The docs for the (awesome) Click package suggest a few reasons to use entry points instead of scripts, including
- cross-platform compatibility and
- avoiding having the interpreter assign
__name__to__main__, which could cause code to be imported twice (if another module imports your script)
Click is a nice way to implement functions for use as entry_points, btw.
The docs for the (awesome) Click package suggest a few reasons to use entry points instead of scripts, including
- cross-platform compatibility and
- avoiding having the interpreter assign
__name__to__main__, which could cause code to be imported twice (if another module imports your script)
Click is a nice way to implement functions for use as entry_points, btw.
One key difference between these two ways of creating command line executables is that with the setuptools approach (your first example), you have to call a function inside of the script -- in your case this is the func inside of your module. However, in the distutils approach (your second example) you call the script directly (which allows being listed with or without an extension).
Console_scripts entrypoints hidden behind extras are always installed
deployment - How can I use setuptools to generate a console_scripts entry point which calls `python -m mypackage`? - Stack Overflow
python - How to set the bin scripts entry point in `setup.py`? - Stack Overflow
entry_points console_scripts not working in setuptools >=69 if `dynamic` is not configured
If there is not a good reason to not do so, I would definitely advocate a spin on option 3. As @jonsharp mentions, breaking up your utility into clean units of functionality is a good way to ensure testability. Even the smallest scripts can eventually morph into a much larger program and making sure that you have an extensible API sooner rather than later will alleviate much headache down the road.
The way I'd approach this is:
- Break up your code into logical methods with clean and clear I/O
- Add unit tests. Having them is never a bad thing.
- Rather than using
if __name__ == '__main__', create amain()(or similar) method containing your entry point - Use setuptool's
setup()function to define the script entry point in your setup.py file.
For example:
from setuptools import setup
setup(
name='mypackage',
version='0.1',
entry_points={
'console_scripts': [ 'myscript = mypackage.mymodule:main' ],
}
)
Now, not only is all of your code (including main()) is easily unit testable, but you can still have your console entry point once you've done a python setup.py install|develop.
Any validation I do using, say, ArgumentParser.add_mutually_exclusive_group will not apply when my tool is run as a library instead of a command line script.
Depending on how your API is designed, you may need to add some extra validation to input parameters, but that should likely be there to prevent unexpected input anyways.
Edit: The only time I would use generally use subprocess is when I'm calling into a non-Python application or another Python script that I don't own or have the time to refactor, but the latter only being as a last resort. Most well-written Python utilities will expose both command line utilities and internal API.
My guess is that you can pursue #3 by way of #2 and that it won't require anything close to a total refactor. Often in these situations, you just need to make a few adjustments at the entry points and (sometimes) exit points.
If needed, wrap your current script in a
main()function.def main(args, stdin = None): if stdin is None: stdin = sys.stdin # full script here if __name__ == '__main__': main(sys.argv[1:])That function should take a list of strings. When you parse command-line options, operate on
args, notsys.argv.Similarly, adjust other parts of your script to avoid direct operations on
sys.stdinand related streams. Instead, use the arguments passed to yourmain()function.
Programatic users (as opposed to command-line users) will pass in a list of strings and any open file handles (or iterables) they want the code to use. Later, if needed, you can make the programatic use more natural by writing a function that knows how to convert typical positional and keyword-style arguments into the list of strings that your option parser needs.
An "entry point" is typically a function (or other callable function-like object) that a developer or user of your Python package might want to use, though a non-callable object can be supplied as an entry point as well (as correctly pointed out in the comments!).
The most popular kind of entry point is the console_scripts entry point, which points to a function that you want made available as a command-line tool to whoever installs your package. This goes into your setup.py script like:
entry_points={
'console_scripts': [
'cursive = cursive.tools.cmd:cursive_command',
],
},
I have a package I've just deployed called cursive.tools, and I wanted it to make available a "cursive" command that someone could run from the command line, like:
$ cursive --help
usage: cursive ...
The way to do this is define a function, like maybe a cursive_command function in the file cursive/tools/cmd.py that looks like:
def cursive_command():
args = sys.argv[1:]
if len(args) < 1:
print "usage: ..."
and so forth; it should assume that it's been called from the command line, parse the arguments that the user has provided, and ... well, do whatever the command is designed to do.
Install the docutils package for a great example of entry-point use: it will install something like a half-dozen useful commands for converting Python documentation to other formats.
EntryPoints provide a persistent, filesystem-based object name registration and name-based direct object import mechanism (implemented by the setuptools package).
They associate names of Python objects with free-form identifiers. So any other code using the same Python installation and knowing the identifier can access an object with the associated name, no matter where the object is defined. The associated names can be any names existing in a Python module; for example name of a class, function or variable. The entry point mechanism does not care what the name refers to, as long as it is importable.
As an example, let's use (the name of) a function, and an imaginary python module with a fully-qualified name 'myns.mypkg.mymodule':
def the_function():
"function whose name is 'the_function', in 'mymodule' module"
print "hello from the_function"
Entry points are registered via an entry points declaration in setup.py. To register the_function under entrypoint called 'my_ep_func':
entry_points = {
'my_ep_group_id': [
'my_ep_func = myns.mypkg.mymodule:the_function'
]
},
As the example shows, entry points are grouped; there's corresponding API to look up all entry points belonging to a group (example below).
Upon a package installation (ie. running 'python setup.py install'), the above declaration is parsed by setuptools. It then writes the parsed information in special file. After that, the pkg_resources API (part of setuptools) can be used to look up the entry point and access the object(s) with the associated name(s):
import pkg_resources
named_objects = {}
for ep in pkg_resources.iter_entry_points(group='my_ep_group_id'):
named_objects.update({ep.name: ep.load()})
Here, setuptools read the entry point information that was written in special files. It found the entry point, imported the module (myns.mypkg.mymodule), and retrieved the_function defined there, upon call to pkg_resources.load().
Calling the_function would then be simple:
>>> named_objects'my_ep_func'
hello from the_function
Thus, while perhaps a bit difficult to grasp at first, the entry point mechanism is actually quite simple to use. It provides an useful tool for pluggable Python software development.