Just to give a full picture of what megazord.py would look like, using @Jeffrey Harris suggestion to use a nice library for parsing the inputs.
import argparse
def main():
''' Example of taking inputs for megazord bin'''
parser = argparse.ArgumentParser(prog='my_megazord_program')
parser.add_argument('-i', nargs='?', help='help for -i blah')
parser.add_argument('-d', nargs='?', help='help for -d blah')
parser.add_argument('-v', nargs='?', help='help for -v blah')
parser.add_argument('-w', nargs='?', help='help for -w blah')
args = parser.parse_args()
collected_inputs = {'i': args.i,
'd': args.d,
'v': args.v,
'w': args.w}
print 'got input: ', collected_inputs
And with using it like in the above, one would get
$ megazord -i input -d database -v xx-xx -w yy-yy
got input: {'i': 'input', 'd': 'database', 'w': 'yy-yy', 'v': 'xx-xx'}
And since they are all optional arguments,
$ megazord
got input: {'i': None, 'd': None, 'w': None, 'v': None}
Answer from HeyWatchThis on Stack Overflowpython - Setuptools not passing arguments for entry_points - Stack Overflow
setuptools - Explain Python entry points? - Stack Overflow
argparse - Preferred way to expand a command line script to be used as a library in Python - Software Engineering Stack Exchange
python - How to add an entry_point/console script to setup.py that uses arguments - Stack Overflow
Just to give a full picture of what megazord.py would look like, using @Jeffrey Harris suggestion to use a nice library for parsing the inputs.
import argparse
def main():
''' Example of taking inputs for megazord bin'''
parser = argparse.ArgumentParser(prog='my_megazord_program')
parser.add_argument('-i', nargs='?', help='help for -i blah')
parser.add_argument('-d', nargs='?', help='help for -d blah')
parser.add_argument('-v', nargs='?', help='help for -v blah')
parser.add_argument('-w', nargs='?', help='help for -w blah')
args = parser.parse_args()
collected_inputs = {'i': args.i,
'd': args.d,
'v': args.v,
'w': args.w}
print 'got input: ', collected_inputs
And with using it like in the above, one would get
$ megazord -i input -d database -v xx-xx -w yy-yy
got input: {'i': 'input', 'd': 'database', 'w': 'yy-yy', 'v': 'xx-xx'}
And since they are all optional arguments,
$ megazord
got input: {'i': None, 'd': None, 'w': None, 'v': None}
The setuptools console_scripts entry point wants a function of no arguments.
Happily, optparse (Parser for command line options) doesn't need to be passed any arguments, it will read in sys.argv[1:] and use that as it's input.
An "entry point" is typically a function (or other callable function-like object) that a developer or user of your Python package might want to use, though a non-callable object can be supplied as an entry point as well (as correctly pointed out in the comments!).
The most popular kind of entry point is the console_scripts entry point, which points to a function that you want made available as a command-line tool to whoever installs your package. This goes into your setup.py script like:
entry_points={
'console_scripts': [
'cursive = cursive.tools.cmd:cursive_command',
],
},
I have a package I've just deployed called cursive.tools, and I wanted it to make available a "cursive" command that someone could run from the command line, like:
$ cursive --help
usage: cursive ...
The way to do this is define a function, like maybe a cursive_command function in the file cursive/tools/cmd.py that looks like:
def cursive_command():
args = sys.argv[1:]
if len(args) < 1:
print "usage: ..."
and so forth; it should assume that it's been called from the command line, parse the arguments that the user has provided, and ... well, do whatever the command is designed to do.
Install the docutils package for a great example of entry-point use: it will install something like a half-dozen useful commands for converting Python documentation to other formats.
EntryPoints provide a persistent, filesystem-based object name registration and name-based direct object import mechanism (implemented by the setuptools package).
They associate names of Python objects with free-form identifiers. So any other code using the same Python installation and knowing the identifier can access an object with the associated name, no matter where the object is defined. The associated names can be any names existing in a Python module; for example name of a class, function or variable. The entry point mechanism does not care what the name refers to, as long as it is importable.
As an example, let's use (the name of) a function, and an imaginary python module with a fully-qualified name 'myns.mypkg.mymodule':
def the_function():
"function whose name is 'the_function', in 'mymodule' module"
print "hello from the_function"
Entry points are registered via an entry points declaration in setup.py. To register the_function under entrypoint called 'my_ep_func':
entry_points = {
'my_ep_group_id': [
'my_ep_func = myns.mypkg.mymodule:the_function'
]
},
As the example shows, entry points are grouped; there's corresponding API to look up all entry points belonging to a group (example below).
Upon a package installation (ie. running 'python setup.py install'), the above declaration is parsed by setuptools. It then writes the parsed information in special file. After that, the pkg_resources API (part of setuptools) can be used to look up the entry point and access the object(s) with the associated name(s):
import pkg_resources
named_objects = {}
for ep in pkg_resources.iter_entry_points(group='my_ep_group_id'):
named_objects.update({ep.name: ep.load()})
Here, setuptools read the entry point information that was written in special files. It found the entry point, imported the module (myns.mypkg.mymodule), and retrieved the_function defined there, upon call to pkg_resources.load().
Calling the_function would then be simple:
>>> named_objects'my_ep_func'
hello from the_function
Thus, while perhaps a bit difficult to grasp at first, the entry point mechanism is actually quite simple to use. It provides an useful tool for pluggable Python software development.
If there is not a good reason to not do so, I would definitely advocate a spin on option 3. As @jonsharp mentions, breaking up your utility into clean units of functionality is a good way to ensure testability. Even the smallest scripts can eventually morph into a much larger program and making sure that you have an extensible API sooner rather than later will alleviate much headache down the road.
The way I'd approach this is:
- Break up your code into logical methods with clean and clear I/O
- Add unit tests. Having them is never a bad thing.
- Rather than using
if __name__ == '__main__', create amain()(or similar) method containing your entry point - Use setuptool's
setup()function to define the script entry point in your setup.py file.
For example:
from setuptools import setup
setup(
name='mypackage',
version='0.1',
entry_points={
'console_scripts': [ 'myscript = mypackage.mymodule:main' ],
}
)
Now, not only is all of your code (including main()) is easily unit testable, but you can still have your console entry point once you've done a python setup.py install|develop.
Any validation I do using, say, ArgumentParser.add_mutually_exclusive_group will not apply when my tool is run as a library instead of a command line script.
Depending on how your API is designed, you may need to add some extra validation to input parameters, but that should likely be there to prevent unexpected input anyways.
Edit: The only time I would use generally use subprocess is when I'm calling into a non-Python application or another Python script that I don't own or have the time to refactor, but the latter only being as a last resort. Most well-written Python utilities will expose both command line utilities and internal API.
My guess is that you can pursue #3 by way of #2 and that it won't require anything close to a total refactor. Often in these situations, you just need to make a few adjustments at the entry points and (sometimes) exit points.
If needed, wrap your current script in a
main()function.def main(args, stdin = None): if stdin is None: stdin = sys.stdin # full script here if __name__ == '__main__': main(sys.argv[1:])That function should take a list of strings. When you parse command-line options, operate on
args, notsys.argv.Similarly, adjust other parts of your script to avoid direct operations on
sys.stdinand related streams. Instead, use the arguments passed to yourmain()function.
Programatic users (as opposed to command-line users) will pass in a list of strings and any open file handles (or iterables) they want the code to use. Later, if needed, you can make the programatic use more natural by writing a function that knows how to convert typical positional and keyword-style arguments into the list of strings that your option parser needs.
The section must be [options.entry_points]. See an example at https://github.com/github/octodns/blob/4b44ab14b1f0a52f1051c67656d6e3dd6f0ba903/setup.cfg#L34
[options.entry_points]
console_scripts =
octodns-compare = octodns.cmds.compare:main
octodns-dump = octodns.cmds.dump:main
octodns-report = octodns.cmds.report:main
octodns-sync = octodns.cmds.sync:main
octodns-validate = octodns.cmds.validate:main
With python 3.6 and setuptools 39.0.1, i had like you to move entry points from .py to .cfg.
My setup.py, focused on the entry points declaration:
from setuptools import setup
setup(
entry_points={
'pytest11': [
'mytest = mytest.plugin',
],
},
)
I ended up with this working setup.cfg:
... # other standard declarations
[options.entry_points]
pytest11 =
mytest = mytest.plugin
Instead of using the entry_points argument to setuptools.setup(), I'm as of recently supposed to put the file entry_points.txt into the .dist-info directory. But I don't understand how to do that because that directory gets automatically created during the package installation.
[EDIT]
I found that I have to make a [project.scripts] section in pyproject.toml, which upon installation creates the appropriate .dist-info/entry_points.txt. But it still doesn't install the actual callable command.
[EDIT 2]
I found that I made a totally stupid mistake and my executable ended up being named console_scripts. Please don't answer any more. I've downvoted my own OP.