Instead of .encode('utf-8'), use .encode('latin-1').
Can you check encoding by following method:
>>> import sys
>>> sys.getdefaultencoding()
'utf-8'
>>>
If encoding is ascii then set to utf-8
open following file(I am using Python 2.7):
/usr/lib/python2.7/sitecustomize.pythen update following to
utf-8sys.setdefaultencoding("utf-8")
[Edit 2]
Can you add following in tour code(at start) and then check:-
>>> try:
... import apport_python_hook
... except ImportError:
... pass
... else:
... apport_python_hook.install()
...
>>> import sys
>>>
>>> sys.setdefaultencoding("utf-8")
>>>
>>>
This error means that your message is already a unicode object, no decoding needed.
When you are doing:
truestr = unicode(string, 'utf-8')
your variable string is first implicitly converted to str type using default 'ascii' codec. And of course, it fails because your string contains non-ascii characters.
If you want to write string somewhere as UTF-8, use string.encode('utf-8').
Note: I've renamed your str variable to string because of name clash with built-in str type. Naming variable str (or int, or float, etc.) is a very bad style.
In Python 2
>>> plain_string = "Hi!"
>>> unicode_string = u"Hi!"
>>> type(plain_string), type(unicode_string)
(<type 'str'>, <type 'unicode'>)
^ This is the difference between a byte string (plain_string) and a unicode string.
>>> s = "Hello!"
>>> u = unicode(s, "utf-8")
^ Converting to unicode and specifying the encoding.
In Python 3
All strings are unicode. The unicode function does not exist anymore. See answer from @Noumenon
If the methods above don't work, you can also tell Python to ignore portions of a string that it can't convert to utf-8:
stringnamehere.decode('utf-8', 'ignore')
Here is a simpler method (hack) that gives you back the setdefaultencoding() function that was deleted from sys:
import sys
# sys.setdefaultencoding() does not exist, here!
reload(sys) # Reload does the trick!
sys.setdefaultencoding('UTF8')
(Note for Python 3.4+: reload() is in the importlib library.)
This is not a safe thing to do, though: this is obviously a hack, since sys.setdefaultencoding() is purposely removed from sys when Python starts. Reenabling it and changing the default encoding can break code that relies on ASCII being the default (this code can be third-party, which would generally make fixing it impossible or dangerous).
PS: This hack doesn't seem to work with Python 3.9 anymore.
If you get this error when you try to pipe/redirect output of your script
UnicodeEncodeError: 'ascii' codec can't encode characters in position 0-5: ordinal not in range(128)
Just export PYTHONIOENCODING in console and then run your code.
export PYTHONIOENCODING=utf8