Use chardet library. It is super easy
import chardet
the_encoding = chardet.detect('your string')['encoding']
and that's it!
in python3 you need to provide type bytes or bytearray so:
import chardet
the_encoding = chardet.detect(b'your string')['encoding']
Answer from george on Stack OverflowUse chardet library. It is super easy
import chardet
the_encoding = chardet.detect('your string')['encoding']
and that's it!
in python3 you need to provide type bytes or bytearray so:
import chardet
the_encoding = chardet.detect(b'your string')['encoding']
if your files either in cp1252 and utf-8, then there is an easy way.
import logging
def force_decode(string, codecs=['utf8', 'cp1252']):
for i in codecs:
try:
return string.decode(i)
except UnicodeDecodeError:
pass
logging.warn("cannot decode url %s" % ([string]))
for item in os.listdir(rootPath):
#Convert to Unicode
if isinstance(item, str):
item = force_decode(item)
print item
otherwise, there is a charset detect lib.
Python - detect charset and convert to utf-8
https://pypi.python.org/pypi/chardet
EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative
Correctly detecting the encoding all times is impossible.
(From chardet FAQ:)
However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.
There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.
You can also use UnicodeDammit. It will try the following methods:
- An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
- An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
- An encoding sniffed by the chardet library, if you have it installed.
- UTF-8
- Windows-1252
Another option for working out the encoding is to use libmagic (which is the code behind the file command). There are a profusion of Python bindings available.
The Python bindings that live in the file source tree are available as the python-magic (or python3-magic) debian package. It can determine the encoding of a file by doing:
import magic
blob = open('unknown-file', 'rb').read()
m = magic.open(magic.MAGIC_MIME_ENCODING)
m.load()
encoding = m.buffer(blob) # "utf-8", "us-ascii", etc.
There is an identically named, but incompatible, python-magic pip package on PyPI that also uses libmagic. It can also get the encoding, by doing:
import magic
m = magic.Magic(mime_encoding=True)
encoding = m.from_file('unknown-file')
» pip install lib-detect-encoding
You can use the chardet package. Read this tutorial.
If you are using Ubuntu:
sudo apt-get install python3-chardet
If you are using pip:
pip install chardet2
Since you've entered it from the console, the encoding will be sys.stdin.encoding
>>> name = '深入 damon'
>>> import sys
>>> sys.stdin.encoding
'UTF-8'
>>> b1 = name.decode(sys.stdin.encoding)
>>> b1
u'\u6df1\u5165 damon'
>>> b1.encode(sys.stdin.encoding)
'\xe6\xb7\xb1\xe5\x85\xa5 damon'
>>> print b1.encode(sys.stdin.encoding)
深入 damon
Files generally indicate their encoding with a file header. There are many examples here. However, even reading the header you can never be sure what encoding a file is really using.
For example, a file with the first three bytes 0xEF,0xBB,0xBF is probably a UTF-8 encoded file. However, it might be an ISO-8859-1 file which happens to start with the characters . Or it might be a different file type entirely.
Notepad++ does its best to guess what encoding a file is using, and most of the time it gets it right. Sometimes it does get it wrong though - that's why that 'Encoding' menu is there, so you can override its best guess.
For the two encodings you mention:
- The "UCS-2 Little Endian" files are UTF-16 files (based on what I understand from the info here) so probably start with
0xFF,0xFEas the first 2 bytes. From what I can tell, Notepad++ describes them as "UCS-2" since it doesn't support certain facets of UTF-16. - The "UTF-8 without BOM" files don't have any header bytes. That's what the "without BOM" bit means.
You cannot. If you could do that, there would not be so many web sites or text files with “random gibberish” out there. That's why the encoding is usually sent along with the payload as meta data.
In case it's not, all you can do is a “smart guess” but the result is often ambiguous since the same byte sequence might be valid in several encodings.