EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
Answer from nosklo on Stack Overflow
Top answer
1 of 16
292

EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
2 of 16
104

Another option for working out the encoding is to use libmagic (which is the code behind the file command). There are a profusion of Python bindings available.

The Python bindings that live in the file source tree are available as the python-magic (or python3-magic) debian package. It can determine the encoding of a file by doing:

import magic

blob = open('unknown-file', 'rb').read()
m = magic.open(magic.MAGIC_MIME_ENCODING)
m.load()
encoding = m.buffer(blob)  # "utf-8", "us-ascii", etc.

There is an identically named, but incompatible, python-magic pip package on PyPI that also uses libmagic. It can also get the encoding, by doing:

import magic

m = magic.Magic(mime_encoding=True)
encoding = m.from_file('unknown-file')
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-a-text-file-with-python
Detect Encoding of a Text file with Python - GeeksforGeeks
April 8, 2026 - Python provides the chardet library, which can automatically detect a file’s encoding.
Discussions

How to know the encoding of a file in Python? - Stack Overflow
Does anybody know how to get the encoding of a file in Python. I know that you can use the codecs module to open a file with a specific encoding but you have to know it in advance. import codecs f = More on stackoverflow.com
🌐 stackoverflow.com
How to determine encoding of text file?
No it's the em dash before the copyright symbol. The encoding that you're trying to decode is actually cp1252. Strangely enough, the encoding of the page you linked to is actually utf-8: from urllib import urlopen x = urlopen("http://www.thrivenotes.com/the-last-answer/").read().decode("utf-8") If the content of that URI wasn't utf-8, that code would raise an error. What probably happened was when you copied and pasted the text, your text editor probably saved the text file using windows-1252. There is the chardet but generally it's pretty easy to guess what the encoding is. If it's western content, it's going to most likely be two different encodings, utf-8 or windows-1252. With utf-8, all bytes with the high bit set signal the start of a multi-byte sequence. Those bytes are reserved in both latin-1 (iso-8859-1) and the corresponding Unicode code points are reserved. Unfortunately for developers everywhere the other most common encoding, windows-1252 used those reserved bytes for really common characters in publishing such as en and em dashes. More on reddit.com
🌐 r/Python
9
2
November 30, 2010
python - How do I detect if a file is encoded using UTF-8? - Stack Overflow
Is there a way to recognize if text file is UTF-8 in Python? I would really like to get if the file is UTF-8 or not. I don't need to detect other encodings. More on stackoverflow.com
🌐 stackoverflow.com
The use of open(encoding="utf-8")
open defaults to whatever encoding your system uses by default, so it can be anything from ASCII to ISO-8859-1. You can, however, make it use a specific encoding instead. utf-8 is useful as it has most characters. More on reddit.com
🌐 r/learnpython
16
3
November 2, 2020
🌐
Readthedocs
unicodebook.readthedocs.io › guess_encoding.html
8. How to guess the encoding of a document? — Programming with Unicode
Example in Python getting the BOMs from the codecs library: from codecs import BOM_UTF8, BOM_UTF16_BE, BOM_UTF16_LE, BOM_UTF32_BE, BOM_UTF32_LE BOMS = ( (BOM_UTF8, "UTF-8"), (BOM_UTF32_BE, "UTF-32-BE"), (BOM_UTF32_LE, "UTF-32-LE"), (BOM_UTF16_BE, "UTF-16-BE"), (BOM_UTF16_LE, "UTF-16-LE"), ) def check_bom(data): return [encoding for bom, encoding in BOMS if data.startswith(bom)]
🌐
Kaggle
kaggle.com › code › rtatman › automatically-detecting-character-encodings
Automatically detecting character encodings | Kaggle
December 16, 2017 - Explore and run AI code with Kaggle Notebooks | Using data from Character Encoding Examples
🌐
Program Creek
programcreek.com › python
Python detect encoding
def open_file_detect_encoding(...th).get('encoding', 'utf-8')) else: return open(file_path, "r") ... def detect_encoding(readline): """ The detect_encoding() function is used to detect the encoding that should be used to decode a Python source file....
Find elsewhere
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Approach: first open the file in binary mode to look for the encoding marker, then reopen in text mode with the identified encoding. Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
🌐
Rip Tutorial
riptutorial.com › how to detect the encoding of a text file with python?
encoding Tutorial => How to detect the encoding of a text file with...
Chardet can detect following encodings: ASCII, UTF-8, UTF-16 (2 variants), UTF-32 (4 variants) Big5, GB2312, EUC-TW, HZ-GB-2312, ISO-2022-CN (Traditional and Simplified Chinese) ...
🌐
Alamot
alamot.github.io › test_encodings
How to determine the character encoding of a text file - The Portal of Knowledge
October 21, 2016 - There is an automatic universal encoding detector for Python 2 and 3 here: https://pypi.python.org/pypi/chardet. The chardet detector comes with a command-line script which reports on the encodings of one or more files:
🌐
GitHub
gist.github.com › a8d048083c0634021487eb43c6b17a02
Detect file encoding in python3 with chardet · GitHub
Detect file encoding in python3 with chardet · Raw · detect_encode.py · This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
🌐
Reddit
reddit.com › r/python › how to determine encoding of text file?
r/Python on Reddit: How to determine encoding of text file?
November 30, 2010 -

I'm trying to open a file and read it line by line, however when I call readline() I get UnicodeDecodeError's such as this:

UnicodeDecodeError: 'charmap' codec can't decode byte 0x96 in position 430: unexpected code bytes.

when trying to read this line:

The Last Answer by Isaac Asimov — © 1980

I'm guessing it's because of the ©, so I figured opening it with encoding='utf-8' would help, but I get the same error (except w/utf-8 codec instead obviously). So.. I'm assuming I just need to find out what this is encoded in, so how to do that? This is just something I copied and pasted off of the web into a plain txt file btw, not sure if that matters. It's the last answer if anyone's interested. Great read.

Top answer
1 of 2
24

You mentioned in a comment you only need to detect UTF-8. If you know the alternative consists of only single byte encodings, then there is a solution that often works.

If you know it's either UTF-8 or single byte encoding like latin-1, then try opening it first in UTF-8 and then in the other encoding. If the file contains only ASCII characters, it will end up opened in UTF-8 even if it was intended as the other encoding. If it contains any non-ASCII characters, this will almost always correctly detect the right character set between the two.

try:
    # or codecs.open on Python <= 2.5
    # or io.open on Python > 2.5 and <= 2.7
    filedata = open(filename, encoding='UTF-8').read() 
except:
    filedata = open(filename, encoding='other-single-byte-encoding').read() 

Your best bet is to use the chardet package from PyPI, either directly or through UnicodeDamnit from BeautifulSoup:

chardet 1.0.1

Universal encoding detector

Detects:

  • ASCII, UTF-8, UTF-16 (2 variants), UTF-32 (4 variants)
  • Big5, GB2312, EUC-TW, HZ-GB-2312, ISO-2022-CN (Traditional and Simplified Chinese)
  • EUC-JP, SHIFT_JIS, ISO-2022-JP (Japanese)
  • EUC-KR, ISO-2022-KR (Korean)
  • KOI8-R, MacCyrillic, IBM855, IBM866, ISO-8859-5, windows-1251 (Cyrillic)
  • ISO-8859-2, windows-1250 (Hungarian)
  • ISO-8859-5, windows-1251 (Bulgarian)
  • windows-1252 (English)
  • ISO-8859-7, windows-1253 (Greek)
  • ISO-8859-8, windows-1255 (Visual and Logical Hebrew)
  • TIS-620 (Thai)

Requires Python 2.1 or later

However, some files will be valid in multiple encodings, so chardet is not a panacea.

2 of 2
3

Reliably? No.

In general, a byte sequence has no meaning unless you know how to interpret it -- this goes for text files, but also integers, floating point numbers, etc.

But, there are ways of guessing the encoding of a file, by looking at the byte order mark (if there is one) and the first chunk of the file (to see which encoding yields the most sensible characters). The chardet library is pretty good at this, but be aware it's only a heuristic, albeit a rather powerful one.

🌐
Song Genius API
melaniewalsh.github.io › Intro-Cultural-Analytics › 02-Python › 07-Files-Character-Encoding.html
Files & Character Encoding — Introduction to Cultural Analytics & Python
Character encodings are systems that map characters to numbers. Each character is given a specific ID number. This way, computers can actually read and understand characters. You can check any characters’ “code point,” or place in the Unicode universe, with the function ord()
🌐
YouTube
youtube.com › watch
Python Tutorial - 20 : How to find file encoding? | Detect file encoding | Python chardet library - YouTube
This tutorial talks about finding which encoding is applied on the file using python chardet library
Published: March 19, 2024
🌐
Krinkere
krinkere.github.io › krinkersite › encoding_csv_file_python.html
How to detect encoding of CSV file in python - Cloud. Big Data. Analytics... and so on
March 30, 2018 - # look at the first ten thousand bytes to guess the character encoding with open("my_data.csv", 'rb') as rawdata: result = chardet.detect(rawdata.read(10000)) # check what the character encoding might be print(result) ... So chardet is 73% confidence that the right encoding is "Windows-1252". Now we can use this data to specify encoding type as we trying to read the file