EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
Answer from nosklo on Stack Overflow
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-a-text-file-with-python
Detect Encoding of a Text file with Python - GeeksforGeeks
April 8, 2026 - Below Python code defines a function, 'detect_encoding(file_path), that uses the 'chardet' library to automatically determine the encoding of a text file specified by its path. It reads the file in binary mode, feeds each line to a universal ...
Top answer
1 of 16
292

EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
2 of 16
104

Another option for working out the encoding is to use libmagic (which is the code behind the file command). There are a profusion of Python bindings available.

The Python bindings that live in the file source tree are available as the python-magic (or python3-magic) debian package. It can determine the encoding of a file by doing:

import magic

blob = open('unknown-file', 'rb').read()
m = magic.open(magic.MAGIC_MIME_ENCODING)
m.load()
encoding = m.buffer(blob)  # "utf-8", "us-ascii", etc.

There is an identically named, but incompatible, python-magic pip package on PyPI that also uses libmagic. It can also get the encoding, by doing:

import magic

m = magic.Magic(mime_encoding=True)
encoding = m.from_file('unknown-file')
🌐
Reddit
reddit.com › r/python › how to determine encoding of text file?
r/Python on Reddit: How to determine encoding of text file?
November 30, 2010 -

I'm trying to open a file and read it line by line, however when I call readline() I get UnicodeDecodeError's such as this:

UnicodeDecodeError: 'charmap' codec can't decode byte 0x96 in position 430: unexpected code bytes.

when trying to read this line:

The Last Answer by Isaac Asimov — © 1980

I'm guessing it's because of the ©, so I figured opening it with encoding='utf-8' would help, but I get the same error (except w/utf-8 codec instead obviously). So.. I'm assuming I just need to find out what this is encoded in, so how to do that? This is just something I copied and pasted off of the web into a plain txt file btw, not sure if that matters. It's the last answer if anyone's interested. Great read.

🌐
Krinkere
krinkere.github.io › krinkersite › encoding_csv_file_python.html
How to detect encoding of CSV file in python - Cloud. Big Data. Analytics... and so on
March 30, 2018 - # look at the first ten thousand bytes to guess the character encoding with open("my_data.csv", 'rb') as rawdata: result = chardet.detect(rawdata.read(10000)) # check what the character encoding might be print(result) ... So chardet is 73% confidence that the right encoding is "Windows-1252". Now we can use this data to specify encoding type as we trying to read the file · data = pd.read_csv("my_data.csv", encoding='Windows-1252'))
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-csv-file-in-python
Detect Encoding of CSV File in Python - GeeksforGeeks
July 23, 2025 - Incorrect encoding can lead to data corruption and misinterpretation. By using the chardet library, you can automatically detect the encoding of a CSV file and ensure that it is properly handled during file operations.
🌐
Rip Tutorial
riptutorial.com › how to detect the encoding of a text file with python?
encoding Tutorial => How to detect the encoding of a text file with...
Chardet can detect following encodings: ASCII, UTF-8, UTF-16 (2 variants), UTF-32 (4 variants) Big5, GB2312, EUC-TW, HZ-GB-2312, ISO-2022-CN (Traditional and Simplified Chinese) ...
Find elsewhere
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Approach: first open the file in binary mode to look for the encoding marker, then reopen in text mode with the identified encoding. Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
🌐
DEV Community
dev.to › bowmanjd › character-encodings-and-detection-with-python-chardet-and-cchardet-4hj7
Character Encodings and Detection with Python, chardet, and cchardet - DEV Community
February 9, 2026 - First, I could use chardet.detect() in a one-off fashion on a text file, to determine the first time what the character encoding will be on subsequent engagements. Let's say there is a source system that always exports a CSV file with the same character encoding.
🌐
Baeldung
baeldung.com › home › files › how to auto-detect encoding of a text file in linux
How to Auto-Detect Encoding of a Text File in Linux | Baeldung on Linux
March 18, 2024 - In Linux, we can use the chardet library to auto-detect the encoding of a text file. It’s a character encoding auto-detection tool in Python.
🌐
sqlpey
sqlpey.com › python › detect-file-encoding-python
How to Detect File Encoding in Python: Methods and Solutions …
July 22, 2025 - In cases where automatic tools fail or for more granular control, you can implement manual detection. This involves trying common encodings sequentially. import codecs def guess_encoding_manually(file_path: str, encodings_to_try: list[str] = None) -> str | None: '''Manually guess encoding by trying a list of common codecs.''' if encodings_to_try is None: encodings_to_try = ['utf-8', 'windows-1250', 'windows-1252', 'iso-8859-1'] # Add more as needed for encoding in encodings_to_try: try: with codecs.open(file_path, 'r', encoding=encoding) as f: f.read() # Attempt to read the file return encoding # Success!
🌐
LabEx
labex.io › tutorials › python-how-to-read-files-with-different-encodings-434794
How to read files with different encodings | LabEx
Python 3 natively supports multiple encodings through built-in functions and methods. The open() function allows specifying encoding when reading or writing files. ## Check file encoding import chardet def detect_file_encoding(filename): with ...
🌐
Readthedocs
chardet.readthedocs.io › en › latest › usage.html
Usage - chardet 7.6.1.dev26+gf88488a2d documentation
from chardet import detect, EncodingEra # Default: all encodings considered result = detect(data) # Restrict to modern web encodings only result = detect(data, encoding_era=EncodingEra.MODERN_WEB) # Only legacy ISO encodings result = detect(data, encoding_era=EncodingEra.LEGACY_ISO) ... By default, chardet returns encoding names compatible with chardet 5.x/6.x (e.g., "utf-8", "ascii", "SHIFT_JIS"). Two parameters control how encoding names are returned: compat_names (default True) — map internal Python codec names to chardet 5.x/6.x compatible display names. Set to False to get raw Python codec names (e.g., "shift_jis_2004" instead of "SHIFT_JIS").
🌐
Alamot
alamot.github.io › test_encodings
How to determine the character encoding of a text file - The Portal of Knowledge
October 21, 2016 - There is an automatic universal encoding detector for Python 2 and 3 here: https://pypi.python.org/pypi/chardet. The chardet detector comes with a command-line script which reports on the encodings of one or more files:
🌐
YouTube
youtube.com › watch
Python Tutorial - 20 : How to find file encoding? | Detect file encoding | Python chardet library - YouTube
This tutorial talks about finding which encoding is applied on the file using python chardet library
Published: March 19, 2024
🌐
Runebook.dev
runebook.dev › en › docs › python › library › tokenize › tokenize.detect_encoding
The Right Way to Detect Python Source Encoding (and When to Use chardet Instead)
The tokenize.detect_encoding(readline) function, part of Python's built-in tokenize module, is specifically designed to find the source file encoding of Python code.
🌐
Kaggle
kaggle.com › rtatman › automatically-detecting-character-encodings
Automatically detecting character encodings
Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
pythontutorials
pythontutorials.net › blog › how-to-know-the-encoding-of-a-file-in-python
How to Automatically Detect File Encoding in Python: A Practical Guide — pythontutorials.net
Scale: Processing hundreds of files (e.g., in data pipelines) requires automation. Global Data: Encodings vary by region (e.g., Shift-JIS for Japanese, GB2312 for Chinese). Automatic detection tools analyze byte patterns to guess the encoding, saving time and reducing errors. Python offers three leading libraries for encoding detection.