EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
Answer from nosklo on Stack Overflow
Top answer
1 of 16
292

EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
2 of 16
104

Another option for working out the encoding is to use libmagic (which is the code behind the file command). There are a profusion of Python bindings available.

The Python bindings that live in the file source tree are available as the python-magic (or python3-magic) debian package. It can determine the encoding of a file by doing:

import magic

blob = open('unknown-file', 'rb').read()
m = magic.open(magic.MAGIC_MIME_ENCODING)
m.load()
encoding = m.buffer(blob)  # "utf-8", "us-ascii", etc.

There is an identically named, but incompatible, python-magic pip package on PyPI that also uses libmagic. It can also get the encoding, by doing:

import magic

m = magic.Magic(mime_encoding=True)
encoding = m.from_file('unknown-file')
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-a-text-file-with-python
Detect Encoding of a Text file with Python - GeeksforGeeks
April 8, 2026 - Python provides the chardet library, which can automatically detect a file’s encoding.
🌐
DEV Community
dev.to › bowmanjd › character-encodings-and-detection-with-python-chardet-and-cchardet-4hj7
Character Encodings and Detection with Python, chardet, and cchardet - DEV Community
February 9, 2026 - If you do not know what the character encoding is for a file you need to handle in Python, then try chardet. ... Use something like the above to install it in your Python virtual environment. Character detection with chardet works something like this:
🌐
Readthedocs
unicodebook.readthedocs.io › guess_encoding.html
8. How to guess the encoding of a document? — Programming with Unicode
If you would like to reject surrogate characters in Python 2, use the following strict function: def isUTF8Strict(data): try: decoded = data.decode('UTF-8') except UnicodeDecodeError: return False else: for ch in decoded: if 0xD800 <= ord(ch) <= 0xDFFF: return False return True · PHP has a builtin function to detect the encoding of a byte string: mb_detect_encoding().
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-csv-file-in-python
Detect Encoding of CSV File in Python - GeeksforGeeks
July 23, 2025 - In this example, below Python code below utilizes the chardet library to automatically detect the encoding of a CSV file. It opens the file in binary mode, reads its content, and employs chardet.detect() to determine the encoding.
🌐
Krinkere
krinkere.github.io › krinkersite › encoding_csv_file_python.html
How to detect encoding of CSV file in python - Cloud. Big Data. Analytics... and so on
March 30, 2018 - # look at the first ten thousand bytes to guess the character encoding with open("my_data.csv", 'rb') as rawdata: result = chardet.detect(rawdata.read(10000)) # check what the character encoding might be print(result)
Find elsewhere
🌐
GeeksforGeeks
geeksforgeeks.org › python › python-character-encoding
Python | Character Encoding - GeeksforGeeks
May 29, 2021 - This module can be simply installed using sudo easy_install charade or pip install charade. Let's see the wrapper function around the charade module. Code : encoding.detect(string), to detect the encoding
🌐
Kaggle
kaggle.com › code › rtatman › automatically-detecting-character-encodings
Automatically detecting character encodings | Kaggle
December 16, 2017 - Explore and run AI code with Kaggle Notebooks | Using data from Character Encoding Examples
🌐
Rip Tutorial
riptutorial.com › how to detect the encoding of a text file with python?
encoding Tutorial => How to detect the encoding of a text file with...
import chardet rawdata = open(file, "r").read() result = chardet.detect(rawdata) charenc = result['encoding']
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
🌐
LinkedIn
linkedin.com › learning › effective-serialization-with-python › detect-encoding
Detect encoding - Python Video Tutorial | LinkedIn Learning, formerly Lynda.com
August 25, 2020 - And we see that LinkedIn is sending ... not be that lucky, and you need to guess the encoding. You can use the external chardet package to guess the encoding. So python-m pip install chardet....
🌐
Plain English
python.plainenglish.io › character-encoding-detection-made-easy-with-chardet-in-python-dab2ac258840
Character Encoding Detection Made Easy with Chardet in Python | by Yancy Dennis | Python in Plain English
February 7, 2023 - Chardet is an essential tool for data analysis and data processing in Python. Using Chardet is straightforward. The library can be installed using the pip package manager with the following command: ... import chardet def detect_encoding(file): detector = chardet.universaldetector.UniversalDetector() with open(file, "rb") as f: for line in f: detector.feed(line) if detector.done: break…
🌐
Medium
cloudmersive.medium.com › how-to-detect-the-encoding-of-a-text-file-in-python-7a21989acd5b
How to Detect the Encoding of a Text File in Python | by Cloudmersive | Medium
January 9, 2023 - from __future__ import print_function import time import cloudmersive_convert_api_client from cloudmersive_convert_api_client.rest import ApiException from pprint import pprint # Configure API key authorization: Apikey configuration = cloudmersive_convert_api_client.Configuration() configuration.api_key['Apikey'] = 'YOUR_API_KEY' # create an instance of the API class api_instance = cloudmersive_convert_api_client.EditTextApi(cloudmersive_convert_api_client.ApiClient(configuration)) input_file = '/path/to/inputfile' # file | Input file to perform the operation on. try: # Detect text encoding of file api_response = api_instance.edit_text_text_encoding_detect(input_file) pprint(api_response) except ApiException as e: print("Exception when calling EditTextApi->edit_text_text_encoding_detect: %s\n" % e)
🌐
Baeldung
baeldung.com › home › files › how to auto-detect encoding of a text file in linux
How to Auto-Detect Encoding of a Text File in Linux | Baeldung on Linux
March 18, 2024 - In Linux, we can use the chardet library to auto-detect the encoding of a text file. It’s a character encoding auto-detection tool in Python.
🌐
GeeksforGeeks
geeksforgeeks.org › python › character-encoding-detection-with-chardet-in-python
Character Encoding Detection With Chardet in Python - GeeksforGeeks
July 23, 2025 - In this example, the Python script uses the chardet library to detect the character encoding of a given byte sequence (data).
🌐
Alamot
alamot.github.io › test_encodings
How to determine the character encoding of a text file - The Portal of Knowledge
October 21, 2016 - There is an automatic universal encoding detector for Python 2 and 3 here: https://pypi.python.org/pypi/chardet.