Use chardet library. It is super easy

import chardet

the_encoding = chardet.detect('your string')['encoding']

and that's it!

in python3 you need to provide type bytes or bytearray so:

import chardet
the_encoding = chardet.detect(b'your string')['encoding']
Answer from george on Stack Overflow
🌐
Readthedocs
unicodebook.readthedocs.io › guess_encoding.html
8. How to guess the encoding of a document? — Programming with Unicode
If you would like to reject surrogate characters in Python 2, use the following strict function: def isUTF8Strict(data): try: decoded = data.decode('UTF-8') except UnicodeDecodeError: return False else: for ch in decoded: if 0xD800 <= ord(ch) <= 0xDFFF: return False return True · PHP has a builtin function to detect the encoding of a byte string: mb_detect_encoding().
🌐
Program Creek
programcreek.com › python
Python detect encoding
def detect_encoding(sample, encoding=None): """Detect encoding of a byte string sample.
🌐
GitHub
github.com › magnetikonline › py-encoding-detect
GitHub - magnetikonline/py-encoding-detect: Python module for detecting common text file encodings. · GitHub
If odd count below negative threshold and even count above positive threshold, return result of UTF-16LE. A detection of sample files with various encoding formats can be run via test/detect.py.
Author: magnetikonline
Top answer
1 of 16
292

EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
2 of 16
104

Another option for working out the encoding is to use libmagic (which is the code behind the file command). There are a profusion of Python bindings available.

The Python bindings that live in the file source tree are available as the python-magic (or python3-magic) debian package. It can determine the encoding of a file by doing:

import magic

blob = open('unknown-file', 'rb').read()
m = magic.open(magic.MAGIC_MIME_ENCODING)
m.load()
encoding = m.buffer(blob)  # "utf-8", "us-ascii", etc.

There is an identically named, but incompatible, python-magic pip package on PyPI that also uses libmagic. It can also get the encoding, by doing:

import magic

m = magic.Magic(mime_encoding=True)
encoding = m.from_file('unknown-file')
🌐
DEV Community
dev.to › bowmanjd › character-encodings-and-detection-with-python-chardet-and-cchardet-4hj7
Character Encodings and Detection with Python, chardet, and cchardet - DEV Community
February 9, 2026 - Try the above print statement in a Python console or script and you should see our beloved "spam". It was automatically decoded in the Python console, printing the corresponding letters (characters). But let's be more explicit, creating a byte string of the above numbers, and specifying the ASCII encoding:
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
While the Windows cp1252 encoding is also sometimes referred to as “latin-1”, it doesn’t map all possible byte values, and thus needs to be used in combination with the surrogateescape error handler to ensure it never throws UnicodeDecodeError. The latin-1 encoding in Python implements ISO_8859-1:1987 which maps all possible byte values to the first 256 Unicode code points, and thus ensures decoding errors will never occur regardless of the configured error handler.
🌐
Python
docs.python.org › 3 › library › codecs.html
codecs — Codec registry and base classes
However that’s not possible with UTF-8, as UTF-8 byte sequences have a structure that doesn’t allow arbitrary byte sequences. To increase the reliability with which a UTF-8 encoding can be detected, Microsoft invented a variant of UTF-8 (that Python calls "utf-8-sig") for its Notepad program: Before any of the Unicode characters is written to the file, a UTF-8 encoded BOM (which looks like this as a byte sequence: 0xef, 0xbb, 0xbf) is written.
🌐
GeeksforGeeks
geeksforgeeks.org › python › python-character-encoding
Python | Character Encoding - GeeksforGeeks
May 29, 2021 - detect() : It is a charade.detect() wrapper. It encodes the strings and handles the UnicodeDecodeError exceptions. It expects a bytes object so therefore the string is encoded before trying to detect the encoding. convert() : It is a ...
Find elsewhere
🌐
LinkedIn
linkedin.com › learning › effective-serialization-with-python › detect-encoding
Detect encoding - Python Video Tutorial | LinkedIn Learning, formerly Lynda.com
August 25, 2020 - And we see that LinkedIn is sending ... not be that lucky, and you need to guess the encoding. You can use the external chardet package to guess the encoding. So python-m pip install chardet....
🌐
Runebook.dev
runebook.dev › en › docs › python › library › tokenize › tokenize.detect_encoding
The Right Way to Detect Python Source Encoding (and When to Use chardet Instead)
The tokenize.detect_encoding(readline) function, part of Python's built-in tokenize module, is specifically designed to find the source file encoding of Python code. It does this by checking for a UTF-8 Byte Order Mark (BOM) or an encoding cookie ...
🌐
sqlpey
sqlpey.com › python › detect-file-encoding-python
How to Detect File Encoding in Python: Methods and Solutions …
July 22, 2025 - A BOM is a special sequence of bytes that indicates the encoding and byte order. import sys import codecs def detect_bom_encoding(file_path: str) -> str | None: """ Returns the encoding string for the open() function based on BOM detection.
🌐
PyPI
pypi.org › project › lib-detect-encoding
lib-detect-encoding · PyPI
October 13, 2023 - Works on posix, windows and WINE ... https://docs.python.org/3/library/codecs.html#standard-encodings """ def get_file_encoding(raw_bytes: bytes) -> str: """ returns the encoding for the raw_bytes passed. if the confidence of the ...
      » pip install lib-detect-encoding
    
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-a-text-file-with-python
Detect Encoding of a Text file with Python - GeeksforGeeks
April 8, 2026 - Python provides the chardet library, which can automatically detect a file’s encoding. It works by analyzing the statistical patterns of byte sequences to estimate the most likely encoding.
🌐
WebLearnModernPython
web.learnmodernpython.com › home › detect file encoding: build your own tool
Detect File Encoding: Build Your Own Tool
April 23, 2026 - Use Libraries: For robust encoding detection, leverage libraries like `chardet` or `cchardet` in Python. These libraries use sophisticated statistical analysis and larger sample sizes to achieve higher accuracy.
🌐
pythontutorials
pythontutorials.net › blog › how-to-know-the-encoding-of-a-file-in-python
How to Automatically Detect File Encoding in Python: A Practical Guide — pythontutorials.net
Scale: Processing hundreds of files (e.g., in data pipelines) requires automation. Global Data: Encodings vary by region (e.g., Shift-JIS for Japanese, GB2312 for Chinese). Automatic detection tools analyze byte patterns to guess the encoding, saving time and reducing errors. Python offers three leading libraries for encoding detection.
🌐
Python Module of the Week
pymotw.com › 2 › codecs
codecs – String encoding and decoding - Python Module of the Week
$ python codecs_bom.py BOM : fffe BOM_BE : feff BOM_LE : fffe BOM_UTF8 : efbb bf BOM_UTF16 : fffe BOM_UTF16_BE : feff BOM_UTF16_LE : fffe BOM_UTF32 : fffe 0000 BOM_UTF32_BE : 0000 feff BOM_UTF32_LE : fffe 0000 · Byte ordering is detected and handled automatically by the decoders in codecs, but ...
🌐
Bytes
bytes.com › home › forum › topic › python
Detect character encoding - Post.Byes
December 4, 2005 - This doesn't really help much by > itself, though.[/color] ----- test.py for enc in ["cp1250", "latin1", "iso-8859-2"]: print enc try: str.decode("".j oin([chr(i) for i in xrange(256)]), enc) except UnicodeDecodeEr ror, e: print e ----- 192:~ deets$ python2.4 /tmp/test.py cp1250 'charmap' codec can't decode byte 0x81 in position 129: character maps to <undefined> latin1 iso-8859-2 So cp1250 doesn't have all codepoints defined - but the others have. Sure, this helps you to eliminate 1 of the three choices the OP wanted to choose between - but how many texts you have that have a 129 in them? Regards, Diez ... Re: Detect character encoding [Diez B.