Use chardet library. It is super easy

import chardet

the_encoding = chardet.detect('your string')['encoding']

and that's it!

in python3 you need to provide type bytes or bytearray so:

import chardet
the_encoding = chardet.detect(b'your string')['encoding']
Answer from george on Stack Overflow
Top answer
1 of 16
292

EDIT: chardet seems to be unmantained but most of the answer applies. Check https://pypi.org/project/charset-normalizer/ for an alternative

Correctly detecting the encoding all times is impossible.

(From chardet FAQ:)

However, some encodings are optimized for specific languages, and languages are not random. Some character sequences pop up all the time, while other sequences make no sense. A person fluent in English who opens a newspaper and finds “txzqJv 2!dasd0a QqdKjvz” will instantly recognize that that isn't English (even though it is composed entirely of English letters). By studying lots of “typical” text, a computer algorithm can simulate this kind of fluency and make an educated guess about a text's language.

There is the chardet library that uses that study to try to detect encoding. chardet is a port of the auto-detection code in Mozilla.

You can also use UnicodeDammit. It will try the following methods:

  • An encoding discovered in the document itself: for instance, in an XML declaration or (for HTML documents) an http-equiv META tag. If Beautiful Soup finds this kind of encoding within the document, it parses the document again from the beginning and gives the new encoding a try. The only exception is if you explicitly specified an encoding, and that encoding actually worked: then it will ignore any encoding it finds in the document.
  • An encoding sniffed by looking at the first few bytes of the file. If an encoding is detected at this stage, it will be one of the UTF-* encodings, EBCDIC, or ASCII.
  • An encoding sniffed by the chardet library, if you have it installed.
  • UTF-8
  • Windows-1252
2 of 16
104

Another option for working out the encoding is to use libmagic (which is the code behind the file command). There are a profusion of Python bindings available.

The Python bindings that live in the file source tree are available as the python-magic (or python3-magic) debian package. It can determine the encoding of a file by doing:

import magic

blob = open('unknown-file', 'rb').read()
m = magic.open(magic.MAGIC_MIME_ENCODING)
m.load()
encoding = m.buffer(blob)  # "utf-8", "us-ascii", etc.

There is an identically named, but incompatible, python-magic pip package on PyPI that also uses libmagic. It can also get the encoding, by doing:

import magic

m = magic.Magic(mime_encoding=True)
encoding = m.from_file('unknown-file')
🌐
GeeksforGeeks
geeksforgeeks.org › python › detect-encoding-of-a-text-file-with-python
Detect Encoding of a Text file with Python - GeeksforGeeks
April 8, 2026 - Below Python code defines a function, 'detect_encoding(file_path), that uses the 'chardet' library to automatically determine the encoding of a text file specified by its path. It reads the file in binary mode, feeds each line to a universal ...
Top answer
1 of 13
325

In Python 3, all strings are sequences of Unicode characters. There is a bytes type that holds raw bytes.

In Python 2, a string may be of type str or of type unicode. You can tell which using code something like this:

def whatisthis(s):
    if isinstance(s, str):
        print "ordinary string"
    elif isinstance(s, unicode):
        print "unicode string"
    else:
        print "not a string"

This does not distinguish "Unicode or ASCII"; it only distinguishes Python types. A Unicode string may consist of purely characters in the ASCII range, and a bytestring may contain ASCII, encoded Unicode, or even non-textual data.

2 of 13
136

How to tell if an object is a unicode string or a byte string

You can use type or isinstance.

In Python 2:

>>> type(u'abc')  # Python 2 unicode string literal
<type 'unicode'>
>>> type('abc')   # Python 2 byte string literal
<type 'str'>

In Python 2, str is just a sequence of bytes. Python doesn't know what its encoding is. The unicode type is the safer way to store text. If you want to understand this more, I recommend http://farmdev.com/talks/unicode/.

In Python 3:

>>> type('abc')   # Python 3 unicode string literal
<class 'str'>
>>> type(b'abc')  # Python 3 byte string literal
<class 'bytes'>

In Python 3, str is like Python 2's unicode, and is used to store text. What was called str in Python 2 is called bytes in Python 3.


How to tell if a byte string is valid utf-8 or ascii

You can call decode. If it raises a UnicodeDecodeError exception, it wasn't valid.

>>> u_umlaut = b'\xc3\x9c'   # UTF-8 representation of the letter 'Ü'
>>> u_umlaut.decode('utf-8')
u'\xdc'
>>> u_umlaut.decode('ascii')
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0: ordinal not in range(128)
🌐
Program Creek
programcreek.com › python
Python detect encoding
def detect_encoding(sample, encoding=None): """Detect encoding of a byte string sample.
🌐
GeeksforGeeks
geeksforgeeks.org › python › python-character-encoding
Python | Character Encoding - GeeksforGeeks
May 29, 2021 - ... # -*- coding: utf-8 -*- import charade def convert(s): # if in the charade instance if isinstance(s, str): s = s.encode() # retrieving the encoding information # from the detect() output encode = detect(s)['encoding'] if encode == 'utf-8': ...
🌐
GitHub
github.com › magnetikonline › py-encoding-detect
GitHub - magnetikonline/py-encoding-detect: Python module for detecting common text file encodings. · GitHub
from encdect import EncodingDetectFile UPPERCASE_WORD = 'string' detect = EncodingDetectFile() result = detect.load('./test/file/utf-8.txt') if (result): encoding,bom_marker,file_decode = result print(type(file_decode)) # <type 'unicode'> file_decode = file_decode.replace( UPPERCASE_WORD, UPPERCASE_WORD.upper() ) fh = open('./output.txt','w') if (bom_marker): fh.write(bom_marker) fh.write(file_decode.encode(encoding)) fh.close() Routines used are based on the work of the C++/C# library https://github.com/AutoIt/text-encoding-detect, with minor tweaks/optimizations.
Author: magnetikonline
🌐
Readthedocs
unicodebook.readthedocs.io › guess_encoding.html
8. How to guess the encoding of a document? — Programming with Unicode
If you would like to reject surrogate characters in Python 2, use the following strict function: def isUTF8Strict(data): try: decoded = data.decode('UTF-8') except UnicodeDecodeError: return False else: for ch in decoded: if 0xD800 <= ord(ch) <= 0xDFFF: return False return True · PHP has a builtin function to detect the encoding of a byte string: mb_detect_encoding().
🌐
Rip Tutorial
riptutorial.com › how to detect the encoding of a text file with python?
encoding Tutorial => How to detect the encoding of a text file with...
Chardet can detect following encodings: ASCII, UTF-8, UTF-16 (2 variants), UTF-32 (4 variants) Big5, GB2312, EUC-TW, HZ-GB-2312, ISO-2022-CN (Traditional and Simplified Chinese) ...
Find elsewhere
🌐
DEV Community
dev.to › bowmanjd › character-encodings-and-detection-with-python-chardet-and-cchardet-4hj7
Character Encodings and Detection with Python, chardet, and cchardet - DEV Community
February 9, 2026 - Try the above print statement in a Python console or script and you should see our beloved "spam". It was automatically decoded in the Python console, printing the corresponding letters (characters). But let's be more explicit, creating a byte string of the above numbers, and specifying the ASCII encoding:
Top answer
1 of 2
3

In my experience enca commandline tool is pretty good at guessing encoding correctly:

http://linux.die.net/man/1/enca

In Python, there's chardet:

https://github.com/chardet/chardet

2 of 2
1

I just think of a way, you can decode the string in every possible encodings.

The following encodings are borrowed from Python's Standard Encodings

code_list = ["ascii", "big5", "big5hkscs", "cp037", "cp424", "cp437", "cp500",
 "cp720", "cp737", "cp775", "cp850", "cp852", "cp855", "cp856", "cp857", "cp858",
 "cp860", "cp861", "cp862", "cp863", "cp864", "cp865", "cp866", "cp869", "cp874",
 "cp875", "cp932", "cp949", "cp950", "cp1006", "cp1026", "cp1140", "cp1250", "cp1251",
 "cp1252", "cp1253", "cp1254", "cp1255", "cp1256", "cp1257", "cp1258", "euc_jp",
 "euc_jis_2004", "euc_jisx0213", "euc_kr", "gb2312", "gbk", "gb18030", "hz", "iso2022_jp",
 "iso2022_jp_1", "iso2022_jp_2", "iso2022_jp_2004", "iso2022_jp_3", "iso2022_jp_ext",
 "iso2022_kr", "latin_1", "iso8859_2", "iso8859_3", "iso8859_4", "iso8859_5", "iso8859_6",
 "iso8859_7", "iso8859_8", "iso8859_9", "iso8859_10", "iso8859_13", "iso8859_14",
 "iso8859_15", "iso8859_16", "johab", "koi8_r", "koi8_u", "mac_cyrillic", "mac_greek",
 "mac_iceland", "mac_latin2", "mac_roman", "mac_turkish", "ptcp154", "shift_jis",
 "shift_jis_2004", "shift_jisx0213", "utf_32", "utf_32_be", "utf_32_le", "utf_16",
 "utf_16_be", "utf_16_le", "utf_7", "utf_8", "utf_8_sig", "idna", "mbcs", "palmos",
 "punycode", "raw_unicode_escape", "rot_13", "undefined", "unicode_escape",
 "unicode_internal", "base64_codec", "bz2_codec", "hex_codec", "quopri_codec",
 "string_escape", "uu_codec", "zlib_codec"]

s = '\x86\x9cG<!\xd9F@\xb4\n\xd6\xd4(\x9cb\xfe'


for i in code_list:
    try:
        print 'Using {0} to decode......{1:<30}'.format(i,s.decode(i).encode('utf-8'))
    except Exception as e:
#         pass
        print e
🌐
Kaggle
kaggle.com › code › rtatman › automatically-detecting-character-encodings
Automatically detecting character encodings | Kaggle
December 16, 2017 - Explore and run AI code with Kaggle Notebooks | Using data from Character Encoding Examples
🌐
Bytes
bytes.com › home › forum › topic › python
Detect character encoding - Post.Byes
December 4, 2005 - Recode might be of help here, it has such heuristics built in AFAIK. But there is _no_ way to be absolutely sure. 8bit are 8bit, so each file is "legal" in all encodings. Diez ... Re: Detect character encoding "Diez B. Roggisch" <deets@nospam.w eb.de> writes:[color=blue] > Michal wrote:[color=green] >> is there any way how to detect string encoding in Python?
🌐
YouTube
youtube.com › codetwist
python detect encoding of string - YouTube
Download this code from https://codegive.com Title: A Guide to Detecting the Encoding of a String in PythonIntroduction:Character encoding is a critical aspe...
Published: December 27, 2023
Views: 6
🌐
Medium
cloudmersive.medium.com › how-to-detect-the-encoding-of-a-text-file-in-python-7a21989acd5b
How to Detect the Encoding of a Text File in Python | by Cloudmersive | Medium
January 9, 2023 - from __future__ import print_function import time import cloudmersive_convert_api_client from cloudmersive_convert_api_client.rest import ApiException from pprint import pprint # Configure API key authorization: Apikey configuration = cloudmersive_convert_api_client.Configuration() configuration.api_key['Apikey'] = 'YOUR_API_KEY' # create an instance of the API class api_instance = cloudmersive_convert_api_client.EditTextApi(cloudmersive_convert_api_client.ApiClient(configuration)) input_file = '/path/to/inputfile' # file | Input file to perform the operation on. try: # Detect text encoding of file api_response = api_instance.edit_text_text_encoding_detect(input_file) pprint(api_response) except ApiException as e: print("Exception when calling EditTextApi->edit_text_text_encoding_detect: %s\n" % e)
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
data corruption is also still possible if the escaped portions of the string are modified directly · Use case: the files to be processed are in a consistent encoding, the encoding can be determined from the OS details and locale settings and it is acceptable to refuse to process files that are not properly encoded. Approach: simply open the file in text mode. This use case describes the default behaviour in Python 3.
🌐
Chatnoir
resiliparse.chatnoir.eu › en › stable › man › parse › encoding.html
Character Encoding — ChatNoir Resiliparse 1.0.9 documentation
With WHATWG remapping enabled, unknown encodings are mapped to UTF-8. This is convenient most of the time, but it also means that the source text is not necessarily decodable without errors, so care should be taken here. You can use bytes_to_str() to avoid decoding errors (see Convert Byte String to Unicode for details). As a convenience shortcut, Resiliparse also provides detect_encoding() which creates and maintains a single global EncodingDetector instance:
🌐
GeeksforGeeks
geeksforgeeks.org › python › character-encoding-detection-with-chardet-in-python
Character Encoding Detection With Chardet in Python - GeeksforGeeks
July 23, 2025 - In this example, the Python script uses the chardet library to detect the character encoding of a given byte sequence (data). The detected encoding and its confidence level are printed, revealing information about the encoding scheme of the ...