Given a file object, and a number of characters, you can use:

# build a table mapping lead byte to expected follow-byte count
# bytes 00-BF have 0 follow bytes, F5-FF is not legal UTF8
# C0-DF: 1, E0-EF: 2 and F0-F4: 3 follow bytes.
# leave F5-FF set to 0 to minimize reading broken data.
_lead_byte_to_count = []
for i in range(256):
    _lead_byte_to_count.append(
        1 + (i >= 0xe0) + (i >= 0xf0) if 0xbf < i < 0xf5 else 0)

def readUTF8(f, count):
    """Read `count` UTF-8 bytes from file `f`, return as unicode"""
    # Assumes UTF-8 data is valid; leaves it up to the `.decode()` call to validate
    res = []
    while count:
        count -= 1
        lead = f.read(1)
        res.append(lead)
        readcount = _lead_byte_to_count[ord(lead)]
        if readcount:
            res.append(f.read(readcount))
    return (''.join(res)).decode('utf8')

Result of a test:

>>> test = StringIO(u'This is a test containing Unicode data: \ua000'.encode('utf8'))
>>> readUTF8(test, 41)
u'This is a test containing Unicode data: \ua000'

In Python 3, it is of course much, much easier to just wrap the file object in a io.TextIOWrapper() object and leave decoding to the native and efficient Python UTF-8 implementation.

Answer from Martijn Pieters on Stack Overflow
🌐
UW PCE
uwpce-pythoncert.github.io › ProgrammingInPython › modules › Files.html
File Reading and Writing — Programming in Python 7.0 documentation
But this too is complicated – ... These days, that mostly “just works”. But if you read a binary file as text, then Python will try to interpret the bytes as utf-8 encoded text – and this will likely fail:...
🌐
Computer Science Atlas
csatlas.com › python-read-file
Python 3: Read a File — Computer Science Atlas
February 2, 2021 - .read() first loads the file in binary format, then .decode() converts it to a string using Unicode UTF-8 decoding rules. Python automatically closes the file f after running all the code inside the with ...
🌐
AskPython
askpython.com › python › examples › binary-to-utf8-conversion
How to Convert Binary Data to UTF-8 in Python - AskPython
April 10, 2025 - If your binary data is already UTF-8 encoded, you can simply decode it to a text string: data = b"Some binary data" text = data.decode("utf-8") Note that UTF-8 is the default encoding for decode(), so you can omit it: ... This handles the majority ...
🌐
Python Morsels
pythonmorsels.com › reading-binary-files-in-python
Reading binary files in Python - Python Morsels
May 16, 2022 - This function reads all of the binary data within this file. We're reading bytes because Python's hashlib module requires us to work with bytes.
🌐
GitHub
github.com › construct › construct › issues › 517
Trying to read binary file in utf8? - Failing in Python3 · Issue #517 · construct/construct
January 28, 2018 - Hi all!, i'm here again, testing and testing, i found this, i don't know if is a problem in the code or in the lib, the weird here is, the code works fine in Python2, i can execute, use all, anyway there no problem, but now i test it with Python3, and show an error, i have no idea why, here the outputs:
Author: construct
🌐
Python
docs.python.org › 3 › library › codecs.html
codecs — Codec registry and base classes
Creates a StreamRecoder instance which implements a two-way conversion: encode and decode work on the frontend — the data visible to code calling read() and write(), while Reader and Writer work on the backend — the data in stream. You can use these objects to do transparent transcodings, e.g., from Latin-1 to UTF-8 and back. The stream argument must be a file-like object.
Top answer
1 of 14
914

Rather than mess with .encode and .decode, specify the encoding when opening the file. The io module, added in Python 2.6, provides an io.open function, which allows specifying the file's encoding.

Supposing the file is encoded in UTF-8, we can use:

>>> import io
>>> f = io.open("test", mode="r", encoding="utf-8")

Then f.read returns a decoded Unicode object:

>>> f.read()
u'Capit\xe1l\n\n'

In 3.x, the io.open function is an alias for the built-in open function, which supports the encoding argument (it does not in 2.x).

We can also use open from the codecs standard library module:

>>> import codecs
>>> f = codecs.open("test", "r", "utf-8")
>>> f.read()
u'Capit\xe1l\n\n'

Note, however, that this can cause problems when mixing read() and readline().

2 of 14
126

In the notation u'Capit\xe1n\n' (should be just 'Capit\xe1n\n' in 3.x, and must be in 3.0 and 3.1), the \xe1 represents just one character. \x is an escape sequence, indicating that e1 is in hexadecimal.

Writing Capit\xc3\xa1n into the file in a text editor means that it actually contains \xc3\xa1. Those are 8 bytes and the code reads them all. We can see this by displaying the result:

# Python 3.x - reading the file as bytes rather than text,
# to ensure we see the raw data
>>> open('f2', 'rb').read()
b'Capit\\xc3\\xa1n\n'

# Python 2.x
>>> open('f2').read()
'Capit\\xc3\\xa1n\n'

Instead, just input characters like á in the editor, which should then handle the conversion to UTF-8 and save it.

In 2.x, a string that actually contains these backslash-escape sequences can be decoded using the string_escape codec:

# Python 2.x
>>> print 'Capit\\xc3\\xa1n\n'.decode('string_escape')
Capitán

The result is a str that is encoded in UTF-8 where the accented character is represented by the two bytes that were written \\xc3\\xa1 in the original string. To get a unicode result, decode again with UTF-8.

In 3.x, the string_escape codec is replaced with unicode_escape, and it is strictly enforced that we can only encode from a str to bytes, and decode from bytes to str. unicode_escape needs to start with a bytes in order to process the escape sequences (the other way around, it adds them); and then it will treat the resulting \xc3 and \xa1 as character escapes rather than byte escapes. As a result, we have to do a bit more work:

# Python 3.x
>>> 'Capit\\xc3\\xa1n\n'.encode('ascii').decode('unicode_escape').encode('latin-1').decode('utf-8')
'Capitán\n'
Find elsewhere
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Approach: first open the file in binary mode to look for the encoding marker, then reopen in text mode with the identified encoding. Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
Top answer
1 of 1
1

One way might be to use Hachoir to define a file parsing protocol.

The simple alternative is to open the file in binary mode and manually initialise a buffer and text wrapper around it. You can then switch in and out of binary pretty neatly:

my_file = io.open("myfile.txt", "rb")
my_file_buffer = io.BufferedReader(my_file, buffer_size=1) # Not as performant but a larger buffer will "eat" into the binary data 
my_file_text_reader = io.TextIOWrapper(my_file_buffer, encoding="utf-8")
string_buffer = ""

while True:
    while "near the end" not in string_buffer:
        string_buffer += my_file_text_reader.read(1) # read one Unicode char at a time

    # binary data must be next. Where do we get the binary length from?
    print string_buffer
    data = my_file_buffer.read(3)

    print data
    string_buffer = ""

A quicker, less extensible way might be to use the approach you've suggested in your question by intelligently parsing the text portions, reading each UTF-8 sequence of bytes at a time. The following code (from http://rosettacode.org/wiki/Read_a_file_character_by_character/UTF8#Python), seems to be a neat way to conservatively read UTF-8 bytes into characters from a binary file:

 def get_next_character(f):
     # note: assumes valid utf-8
     c = f.read(1)
     while c:
         while True:
             try:
                 yield c.decode('utf-8')
             except UnicodeDecodeError:
                 # we've encountered a multibyte character
                 # read another byte and try again
                 c += f.read(1)
             else:
                 # c was a valid char, and was yielded, continue
                 c = f.read(1)
                 break

# Usage:
with open("input.txt","rb") as f:
    my_unicode_str = ""
    for c in get_next_character(f):
        my_unicode_str += c
🌐
Python.org
discuss.python.org › python help
How to fix utf-8 error when reading text file? - Python Help - Discussions on Python.org
April 3, 2024 - I have Python 3.12 on Windows 10. I have a program to find a string in a 12MB file .dat file which was exported from Excel to be a tab-delimited file. However when the file is read I get this error: UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in byte position 7997: invalid continuation byte When I open the file in my text editor (Notepad++) and go to position 7997 I don’t see any special characters when I turn on “Show special characters”. The cursor is between 2 normal letters: H...
🌐
CodeRivers
coderivers.org › blog › python-read-binary-file
Python Read Binary File: A Comprehensive Guide - CodeRivers
February 22, 2026 - In this example, the decode('utf-8') method decodes the bytes object data using the UTF-8 encoding and stores the resulting text in the text variable. The print(text) statement prints the decoded text. When working with specific binary formats, such as images, audio, or video, you may need ...
🌐
Accuweb
accuweb.cloud › home › how to read binary file in python?
How to Read Binary File in Python?
May 2, 2024 - text = data.decode("utf-8")Copy · Only do this if the content is actually encoded text. Q) How do I read a specific byte range from a binary file? A) Use seek() to move the file pointer and read() the required bytes: with open("file.bin", "rb") as f: f.seek(100) # Move to byte 100 data = f.read(20) # Read next 20 bytesCopy · Q) How do I detect the size of a binary file in Python?
🌐
Stack Overflow
stackoverflow.com › questions › 48769227 › how-to-read-binary-file-on-python-from-unknown-encoding
How to read Binary file on PYTHON from unknown encoding? - Stack Overflow
February 13, 2018 - I have a problem when I try to read a binary file in python. ... file_id = fopen("file.bin", "r"); file = fread(file_id); fclose(file_id); dataA = typecast (uint8(file),"uint16"); data = double(dataA); But I now have to read it with python, I tried several methods : import numpy as np import base64 file_id = open("file.bin", "rb"); file = file_id.read(); ... with codecs.open('file.bin', 'r', encoding='utf-8', errors='strict') as fdata: A = codecs.lookup(fdata);
🌐
CodeRivers
coderivers.org › blog › python-read-binary-file-as-string
Reading Binary Files as Strings in Python - CodeRivers
April 16, 2025 - In this example, we open the binary file binary_file.bin in read - binary mode. The read() method reads the entire file content as a bytes object. Then, we try to decode the bytes into a string using the UTF - 8 encoding. The codecs module in Python provides a more flexible way to work with ...
🌐
Learning Python
learning-python.com › strings30.html
Python 3.X Strings Tutorial by Mark Lutz
If you fall into neither of the prior two categories, you can generally use strings in 3.0 much as you would in 2.6: with the general str string type, text files, and all the familiar string operations. Your strings will be encoded and decoded using your platform's default encoding (e.g., ASCII, ...