The method str.encode turns a unicode string into a bytes object:

str.encode(encoding="utf-8", errors="strict")
Return an encoded version of the string as a bytes object. Default encoding is 'utf-8'. errors may be given to set a different error handling scheme. The default for errors is 'strict', meaning that encoding errors raise a UnicodeError. Other possible values are 'ignore', 'replace', 'xmlcharrefreplace', 'backslashreplace' and any other name registered via codecs.register_error(), see section Error Handlers. For a list of possible encodings, see section Standard Encodings.

So what you get is exactly what is expected.

On most machines, you can just open the files and read. If the file encoding is not the system default, you can pass it as keyword argument:

with open(filename, encoding='utf8') as f:
    line = f.readline()
Answer from MaxNoe on Stack Overflow
🌐
CompuTicket
computicket.co.za › home › python › python encoding error: readline () when reading utf-8 file swears: 'charmap' codec...
python - Python encoding error: readline () when reading utf-8 file swears: ‘charmap’ codec can’t decode byte
January 14, 2022 - try: file = open (path, 'r') while ... read line by line fileObj.close () To read a text file encoded using utf-8 encoding in Python, you can use the io.open () function, which is available as built-in open () in Python 3 :...
Discussions

interactive() readline history chokes on UTF-8 characters
File "/lib/python2.7/site-packages/pwnlib/term/readline.py", line 131, in set_buffer buffer_left = unicode(left) UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0: ordinal not in range(128) More on github.com
🌐 github.com
6
July 27, 2017
dataset - Python read huge file line by line with utf-8 encoding - Stack Overflow
I want to read some quite huge files(to be precise: the google ngram 1 word dataset) and count how many times a character occurs. Now I wrote this script: import fileinput files = ['../../datasets/ More on stackoverflow.com
🌐 stackoverflow.com
utf 8 - Python read() works with UTF-8 but readlines() "doesn't" - Stack Overflow
So, I am working with a (huge) UTF-8 encoded file. The first thing I do with it it's get it's lines in a list using the File Object readlines() method. However when I use the print command for debu... More on stackoverflow.com
🌐 stackoverflow.com
September 3, 2013
python - Python3 UnicodeDecodeError with readlines() method - Stack Overflow
You need to know the real encoding, not just guess; UTF-8 is mostly self-checking, so it's unlikely to decode binary gibberish, but latin-1 will happily decode binary gibberish to text gibberish and never whisper a word of complaint. ... Save this answer. Show activity on this post. I think the best answer (in Python 3) is to use the errors= parameter: Copywith open('evil_unicode.txt', 'r', errors='replace') as f: lines = f.readlines... More on stackoverflow.com
🌐 stackoverflow.com
January 15, 2017
🌐
Stanford
web.stanford.edu › class › archive › cs › cs106a › cs106a.1204 › handouts › py-file.html
Python File Reading
The form open(filename, encoding='utf-8') can specify the encoding to use to interpret the text file as unicode.
🌐
GitHub
github.com › Gallopsled › pwntools › issues › 1004
interactive() readline history chokes on UTF-8 characters · Issue #1004 · Gallopsled/pwntools
July 27, 2017 - test.py (doesn't work from inside REPL): from pwn import * r = process('cat', shell=True) r.interactive() enter a string containing UTF-8 characters such löl . hit cursor-up to access readline history encounter exception: File "/l...
Author: Gallopsled
🌐
TutorialsPoint
tutorialspoint.com › article › how-to-read-and-write-unicode-utf-8-files-in-python
How to read and write unicode (UTF-8) files in Python?
Python provides built-in support for reading and writing Unicode (UTF-8) files through the open() function.
🌐
ZetCode
zetcode.com › python › readlines
Python readlines Function - Complete Guide
The readlines function works with different file encodings when specified during file opening. This is crucial for international text files. ... # Reading a UTF-8 encoded file try: with open('multilingual.txt', 'r', encoding='utf-8') as file: lines = file.readlines() for line in lines: ...
Find elsewhere
🌐
University of Pittsburgh
sites.pitt.edu › ~naraehan › python3 › mbb12.html
Python 3 Notes: Reading Text
Python 3 Notes [ HOME | LING 1330/2330 ] Tutorial 12: Reading Text << Previous Tutorial Next Tutorial >> On this page: reading from a file, open(), .readlines(). Video Tutorial Python 3 Changes print(x,y) instead of print x, y You may need open(filename, encoding="utf-8") instead of open(filename).
Top answer
1 of 3
76

I think the best answer (in Python 3) is to use the errors= parameter:

with open('evil_unicode.txt', 'r', errors='replace') as f:
    lines = f.readlines()

Proof:

>>> s = b'\xe5abc\nline2\nline3'
>>> with open('evil_unicode.txt','wb') as f:
...     f.write(s)
...
16
>>> with open('evil_unicode.txt', 'r') as f:
...     lines = f.readlines()
...
Traceback (most recent call last):
  File "<stdin>", line 2, in <module>
  File "/Library/Frameworks/Python.framework/Versions/3.4/lib/python3.4/codecs.py", line 319, in decode
    (result, consumed) = self._buffer_decode(data, self.errors, final)
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe5 in position 0: invalid continuation byte
>>> with open('evil_unicode.txt', 'r', errors='replace') as f:
...     lines = f.readlines()
...
>>> lines
['�abc\n', 'line2\n', 'line3']
>>>

Note that the errors= can be replace or ignore. Here's what ignore looks like:

>>> with open('evil_unicode.txt', 'r', errors='ignore') as f:
...     lines = f.readlines()
...
>>> lines
['abc\n', 'line2\n', 'line3']
2 of 3
24

Your default encoding appears to be ASCII, where the input is more than likely UTF-8. When you hit non-ASCII bytes in the input, it's throwing the exception. It's not so much that readlines itself is responsible for the problem; rather, it's causing the read+decode to occur, and the decode is failing.

It's an easy fix though; the default open in Python 3 allows you to provide the known encoding of an input, replacing the default (ASCII in your case) with any other recognized encoding. Providing it allows you to keep reading as str (rather than the significantly different raw binary data bytes objects), while letting Python do the work of converting from raw disk bytes to true text data:

# Using with statement closes the file for us without needing to remember to close
# explicitly, and closes even when exceptions occur
with open(argfile, encoding='utf-8') as inf:
    f = inf.readlines()

If the file is some other encoding, you'd change encoding='utf-8' to the appropriate argument. Note that while some people will tell you to "Just use 'latin-1'" here if 'utf-8' doesn't work":

  1. That's often wrong (modern text editors tend to produce UTF-8 or UTF-16, with latin-1 being much less common; frankly, you're more likely to see Microsoft's 'latin-1' variant, 'cp1252', that's mostly the same but remaps some characters to support stuff like smart quotes), and
  2. Unlike the UTF encodings, the various byte-per-character ASCII superset encodings (including 'latin-1', 'cp1252', 'cp437', and many others) are not self-checking; if the data isn't in the encoding specified, they'll still happily decode it, it will just produce gibberish for stuff above the ASCII range.

In short, if your data isn't a UTF encoding (or one of the rare non-UTF self-checking encodings), you need to know the encoding used, or you're stuck guessing and checking the result to see if it makes sense (and for stuff like a source that might be latin-1 or cp1252, you'll never be sure unless it eventually contains a cp1252-specific character).

🌐
University of Pittsburgh
sites.pitt.edu › ~naraehan › python3 › reading_writing_methods.html
Python 3 Notes: Reading and Writing Methods
Python 3 Notes [ HOME | LING 1330/2330 ] File Reading and Writing Methods << Previous Note Next Note >> On this page: open(), file.read(), file.readlines(), file.write(), file.writelines(), with open() as f:. Before proceeding, make sure you understand the concepts of file path and CWD.
🌐
Hyperskill
hyperskill.org › university › python › readline-method-in-python
Python readline(): Read Files Line by Line with Examples
June 5, 2026 - When working with Python there are methods available to transform binary data into a readable form. A popular method involves utilizing the decode() function. ... binary_data = b'\x68\x65\x6c\x6c\x6f' readable_data = binary_data.decode('utf-8') print(readable_data) # Output: hello · To handle the information in a file with the readlines() function here are the steps you should take;
🌐
Python Tutorial
pythontutorial.net › home › python basics › python read text file
How to Read a Text file In Python Effectively
October 22, 2020 - Always close a file after completing reading it using the close() method or the with statement. Use the encoding='utf-8' to read the UTF-8 text file. Quiz · 5 questions · To help you understand various ways to read text files in Python.Start ...
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
🌐
Python.org
discuss.python.org › python help
How to fix utf-8 error when reading text file? - Python Help - Discussions on Python.org
April 3, 2024 - I have Python 3.12 on Windows 10. I have a program to find a string in a 12MB file .dat file which was exported from Excel to be a tab-delimited file. However when the file is read I get this error: UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in byte position 7997: invalid continuation byte When I open the file in my text editor (Notepad++) and go to position 7997 I don’t see any special characters when I turn on “Show special characters”. The cursor is between 2 normal letters: H...
🌐
Reddit
reddit.com › r/learnpython › i have a unicode ß that i encoded to 'utf-8', and wrote it in test.txt as a string. how do i read it back again to python as a unicode?
r/learnpython on Reddit: I have a Unicode ß that I encoded to 'utf-8', and wrote it in test.txt as a string. How do I read it back again to Python as a Unicode?
February 16, 2022 -
with open("test.txt", "w") as f:
    f.write(str("ß".encode("utf-8")))

In my text file, this is written: b'\xc3\x9f'

How do I read that back in Python as a Unicode?

I have tried reading it with:

with open("test.txt", "r") as f:
    tmp = f.readline()

But it returns a string "b'\\xc3\\x9f'" and I can't figure out how to decode it

🌐
Python Forum
python-forum.io › thread-22047.html
Problem with readlines() assignment
October 26, 2019 - Hello there, I have an assignment which I need to make a loop that reads each line in a text file and then run a condition against the line of text. If it has special characters, it is rejected. If it is alphanumeric, then it is valid: filename = op...