1. To get an encoding parameter in Python 2:

If you only need to support Python 2.6 and 2.7 you can use io.open instead of open. io is the new io subsystem for Python 3, and it exists in Python 2,6 ans 2.7 as well. Please be aware that in Python 2.6 (as well as 3.0) it's implemented purely in python and very slow, so if you need speed in reading files, it's not a good option.

If you need speed, and you need to support Python 2.6 or earlier, you can use codecs.open instead. It also has an encoding parameter, and is quite similar to io.open except it handles line-endings differently.

2. To get a Python 3 open() style file handler which streams bytestrings:

open(filename, 'rb')

Note the 'b', meaning 'binary'.

Answer from Lennart Regebro on Stack Overflow
Discussions

Unicode (UTF-8) reading and writing to files in Python - Stack Overflow
Encoding is the name of the encoding ... the file. This should only be used in text mode. The default encoding is platform dependent (whatever locale.getpreferredencoding() returns), but any text encoding supported by Python can be used. See the codecs module for the list of supported encodings. So by adding encoding='utf-8' as a parameter to the open function, ... More on stackoverflow.com
🌐 stackoverflow.com
python - What encoding does open() use by default? - Stack Overflow
The default UTF-8 encoding of Python 3 only extends to conversions between bytes and str types. open() instead chooses an appropriate default encoding based on the environment: encoding is the name of the encoding used to decode or encode the file. This should only be used in text mode. More on stackoverflow.com
🌐 stackoverflow.com
How to get python to read files encoded in 'latin-1'
Are you using Python2X. Enter python -V on the command line and upgrade if necessary. You can use a "magic line" #!/usr/bin/python # -*- coding: -*- You also can specify the encoding when opening https://www.folkstalk.com/2022/10/python-open-encoding-utf-8-with-code-examples.html and many other pages on the web. More on reddit.com
🌐 r/learnpython
2
1
November 30, 2022
The use of open(encoding="utf-8")
open defaults to whatever encoding your system uses by default, so it can be anything from ASCII to ISO-8859-1. You can, however, make it use a specific encoding instead. utf-8 is useful as it has most characters. More on reddit.com
🌐 r/learnpython
16
3
November 2, 2020
🌐
Python
docs.python.org › 3 › library › codecs.html
codecs — Codec registry and base classes
Open an encoded file using the given mode and return an instance of StreamReaderWriter, providing transparent encoding/decoding.
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Approach: first open the file in binary mode to look for the encoding marker, then reopen in text mode with the identified encoding. Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
🌐
Learn By Example
learnbyexample.org › python-open-function
Python open() Function - Learn By Example
April 21, 2020 - # Read a file in 'UTF-8' encoding f = open('myfile.txt', encoding='UTF-8')
🌐
Honeybadger
honeybadger.io › blog › python-character-encoding
Python developer's guide to character encoding - Honeybadger Developer Blog
March 6, 2023 - This is because, when working with ... Usually, when you open a file using the open() method, Python automatically treats it as a text file to convert the bytes in the text file to a string with the encoding you want....
🌐
Stanford
web.stanford.edu › class › archive › cs › cs106a › cs106a.1204 › handouts › py-file.html
Python File Reading
The form open(filename, encoding='utf-8') can specify the encoding to use to interpret the text file as unicode. If reading a file crashes with a "UnicodeDecodeError", probably the reading code needs to specify an encoding as above. Try the 'utf-8' encoding first, as many files are encoded with it.
Find elsewhere
🌐
University of Pittsburgh
sites.pitt.edu › ~naraehan › python3 › reading_writing_methods.html
Python 3 Notes: Reading and Writing Methods
Python 3 Notes [ HOME | LING 1330/2330 ] File Reading and Writing Methods << Previous Note Next Note >> On this page: open(), file.read(), file.readlines(), file.write(), file.writelines(), with open() as f:. Before proceeding, make sure you understand the concepts of file path and CWD.
🌐
Python Morsels
pythonmorsels.com › unicode-character-encodings-in-python
Unicode character encodings - Python Morsels
May 2, 2022 - On my machine, the default character encoding is utf-8. But on Windows, the default character encoding is usually cp1252. Note: Since Python 3.6, all files are read and written by Python using utf-8 by default, even on Windows.
🌐
Python Module of the Week
pymotw.com › 2 › codecs
codecs – String encoding and decoding - Python Module of the Week
$ python codecs_socket.py Sending ... to that intermediate data format is useful. EncodedFile() takes an open file handle using one encoding and wraps it with a class that translates the data to another encoding as the I/O occurs....
Top answer
1 of 14
914

Rather than mess with .encode and .decode, specify the encoding when opening the file. The io module, added in Python 2.6, provides an io.open function, which allows specifying the file's encoding.

Supposing the file is encoded in UTF-8, we can use:

>>> import io
>>> f = io.open("test", mode="r", encoding="utf-8")

Then f.read returns a decoded Unicode object:

>>> f.read()
u'Capit\xe1l\n\n'

In 3.x, the io.open function is an alias for the built-in open function, which supports the encoding argument (it does not in 2.x).

We can also use open from the codecs standard library module:

>>> import codecs
>>> f = codecs.open("test", "r", "utf-8")
>>> f.read()
u'Capit\xe1l\n\n'

Note, however, that this can cause problems when mixing read() and readline().

2 of 14
126

In the notation u'Capit\xe1n\n' (should be just 'Capit\xe1n\n' in 3.x, and must be in 3.0 and 3.1), the \xe1 represents just one character. \x is an escape sequence, indicating that e1 is in hexadecimal.

Writing Capit\xc3\xa1n into the file in a text editor means that it actually contains \xc3\xa1. Those are 8 bytes and the code reads them all. We can see this by displaying the result:

# Python 3.x - reading the file as bytes rather than text,
# to ensure we see the raw data
>>> open('f2', 'rb').read()
b'Capit\\xc3\\xa1n\n'

# Python 2.x
>>> open('f2').read()
'Capit\\xc3\\xa1n\n'

Instead, just input characters like á in the editor, which should then handle the conversion to UTF-8 and save it.

In 2.x, a string that actually contains these backslash-escape sequences can be decoded using the string_escape codec:

# Python 2.x
>>> print 'Capit\\xc3\\xa1n\n'.decode('string_escape')
Capitán

The result is a str that is encoded in UTF-8 where the accented character is represented by the two bytes that were written \\xc3\\xa1 in the original string. To get a unicode result, decode again with UTF-8.

In 3.x, the string_escape codec is replaced with unicode_escape, and it is strictly enforced that we can only encode from a str to bytes, and decode from bytes to str. unicode_escape needs to start with a bytes in order to process the escape sequences (the other way around, it adds them); and then it will treat the resulting \xc3 and \xa1 as character escapes rather than byte escapes. As a result, we have to do a bit more work:

# Python 3.x
>>> 'Capit\\xc3\\xa1n\n'.encode('ascii').decode('unicode_escape').encode('latin-1').decode('utf-8')
'Capitán\n'
🌐
Python Forum
python-forum.io › thread-42020.html
[SOLVED] Right way to open files with different encodings?
April 23, 2024 - Hello, Some of the files could be Windows (latin1, iso9959-1, cp1252), others could be utf-8. Is try/except the right way to do it? #with open(file, 'r') as f: #with open(file, 'r',encoding='utf-8') as f: #latin1, iso9959-1, cp1252 with open(file,...
🌐
LabEx
labex.io › tutorials › python-how-to-read-files-with-different-encodings-434794
How to read files with different encodings | LabEx
graph LR A[Human-Readable Text] ... encodings through built-in functions and methods. The open() function allows specifying encoding when reading or writing files....
🌐
Code with C
codewithc.com › code with c › blog › python with open encoding: specifying file encoding
Python With Open Encoding: Specifying File Encoding - Code With C
January 27, 2024 - So, what are your thoughts on file encoding in Python? Have you encountered any fascinating encoding puzzles? Share your experiences below! Let’s spark a lively discussion. 🤓 ... # Importing necessary modules import io # Main function demonstrating the use of open with specific encoding def main(): # Define the file path and the encoding file_path = 'example.txt' encoding_type = 'utf-8' # Let's begin by writing some data to a file using our specified encoding data_to_write = 'This is a line of text with emojis 🚀🐍.' # Opening the file with the 'write' mode and specific encoding with
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
The default encoding for Python source code is UTF-8, so you can simply include a Unicode character in a string literal: try: with open('/tmp/input.txt', 'r') as f: ... except OSError: # 'File not found' error message.
🌐
Plain English
python.plainenglish.io › working-with-encoded-files-in-python-e2442ec604af
Working with Encoded Files in Python | Python in Plain English
January 13, 2025 - It can represent characters from virtually any language and is the default encoding in Python 3. To open and read a UTF-8 encoded file in Python, we can use the built-in open() function.
🌐
Diveintopython3
diveintopython3.net › files.html
Files - Dive Into Python 3
Before you can read from a file, you need to open it. Opening a file in Python couldn’t be easier: a_file = open('examples/chinese.txt', encoding='utf-8')