1. To get an encoding parameter in Python 2:

If you only need to support Python 2.6 and 2.7 you can use io.open instead of open. io is the new io subsystem for Python 3, and it exists in Python 2,6 ans 2.7 as well. Please be aware that in Python 2.6 (as well as 3.0) it's implemented purely in python and very slow, so if you need speed in reading files, it's not a good option.

If you need speed, and you need to support Python 2.6 or earlier, you can use codecs.open instead. It also has an encoding parameter, and is quite similar to io.open except it handles line-endings differently.

2. To get a Python 3 open() style file handler which streams bytestrings:

open(filename, 'rb')

Note the 'b', meaning 'binary'.

Answer from Lennart Regebro on Stack Overflow
Discussions

Unicode (UTF-8) reading and writing to files in Python - Stack Overflow
Encoding is the name of the encoding ... the file. This should only be used in text mode. The default encoding is platform dependent (whatever locale.getpreferredencoding() returns), but any text encoding supported by Python can be used. See the codecs module for the list of supported encodings. So by adding encoding='utf-8' as a parameter to the open function, ... More on stackoverflow.com
🌐 stackoverflow.com
python - What encoding does open() use by default? - Stack Overflow
The default UTF-8 encoding of Python 3 only extends to conversions between bytes and str types. open() instead chooses an appropriate default encoding based on the environment: encoding is the name of the encoding used to decode or encode the file. This should only be used in text mode. More on stackoverflow.com
🌐 stackoverflow.com
File.open to support different encodings - Django Internals - Django Forum
Currently the ​File.open method has only a single parameter mode. mode can be set to suggest opening a file in text mode. Then, Python’s open will by default attempt to use the common utf-8 encoding. It is a common pattern in a code base I work on to use the utf-8-sig encoding for CSV files, ... More on forum.djangoproject.com
🌐 forum.djangoproject.com
1
June 9, 2023
How to get python to read files encoded in 'latin-1'
Are you using Python2X. Enter python -V on the command line and upgrade if necessary. You can use a "magic line" #!/usr/bin/python # -*- coding: -*- You also can specify the encoding when opening https://www.folkstalk.com/2022/10/python-open-encoding-utf-8-with-code-examples.html and many other pages on the web. More on reddit.com
🌐 r/learnpython
2
1
November 30, 2022
🌐
Python
docs.python.org › 3 › library › codecs.html
codecs — Codec registry and base classes
Open an encoded file using the given mode and return an instance of StreamReaderWriter, providing transparent encoding/decoding.
🌐
Curiousefficiency
python-notes.curiousefficiency.org › en › latest › python3 › text_file_processing.html
Processing Text Files in Python 3 - Alyssa Coghlan's Python Notes
Approach: first open the file in binary mode to look for the encoding marker, then reopen in text mode with the identified encoding. Example: f = tokenize.open(fname) uses PEP 263 encoding markers to detect the encoding of Python source files (defaulting to UTF-8 if no encoding marker is detected)
🌐
Learn By Example
learnbyexample.org › python-open-function
Python open() Function - Learn By Example
April 21, 2020 - # Read a file in 'UTF-8' encoding f = open('myfile.txt', encoding='UTF-8')
🌐
Stanford
web.stanford.edu › class › archive › cs › cs106a › cs106a.1204 › handouts › py-file.html
Python File Reading
The form open(filename, encoding='utf-8') can specify the encoding to use to interpret the text file as unicode. If reading a file crashes with a "UnicodeDecodeError", probably the reading code needs to specify an encoding as above. Try the 'utf-8' encoding first, as many files are encoded with it.
Find elsewhere
🌐
Honeybadger
honeybadger.io › blog › python-character-encoding
Python developer's guide to character encoding - Honeybadger Developer Blog
March 6, 2023 - This is because, when working with ... Usually, when you open a file using the open() method, Python automatically treats it as a text file to convert the bytes in the text file to a string with the encoding you want....
🌐
University of Pittsburgh
sites.pitt.edu › ~naraehan › python3 › reading_writing_methods.html
Python 3 Notes: Reading and Writing Methods
Python 3 Notes [ HOME | LING 1330/2330 ] File Reading and Writing Methods << Previous Note Next Note >> On this page: open(), file.read(), file.readlines(), file.write(), file.writelines(), with open() as f:. Before proceeding, make sure you understand the concepts of file path and CWD.
🌐
Python Module of the Week
pymotw.com › 2 › codecs
codecs – String encoding and decoding - Python Module of the Week
$ python codecs_socket.py Sending ... to that intermediate data format is useful. EncodedFile() takes an open file handle using one encoding and wraps it with a class that translates the data to another encoding as the I/O occurs....
🌐
Python Morsels
pythonmorsels.com › unicode-character-encodings-in-python
Unicode character encodings - Python Morsels
May 2, 2022 - On my machine, the default character encoding is utf-8. But on Windows, the default character encoding is usually cp1252. Note: Since Python 3.6, all files are read and written by Python using utf-8 by default, even on Windows.
🌐
Code with C
codewithc.com › code with c › blog › python with open encoding: specifying file encoding
Python With Open Encoding: Specifying File Encoding - Code With C
January 27, 2024 - You simply include the encoding parameter along with the desired encoding format when opening a file. The encoding parameter is your gateway to a world of text encoding bliss. By using this parameter, you communicate to Python the specific encoding ...
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
The default encoding for Python source code is UTF-8, so you can simply include a Unicode character in a string literal: try: with open('/tmp/input.txt', 'r') as f: ... except OSError: # 'File not found' error message.
Top answer
1 of 14
914

Rather than mess with .encode and .decode, specify the encoding when opening the file. The io module, added in Python 2.6, provides an io.open function, which allows specifying the file's encoding.

Supposing the file is encoded in UTF-8, we can use:

>>> import io
>>> f = io.open("test", mode="r", encoding="utf-8")

Then f.read returns a decoded Unicode object:

>>> f.read()
u'Capit\xe1l\n\n'

In 3.x, the io.open function is an alias for the built-in open function, which supports the encoding argument (it does not in 2.x).

We can also use open from the codecs standard library module:

>>> import codecs
>>> f = codecs.open("test", "r", "utf-8")
>>> f.read()
u'Capit\xe1l\n\n'

Note, however, that this can cause problems when mixing read() and readline().

2 of 14
126

In the notation u'Capit\xe1n\n' (should be just 'Capit\xe1n\n' in 3.x, and must be in 3.0 and 3.1), the \xe1 represents just one character. \x is an escape sequence, indicating that e1 is in hexadecimal.

Writing Capit\xc3\xa1n into the file in a text editor means that it actually contains \xc3\xa1. Those are 8 bytes and the code reads them all. We can see this by displaying the result:

# Python 3.x - reading the file as bytes rather than text,
# to ensure we see the raw data
>>> open('f2', 'rb').read()
b'Capit\\xc3\\xa1n\n'

# Python 2.x
>>> open('f2').read()
'Capit\\xc3\\xa1n\n'

Instead, just input characters like á in the editor, which should then handle the conversion to UTF-8 and save it.

In 2.x, a string that actually contains these backslash-escape sequences can be decoded using the string_escape codec:

# Python 2.x
>>> print 'Capit\\xc3\\xa1n\n'.decode('string_escape')
Capitán

The result is a str that is encoded in UTF-8 where the accented character is represented by the two bytes that were written \\xc3\\xa1 in the original string. To get a unicode result, decode again with UTF-8.

In 3.x, the string_escape codec is replaced with unicode_escape, and it is strictly enforced that we can only encode from a str to bytes, and decode from bytes to str. unicode_escape needs to start with a bytes in order to process the escape sequences (the other way around, it adds them); and then it will treat the resulting \xc3 and \xa1 as character escapes rather than byte escapes. As a result, we have to do a bit more work:

# Python 3.x
>>> 'Capit\\xc3\\xa1n\n'.encode('ascii').decode('unicode_escape').encode('latin-1').decode('utf-8')
'Capitán\n'
🌐
LabEx
labex.io › tutorials › python-how-to-read-files-with-different-encodings-434794
How to read files with different encodings | LabEx
graph LR A[Human-Readable Text] ... encodings through built-in functions and methods. The open() function allows specifying encoding when reading or writing files....
🌐
Plain English
python.plainenglish.io › working-with-encoded-files-in-python-e2442ec604af
Working with Encoded Files in Python | Python in Plain English
January 13, 2025 - It can represent characters from virtually any language and is the default encoding in Python 3. To open and read a UTF-8 encoded file in Python, we can use the built-in open() function.
🌐
Python Forum
python-forum.io › thread-42020.html
[SOLVED] Right way to open files with different encodings?
April 23, 2024 - Hello, Some of the files could be Windows (latin1, iso9959-1, cp1252), others could be utf-8. Is try/except the right way to do it? #with open(file, 'r') as f: #with open(file, 'r',encoding='utf-8') a
🌐
Django Forum
forum.djangoproject.com › django internals
File.open to support different encodings - Django Internals - Django Forum
June 9, 2023 - Currently the ​File.open method has only a single parameter mode. mode can be set to suggest opening a file in text mode. Then, Python’s open will by default attempt to use the common utf-8 encoding.