The method str.encode turns a unicode string into a bytes object:
str.encode(encoding="utf-8", errors="strict")
Return an encoded version of the string as a bytes object. Default encoding is 'utf-8'. errors may be given to set a different error handling scheme. The default for errors is 'strict', meaning that encoding errors raise a UnicodeError. Other possible values are 'ignore', 'replace', 'xmlcharrefreplace', 'backslashreplace' and any other name registered via codecs.register_error(), see section Error Handlers. For a list of possible encodings, see section Standard Encodings.
So what you get is exactly what is expected.
On most machines, you can just open the files and read. If the file encoding is not the system default, you can pass it as keyword argument:
with open(filename, encoding='utf8') as f:
line = f.readline()
Answer from MaxNoe on Stack Overflowinteractive() readline history chokes on UTF-8 characters
dataset - Python read huge file line by line with utf-8 encoding - Stack Overflow
utf 8 - Python read() works with UTF-8 but readlines() "doesn't" - Stack Overflow
python - Python3 UnicodeDecodeError with readlines() method - Stack Overflow
In Python 3, pass an appropriate errors= value (such as errors=ignore or errors=replace) on creating your file object (presuming it to be a subclass of io.TextIOWrapper -- and if it isn't, consider wrapping it in one!); also, consider passing a more likely encoding than charmap (when you aren't sure, utf-8 is always a good place to start).
For instance:
f = open('misc-notes.txt', encoding='utf-8', errors='ignore')
In Python 2, the read() operation simply returns bytes; the trick, then, is decoding them to get them into a string (if you do, in fact, want characters as opposed to bytes). If you don't have a better guess for their real encoding:
your_string.decode('utf-8', 'replace')
...to replace unhandled characters, or
your_string.decode('utf-8', 'ignore')
to simply ignore them.
That said, finding and using their real encoding (rather than guessing utf-8) would be preferred.
You should open the file with a codecs to make sure that the file gets interpreted as UTF8.
import codecs fd = codecs.open(filename,'r',encoding='utf-8') data = fd.read()
I do not know why fileinput does not work as expected.
I suggest you use the open function instead. The return value can be iterated over and will return lines, just like fileinput.
The code will then be something like:
for filename in files:
print(filename)
for filelineno, line in enumerate(open(filename, encoding="utf-8")):
line = line.strip()
data = line.split('\t')
# ...
Some documentation links: enumerate, open, io.TextIOWrapper (open returns an instance of TextIOWrapper).
The problem is that fileinput doesn't use file.xreadlines(), which reads line by line, but file.readline(bufsize), which reads bufsize bytes at once (and turns that into a list of lines). You are providing 0 for the bufsize parameter of fileinput.input() (which is also the default value). Bufsize 0 means that the whole file is buffered.
Solution: provide a reasonable bufsize.
I think the best answer (in Python 3) is to use the errors= parameter:
with open('evil_unicode.txt', 'r', errors='replace') as f:
lines = f.readlines()
Proof:
>>> s = b'\xe5abc\nline2\nline3'
>>> with open('evil_unicode.txt','wb') as f:
... f.write(s)
...
16
>>> with open('evil_unicode.txt', 'r') as f:
... lines = f.readlines()
...
Traceback (most recent call last):
File "<stdin>", line 2, in <module>
File "/Library/Frameworks/Python.framework/Versions/3.4/lib/python3.4/codecs.py", line 319, in decode
(result, consumed) = self._buffer_decode(data, self.errors, final)
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe5 in position 0: invalid continuation byte
>>> with open('evil_unicode.txt', 'r', errors='replace') as f:
... lines = f.readlines()
...
>>> lines
['�abc\n', 'line2\n', 'line3']
>>>
Note that the errors= can be replace or ignore. Here's what ignore looks like:
>>> with open('evil_unicode.txt', 'r', errors='ignore') as f:
... lines = f.readlines()
...
>>> lines
['abc\n', 'line2\n', 'line3']
Your default encoding appears to be ASCII, where the input is more than likely UTF-8. When you hit non-ASCII bytes in the input, it's throwing the exception. It's not so much that readlines itself is responsible for the problem; rather, it's causing the read+decode to occur, and the decode is failing.
It's an easy fix though; the default open in Python 3 allows you to provide the known encoding of an input, replacing the default (ASCII in your case) with any other recognized encoding. Providing it allows you to keep reading as str (rather than the significantly different raw binary data bytes objects), while letting Python do the work of converting from raw disk bytes to true text data:
# Using with statement closes the file for us without needing to remember to close
# explicitly, and closes even when exceptions occur
with open(argfile, encoding='utf-8') as inf:
f = inf.readlines()
If the file is some other encoding, you'd change encoding='utf-8' to the appropriate argument. Note that while some people will tell you to "Just use 'latin-1'" here if 'utf-8' doesn't work":
- That's often wrong (modern text editors tend to produce UTF-8 or UTF-16, with latin-1 being much less common; frankly, you're more likely to see Microsoft's
'latin-1'variant,'cp1252', that's mostly the same but remaps some characters to support stuff like smart quotes), and - Unlike the UTF encodings, the various byte-per-character ASCII superset encodings (including
'latin-1','cp1252','cp437', and many others) are not self-checking; if the data isn't in the encoding specified, they'll still happily decode it, it will just produce gibberish for stuff above the ASCII range.
In short, if your data isn't a UTF encoding (or one of the rare non-UTF self-checking encodings), you need to know the encoding used, or you're stuck guessing and checking the result to see if it makes sense (and for stuff like a source that might be latin-1 or cp1252, you'll never be sure unless it eventually contains a cp1252-specific character).
with open("test.txt", "w") as f:
f.write(str("ß".encode("utf-8")))In my text file, this is written: b'\xc3\x9f'
How do I read that back in Python as a Unicode?
I have tried reading it with:
with open("test.txt", "r") as f:
tmp = f.readline()But it returns a string "b'\\xc3\\x9f'" and I can't figure out how to decode it