As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Answer from Sam on Stack OverflowAs the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
csv - Encoding Error in Panda read_csv - Stack Overflow
BUG: read_csv is failing with an encoding different that UTF-8 and memory_map set to True in version 1.2.4
pandas.read_csv() encoding issue
python - UnicodeDecodeError when reading CSV file in Pandas - Stack Overflow
I'm trying to read in a CSV file using pandas.read\_file(report), but I'm hitting this error message:
'utf-8' codec can't decode byte 0x93 in position 28: invalid start byte
I've also tried this with various encoding types, pandas.read_file(report, encoding='UTF-16')
but for some reason it has no effect on the error, it always shows as 'utf-8' The encoding line shows in my traceback so I know that it's there. Does anyone know why this is?
read_csv takes an encoding option to deal with files in different formats. I mostly use read_csv('file', encoding = "ISO-8859-1"), or alternatively encoding = "utf-8" for reading, and generally utf-8 for to_csv.
You can also use one of several alias options like 'latin' or 'cp1252' (Windows) instead of 'ISO-8859-1' (see python docs, also for numerous other encodings you may encounter).
See relevant Pandas documentation, python docs examples on csv files, and plenty of related questions here on SO. A good background resource is What every developer should know about unicode and character sets.
To detect the encoding (assuming the file contains non-ascii characters), you can use enca (see man page) or file -i (linux) or file -I (osx) (see man page).
Simplest of all Solutions:
import pandas as pd
df = pd.read_csv('file_name.csv', engine='python')
Alternate Solution:
Sublime Text:
- Open the csv file in Sublime text editor or VS Code.
- Save the file in utf-8 format.
- In sublime, Click File -> Save with encoding -> UTF-8
VS Code:
In the bottom bar of VSCode, you'll see the label UTF-8. Click it. A popup opens. Click Save with encoding. You can now pick a new encoding for that file.
Then, you could read your file as usual:
import pandas as pd
data = pd.read_csv('file_name.csv', encoding='utf-8')
and the other different encoding types are:
encoding = "cp1252"
encoding = "ISO-8859-1"
You can pass encoding_errors='ignore', but I would advice to try different encoding first.
tl;dr
For your situation, you probably need one of these:
pd.read_csv(file, encoding='utf-8', encoding_errors='replace')
# or
pd.read_csv(file, encoding='utf-8', encoding_errors='ignore')
Longer answer
As of version 1.3+ of pandas.read_csv() you can pass argument encoding_errors.
Possible values:
strict: Raise UnicodeError (or a subclass); this is the default.ignore: Ignore the malformed data and continue without further notice.replace: Replace with a suitable replacement marker; Python will use the official U+FFFD REPLACEMENT CHARACTER for the built-in codecs on decoding, and ‘?’ on encoding.xmlcharrefreplace: Replace with the appropriate XML character reference (only for encoding).backslashreplace: Replace with backslashed escape sequences.namereplace: Replace with \N{...} escape sequences (only for encoding).surrogateescape: On decoding, replace byte with individual surrogate code ranging from U+DC80 to U+DCFF. This code will then be turned back into the same byte when the 'surrogateescape' error handler is used when encoding the data.surrogatepass: Allow encoding and decoding of surrogate codes. These codecs normally treat the presence of surrogates as an error.