As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Answer from Sam on Stack OverflowAs the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
python - Pandas convert dataframe to Utf-8 - Stack Overflow
to_csv() enconding option "utf-8-BOM"
python - Pandas df.to_csv("file.csv" encode="utf-8") still gives trash characters for minus sign - Stack Overflow
Pandas not reading UTF-8
I'm trying to read in a CSV file using pandas.read\_file(report), but I'm hitting this error message:
'utf-8' codec can't decode byte 0x93 in position 28: invalid start byte
I've also tried this with various encoding types, pandas.read_file(report, encoding='UTF-16')
but for some reason it has no effect on the error, it always shows as 'utf-8' The encoding line shows in my traceback so I know that it's there. Does anyone know why this is?
Your "bad" output is UTF-8 displayed as CP1252.
On Windows, many editors assume the default ANSI encoding (CP1252 on US Windows) instead of UTF-8 if there is no byte order mark (BOM) character at the start of the file. While a BOM is meaningless to the UTF-8 encoding, its UTF-8-encoded presence serves as a signature for some programs. For example, Microsoft Office's Excel requires it even on non-Windows OSes. Try:
df.to_csv('file.csv',encoding='utf-8-sig')
That encoder will add the BOM.
encoding='utf-8-sig does not work for me. Excel reads the special characters fine now, but the Tab separators are gone! However, encoding='utf-16 does work correctly: special characters OK and Tab separators work. This is the solution for me.
Hi. I'm using Python + Camelot (OCR library) to read a PDF, clean up, and write to output. There are some non-standard dashes that print out a weird character.
For example, a value that is supposed to be "1-4" prints out as 1–4 when saving the output as csv and opening in Excel. It renders fine if I open the output in a text editor or save as .xlsx.
I realize that I could handle manually in Excel when I open the file, but I will need to automate this. A colleague suggested standardizing to UTF-8. I tried to do that for the header like this:
header = df.iloc[0, 1:].str.encode('utf-8')But then the output looks like b'1\xe2\x80\x934'.
What am I missing? What's the best way to standardize to UTF-8?