As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Answer from Sam on Stack OverflowAs the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
weird characters in dataframe - how to standardize to UTF-8?
utf 8 - How to read a dataframe of encoded strings from csv in python - Stack Overflow
python 2.7 - Dataframe encoding - Stack Overflow
python - Pandas df.to_csv("file.csv" encode="utf-8") still gives trash characters for minus sign - Stack Overflow
Hi. I'm using Python + Camelot (OCR library) to read a PDF, clean up, and write to output. There are some non-standard dashes that print out a weird character.
For example, a value that is supposed to be "1-4" prints out as 1–4 when saving the output as csv and opening in Excel. It renders fine if I open the output in a text editor or save as .xlsx.
I realize that I could handle manually in Excel when I open the file, but I will need to automate this. A colleague suggested standardizing to UTF-8. I tried to do that for the header like this:
header = df.iloc[0, 1:].str.encode('utf-8')But then the output looks like b'1\xe2\x80\x934'.
What am I missing? What's the best way to standardize to UTF-8?
If you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Setting up the encoding should be treated when reading the input file, using the option encoding
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8")
or if the file uses BOM,
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8-sig")
Your "bad" output is UTF-8 displayed as CP1252.
On Windows, many editors assume the default ANSI encoding (CP1252 on US Windows) instead of UTF-8 if there is no byte order mark (BOM) character at the start of the file. While a BOM is meaningless to the UTF-8 encoding, its UTF-8-encoded presence serves as a signature for some programs. For example, Microsoft Office's Excel requires it even on non-Windows OSes. Try:
df.to_csv('file.csv',encoding='utf-8-sig')
That encoder will add the BOM.
encoding='utf-8-sig does not work for me. Excel reads the special characters fine now, but the Tab separators are gone! However, encoding='utf-16 does work correctly: special characters OK and Tab separators work. This is the solution for me.