Your "bad" output is UTF-8 displayed as CP1252.
On Windows, many editors assume the default ANSI encoding (CP1252 on US Windows) instead of UTF-8 if there is no byte order mark (BOM) character at the start of the file. While a BOM is meaningless to the UTF-8 encoding, its UTF-8-encoded presence serves as a signature for some programs. For example, Microsoft Office's Excel requires it even on non-Windows OSes. Try:
df.to_csv('file.csv',encoding='utf-8-sig')
That encoder will add the BOM.
Answer from Mark Tolonen on Stack OverflowYour "bad" output is UTF-8 displayed as CP1252.
On Windows, many editors assume the default ANSI encoding (CP1252 on US Windows) instead of UTF-8 if there is no byte order mark (BOM) character at the start of the file. While a BOM is meaningless to the UTF-8 encoding, its UTF-8-encoded presence serves as a signature for some programs. For example, Microsoft Office's Excel requires it even on non-Windows OSes. Try:
df.to_csv('file.csv',encoding='utf-8-sig')
That encoder will add the BOM.
encoding='utf-8-sig does not work for me. Excel reads the special characters fine now, but the Tab separators are gone! However, encoding='utf-16 does work correctly: special characters OK and Tab separators work. This is the solution for me.
Heya,
I got a DataFrame filled with strings in the first column and the rest consisting of integers (except for the headers).
Now when I export this dataframe to a csv file, and the strings contain German Umlauts (ä,ö,ü or something like ß), the exported csv file has weird looking strings at these indices.
Like "für" became "für".
As far as I know the default encoding for to_csv() is utf-8, which means it should work fine? But I also tried the parameter encoding='utf-8', same results.
What am I doing wrong here?
to_csv() enconding option "utf-8-BOM"
Python 3 writing to_csv file ignores encoding argument.
Trying to Convert Text File to CSV using Python Pandas but having 'utf-8' codec problems
Bug: On Python 3 to_csv() encoding defaults to ascii if the dataframe contains special characters.
I get the error 'utf-8' codec can't decode byte 0xe9 in position 5070: invalid continuation byte. Nothing exports.
my code in question
# importing pandas library
import pandas as pd
# reading the given csv file
# and creating dataframe
account = pd.read_csv("Output Configured 3.txt",
delimiter = '/')
# store dataframe into csv file
account.to_csv('Output Configured 3.csv',
index = None)
What do I need to do to have it actually implement?As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')