As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Answer from Sam on Stack OverflowAs the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
python 2.7 - Dataframe encoding - Stack Overflow
Problem with Pandas read_csv always trying to read as UTF-8
python 3.x - convert pandas dataframe to utf8 - Stack Overflow
python - Pandas df.to_csv("file.csv" encode="utf-8") still gives trash characters for minus sign - Stack Overflow
If you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Setting up the encoding should be treated when reading the input file, using the option encoding
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8")
or if the file uses BOM,
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8-sig")
I'm trying to read in a CSV file using pandas.read\_file(report), but I'm hitting this error message:
'utf-8' codec can't decode byte 0x93 in position 28: invalid start byte
I've also tried this with various encoding types, pandas.read_file(report, encoding='UTF-16')
but for some reason it has no effect on the error, it always shows as 'utf-8' The encoding line shows in my traceback so I know that it's there. Does anyone know why this is?
Df.x.str.encode('utf-8')
Will fix your problems.
http://pandas.pydata.org/pandas-docs/stable/generated/pandas.Series.str.encode.html
Change the code
messages.head().apply(split_into_tokens(messages))
to
messages.head().apply(split_into_tokens)
while using 'apply' with a funtion like in your case passing parameters is not required, as your code shows it is passing a dataframe which is giving error on execution.
Your "bad" output is UTF-8 displayed as CP1252.
On Windows, many editors assume the default ANSI encoding (CP1252 on US Windows) instead of UTF-8 if there is no byte order mark (BOM) character at the start of the file. While a BOM is meaningless to the UTF-8 encoding, its UTF-8-encoded presence serves as a signature for some programs. For example, Microsoft Office's Excel requires it even on non-Windows OSes. Try:
df.to_csv('file.csv',encoding='utf-8-sig')
That encoder will add the BOM.
encoding='utf-8-sig does not work for me. Excel reads the special characters fine now, but the Tab separators are gone! However, encoding='utf-16 does work correctly: special characters OK and Tab separators work. This is the solution for me.