As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Answer from Sam on Stack OverflowAs the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
python - Pandas convert dataframe to Utf-8 - Stack Overflow
Encoding issues when converting a dataframe from pd.read_pickle using dd.from_pandas
python 3.x - convert pandas dataframe to utf8 - Stack Overflow
weird characters in dataframe - how to standardize to UTF-8?
If you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Setting up the encoding should be treated when reading the input file, using the option encoding
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8")
or if the file uses BOM,
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8-sig")
Df.x.str.encode('utf-8')
Will fix your problems.
http://pandas.pydata.org/pandas-docs/stable/generated/pandas.Series.str.encode.html
Change the code
messages.head().apply(split_into_tokens(messages))
to
messages.head().apply(split_into_tokens)
while using 'apply' with a funtion like in your case passing parameters is not required, as your code shows it is passing a dataframe which is giving error on execution.
Hi. I'm using Python + Camelot (OCR library) to read a PDF, clean up, and write to output. There are some non-standard dashes that print out a weird character.
For example, a value that is supposed to be "1-4" prints out as 1–4 when saving the output as csv and opening in Excel. It renders fine if I open the output in a text editor or save as .xlsx.
I realize that I could handle manually in Excel when I open the file, but I will need to automate this. A colleague suggested standardizing to UTF-8. I tried to do that for the header like this:
header = df.iloc[0, 1:].str.encode('utf-8')But then the output looks like b'1\xe2\x80\x934'.
What am I missing? What's the best way to standardize to UTF-8?
I'm trying to read in a CSV file using pandas.read\_file(report), but I'm hitting this error message:
'utf-8' codec can't decode byte 0x93 in position 28: invalid start byte
I've also tried this with various encoding types, pandas.read_file(report, encoding='UTF-16')
but for some reason it has no effect on the error, it always shows as 'utf-8' The encoding line shows in my traceback so I know that it's there. Does anyone know why this is?