Seems your roundtripping IS causing some unicode. Not sure why that is, but easy to fix. You cannot store unicode in a HDFStore Table in python 2, (this works correctly in python 3 however). You could do it as a Fixed format if you want though (it would be pickled). See here.

In [33]: df = pd.read_json(s)

In [25]: df
Out[25]: 
  args                date            host kwargs     operation  status   thingy      time
0   [] 2013-12-02 00:33:59  yy38.segm1.org     {}       x_gbinf    -101  a13yy38  0.000801
1   [] 2013-12-02 00:33:59  kyy1.segm1.org     {}     x_initobj       1  a19kyy1  0.003244
2   [] 2013-12-02 00:34:00  yy10.segm1.org     {}  x_gobjParams    -101  a14yy10  0.002247
3   [] 2013-12-02 00:34:00  yy24.segm1.org     {}        gtfull    -101  a14yy24  0.002787
4   [] 2013-12-02 00:34:00  yy24.segm1.org     {}       x_gbinf    -101  a14yy24  0.001067
5   [] 2013-12-02 00:34:00  yy34.segm1.org     {}       gxyzinf    -101  a12yy34  0.002652
6   [] 2013-12-02 00:34:00  yy15.segm1.org     {}     deletemfg       1  a15yy15  0.004371
7   [] 2013-12-02 00:34:00  yy15.segm1.org     {}       gxyzinf    -101  a15yy15  0.000602

[8 rows x 8 columns]

In [26]: df.dtypes
Out[26]: 
args                 object
date         datetime64[ns]
host                 object
kwargs               object
operation            object
status                int64
thingy               object
time                float64
dtype: object

This is inferring the actual type of the object dtyped Series. They will come out as unicode only if at least 1 string is unicode (otherwise they would be inferred as string)

In [27]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[27]: 
args            unicode
date         datetime64
host            unicode
kwargs          unicode
operation       unicode
status          integer
thingy          unicode
time           floating
dtype: object

Here's how to 'fix' it

In [28]: types = df.apply(lambda x: pd.lib.infer_dtype(x.values))

In [29]: types[types=='unicode']
Out[29]: 
args         unicode
host         unicode
kwargs       unicode
operation    unicode
thingy       unicode
dtype: object

In [30]: for col in types[types=='unicode'].index:
   ....:     df[col] = df[col].astype(str)
   ....:     

Looks the same

In [31]: df
Out[31]: 
  args                date            host kwargs     operation  status   thingy      time
0   [] 2013-12-02 00:33:59  yy38.segm1.org     {}       x_gbinf    -101  a13yy38  0.000801
1   [] 2013-12-02 00:33:59  kyy1.segm1.org     {}     x_initobj       1  a19kyy1  0.003244
2   [] 2013-12-02 00:34:00  yy10.segm1.org     {}  x_gobjParams    -101  a14yy10  0.002247
3   [] 2013-12-02 00:34:00  yy24.segm1.org     {}        gtfull    -101  a14yy24  0.002787
4   [] 2013-12-02 00:34:00  yy24.segm1.org     {}       x_gbinf    -101  a14yy24  0.001067
5   [] 2013-12-02 00:34:00  yy34.segm1.org     {}       gxyzinf    -101  a12yy34  0.002652
6   [] 2013-12-02 00:34:00  yy15.segm1.org     {}     deletemfg       1  a15yy15  0.004371
7   [] 2013-12-02 00:34:00  yy15.segm1.org     {}       gxyzinf    -101  a15yy15  0.000602

[8 rows x 8 columns]

But now infers correctly.

In [32]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[32]: 
args             string
date         datetime64
host             string
kwargs           string
operation        string
status          integer
thingy           string
time           floating
dtype: object
Answer from Jeff on Stack Overflow
Top answer
1 of 2
28

Seems your roundtripping IS causing some unicode. Not sure why that is, but easy to fix. You cannot store unicode in a HDFStore Table in python 2, (this works correctly in python 3 however). You could do it as a Fixed format if you want though (it would be pickled). See here.

In [33]: df = pd.read_json(s)

In [25]: df
Out[25]: 
  args                date            host kwargs     operation  status   thingy      time
0   [] 2013-12-02 00:33:59  yy38.segm1.org     {}       x_gbinf    -101  a13yy38  0.000801
1   [] 2013-12-02 00:33:59  kyy1.segm1.org     {}     x_initobj       1  a19kyy1  0.003244
2   [] 2013-12-02 00:34:00  yy10.segm1.org     {}  x_gobjParams    -101  a14yy10  0.002247
3   [] 2013-12-02 00:34:00  yy24.segm1.org     {}        gtfull    -101  a14yy24  0.002787
4   [] 2013-12-02 00:34:00  yy24.segm1.org     {}       x_gbinf    -101  a14yy24  0.001067
5   [] 2013-12-02 00:34:00  yy34.segm1.org     {}       gxyzinf    -101  a12yy34  0.002652
6   [] 2013-12-02 00:34:00  yy15.segm1.org     {}     deletemfg       1  a15yy15  0.004371
7   [] 2013-12-02 00:34:00  yy15.segm1.org     {}       gxyzinf    -101  a15yy15  0.000602

[8 rows x 8 columns]

In [26]: df.dtypes
Out[26]: 
args                 object
date         datetime64[ns]
host                 object
kwargs               object
operation            object
status                int64
thingy               object
time                float64
dtype: object

This is inferring the actual type of the object dtyped Series. They will come out as unicode only if at least 1 string is unicode (otherwise they would be inferred as string)

In [27]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[27]: 
args            unicode
date         datetime64
host            unicode
kwargs          unicode
operation       unicode
status          integer
thingy          unicode
time           floating
dtype: object

Here's how to 'fix' it

In [28]: types = df.apply(lambda x: pd.lib.infer_dtype(x.values))

In [29]: types[types=='unicode']
Out[29]: 
args         unicode
host         unicode
kwargs       unicode
operation    unicode
thingy       unicode
dtype: object

In [30]: for col in types[types=='unicode'].index:
   ....:     df[col] = df[col].astype(str)
   ....:     

Looks the same

In [31]: df
Out[31]: 
  args                date            host kwargs     operation  status   thingy      time
0   [] 2013-12-02 00:33:59  yy38.segm1.org     {}       x_gbinf    -101  a13yy38  0.000801
1   [] 2013-12-02 00:33:59  kyy1.segm1.org     {}     x_initobj       1  a19kyy1  0.003244
2   [] 2013-12-02 00:34:00  yy10.segm1.org     {}  x_gobjParams    -101  a14yy10  0.002247
3   [] 2013-12-02 00:34:00  yy24.segm1.org     {}        gtfull    -101  a14yy24  0.002787
4   [] 2013-12-02 00:34:00  yy24.segm1.org     {}       x_gbinf    -101  a14yy24  0.001067
5   [] 2013-12-02 00:34:00  yy34.segm1.org     {}       gxyzinf    -101  a12yy34  0.002652
6   [] 2013-12-02 00:34:00  yy15.segm1.org     {}     deletemfg       1  a15yy15  0.004371
7   [] 2013-12-02 00:34:00  yy15.segm1.org     {}       gxyzinf    -101  a15yy15  0.000602

[8 rows x 8 columns]

But now infers correctly.

In [32]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[32]: 
args             string
date         datetime64
host             string
kwargs           string
operation        string
status          integer
thingy           string
time           floating
dtype: object
2 of 2
8

The above solution may cause some errors with unicode special characters. A similar solution to convert unicode to string that will not get hung up on unicode special characters:

for col in types[types=='unicode'].index:
     df[col] = df[col].apply(lambda x: x.encode('utf-8').strip())

This is due in part to how python handles unicode. More info on that in the Python Unicode How-To.

🌐
Pandas
pandas.pydata.org › pandas-docs › version › 1.0 › user_guide › options.html
Options and settings — pandas 1.0.5 documentation
In addition, Unicode characters whose width is “Ambiguous” can either be 1 or 2 characters wide depending on the terminal setting or encoding.
Discussions

Unicode error while importing CSV using pandas
import pandas as pd df = pd.read_csv ('C:\Users\Acer\Desktop\fbreb.csv', encoding="utf8") print(df) More on reddit.com
🌐 r/learnpython
4
1
September 11, 2020
python - How to read UTF-8 files with Pandas? - Stack Overflow
Pandas stores strings in objects. In python 3, all string are in unicode by default. More on stackoverflow.com
🌐 stackoverflow.com
astype(unicode) does not work as expected
astype unicode seems to call str, so that the following code throws import pandas df = pandas.DataFrame({"somecol": [u"適当"]}) df["somecol"].astype("unicode") raises : UnicodeEncodeError: 'ascii' co... More on github.com
🌐 github.com
11
July 15, 2014
Support writing unicode characters in df.to_stata()
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more - Issue · pandas-dev/p... More on github.com
🌐 github.com
13
November 8, 2018
🌐
DEV Community
dev.to › _aadidev › 3-ways-to-handle-non-utf-8-characters-in-pandas-242
3 Ways to Handle non UTF-8 Characters in Pandas - DEV Community
January 20, 2022 - The number of unique characters that ASCII can handle is limited by the number of unique bytes (combinations of 1 and 0) available. However, to summarize: using 8 bits allows for 256 unique characters which is NO where close in handling every single character from every single language. This is where Unicode comes in; unicode assigns a "code points" in hexadecimal to each character.
🌐
Pandas
pandas.pydata.org › docs › dev › reference › api › pandas.DataFrame.to_string.html
pandas.DataFrame.to_string — pandas 3.0.0.dev0+2555.gf7447cc05e documentation
Formatter function to apply to columns’ elements if they are floats. This function must return a unicode string and will be applied only to the non-NaN elements, with NaN being handled by na_rep.
🌐
Reddit
reddit.com › r/learnpython › unicode error while importing csv using pandas
r/learnpython on Reddit: Unicode error while importing CSV using pandas
September 11, 2020 -

I tried to import a CSV file into my Jupyter notebook using pandas read_csv option. And I am getting Unicode error.

I'm using Windows 10, and the file name comes with back slashes and from what I understand its the cause of the error. I tired encoding="utf8" still not working.

I have added the code as a comment below.

How can I resolve this issue?

Would be really helpful. Have a nice day.

🌐
Medium
medium.com › @anala007 › dealing-with-the-unicodedecodeerror-in-pandas-when-reading-csv-files-edc4987bf68b
Dealing with the UnicodeDecodeError in Pandas When Reading CSV Files | by Arun | Medium
June 12, 2023 - import pandas as pd df = pd.read_csv('file.csv', encoding='utf-8-sig') This could be a handy trick when dealing with files that have minor issues with their encoding. While encountering a UnicodeDecodeError when reading a CSV file with Pandas can be quite annoying, several strategies can help you overcome this hurdle.
Find elsewhere
🌐
Apache JIRA
issues.apache.org › jira › browse › ARROW-1976
[ARROW-1976] [Python] Handling unicode pandas columns on parquet.read_table - ASF Jira
January 11, 2023 - Unicode columns in pandas DataFrames aren't being handled correctly for some datasets when reading a parquet file into a pandas DataFrame, leading to the common Python ASCII encoding error.
🌐
Quora
quora.com › How-do-I-fix-a-Unicode-error-while-reading-a-CSV-file-with-a-pandas-library-in-Python-3-6
How to fix a Unicode error while reading a CSV file with a pandas library in Python 3.6 - Quora
Quora is a place to gain and share knowledge. It's a platform to ask questions and connect with people who contribute unique insights and quality answers.
🌐
GitHub
github.com › pandas-dev › pandas › issues › 7758
astype(unicode) does not work as expected · Issue #7758 · pandas-dev/pandas
July 15, 2014 - astype unicode seems to call str, so that the following code throws import pandas df = pandas.DataFrame({"somecol": [u"適当"]}) df["somecol"].astype("unicode") raises : UnicodeEncodeError: 'ascii' codec can't encode ch aracters in position...
Author: pandas-dev
🌐
Pandas How To
pandashowto.com › pandas how to › data input and output › how to handle different encodings (utf-8, latin-1, etc.) in pandas • pandas how to
How To Handle Different Encodings (UTF-8, Latin-1, Etc.) In Pandas • Pandas How To
March 30, 2025 - import pandas as pd df = pd.read_csv('your_file.csv', encoding='latin-1') print(df) In this example, the encoding parameter is set to ‘latin-1′. If you were working with a UTF-8 file, you would use encoding=’utf-8′. If you encounter an UnicodeDecodeError, it usually means Pandas is trying to decode the file using the wrong encoding.
🌐
GitHub
github.com › pandas-dev › pandas › issues › 23573
Support writing unicode characters in df.to_stata() #23573
November 8, 2018 - import pandas as pd df = pd.DataFrame({'a': ['丆']}) df.to_stata('test.dta') # UnicodeEncodeError: 'latin-1' codec can't encode character '\u4e06' in position 0: ordinal not in range(256) I picked an arbitrary CJK character to test this with. It would be possible to write Unicode strings to a Stata file by implementing a writer according to version 118 of the dta format.
Author: pandas-dev
🌐
Medium
medium.com › @deeptibhatia › how-to-read-utf-8-characters-using-pandas-in-python-machine-learning-course-by-hackveda-88812b1433a1
How to read utf-8 characters using pandas in python Machine Learning course by Hackveda ! | by Deepti Bhatia | Medium
September 9, 2018 - Pandas help us to read any file just by copying the link from where it is downloaded ending with the file name with its extension . ... Encoding to use for UTF when reading/writing (ex. ‘utf-8’). List of Python standard encodings . UTF-8 is a compromise character encoding that can be as compact as ASCII (if the file is just plain English text) but can also contain any unicode ...
🌐
Unicode Compart
compart.com › en › unicode › U+1F43C
“🐼” U+1F43C Panda Face Unicode Character
U+1F43C is the unicode hex value of the character Panda Face. Char U+1F43C, Encodings, HTML Entitys:🐼,🐼, UTF-8 (hex), UTF-16 (hex), UTF-32 (hex)
🌐
pythontutorials
pythontutorials.net › blog › writing-pandas-dataframe-to-json-in-unicode
How to Write Pandas DataFrame to JSON with Unicode Characters Without Escaping (Fixing force_ascii=False Error) — pythontutorials.net
By default, JSON allows Unicode characters, but many tools (including Pandas) escape non-ASCII characters into \uXXXX sequences for compatibility with systems that only support ASCII.
🌐
Plain English
python.plainenglish.io › easy-ways-to-handle-unicodedecodeerrors-when-reading-csv-files-in-pandas-84c7aad1d5ac
Easy Ways To Handle UnicodeDecodeErrors When Reading CSV Files in Pandas | by Gabriel Ejiro | Python in Plain English
November 10, 2023 - Pandas usually reads CSV files with ‘utf-8’ encoding by default. However, if your file uses a different encoding, you might meet the formidable ‘UnicodeDecodeError.’ To avoid this, a simple fix is to determine/identify your file’s encoding and explicitly mention it in your code.
🌐
GitHub
github.com › pandas-dev › pandas › issues › 25287
Python 2: pandas.Series.eq() crashes with unicode data · Issue #25287 · pandas-dev/pandas
February 12, 2019 - Code Sample #!/usr/bin/env python # -*- coding: utf-8 -*- import pandas as pd weird_string = u"føøbår" x = pd.DataFrame({"a" : [weird_string, "xx"]}) bool_series = (x["a"].eq(weird_string)) print(x[bool_series].to_string()) assert len(x[...
Author: pandas-dev