Seems your roundtripping IS causing some unicode. Not sure why that is, but easy to fix. You cannot store unicode in a HDFStore Table in python 2, (this works correctly in python 3 however). You could do it as a Fixed format if you want though (it would be pickled). See here.
In [33]: df = pd.read_json(s)
In [25]: df
Out[25]:
args date host kwargs operation status thingy time
0 [] 2013-12-02 00:33:59 yy38.segm1.org {} x_gbinf -101 a13yy38 0.000801
1 [] 2013-12-02 00:33:59 kyy1.segm1.org {} x_initobj 1 a19kyy1 0.003244
2 [] 2013-12-02 00:34:00 yy10.segm1.org {} x_gobjParams -101 a14yy10 0.002247
3 [] 2013-12-02 00:34:00 yy24.segm1.org {} gtfull -101 a14yy24 0.002787
4 [] 2013-12-02 00:34:00 yy24.segm1.org {} x_gbinf -101 a14yy24 0.001067
5 [] 2013-12-02 00:34:00 yy34.segm1.org {} gxyzinf -101 a12yy34 0.002652
6 [] 2013-12-02 00:34:00 yy15.segm1.org {} deletemfg 1 a15yy15 0.004371
7 [] 2013-12-02 00:34:00 yy15.segm1.org {} gxyzinf -101 a15yy15 0.000602
[8 rows x 8 columns]
In [26]: df.dtypes
Out[26]:
args object
date datetime64[ns]
host object
kwargs object
operation object
status int64
thingy object
time float64
dtype: object
This is inferring the actual type of the object dtyped Series. They will come out as unicode only if at least 1 string is unicode (otherwise they would be inferred as string)
In [27]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[27]:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
status integer
thingy unicode
time floating
dtype: object
Here's how to 'fix' it
In [28]: types = df.apply(lambda x: pd.lib.infer_dtype(x.values))
In [29]: types[types=='unicode']
Out[29]:
args unicode
host unicode
kwargs unicode
operation unicode
thingy unicode
dtype: object
In [30]: for col in types[types=='unicode'].index:
....: df[col] = df[col].astype(str)
....:
Looks the same
In [31]: df
Out[31]:
args date host kwargs operation status thingy time
0 [] 2013-12-02 00:33:59 yy38.segm1.org {} x_gbinf -101 a13yy38 0.000801
1 [] 2013-12-02 00:33:59 kyy1.segm1.org {} x_initobj 1 a19kyy1 0.003244
2 [] 2013-12-02 00:34:00 yy10.segm1.org {} x_gobjParams -101 a14yy10 0.002247
3 [] 2013-12-02 00:34:00 yy24.segm1.org {} gtfull -101 a14yy24 0.002787
4 [] 2013-12-02 00:34:00 yy24.segm1.org {} x_gbinf -101 a14yy24 0.001067
5 [] 2013-12-02 00:34:00 yy34.segm1.org {} gxyzinf -101 a12yy34 0.002652
6 [] 2013-12-02 00:34:00 yy15.segm1.org {} deletemfg 1 a15yy15 0.004371
7 [] 2013-12-02 00:34:00 yy15.segm1.org {} gxyzinf -101 a15yy15 0.000602
[8 rows x 8 columns]
But now infers correctly.
In [32]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[32]:
args string
date datetime64
host string
kwargs string
operation string
status integer
thingy string
time floating
dtype: object
Answer from Jeff on Stack OverflowSeems your roundtripping IS causing some unicode. Not sure why that is, but easy to fix. You cannot store unicode in a HDFStore Table in python 2, (this works correctly in python 3 however). You could do it as a Fixed format if you want though (it would be pickled). See here.
In [33]: df = pd.read_json(s)
In [25]: df
Out[25]:
args date host kwargs operation status thingy time
0 [] 2013-12-02 00:33:59 yy38.segm1.org {} x_gbinf -101 a13yy38 0.000801
1 [] 2013-12-02 00:33:59 kyy1.segm1.org {} x_initobj 1 a19kyy1 0.003244
2 [] 2013-12-02 00:34:00 yy10.segm1.org {} x_gobjParams -101 a14yy10 0.002247
3 [] 2013-12-02 00:34:00 yy24.segm1.org {} gtfull -101 a14yy24 0.002787
4 [] 2013-12-02 00:34:00 yy24.segm1.org {} x_gbinf -101 a14yy24 0.001067
5 [] 2013-12-02 00:34:00 yy34.segm1.org {} gxyzinf -101 a12yy34 0.002652
6 [] 2013-12-02 00:34:00 yy15.segm1.org {} deletemfg 1 a15yy15 0.004371
7 [] 2013-12-02 00:34:00 yy15.segm1.org {} gxyzinf -101 a15yy15 0.000602
[8 rows x 8 columns]
In [26]: df.dtypes
Out[26]:
args object
date datetime64[ns]
host object
kwargs object
operation object
status int64
thingy object
time float64
dtype: object
This is inferring the actual type of the object dtyped Series. They will come out as unicode only if at least 1 string is unicode (otherwise they would be inferred as string)
In [27]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[27]:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
status integer
thingy unicode
time floating
dtype: object
Here's how to 'fix' it
In [28]: types = df.apply(lambda x: pd.lib.infer_dtype(x.values))
In [29]: types[types=='unicode']
Out[29]:
args unicode
host unicode
kwargs unicode
operation unicode
thingy unicode
dtype: object
In [30]: for col in types[types=='unicode'].index:
....: df[col] = df[col].astype(str)
....:
Looks the same
In [31]: df
Out[31]:
args date host kwargs operation status thingy time
0 [] 2013-12-02 00:33:59 yy38.segm1.org {} x_gbinf -101 a13yy38 0.000801
1 [] 2013-12-02 00:33:59 kyy1.segm1.org {} x_initobj 1 a19kyy1 0.003244
2 [] 2013-12-02 00:34:00 yy10.segm1.org {} x_gobjParams -101 a14yy10 0.002247
3 [] 2013-12-02 00:34:00 yy24.segm1.org {} gtfull -101 a14yy24 0.002787
4 [] 2013-12-02 00:34:00 yy24.segm1.org {} x_gbinf -101 a14yy24 0.001067
5 [] 2013-12-02 00:34:00 yy34.segm1.org {} gxyzinf -101 a12yy34 0.002652
6 [] 2013-12-02 00:34:00 yy15.segm1.org {} deletemfg 1 a15yy15 0.004371
7 [] 2013-12-02 00:34:00 yy15.segm1.org {} gxyzinf -101 a15yy15 0.000602
[8 rows x 8 columns]
But now infers correctly.
In [32]: df.apply(lambda x: pd.lib.infer_dtype(x.values))
Out[32]:
args string
date datetime64
host string
kwargs string
operation string
status integer
thingy string
time floating
dtype: object
The above solution may cause some errors with unicode special characters. A similar solution to convert unicode to string that will not get hung up on unicode special characters:
for col in types[types=='unicode'].index:
df[col] = df[col].apply(lambda x: x.encode('utf-8').strip())
This is due in part to how python handles unicode. More info on that in the Python Unicode How-To.
Unicode error while importing CSV using pandas
python - How to read UTF-8 files with Pandas? - Stack Overflow
astype(unicode) does not work as expected
Support writing unicode characters in df.to_stata()
I tried to import a CSV file into my Jupyter notebook using pandas read_csv option. And I am getting Unicode error.
I'm using Windows 10, and the file name comes with back slashes and from what I understand its the cause of the error. I tired encoding="utf8" still not working.
I have added the code as a comment below.
How can I resolve this issue?
Would be really helpful. Have a nice day.
As the other poster mentioned, you might try:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
However this could still leave you looking at 'object' when you print the dtypes. To confirm they are utf8, try this line after reading the CSV:
df.apply(lambda x: pd.lib.infer_dtype(x.values))
Example output:
args unicode
date datetime64
host unicode
kwargs unicode
operation unicode
Use the encoding keyword with the appropriate parameter:
df = pd.read_csv('1459966468_324.csv', encoding='utf8')
You have unicode values in your DataFrame. Files store bytes, which means all unicode have to be encoded into bytes before they can be stored in a file. You have to specify an encoding, such as utf-8. For example,
df.to_csv('path', header=True, index=False, encoding='utf-8')
If you don't specify an encoding, then the encoding used by df.to_csv defaults to ascii in Python2, or utf-8 in Python3.
Adding an answer to help myself google it later:
One trick that helped me is to encode a problematic series first, then decode it back to utf-8. Like:
df['crumbs'] = df['crumbs'].map(lambda x: x.encode('unicode-escape').decode('utf-8'))
This would get the dataframe to print correctly too.