If you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Answer from BKS on Stack OverflowIf you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Setting up the encoding should be treated when reading the input file, using the option encoding
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8")
or if the file uses BOM,
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8-sig")
Encoding issues when converting a dataframe from pd.read_pickle using dd.from_pandas
Pandas DataFrame.to_csv() struggling with encoding
python - How to achieve this encoding in pandas dataframe - Stack Overflow
python - Pandas convert dataframe to Utf-8 - Stack Overflow
Heya,
I got a DataFrame filled with strings in the first column and the rest consisting of integers (except for the headers).
Now when I export this dataframe to a csv file, and the strings contain German Umlauts (ä,ö,ü or something like ß), the exported csv file has weird looking strings at these indices.
Like "für" became "für".
As far as I know the default encoding for to_csv() is utf-8, which means it should work fine? But I also tried the parameter encoding='utf-8', same results.
What am I doing wrong here?
Create dictionary for top N values by counts in column Master by Series.value_counts with Series.head and use them for Series.map with replace not matched values to N + 1 in Series.fillna:
N = 4
d = {v: k+1 for k, v in enumerate(df['Master'].value_counts().head(N).index)}
print (d)
df['Master'] = df['Master'].map(d).fillna(N + 1).astype(int)
If you have list of top values by list:
L = ['Apple','Pineapple','Orange','Strawberrry']
d = {v: k+1 for k, v in enumerate(L)}
print (d)
{'Apple': 1, 'Pineapple': 2, 'Orange': 3, 'Strawberrry': 4}
df['Master'] = df['Master'].map(d).fillna(len(L) + 1).astype(int)
print (df)
Number Master
0 1 1
1 2 3
2 3 2
3 4 4
4 5 5
5 6 5
6 7 5
7 8 5
8 9 5
9 10 5
You can use dict and Series.map and fillna(5) for keys that don't exist in dct.
dct = {'Apple':1, 'Pineapple':2,'Orange':3 , 'Strawberrry':4}
df['Master'] = df['Master'].map(dct).fillna(5).astype(int)
print(df)
Number Master
0 1 1
1 2 3
2 3 2
3 4 4
4 5 5
5 6 5
6 7 5
7 8 5
8 9 5
9 10 5
I'm trying to read in a CSV file using pandas.read\_file(report), but I'm hitting this error message:
'utf-8' codec can't decode byte 0x93 in position 28: invalid start byte
I've also tried this with various encoding types, pandas.read_file(report, encoding='UTF-16')
but for some reason it has no effect on the error, it always shows as 'utf-8' The encoding line shows in my traceback so I know that it's there. Does anyone know why this is?
If you have a look at the documentation for read_csv() you'll find that you can use the argument encoding_errors='ignore' to ignore those encoding errors and move on with the import. This should allow you to open the file with the most appropriate codec.
Other suitable values for this argument can be found in the python codecs documentation.
I suppose if your Excel opens the data the case isn't with encoding. You can get its's encoding importing 'locale' and asking locale.getpreferredencoding(). Look at header row while opening data in text redactor to find out if any field has escaped characters like '\t' and also look at a csv delimiter (the default in read_csv is ',', your may have another)