If you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Answer from BKS on Stack OverflowIf you're not using csv, and you want to encode your string index, this is what worked for me:
df.index = df.index.str.encode('utf-8')
Setting up the encoding should be treated when reading the input file, using the option encoding
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8")
or if the file uses BOM,
df = pd.read_csv('bibliography.csv', delimiter=',', encoding="utf-8-sig")
Pandas DataFrame.to_csv() struggling with encoding
Encoding issues when converting a dataframe from pd.read_pickle using dd.from_pandas
python - How to achieve this encoding in pandas dataframe - Stack Overflow
csv - Encoding Error in Panda read_csv - Stack Overflow
Heya,
I got a DataFrame filled with strings in the first column and the rest consisting of integers (except for the headers).
Now when I export this dataframe to a csv file, and the strings contain German Umlauts (ä,ö,ü or something like ß), the exported csv file has weird looking strings at these indices.
Like "für" became "für".
As far as I know the default encoding for to_csv() is utf-8, which means it should work fine? But I also tried the parameter encoding='utf-8', same results.
What am I doing wrong here?
Create dictionary for top N values by counts in column Master by Series.value_counts with Series.head and use them for Series.map with replace not matched values to N + 1 in Series.fillna:
N = 4
d = {v: k+1 for k, v in enumerate(df['Master'].value_counts().head(N).index)}
print (d)
df['Master'] = df['Master'].map(d).fillna(N + 1).astype(int)
If you have list of top values by list:
L = ['Apple','Pineapple','Orange','Strawberrry']
d = {v: k+1 for k, v in enumerate(L)}
print (d)
{'Apple': 1, 'Pineapple': 2, 'Orange': 3, 'Strawberrry': 4}
df['Master'] = df['Master'].map(d).fillna(len(L) + 1).astype(int)
print (df)
Number Master
0 1 1
1 2 3
2 3 2
3 4 4
4 5 5
5 6 5
6 7 5
7 8 5
8 9 5
9 10 5
You can use dict and Series.map and fillna(5) for keys that don't exist in dct.
dct = {'Apple':1, 'Pineapple':2,'Orange':3 , 'Strawberrry':4}
df['Master'] = df['Master'].map(dct).fillna(5).astype(int)
print(df)
Number Master
0 1 1
1 2 3
2 3 2
3 4 4
4 5 5
5 6 5
6 7 5
7 8 5
8 9 5
9 10 5