Here's a list of available python 3 encodings -
https://docs.python.org/3/library/codecs.html#standard-encodings
I don't think pandas includes or excludes any additional encodings.
Answer from Shashank Agarwal on Stack OverflowHere's a list of available python 3 encodings -
https://docs.python.org/3/library/codecs.html#standard-encodings
I don't think pandas includes or excludes any additional encodings.
I wrote a simple checker for all encodings types. My data included problematic signs in headers so I use
df.info()
function to check if everything is correct.
import pandas as pd
codecs = ['ascii','big5','big5hkscs','cp037','cp273','cp424','cp437','cp500','cp720','cp737','cp775','cp850','cp852','cp855',
'cp856','cp857','cp858','cp860','cp861','cp862','cp863','cp864','cp865','cp866','cp869','cp874','cp875','cp932','cp949',
'cp950','cp1006','cp1026','cp1125','cp1140','cp1250','cp1251','cp1252','cp1253','cp1254','cp1255','cp1256','cp1257','cp1258',
'euc_jp','euc_jis_2004','euc_jisx0213','euc_kr','gb2312','gbk','gb18030','hz','iso2022_jp','iso2022_jp_1','iso2022_jp_2',
'iso2022_jp_2004','iso2022_jp_3','iso2022_jp_ext','iso2022_kr','latin_1','iso8859_2','iso8859_3','iso8859_4','iso8859_5','iso8859_6',
'iso8859_7','iso8859_8','iso8859_9','iso8859_10','iso8859_11','iso8859_13','iso8859_14','iso8859_15','iso8859_16','johab','koi8_r','koi8_t',
'koi8_u','kz1048','mac_cyrillic','mac_greek','mac_iceland','mac_latin2','mac_roman','mac_turkish','ptcp154','shift_jis','shift_jis_2004',
'shift_jisx0213','utf_32','utf_32_be','utf_32_le','utf_16','utf_16_be','utf_16_le','utf_7','utf_8','utf_8_sig']
for x in range(len(codecs)):
print(x,': Now checking use of:', codecs[x])
try:
df = pd.read_csv('*your_csv_file*.csv', header = 0, encoding = (codecs[x]), sep=';')
print(df.info())
print(input('Press any key...'))
except:
print('I can\'t load data for', codecs[x], '\n')
print(input('Press any key...'))
Remember about giving a sep parameter it also helps.
to_csv with lists of strings and unicode encoding produces wrong output
python - How to find out which encoding to use in Pandas - Stack Overflow
python - LabelEncoding in Pandas on a column with list of strings across rows - Stack Overflow
python - How to achieve this encoding in pandas dataframe - Stack Overflow
If you have a look at the documentation for read_csv() you'll find that you can use the argument encoding_errors='ignore' to ignore those encoding errors and move on with the import. This should allow you to open the file with the most appropriate codec.
Other suitable values for this argument can be found in the python codecs documentation.
I suppose if your Excel opens the data the case isn't with encoding. You can get its's encoding importing 'locale' and asking locale.getpreferredencoding(). Look at header row while opening data in text redactor to find out if any field has escaped characters like '\t' and also look at a csv delimiter (the default in read_csv is ',', your may have another)
Create dictionary for top N values by counts in column Master by Series.value_counts with Series.head and use them for Series.map with replace not matched values to N + 1 in Series.fillna:
N = 4
d = {v: k+1 for k, v in enumerate(df['Master'].value_counts().head(N).index)}
print (d)
df['Master'] = df['Master'].map(d).fillna(N + 1).astype(int)
If you have list of top values by list:
L = ['Apple','Pineapple','Orange','Strawberrry']
d = {v: k+1 for k, v in enumerate(L)}
print (d)
{'Apple': 1, 'Pineapple': 2, 'Orange': 3, 'Strawberrry': 4}
df['Master'] = df['Master'].map(d).fillna(len(L) + 1).astype(int)
print (df)
Number Master
0 1 1
1 2 3
2 3 2
3 4 4
4 5 5
5 6 5
6 7 5
7 8 5
8 9 5
9 10 5
You can use dict and Series.map and fillna(5) for keys that don't exist in dct.
dct = {'Apple':1, 'Pineapple':2,'Orange':3 , 'Strawberrry':4}
df['Master'] = df['Master'].map(dct).fillna(5).astype(int)
print(df)
Number Master
0 1 1
1 2 3
2 3 2
3 4 4
4 5 5
5 6 5
6 7 5
7 8 5
8 9 5
9 10 5