Here's a list of available python 3 encodings -

https://docs.python.org/3/library/codecs.html#standard-encodings

I don't think pandas includes or excludes any additional encodings.

Answer from Shashank Agarwal on Stack Overflow
Top answer
1 of 2
10

Here's a list of available python 3 encodings -

https://docs.python.org/3/library/codecs.html#standard-encodings

I don't think pandas includes or excludes any additional encodings.

2 of 2
9

I wrote a simple checker for all encodings types. My data included problematic signs in headers so I use

df.info()

function to check if everything is correct.

import pandas as pd

codecs = ['ascii','big5','big5hkscs','cp037','cp273','cp424','cp437','cp500','cp720','cp737','cp775','cp850','cp852','cp855',
          'cp856','cp857','cp858','cp860','cp861','cp862','cp863','cp864','cp865','cp866','cp869','cp874','cp875','cp932','cp949',
          'cp950','cp1006','cp1026','cp1125','cp1140','cp1250','cp1251','cp1252','cp1253','cp1254','cp1255','cp1256','cp1257','cp1258',
          'euc_jp','euc_jis_2004','euc_jisx0213','euc_kr','gb2312','gbk','gb18030','hz','iso2022_jp','iso2022_jp_1','iso2022_jp_2',
          'iso2022_jp_2004','iso2022_jp_3','iso2022_jp_ext','iso2022_kr','latin_1','iso8859_2','iso8859_3','iso8859_4','iso8859_5','iso8859_6',
          'iso8859_7','iso8859_8','iso8859_9','iso8859_10','iso8859_11','iso8859_13','iso8859_14','iso8859_15','iso8859_16','johab','koi8_r','koi8_t',
          'koi8_u','kz1048','mac_cyrillic','mac_greek','mac_iceland','mac_latin2','mac_roman','mac_turkish','ptcp154','shift_jis','shift_jis_2004',
          'shift_jisx0213','utf_32','utf_32_be','utf_32_le','utf_16','utf_16_be','utf_16_le','utf_7','utf_8','utf_8_sig']


for x in range(len(codecs)):
    print(x,': Now checking use of:', codecs[x])
    try:
        df = pd.read_csv('*your_csv_file*.csv', header = 0, encoding = (codecs[x]), sep=';')
        print(df.info())
        print(input('Press any key...'))
    except:
        print('I can\'t load data for', codecs[x], '\n')
        print(input('Press any key...'))

Remember about giving a sep parameter it also helps.

๐ŸŒ
Saturn Cloud
saturncloud.io โ€บ blog โ€บ a-list-of-pandas-readcsv-encoding-options
A List of Pandas readcsv Encoding Options | Saturn Cloud Blog
May 1, 2026 - UTF-8 is the most widely used encoding format for text data. It supports all characters in the Unicode standard and is compatible with ASCII. To specify UTF-8 encoding in Pandas, use the encoding='utf-8' parameter.
Discussions

to_csv with lists of strings and unicode encoding produces wrong output
IO CSVread_csv, to_csvread_csv, to_csvOutput-Formatting__repr__ of pandas objects, to_string__repr__ of pandas objects, to_string ... If I have a dataframe with cells containing lists of strings (or unicode strings), then these lists are broken when I use to_csv() with the encoding parameter set. More on github.com
๐ŸŒ github.com
11
August 13, 2015
python - How to find out which encoding to use in Pandas - Stack Overflow
I am trying to open a .CSV file in Pandas, but I keep getting an encoding error. I have literally tried all possible encoding codes, but none of them work: encode_list = ['ascii','big5','big5hkscs',' More on stackoverflow.com
๐ŸŒ stackoverflow.com
March 22, 2022
python - LabelEncoding in Pandas on a column with list of strings across rows - Stack Overflow
I would like to LabelEncode a column in pandas where each row contains a list of strings. Since a similar string/text carries a same meaning across rows, encoding should respect that, and ideally e... More on stackoverflow.com
๐ŸŒ stackoverflow.com
python - How to achieve this encoding in pandas dataframe - Stack Overflow
I have a dataframe df : Number Master 1 Apple 2 Orange 3 Pineapple 4 Strawberrry 5 Blueberry 6 Plums 7 Cherry 8 Dragonfruit 9 Iceapple 10 Litchie This is just a sample df . original dataframe has 1... More on stackoverflow.com
๐ŸŒ stackoverflow.com
July 19, 2022
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ one-hot-encoding-from-a-pandas-column-containing-a-list
One-Hot-Encoding from a Pandas Column Containing a List - GeeksforGeeks
August 22, 2024 - It converts categorical variables into a binary matrix representation, where each category is represented by a separate column. This article will guide you through the process of one-hot encoding a Pandas column containing a list of elements, a common scenario in data analysis and machine learning.
๐ŸŒ
GitHub
github.com โ€บ pandas-dev โ€บ pandas โ€บ issues โ€บ 10813
to_csv with lists of strings and unicode encoding produces wrong output ยท Issue #10813 ยท pandas-dev/pandas
August 13, 2015 - IO CSVread_csv, to_csvread_csv, to_csvOutput-Formatting__repr__ of pandas objects, to_string__repr__ of pandas objects, to_string ... If I have a dataframe with cells containing lists of strings (or unicode strings), then these lists are broken when I use to_csv() with the encoding parameter set.
Author: pandas-dev
๐ŸŒ
Pandas
pandas.pydata.org โ€บ docs โ€บ search.html
Search - pandas 3.0.2 documentation
Created using Sphinx 9.0.4 ยท Built with the PyData Sphinx Theme 0.16.1
Find elsewhere
๐ŸŒ
Practical Business Python
pbpython.com โ€บ categorical-encoding.html
Guide to Encoding Categorical Values in Python - Practical Business Python
A common alternative approach is called one hot encoding (but also goes by several different names shown below). Despite the different names, the basic strategy is to convert each category value into a new column and assigns a 1 or 0 (True/False) value to the column. This has the benefit of not weighting a value improperly but does have the downside of adding more columns to the data set. Pandas supports this feature using get_dummies.
๐ŸŒ
Data Science for Everyone
matthew-brett.github.io โ€บ cfd2019 โ€บ chapters โ€บ 07 โ€บ text_encoding
6.9 Text encoding - Coding for Data - 2019 edition
August 14, 2020 - In fact, Pandas assumes that text is in UTF-8 format, because it is so common. In this case, as the filename suggests, the bytes for the text are in Latin 1 encoding.
๐ŸŒ
Stack Abuse
stackabuse.com โ€บ one-hot-encoding-in-python-with-pandas-and-scikit-learn
One-Hot Encoding in Python with Pandas and Scikit-Learn
July 31, 2021 - Let's take a look at a simple example of how we can convert values from a categorical column in our dataset into their numerical counterparts, via the one-hot encoding scheme. We'll be creating a really simple dataset - a list of countries and their ID's: import pandas as pd ids = [11, 22, 33, 44, 55, 66, 77] countries = ['Spain', 'France', 'Spain', 'Germany', 'France'] df = pd.DataFrame(list(zip(ids, countries)), columns=['Ids', 'Countries'])
๐ŸŒ
KDnuggets
kdnuggets.com โ€บ 2023 โ€บ 07 โ€บ pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - For this, we use pandas.get_dummies() method. It takes the following arguments: To better understand the function, let us work on one-hot encoding the dummy dataset. We use the get_dummies method and pass the original data frame as data input. In columns, we pass a list containing only the categorical_column header.
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ pandas โ€บ python-pandas-series-str-encode
Python | Pandas Series.str.encode() - GeeksforGeeks
March 27, 2019 - Use 'raw_unicode_escape' for encoding. ... # importing pandas as pd import pandas as pd # Creating the Series sr = pd.Series(['New_York', 'Lisbon', 'Tokyo', 'Paris', 'Munich']) # Creating the index idx = ['City 1', 'City 2', 'City 3', 'City 4', 'City 5'] # set the index sr.index = idx # Print the series print(sr) Output : Now we will use Series.str.encode() function to encode the character strings present in the underlying data of the given series object.
๐ŸŒ
GitHub
github.com โ€บ jvkersch โ€บ pandas-encoding
GitHub - jvkersch/pandas-encoding: A natural encoding for Giant Pandas
February 17, 2019 - $ PYTHONIOENCODING=pandas python Python 3.6.7 (tags/v3.6.7:6ec5cf24b7, Feb 16 2019, 22:29:37) [GCC 7.3.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> import pandas as ๐Ÿผ >>> df = ๐Ÿผ.DataFrame() >>> df Empty DataFrame Columns: [] Index: [] You can also use this encoding directly, by including the line # -*- coding: pandas -*- at the top of your script:
Author: jvkersch