Here's a list of available python 3 encodings -

https://docs.python.org/3/library/codecs.html#standard-encodings

I don't think pandas includes or excludes any additional encodings.

Answer from Shashank Agarwal on Stack Overflow
Top answer
1 of 2
10

Here's a list of available python 3 encodings -

https://docs.python.org/3/library/codecs.html#standard-encodings

I don't think pandas includes or excludes any additional encodings.

2 of 2
9

I wrote a simple checker for all encodings types. My data included problematic signs in headers so I use

df.info()

function to check if everything is correct.

import pandas as pd

codecs = ['ascii','big5','big5hkscs','cp037','cp273','cp424','cp437','cp500','cp720','cp737','cp775','cp850','cp852','cp855',
          'cp856','cp857','cp858','cp860','cp861','cp862','cp863','cp864','cp865','cp866','cp869','cp874','cp875','cp932','cp949',
          'cp950','cp1006','cp1026','cp1125','cp1140','cp1250','cp1251','cp1252','cp1253','cp1254','cp1255','cp1256','cp1257','cp1258',
          'euc_jp','euc_jis_2004','euc_jisx0213','euc_kr','gb2312','gbk','gb18030','hz','iso2022_jp','iso2022_jp_1','iso2022_jp_2',
          'iso2022_jp_2004','iso2022_jp_3','iso2022_jp_ext','iso2022_kr','latin_1','iso8859_2','iso8859_3','iso8859_4','iso8859_5','iso8859_6',
          'iso8859_7','iso8859_8','iso8859_9','iso8859_10','iso8859_11','iso8859_13','iso8859_14','iso8859_15','iso8859_16','johab','koi8_r','koi8_t',
          'koi8_u','kz1048','mac_cyrillic','mac_greek','mac_iceland','mac_latin2','mac_roman','mac_turkish','ptcp154','shift_jis','shift_jis_2004',
          'shift_jisx0213','utf_32','utf_32_be','utf_32_le','utf_16','utf_16_be','utf_16_le','utf_7','utf_8','utf_8_sig']


for x in range(len(codecs)):
    print(x,': Now checking use of:', codecs[x])
    try:
        df = pd.read_csv('*your_csv_file*.csv', header = 0, encoding = (codecs[x]), sep=';')
        print(df.info())
        print(input('Press any key...'))
    except:
        print('I can\'t load data for', codecs[x], '\n')
        print(input('Press any key...'))

Remember about giving a sep parameter it also helps.

🌐
Pandas
pandas.pydata.org › docs › reference › api › pandas.read_csv.html
pandas.read_csv — pandas 3.0.6 documentation - PyData |
Encoding to use for UTF when reading/writing (ex. 'utf-8'). List of Python standard encodings .
🌐
Saturn Cloud
saturncloud.io › blog › a-list-of-pandas-readcsv-encoding-options
A List of Pandas readcsv Encoding Options | Saturn Cloud Blog
May 1, 2026 - UTF-8 is the most widely used encoding format for text data. It supports all characters in the Unicode standard and is compatible with ASCII. To specify UTF-8 encoding in Pandas, use the encoding='utf-8' parameter.
🌐
GeeksforGeeks
geeksforgeeks.org › one-hot-encoding-from-a-pandas-column-containing-a-list
One-Hot-Encoding from a Pandas Column Containing a List - GeeksforGeeks
August 22, 2024 - It converts categorical variables into a binary matrix representation, where each category is represented by a separate column. This article will guide you through the process of one-hot encoding a Pandas column containing a list of elements, a common scenario in data analysis and machine learning.
🌐
Pandas
pandas.pydata.org › docs › search.html
Search - pandas 3.0.2 documentation
Created using Sphinx 9.0.4 · Built with the PyData Sphinx Theme 0.16.1
🌐
GitHub
github.com › pandas-dev › pandas › issues › 10813
to_csv with lists of strings and unicode encoding produces wrong output · Issue #10813 · pandas-dev/pandas
August 13, 2015 - IO CSVread_csv, to_csvread_csv, to_csvOutput-Formatting__repr__ of pandas objects, to_string__repr__ of pandas objects, to_string ... If I have a dataframe with cells containing lists of strings (or unicode strings), then these lists are broken when I use to_csv() with the encoding parameter set.
Author: pandas-dev
Find elsewhere
🌐
Practical Business Python
pbpython.com › categorical-encoding.html
Guide to Encoding Categorical Values in Python - Practical Business Python
A common alternative approach is called one hot encoding (but also goes by several different names shown below). Despite the different names, the basic strategy is to convert each category value into a new column and assigns a 1 or 0 (True/False) value to the column. This has the benefit of not weighting a value improperly but does have the downside of adding more columns to the data set. Pandas supports this feature using get_dummies.
🌐
Data Science for Everyone
matthew-brett.github.io › cfd2019 › chapters › 07 › text_encoding
6.9 Text encoding - Coding for Data - 2019 edition
August 14, 2020 - In fact, Pandas assumes that text is in UTF-8 format, because it is so common. In this case, as the filename suggests, the bytes for the text are in Latin 1 encoding.
🌐
GitHub
github.com › pandas-dev › pandas › blob › main › pandas › core › reshape › encoding.py
pandas/pandas/core/reshape/encoding.py at main · pandas-dev/pandas
>>> pd.get_dummies(pd.Series(list("abc")), dtype=float) a b c · 0 1.0 0.0 0.0 · 1 0.0 1.0 0.0 · 2 0.0 0.0 1.0 · """ from pandas.core.reshape.concat import concat · · def _is_encodable(arr) -> bool: # The columns get_dummies encodes by default: object, string (any ·
Author: pandas-dev
🌐
Stack Abuse
stackabuse.com › one-hot-encoding-in-python-with-pandas-and-scikit-learn
One-Hot Encoding in Python with Pandas and Scikit-Learn
July 31, 2021 - Let's take a look at a simple example of how we can convert values from a categorical column in our dataset into their numerical counterparts, via the one-hot encoding scheme. We'll be creating a really simple dataset - a list of countries and their ID's: import pandas as pd ids = [11, 22, 33, 44, 55, 66, 77] countries = ['Spain', 'France', 'Spain', 'Germany', 'France'] df = pd.DataFrame(list(zip(ids, countries)), columns=['Ids', 'Countries'])
🌐
Wayama
wayama.io › en › article › library › pandas › 02
[pandas] One-hot encoding list elements in pandas columns
December 30, 2023 - Explains how to one-hot encode list values stored in a pandas column using scikit-learn's MultiLabelBinarizer and merge the result back into the original DataFrame.
🌐
GeeksforGeeks
geeksforgeeks.org › pandas › python-pandas-series-str-encode
Python | Pandas Series.str.encode() - GeeksforGeeks
March 27, 2019 - Use 'raw_unicode_escape' for encoding. ... # importing pandas as pd import pandas as pd # Creating the Series sr = pd.Series(['New_York', 'Lisbon', 'Tokyo', 'Paris', 'Munich']) # Creating the index idx = ['City 1', 'City 2', 'City 3', 'City 4', 'City 5'] # set the index sr.index = idx # Print the series print(sr) Output : Now we will use Series.str.encode() function to encode the character strings present in the underlying data of the given series object.
🌐
GitHub
github.com › jvkersch › pandas-encoding
GitHub - jvkersch/pandas-encoding: A natural encoding for Giant Pandas
February 17, 2019 - $ PYTHONIOENCODING=pandas python Python 3.6.7 (tags/v3.6.7:6ec5cf24b7, Feb 16 2019, 22:29:37) [GCC 7.3.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> import pandas as 🐼 >>> df = 🐼.DataFrame() >>> df Empty DataFrame Columns: [] Index: [] You can also use this encoding directly, by including the line # -*- coding: pandas -*- at the top of your script:
Author: jvkersch
🌐
KDnuggets
kdnuggets.com › 2023 › 07 › pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - For this, we use pandas.get_dummies() method. It takes the following arguments: To better understand the function, let us work on one-hot encoding the dummy dataset. We use the get_dummies method and pass the original data frame as data input. In columns, we pass a list containing only the categorical_column header.