Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

Answer from Sayali Sonawane on Stack Overflow
Top answer
1 of 16
317

Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

2 of 16
149

Much easier to use Pandas for basic one-hot encoding. If you're looking for more options you can use scikit-learn.

For basic one-hot encoding with Pandas you pass your data frame into the get_dummies function.

For example, if I have a dataframe called imdb_movies:

...and I want to one-hot encode the Rated column, I do this:

pd.get_dummies(imdb_movies.Rated)

This returns a new dataframe with a column for every "level" of rating that exists, along with either a 1 or 0 specifying the presence of that rating for a given observation.

Usually, we want this to be part of the original dataframe. In this case, we attach our new dummy coded frame onto the original frame using "column-binding.

We can column-bind by using Pandas concat function:

rated_dummies = pd.get_dummies(imdb_movies.Rated)
pd.concat([imdb_movies, rated_dummies], axis=1)

We can now run an analysis on our full dataframe.

SIMPLE UTILITY FUNCTION

I would recommend making yourself a utility function to do this quickly:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    return(res)

Usage:

encode_and_bind(imdb_movies, 'Rated')

Result:

Also, as per @pmalbu comment, if you would like the function to remove the original feature_to_encode then use this version:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    res = res.drop([feature_to_encode], axis=1)
    return(res) 

You can encode multiple features at the same time as follows:

features_to_encode = ['feature_1', 'feature_2', 'feature_3',
                      'feature_4']
for feature in features_to_encode:
    res = encode_and_bind(train_set, feature)
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
Pandas provides the get_dummies() function to perform one-hot encoding on categorical columns.
Published: May 29, 2026
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are encoded using a one-hot (aka ‘one-of-K’ or ‘dummy’) encoding scheme.
🌐
Train in Data
blog.trainindata.com › one-hot-encoding-categorical-variables
One-hot encoding categorical variables | Train in Data Blog
January 25, 2023 - Pandas, Feature-engine and Category Encoders can automatically identify and encode categorical variables, that is, those of type object or categorical. Scikit-learn’s OneHotEncoder(), on the other hand, will encode all variables in the dataset.
🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
In Python, we can implement one-hot encoding using the get_dummies() function in the pandas module and the OneHotEncoder class in the sklearn module.
🌐
Medium
medium.com › @amit25173 › applying-one-hot-encoding-in-pandas-5a639cb3bc69
Applying One-Hot Encoding in Pandas | by Amit Yadav | Medium
March 6, 2025 - This might surprise you: one-hot encoding can actually cause issues in machine learning models if you’re not careful. This issue is called the dummy variable trap, where one of the encoded columns becomes redundant, causing multicollinearity. To avoid it, just set drop_first=True in pd.get_dummies(): import pandas as pd data = {'Color': ['Red', 'Blue', 'Green']} df = pd.DataFrame(data) # Avoiding the dummy variable trap encoded_df = pd.get_dummies(df, drop_first=True) print(encoded_df)
Find elsewhere
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - Then, we'll explore Scikit-learn's OneHotEncoder, which offers more flexibility and control, particularly useful for more complex encoding needs. Pandas provides a very convenient function, get_dummies(), to create one-hot encoded columns directly from a DataFrame.
🌐
Towards Data Science
towardsdatascience.com › home › data science › pandas for one-hot encoding data preventing high cardinality
Pandas for One-Hot Encoding Data Preventing High Cardinality | Towards Data Science
November 15, 2022 - One Hot Encoding is useful to transform categorical data into numbers. Using OHE in a dataset with too many variable will create a wide dataset. Too wide data can suffer with the Curse of Dimensionality, putting the performance of the model ...
🌐
KDnuggets
kdnuggets.com › 2023 › 07 › pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - Therefore, the single categorical column is converted into 4 new columns where only one of the 4 columns will have a 1 value, and all of the other 3 are encoded 0. This is why it is called One-Hot Encoding.
🌐
Medium
blog.cambridgespark.com › robust-one-hot-encoding-in-python-3e29bfcec77e
Tutorial: (Robust) One Hot Encoding in Python | by Kevin Lemagnen | Cambridge Spark
October 11, 2018 - Our column transport has no value bus but the new value bike. Let's see how we can build one hot encoded features for those datasets! We’ll show two different methods, one using the get_dummies method from pandas, and the other with the OneHotEncoder class from sklearn.
🌐
Towards Data Science
towardsdatascience.com › home › latest › one hot encoding scikit vs pandas
One Hot Encoding scikit vs pandas | Towards Data Science
March 5, 2025 - Both sklearn.preprocessing.OneHotEncoder and pandas.get_dummies are popular choices (well, practically the only choices unless you want want to implement it yourself) to perform One Hot Encoding.
Top answer
1 of 4
21

LabelEncoder is not made to transform the data but the target (also known as labels) as explained here. If you want to encode the data you should use OrdinalEncoder.

If you really need to do it this way:

categorical_cols = ['a', 'b', 'c', 'd'] 

from sklearn.preprocessing import LabelEncoder
# instantiate labelencoder object
le = LabelEncoder()

# apply le on categorical feature columns
data[categorical_cols] = data[categorical_cols].apply(lambda col: le.fit_transform(col))    
from sklearn.preprocessing import OneHotEncoder
ohe = OneHotEncoder()

#One-hot-encode the categorical columns.
#Unfortunately outputs an array instead of dataframe.
array_hot_encoded = ohe.fit_transform(data[categorical_cols])

#Convert it to df
data_hot_encoded = pd.DataFrame(array_hot_encoded, index=data.index)

#Extract only the columns that didnt need to be encoded
data_other_cols = data.drop(columns=categorical_cols)

#Concatenate the two dataframes : 
data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1)

Otherwise:

I suggest you to use pandas.get_dummies if you want to achieve one-hot-encoding from raw data (without having to use OrdinalEncoder before) :

#categorical data
categorical_cols = ['a', 'b', 'c', 'd'] 

#import pandas as pd
df = pd.get_dummies(data, columns = categorical_cols)

You can also use drop_first argument to remove one of the one-hot-encoded columns, as some models require.

2 of 4
8

You can do dummy encoding using Pandas in order to get one-hot encoding as shown below:

import pandas as pd

# Multiple categorical columns
categorical_cols = ['a', 'b', 'c', 'd']

pd.get_dummies(data, columns=categorical_cols)

If you want to do one-hot encoding using sklearn library, you can get it done as shown below:

from sklearn.preprocessing import OneHotEncoder
onehotencoder = OneHotEncoder()

transformed_data = onehotencoder.fit_transform(data[categorical_cols])

# the above transformed_data is an array so convert it to dataframe
encoded_data = pd.DataFrame(transformed_data, index=data.index)

# now concatenate the original data and the encoded data using pandas
concatenated_data = pd.concat([data, encoded_data], axis=1)

If a single column has more than 500 categories, the aforementioned way of one-hot encoding is not a good approach. In this case, we can do one-hot encoding for the top 10 or 20 categories that are occurring most for a particular column. A sample code is shown below:

categorical_cols = ['a', 'b', 'c', 'd']

# Let's say we have a column 'b' which has more than 500 categories.
# Find the top 10 most frequent categories for column 'b'
data.b.value_counts().sort_values(ascending = False).head(20)

# make a list of the most frequent categories of the column
top_10_occurring_cat = [cat for cat in data.b.value_counts().sort_values(ascending = False).head(10).index]

# now make the 10 binary variables
for cat in top_10_occurring_cat:
    data[cat] = np.where(data['b'] == cat, 1, 0) # whenever data['b'] == cat replace it with 1 else 0

# This is done for one categorical column, similarly you can repeat for all categorical columns
🌐
Kaggle
kaggle.com › code › marcinrutecki › one-hot-encoding-everything-you-need-to-know
One Hot Encoding - everything you need to know
February 23, 2023 - 1.1 One Hot Encoder vs get_dummies1.2 The dummy variable trap: drop or not to drop?1.3 Possible drawbacks of dropping a column during one hot encoding1.4 Decision tree-based models vs one hot encoding1.5 One Hot Encoding vs very high number of categorical features1.6 Pipelines and One Hot Encoding1.7 One Hot Encoding - before or after train-test split?1.8 Best practices1.9 Simple examples2.1 Import Libraries2.2 Import Data2.3 Data Set Characteristics2.4 Dataset Attributes3.1 Dealin with missing values in TotalCharges3.2 Dealing with duplicated values3.3 Creating numerical and categorical lists4.1 Train test split - stratified splitting4.2 Feature scaling4.3 One hot Encoding4.4 Feature importance
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › onehot.html
One Hot — Category Encoders 2.11.1 documentation
Onehot (or dummy) coding for categorical features, produces a binary feature per category. ... boolean for whether to return a pandas DataFrame from transform (otherwise it will be a numpy array). ... if True, category values will be included in the encoded column names.
🌐
Saturn Cloud
saturncloud.io › blog › pandas-vs-scikitlearn-onehot-encoding-dataframes
Pandas vs. Scikit-learn: One-Hot Encoding Dataframes | Saturn Cloud Blog
May 1, 2026 - It provides a range of functions for cleaning, transforming, and analyzing data, including one-hot encoding. Pandas provides the get_dummies() function to one-hot encode categorical variables.