Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

Answer from Sayali Sonawane on Stack Overflow
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - Pandas provides a very convenient function, get_dummies(), to create one-hot encoded columns directly from a DataFrame. Here's how you can use it (we’ll explain all the code step-by-step below):
Top answer
1 of 16
317

Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

2 of 16
149

Much easier to use Pandas for basic one-hot encoding. If you're looking for more options you can use scikit-learn.

For basic one-hot encoding with Pandas you pass your data frame into the get_dummies function.

For example, if I have a dataframe called imdb_movies:

...and I want to one-hot encode the Rated column, I do this:

pd.get_dummies(imdb_movies.Rated)

This returns a new dataframe with a column for every "level" of rating that exists, along with either a 1 or 0 specifying the presence of that rating for a given observation.

Usually, we want this to be part of the original dataframe. In this case, we attach our new dummy coded frame onto the original frame using "column-binding.

We can column-bind by using Pandas concat function:

rated_dummies = pd.get_dummies(imdb_movies.Rated)
pd.concat([imdb_movies, rated_dummies], axis=1)

We can now run an analysis on our full dataframe.

SIMPLE UTILITY FUNCTION

I would recommend making yourself a utility function to do this quickly:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    return(res)

Usage:

encode_and_bind(imdb_movies, 'Rated')

Result:

Also, as per @pmalbu comment, if you would like the function to remove the original feature_to_encode then use this version:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    res = res.drop([feature_to_encode], axis=1)
    return(res) 

You can encode multiple features at the same time as follows:

features_to_encode = ['feature_1', 'feature_2', 'feature_3',
                      'feature_4']
for feature in features_to_encode:
    res = encode_and_bind(train_set, feature)
Discussions

Pandas factorize and one hot encoding
Say you use factorize and then apply k means clustering. The algorithm will assume the 1 is closer to 2 than to 10. Does that make sense? If factorize just converts text to numbers, is there any reason to believe that the order has some meaning? With one hot encoding theres no such assumption. Each possible value is a boolean, they're all equally close to each other. So, it depends heavily on what you're doing with the data, but one hot encoding is usually safer. More on reddit.com
🌐 r/learnmachinelearning
8
10
June 1, 2024
What alternatives are there to one hot encoding?
There's something called Target Encoding which I've found to be quite effective. It is basically where you use the target variable itself to inform the encoding of the category. For instance, let's say you're doing the classic home price problem where you're trying to predict a home's value. You've got home style as an input (Crafstman, Modern, etc.). You would order the home styles by their average (or perhaps median) home price for that style and use that as the numeric encoding in your model. It's tricky to get quite right, because there can be high variability (especially among rarer home types, in this example). So you can have a cut-off that says "for instances where there's less than X examples in the training set, use the average/min/max/whatever." You can also remove some variability by doing some k-folding. Just realized as I'm typing this, that articles have been written. so why am I typing this out? https://medium.com/@pouryaayria/k-fold-target-encoding-dfe9a594874b -- not sure if this is the best article, but on a quick skim seems fine. One thing to consider is to try multiple methods at once, like create a frequency-encoded and a target-encoded version of the same feature. They may convey different information. More on reddit.com
🌐 r/datascience
12
6
October 6, 2022
How can I decode one hot vector? (and when can we use it?)
np.argmax(one_hot, axis=1) More on reddit.com
🌐 r/MachineLearning
2
0
September 12, 2016
What is the difference between applying One Hot Encoding to a categorical column and changing the data type to categorical (in pandas)?
To the best of my understanding, the regressor can't read strings, so the string "32" instead of the number 32 cannot be parsed and would result in an error.If you get the 'dummies' you convert each label into a column filled with ones (1) and zeros (0). These numeric values can then be fed into a regressor. As an aside, the regressor doesn't "recognise" the labels (as strings) but it can deal with the distance between the transformed columns for two (or more) different labels and use that as a differentiation in the regression. More on reddit.com
🌐 r/learnmachinelearning
6
2
January 28, 2022
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
One-Hot Encoding can be implemented in Python using libraries such as Pandas and Scikit-learn, which provide simple and efficient methods for converting categorical data into binary columns.
Published: May 29, 2026
🌐
Medium
medium.com › @creatorvision03 › one-hot-encoding-a-comprehensive-guide-with-python-code-and-examples-for-effective-categorical-2fbbc111c320
“One-Hot Encoding: A Comprehensive Guide with Python Code and Examples for Effective Categorical Data Representation” | by Shivang Gupta | Medium
July 2, 2023 - By converting categorical data into binary vectors, it allows algorithms to effectively process and interpret the information. In this article, we discussed the concept of one-hot encoding, and its benefits, and provided a code example for implementation using Python and scikit-learn.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
This encoding is needed for feeding categorical data to many scikit-learn estimators, notably linear models and SVMs with the standard kernels. Note: a one-hot encoding of y labels should use a LabelBinarizer instead.
🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
You can set the sparse_output parameter to False for dense one-hot encoding and True otherwise. ... 'The Codecademy Team, composed of experienced educators and tech experts, is dedicated to making tech skills accessible to all.
🌐
Statology
statology.org › home › how to perform one-hot encoding in python
How to Perform One-Hot Encoding in Python
September 28, 2021 - Next, let’s import the ...eprocessing import OneHotEncoder #creating instance of one-hot-encoder encoder = OneHotEncoder(handle_unknown='ignore') #perform one-hot encoding on 'team' column encoder_df = pd.DataFrame(e...
Find elsewhere
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to one hot encode sequence data in python
How to One Hot Encode Sequence Data in Python - MachineLearningMastery.com
August 14, 2019 - How to calculate an integer encoding and one hot encoding by hand in Python. How to use the scikit-learn and Keras libraries to automatically encode your sequence data in Python. Kick-start your project with my new book Long Short-Term Memory Networks With Python, including step-by-step tutorials and the Python source code files for all examples.
🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding data in machine learning
One Hot Encoding Data in Machine Learning - Analytics vidhya
March 28, 2025 - A. For one-hot encoding in a Python DataFrame, use the get_dummies function from the pandas library.
🌐
Stack Abuse
stackabuse.com › one-hot-encoding-in-python-with-pandas-and-scikit-learn
One-Hot Encoding in Python with Pandas and Scikit-Learn
July 31, 2021 - If we represented these categories in one-hot encoding, we would actually replace the rows with columns. We do this by creating one boolean column for each of our given categories, where only one of these columns could take on the value 1 for each sample: We can see from the tables above that more digits are needed in one-hot representation compared to Binary or Gray code.
🌐
Educative
educative.io › answers › one-hot-encoding-in-python
One-hot encoding in Python
Take a look at the example below. It uses the scikit-learn library to perform one-hot encoding:
🌐
Educative
educative.io › blog › one-hot-encoding
Data Science in 5 Minutes: What is One Hot Encoding?
We don’t have to one hot encode manually. Many data science tools offer easy ways to encode your data. The Python library Pandas provides a function called get_dummies to enable one-hot encoding.
🌐
YouTube
youtube.com › watch
Step-by-Step Guide to One Hot Encoding in Python | Machine Learning Essentials - YouTube
In this tutorial, you'll learn how to implement One Hot Encoding in Python, a key technique for data preprocessing in Machine Learning. This step-by-step gui...
Published: September 17, 2024
Views: 150
🌐
AskPython
askpython.com › python › examples › one-hot-encoding
One hot encoding in Python - A Practical Approach - AskPython
February 16, 2023 - Further, we would pass the same integer data to the OneHotEncoder() to encode the integer values into the binary vectors of the categories. The fit_transform() function applies the particular function to be performed on the data or set of values. In this example, we have pulled a dataset into the Python environment.
🌐
Spot Intelligence
spotintelligence.com › home › how to use one hot encoding in python with 3 tutorials
How To Use One Hot Encoding In Python With 3 Tutorials
November 1, 2023 - In Python, one-hot encoding can be easily performed using the OneHotEncoder class from the sklearn.preprocessing module.
🌐
Medium
blog.cambridgespark.com › robust-one-hot-encoding-in-python-3e29bfcec77e
Tutorial: (Robust) One Hot Encoding in Python | by Kevin Lemagnen | Cambridge Spark
October 11, 2018 - We’ll need to specify handle_unknown as ignore so the OneHotEncoder can work later on with our unseen data. The OneHotEncoder will build a numpy array for our data, replacing our original features by one hot encoding versions.
🌐
IQCode
iqcode.com › code › python › one-hot-encoding
one hot encoding Code Example
y = pd.get_dummies(df.Countries, prefix='Country') print(y.head()) # from here you can merge it onto your main DF
🌐
GitHub
rasbt.github.io › mlxtend › user_guide › preprocessing › one-hot_encoding
One hot encoding - mlxtend
from mlxtend.preprocessing import one_hot y = [0, 1, 2, 1, 2] one_hot(y, dtype='int')
🌐
Medium
medium.com › @heyamit10 › one-hot-encoding-in-python-67e7364995cb
One Hot Encoding in Python. Imagine trying to teach a machine to… | by Hey Amit | Medium
November 22, 2024 - This is especially useful when ... or zip codes with thousands of unique values. For example, if you’re predicting whether a customer will buy a product, you can replace each category (e.g., product type) with the average probability of purchase for that type. This helps reduce dimensionality and avoids the sparse matrix problem that One Hot Encoding can cause. And there you have it — everything you need to know about One Hot Encoding in Python...
🌐
OpenAI
blog.gopenai.com › one-hot-encoding-a-comprehensive-guide-with-python-implementation-e57bc7067ce6
One-Hot Encoding: A Comprehensive Guide with Python Implementation | by Shubham Sangole | GoPenAI
May 21, 2024 - By representing categories as binary vectors, we can ensure that algorithms can effectively process and interpret categorical variables. In this guide, we’ve covered the mathematical foundation of one-hot encoding and provided a practical Python implementation using pandas and scikit-learn.