🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are encoded using a one-hot (aka ‘one-of-K’ or ‘dummy’) encoding scheme.
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
One-Hot Encoding can be implemented in Python using libraries such as Pandas and Scikit-learn, which provide simple and efficient methods for converting categorical data into binary columns.
Published: May 29, 2026
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - Pandas provides a very convenient function, get_dummies(), to create one-hot encoded columns directly from a DataFrame. Here's how you can use it (we’ll explain all the code step-by-step below):
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › one hot encoding: understanding the “hot” in data
One Hot Encoding: Understanding the "Hot" in Data - MachineLearningMastery.com
February 27, 2025 - After understanding the basic premise and application of One Hot Encoding in linear models, the next step in our analysis involves identifying which categorical feature contributes most significantly to predicting our target variable. In the code snippet below, we iterate through each categorical feature in our dataset, apply One Hot Encoding, and evaluate its predictive power using a linear regression model in conjunction with cross-validation.
🌐
Google
developers.google.com › machine learning › categorical data: vocabulary and one-hot encoding
Categorical data: Vocabulary and one-hot encoding | Machine Learning | Google for Developers
Sparse representation efficiently stores one-hot encoded data by only recording the position of the '1' value to reduce memory usage.
bit-vector representation where exactly one bit must be set
In digital circuits and machine learning, a one-hot is a group of bits among which the legal combinations of values are only those with a single high (1) bit and all the … Wikipedia
🌐
Wikipedia
en.wikipedia.org › wiki › One-hot
One-hot - Wikipedia
February 14, 2026 - Requires more flip-flops than other encodings, making it impractical for PAL devices ... In natural language processing, a one-hot vector is a 1 × N matrix (vector) used to distinguish each word in a vocabulary from every other word in the vocabulary. The vector consists of 0s in all cells ...
🌐
Medium
medium.com › @creatorvision03 › one-hot-encoding-a-comprehensive-guide-with-python-code-and-examples-for-effective-categorical-2fbbc111c320
“One-Hot Encoding: A Comprehensive Guide with Python Code and Examples for Effective Categorical Data Representation” | by Shivang Gupta | Medium
July 2, 2023 - By converting categorical data into binary vectors, it allows algorithms to effectively process and interpret the information. In this article, we discussed the concept of one-hot encoding, and its benefits, and provided a code example for implementation using Python and scikit-learn.
Top answer
1 of 16
317

Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

2 of 16
149

Much easier to use Pandas for basic one-hot encoding. If you're looking for more options you can use scikit-learn.

For basic one-hot encoding with Pandas you pass your data frame into the get_dummies function.

For example, if I have a dataframe called imdb_movies:

...and I want to one-hot encode the Rated column, I do this:

pd.get_dummies(imdb_movies.Rated)

This returns a new dataframe with a column for every "level" of rating that exists, along with either a 1 or 0 specifying the presence of that rating for a given observation.

Usually, we want this to be part of the original dataframe. In this case, we attach our new dummy coded frame onto the original frame using "column-binding.

We can column-bind by using Pandas concat function:

rated_dummies = pd.get_dummies(imdb_movies.Rated)
pd.concat([imdb_movies, rated_dummies], axis=1)

We can now run an analysis on our full dataframe.

SIMPLE UTILITY FUNCTION

I would recommend making yourself a utility function to do this quickly:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    return(res)

Usage:

encode_and_bind(imdb_movies, 'Rated')

Result:

Also, as per @pmalbu comment, if you would like the function to remove the original feature_to_encode then use this version:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    res = res.drop([feature_to_encode], axis=1)
    return(res) 

You can encode multiple features at the same time as follows:

features_to_encode = ['feature_1', 'feature_2', 'feature_3',
                      'feature_4']
for feature in features_to_encode:
    res = encode_and_bind(train_set, feature)
Find elsewhere
🌐
Towards Data Science
towardsdatascience.com › home › data science › robust one-hot encoding
Robust One-Hot Encoding | Towards Data Science
April 26, 2024 - We have shown how naive use of one-hot encoding techniques can lead to mistakes and problems with inference data, and we have also seen how to mitigate and resolve those issues using both Python and R. If left unresolved, poor management of one-hot encoding can potentially lead to crashes and problems with your inference, so it is strongly recommended to use more robust techniques—like either sklearn's OneHotEncoder or the R function we developed. Thanks for reading! _All the code presented and used in the article can be found in the following Github repo: https://github.com/hcekne/robust_one_hot_encoding_
🌐
Medium
medium.com › @michaeldelsole › what-is-one-hot-encoding-and-how-to-do-it-f0ae272f1179
What is One Hot Encoding and How to Do It | by Michael DelSole | Medium
April 24, 2018 - This means representing each piece of data in a way that the computer can understand, hence the name encode, which literally means “convert to [computer] code”. There’s many different ways of encoding such as Label Encoding, or as you might of guessed, One Hot Encoding.
🌐
Medium
medium.com › @heyamit10 › one-hot-encoding-explained-0b0130ccd78e
One Hot Encoding Explained
November 26, 2024 - Simple to Implement: This might surprise you: one-hot encoding is incredibly easy to implement using Python libraries like pandas or sklearn. With just one line of code, you can transform categorical data into a format that most machine learning models can work with.
🌐
Kaggle
kaggle.com › code › marcinrutecki › one-hot-encoding-everything-you-need-to-know
One Hot Encoding - everything you need to know
February 23, 2023 - 1.1 One Hot Encoder vs get_dummies1.2 The dummy variable trap: drop or not to drop?1.3 Possible drawbacks of dropping a column during one hot encoding1.4 Decision tree-based models vs one hot encoding1.5 One Hot Encoding vs very high number of categorical features1.6 Pipelines and One Hot Encoding1.7 One Hot Encoding - before or after train-test split?1.8 Best practices1.9 Simple examples2.1 Import Libraries2.2 Import Data2.3 Data Set Characteristics2.4 Dataset Attributes3.1 Dealin with missing values in TotalCharges3.2 Dealing with duplicated values3.3 Creating numerical and categorical lists4.1 Train test split - stratified splitting4.2 Feature scaling4.3 One hot Encoding4.4 Feature importance
🌐
Educative
educative.io › blog › one-hot-encoding
Data Science in 5 Minutes: What is One Hot Encoding?
Here, we are passing the value City for the prefix attribute of the method get_dummies(). If we run the code now, we will print our encoded values: ... We can implement a similar functionality with Sklearn, which provides an object/function for one-hot encoding in the preprocessing module.
🌐
MathWorks
mathworks.com › deep learning toolbox › train deep neural networks › custom training using automatic differentiation
onehotencode - Encode data labels into one-hot vectors - MATLAB
B = onehotencode(A,featureDim) encodes data labels in categorical array A into a one-hot encoded array B. The function replaces each element of A with a numeric vector of length equal to the number of unique classes in A along the dimension specified by featureDim.
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.5.2 documentation
This encoding is needed for feeding categorical data to many scikit-learn estimators, notably linear models and SVMs with the standard kernels. Note: a one-hot encoding of y labels should use a LabelBinarizer instead.
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to one hot encode sequence data in python
How to One Hot Encode Sequence Data in Python - MachineLearningMastery.com
August 14, 2019 - How to calculate an integer encoding and one hot encoding by hand in Python. How to use the scikit-learn and Keras libraries to automatically encode your sequence data in Python. Kick-start your project with my new book Long Short-Term Memory Networks With Python, including step-by-step tutorials and the Python source code files for all examples.
🌐
Kaggle
kaggle.com › code › dansbecker › using-categorical-data-with-one-hot-encoding
Using Categorical Data with One Hot Encoding | Kaggle
January 22, 2018 - Explore and run AI code with Kaggle Notebooks | Using data from House Prices - Advanced Regression Techniques
🌐
ScienceDirect
sciencedirect.com › topics › computer-science › one-hot-encoding
One-Hot Encoding - an overview | ScienceDirect Topics
Although this method is simple to implement, the One-Hot vector usually is sparse, semantically independent, and with a very high dimension. ... Word-count-based encoding. This approach initializes a zero-coded vector with the length of the vocabulary size and then replaces each element with a specific value.
🌐
Towards Data Science
towardsdatascience.com › home › latest › how and why performing one-hot encoding in your data science project
How and Why Performing One-Hot Encoding in Your Data Science Project | Towards Data Science
January 19, 2025 - Image by Author. As simple as that. With just one line of code, you get the columns named with the unique values present in the "Name" column and you get the "Name" column dropped.
🌐
Built In
builtin.com › articles › one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - When writing this transformer we assumed that the relevant columns already have categorical dtypes. But it’s very simple to add a few lines of code to GetDummiesTransformer to allow the specification of the columns in the __init__ function. A tutorial on one hot encoding.