🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
One-Hot Encoding can be implemented in Python using libraries such as Pandas and Scikit-learn, which provide simple and efficient methods for converting categorical data into binary columns.
Published: May 29, 2026
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - One-hot encoding is a powerful and essential technique for transforming categorical data into a numerical format suitable for machine learning algorithms. It enhances the accuracy and efficiency of machine learning models by avoiding the pitfalls of ordinality and facilitating the use of ...
🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
In Python, we can implement one-hot encoding using the get_dummies() function in the pandas module and the OneHotEncoder class in the sklearn module.
🌐
Statology
statology.org › home › how to perform one-hot encoding in python
How to Perform One-Hot Encoding in Python
September 28, 2021 - from sklearn.preprocessing import OneHotEncoder #creating instance of one-hot-encoder encoder = OneHotEncoder(handle_unknown='ignore') #perform one-hot encoding on 'team' column encoder_df = pd.DataFrame(encoder.fit_transform(df[['team']]).toarray()) #merge one-hot encoded columns back with original DataFrame final_df = df.join(encoder_df) #view final df print(final_df) team points 0 1 2 0 A 25 1.0 0.0 0.0 1 A 12 1.0 0.0 0.0 2 B 15 0.0 1.0 0.0 3 B 14 0.0 1.0 0.0 4 B 19 0.0 1.0 0.0 5 B 23 0.0 1.0 0.0 6 C 25 0.0 0.0 1.0 7 C 29 0.0 0.0 1.0
🌐
Stack Abuse
stackabuse.com › one-hot-encoding-in-python-with-pandas-and-scikit-learn
One-Hot Encoding in Python with Pandas and Scikit-Learn
July 31, 2021 - x = [[11, "Spain"], [22, "France"], [33, "Spain"], [44, "Germany"], [55, "France"]] y = OneHotEncoder().fit_transform(x).toarray() print(y)
🌐
scikit-learn
scikit-learn.org › 0.20 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.20.4 documentation
>>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) ...
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to one hot encode sequence data in python
How to One Hot Encode Sequence Data in Python - MachineLearningMastery.com
August 14, 2019 - In this example, we will use the encoders from the scikit-learn library. Specifically, the LabelEncoder of creating an integer encoding of labels and the OneHotEncoder for creating a one hot encoding of integer encoded values.
🌐
Built In
builtin.com › articles › one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - One hot encoding (OHE) is a machine learning technique that encodes categorical data to numerical ones. If you want to perform one hot encoding, both sklearn.preprocessing.OneHotEncoder and pandas.get_dummies are popular choices.
Find elsewhere
🌐
Medium
medium.com › @creatorvision03 › one-hot-encoding-a-comprehensive-guide-with-python-code-and-examples-for-effective-categorical-2fbbc111c320
“One-Hot Encoding: A Comprehensive Guide with Python Code and Examples for Effective Categorical Data Representation” | by Shivang Gupta | Medium
July 2, 2023 - Let’s now dive into a practical implementation of one-hot encoding using Python. We’ll be using the popular scikit-learn library, which provides various tools for machine learning tasks. ... from sklearn.preprocessing import OneHotEncoder import pandas as pd # Create a sample dataframe with categorical variables data = {'Color': ['red', 'blue', 'green', 'blue']} df = pd.DataFrame(data) # Initialize the OneHotEncoder encoder = OneHotEncoder() # Fit and transform the dataframe encoded_data = encoder.fit_transform(df[['Color']]) # Convert the encoded data to a pandas dataframe encoded_df = pd.DataFrame(encoded_data.toarray(), columns=encoder.get_feature_names_out(['Color'])) # Print the encoded dataframe print(encoded_df)
Top answer
1 of 16
317

Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

2 of 16
149

Much easier to use Pandas for basic one-hot encoding. If you're looking for more options you can use scikit-learn.

For basic one-hot encoding with Pandas you pass your data frame into the get_dummies function.

For example, if I have a dataframe called imdb_movies:

...and I want to one-hot encode the Rated column, I do this:

pd.get_dummies(imdb_movies.Rated)

This returns a new dataframe with a column for every "level" of rating that exists, along with either a 1 or 0 specifying the presence of that rating for a given observation.

Usually, we want this to be part of the original dataframe. In this case, we attach our new dummy coded frame onto the original frame using "column-binding.

We can column-bind by using Pandas concat function:

rated_dummies = pd.get_dummies(imdb_movies.Rated)
pd.concat([imdb_movies, rated_dummies], axis=1)

We can now run an analysis on our full dataframe.

SIMPLE UTILITY FUNCTION

I would recommend making yourself a utility function to do this quickly:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    return(res)

Usage:

encode_and_bind(imdb_movies, 'Rated')

Result:

Also, as per @pmalbu comment, if you would like the function to remove the original feature_to_encode then use this version:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    res = res.drop([feature_to_encode], axis=1)
    return(res) 

You can encode multiple features at the same time as follows:

features_to_encode = ['feature_1', 'feature_2', 'feature_3',
                      'feature_4']
for feature in features_to_encode:
    res = encode_and_bind(train_set, feature)
🌐
datagy
datagy.io › home › machine learning › one-hot encoding in machine learning with python
One-Hot Encoding in Machine Learning with Python • datagy
September 6, 2023 - # Understanding the OneHotEncoder Class in Sklearn from sklearn.preprocessing import OneHotEncoder OneHotEncoder( categories='auto', # Categories per feature drop=None, # Whether to drop one of the features sparse=True, # Will return sparse matrix if set True dtype=<class 'numpy.float64'>, # Desired data type of the output handle_unknown='error' # Whether to raise an error )
🌐
Train in Data
blog.trainindata.com › one-hot-encoding-categorical-variables
One-hot encoding categorical variables | Train in Data Blog
January 25, 2023 - Let’s first import the necessary Python libraries and get the dataset ready: import pandas as pd import numpy as np from sklearn.model_selection import train_test_split from feature_engine.encoding import OneHotEncoder
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › api › pyspark.ml.feature.OneHotEncoder.html
OneHotEncoder — PySpark 4.2.0 documentation
>>> from pyspark.ml.linalg import Vectors >>> df = spark.createDataFrame([(0.0,), (1.0,), (2.0,)], ["input"]) >>> ohe = OneHotEncoder() >>> ohe.setInputCols(["input"]) OneHotEncoder... >>> ohe.setOutputCols(["output"]) OneHotEncoder... >>> model = ohe.fit(df) >>> model.setOutputCols(["output"]) OneHotEncoderModel...
🌐
AI Mind
pub.aimind.so › one-hot-encoding-for-machine-learning-with-python-and-scikit-learn-c6d8e1173760
One-Hot Encoding for Machine Learning (with Python and Scikit-Learn) | by Francesco Franco | AI Mind
November 22, 2024 - We can then use Scikit-learn for converting the values into a one-hot encoded array, because it offers the sklearn.preprocessing.OneHotEncoder module. We first import the numpy module for converting a Python list into a NumPy array, and the preprocessing module from Scikit-learn.
🌐
Trainindata
feature-engine.trainindata.com › en › latest › user_guide › encoding › OneHotEncoder.html
OneHotEncoder — 1.9.4 - Feature-engine
OneHotEncoder() can specifically encode binary variables into k-1 variables (that is, 1 dummy) while encoding categorical features of higher cardinality into k dummies. This behaviour is specified by setting the parameter drop_last_binary=True. This will ensure that for every binary variable ...
🌐
Kaggle
kaggle.com › code › marcinrutecki › one-hot-encoding-everything-you-need-to-know
One Hot Encoding - everything you need to know
February 23, 2023 - Explore and run AI code with Kaggle Notebooks | Using data from multiple data sources
🌐
scikit-learn
scikit-learn.org › 0.16 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.16.1 documentation
>>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]]) OneHotEncoder(categorical_features='all', dtype=<... 'float'>, handle_unknown='error', n_values='auto', sparse=True) >>> enc.n_values_ array([2, 3, 4]) >>> enc.feature_indices_ array([0, 2, 5, 9]) >>> enc.transform([[0, 1, 1]]).toarray() array([[ 1., 0., 0., 1., 0., 0., 1., 0., 0.]])