๐ŸŒ
scikit-learn
scikit-learn.org โ€บ stable โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder โ€” scikit-learn 1.9.1 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
๐ŸŒ
Codecademy
codecademy.com โ€บ article โ€บ what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
In Python, we can implement one-hot encoding using the get_dummies() function in the pandas module and the OneHotEncoder class in the sklearn module.
๐ŸŒ
DataCamp
datacamp.com โ€บ tutorial โ€บ one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - For more flexibility and control over the encoding process, Scikit-learn offers the OneHotEncoder class. This class provides advanced options, such as handling unknown categories and fitting the encoder to the training data. from sklearn.preprocessing import OneHotEncoder import numpy as np ...
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ dev โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
One-Hot Encoding in Scikit-learn
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
๐ŸŒ
Built In
builtin.com โ€บ articles โ€บ one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - One hot encoding (OHE) is a machine learning technique that encodes categorical data to numerical ones. If you want to perform one hot encoding, both sklearn.preprocessing.OneHotEncoder and pandas.get_dummies are popular choices.
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.16 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder โ€” scikit-learn 0.16.1 documentation
Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.19 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder โ€” scikit-learn 0.19.2 documentation
Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]]) OneHotEncoder(categorical_features='all', dtype=<... 'numpy.float64'>, handle_unknown='error', n_values='auto', sparse=True) >>> enc.n_values_ array([2, 3, 4]) >>> enc.feature_indices_ array([0, 2, 5, 9]) >>> enc.transform([[0, 1, 1]]).toarray() array([[ 1., 0., 0., 1., 0., 0., 1., 0., 0.]])
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.20 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder โ€” scikit-learn 0.20.4 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', ...
Find elsewhere
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 1.5 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder โ€” scikit-learn 1.5.2 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 1.0 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder โ€” scikit-learn 1.0.2 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ machine learning โ€บ ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
One-Hot Encoding can be implemented in Python using libraries such as Pandas and Scikit-learn, which provide simple and efficient methods for converting categorical data into binary columns.
Published: May 29, 2026
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.18 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder โ€” scikit-learn 0.18.2 documentation
Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
Top answer
1 of 10
49

OneHotEncoder Encodes categorical integer features as a one-hot numeric array. Its Transform method returns a sparse matrix if sparse=True, otherwise it returns a 2-d array.

You can't cast a 2-d array (or sparse matrix) into a Pandas Series. You must create a Pandas Serie (a column in a Pandas dataFrame) for each category.

I would recommend pandas.get_dummies instead:

data = pd.get_dummies(data,prefix=['Profession'], columns = ['Profession'], drop_first=True)

EDIT:

Using Sklearn OneHotEncoder:

transformed = jobs_encoder.transform(data['Profession'].to_numpy().reshape(-1, 1))
#Create a Pandas DataFrame of the hot encoded column
ohe_df = pd.DataFrame(transformed, columns=jobs_encoder.get_feature_names())
#concat with original data
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)

Other Options: If you are doing hyperparameter tuning with GridSearch it's recommanded to use ColumnTransformer and FeatureUnion with Pipeline or directly make_column_transformer

2 of 10
24

So turned out that Scikit-Learns LabelBinarizer gave me better luck in converting the data to one-hot encoded format, with help from Amnie's solution, my final code is as follows

import pandas as pd
from sklearn.preprocessing import LabelBinarizer

jobs_encoder = LabelBinarizer()
jobs_encoder.fit(data['Profession'])
transformed = jobs_encoder.transform(data['Profession'])
ohe_df = pd.DataFrame(transformed)
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.21 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder โ€” scikit-learn 0.21.3 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', ...
๐ŸŒ
Towards Data Science
towardsdatascience.com โ€บ home โ€บ latest โ€บ one hot encoding scikit vs pandas
One Hot Encoding scikit vs pandas | Towards Data Science
March 5, 2025 - Both sklearn.preprocessing.OneHotEncoder and pandas.get_dummies are popular choices (well, practically the only choices unless you want want to implement it yourself) to perform One Hot Encoding.
Top answer
1 of 16
317

Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

2 of 16
149

Much easier to use Pandas for basic one-hot encoding. If you're looking for more options you can use scikit-learn.

For basic one-hot encoding with Pandas you pass your data frame into the get_dummies function.

For example, if I have a dataframe called imdb_movies:

...and I want to one-hot encode the Rated column, I do this:

pd.get_dummies(imdb_movies.Rated)

This returns a new dataframe with a column for every "level" of rating that exists, along with either a 1 or 0 specifying the presence of that rating for a given observation.

Usually, we want this to be part of the original dataframe. In this case, we attach our new dummy coded frame onto the original frame using "column-binding.

We can column-bind by using Pandas concat function:

rated_dummies = pd.get_dummies(imdb_movies.Rated)
pd.concat([imdb_movies, rated_dummies], axis=1)

We can now run an analysis on our full dataframe.

SIMPLE UTILITY FUNCTION

I would recommend making yourself a utility function to do this quickly:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    return(res)

Usage:

encode_and_bind(imdb_movies, 'Rated')

Result:

Also, as per @pmalbu comment, if you would like the function to remove the original feature_to_encode then use this version:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    res = res.drop([feature_to_encode], axis=1)
    return(res) 

You can encode multiple features at the same time as follows:

features_to_encode = ['feature_1', 'feature_2', 'feature_3',
                      'feature_4']
for feature in features_to_encode:
    res = encode_and_bind(train_set, feature)
๐ŸŒ
pythontutorials
pythontutorials.net โ€บ blog โ€บ one-hot-encode-sklearn
One-Hot Encoding with scikit-learn: A Comprehensive Guide โ€” pythontutorials.net
For example, consider a categorical variable "color" with three unique categories: "red", "blue", and "green". The one-hot encoded representation of these categories would be: ... The OneHotEncoder class in sklearn.preprocessing can be used to perform one-hot encoding.
๐ŸŒ
Medium
datasensei.medium.com โ€บ how-to-transform-nominal-data-for-ml-with-onehotencoder-from-scikit-learn-f6febfefb3c6
How to Transform Nominal Data for ML with OneHotEncoder from Scikit-Learn | by Data Seito | Medium
January 18, 2022 - DataFrame containing the one-hot encoded features. Finally, we can join the numerical features with our encoded categorical features. ... The full DataFrame with the encoded categorical features. And thatโ€™s it! You are now one step closer to mastering the art and science of machine learning. Putting together everything above yields a short script that you can use in your own machine learning pipelines. import pandas as pdfrom sklearn.preprocessing import OneHotEncoder # create an example dataframe to work withdf = pd.DataFrame([ [27, 'Sedan', 'Toyota'], [23, 'Hatchback', 'Honda'], [21, 'SUV'