Try:

oneEncoder.fit_transform(df[["Col2"]]).todense()

Suppose we have:

features = pd.DataFrame({"Col2":["a","b","c"]})

Then:

oneEncoder= OneHotEncoder()
oneEncoder.fit_transform(features[["Col2"]]).todense()
matrix([[1., 0., 0.],
        [0., 1., 0.],
        [0., 0., 1.]])

If you're dealing with a Series object, you may wish to reshape it:

oneEncoder.fit_transform(features.Col2.values.reshape(-1,1)).todense()
matrix([[1., 0., 0.],
        [0., 1., 0.],
        [0., 0., 1.]])

Dropping todense() method will leave your transformation in a sparse matrix.

And finally, you may alway decode what the columns of your matrix mean by:

oneEncoder.categories_
[array(['a', 'b', 'c'], dtype=object)]

Not surpisingly, they are your unique inputs ordered alphabetically.

Answer from Sergey Bushmanov on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Fits transformer to X and y with optional parameters fit_params and returns a transformed version of X. ... Input samples. yarray-like of shape (n_samples,) or (n_samples, n_outputs), default=None · Target values (None for unsupervised ...
🌐
datagy
datagy.io › home › python posts › one-hot encoding in scikit-learn with onehotencoder
One-Hot Encoding in Scikit-Learn with OneHotEncoder • datagy
April 14, 2024 - # One-hot encoding a single column from sklearn.preprocessing import OneHotEncoder from seaborn import load_dataset df = load_dataset('penguins') ohe = OneHotEncoder() transformed = ohe.fit_transform(df[['island']]) print(transformed.toarray()) # ...
🌐
scikit-learn
scikit-learn.org › 0.16 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.16.1 documentation
Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
Top answer
1 of 4
7

You can go directly to OneHotEncoding now without using the LabelEncoder, and as we move toward version 0.22 many might want to do things this way to avoid warnings and potential errors (see DOCS and EXAMPLES).


Example code 1 where ALL columns are encoded and where the categories are explicitly specified:

import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]

df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

countries = np.unique(X[:,0])
names = np.unique(X[:,1])

ohe = OneHotEncoder(categories=[countries, names])
X = ohe.fit_transform(X).toarray()

print (X)

Output for code example 1:

[[1. 0. 0. 0. 0. 1. 0.]
 [0. 0. 1. 0. 0. 0. 1.]
 [0. 1. 0. 1. 0. 0. 0.]
 [0. 0. 1. 0. 1. 0. 0.]]

Example code 2 showing the 'auto' option for specification of categories:

The first 3 columns encode the country names, the last four the personal names.

import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]

df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

ohe = OneHotEncoder(categories='auto')
X = ohe.fit_transform(X).toarray()

print (X)

Output for code example 2 (same as for 1):

[[1. 0. 0. 0. 0. 1. 0.]
 [0. 0. 1. 0. 0. 0. 1.]
 [0. 1. 0. 1. 0. 0. 0.]
 [0. 0. 1. 0. 1. 0. 0.]]

Example code 3 where only the first column is one hot encoded:

Now, here's the unique part. What if you only need to One Hot Encode a specific column for your data?

(Note: I've left the last column as strings for easier illustration. In reality it makes more sense to do this WHEN the last column was already numerical).

import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]

df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

countries = np.unique(X[:,0])
names = np.unique(X[:,1])

ohe = OneHotEncoder(categories=[countries]) # specify ONLY unique country names
tmp = ohe.fit_transform(X[:,0].reshape(-1, 1)).toarray()

X = np.append(tmp, names.reshape(-1,1), axis=1)

print (X)

Output for code example 3:

[[1.0 0.0 0.0 'Pechi']
 [0.0 0.0 1.0 'Raj']
 [0.0 1.0 0.0 'Sri']
 [0.0 0.0 1.0 'Vignesh']]
2 of 4
4

Below implementation should work well. Note that the input of onehotencoder fit_transform must not be 1-rank array and also output is sparse and we have used to_array() to expand it.

import pandas as pd
import numpy as np
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]


df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

le = LabelEncoder()
X_num = le.fit_transform(X[:,0]).reshape(-1,1)

ohe = OneHotEncoder()
X_num = ohe.fit_transform(X_num)

print (X_num.toarray())

X[:,0] = X_num

print (X)
🌐
Readthedocs
scikit-survival.readthedocs.io › en › stable › api › generated › sksurv.preprocessing.OneHotEncoder.html
sksurv.preprocessing.OneHotEncoder — scikit-survival 0.28.0
Fits the transformer to X by identifying categorical features and then returns a transformed version of X with categorical features one-hot encoded.
🌐
Medium
datasensei.medium.com › how-to-transform-nominal-data-for-ml-with-onehotencoder-from-scikit-learn-f6febfefb3c6
How to Transform Nominal Data for ML with OneHotEncoder from Scikit-Learn | by Data Seito | Medium
January 18, 2022 - Now let’s import OneHotEncoder from scikit-learn, break apart our dataframe into its numerical and categorical components, and fit our encoder to the example dataframe. from sklearn.preprocessing import OneHotEncoderX_num = df.select_dtypes(exclude='object') X_cat = df.select_dtypes(include='object')encoder = OneHotEncoder(sparse=False, handle_unknown='error')X_encoded = encoder.fit_transform(X_cat)X_encoded
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-4-data-preprocessing-feature-engineering › applying-encoders
Applying Encoders in Scikit-learn
Instantiate the encoder: Create an instance of OneHotEncoder. Fit and transform: Use the fit_transform method on the selected data.
Find elsewhere
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - Scikit-learn's OneHotEncoder can handle unknown categories by ignoring them or assigning them to a dedicated column, ensuring the model can still process new data effectively. This example demonstrates how to fit the encoder on the training data and then transform both training and test data, ...
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
import pandas as pd from sklearn.preprocessing import OneHotEncoder data = { 'Employee_ID': [10, 20, 15, 25, 30], 'Gender': ['M', 'F', 'F', 'M', 'F'], 'Remarks': ['Good', 'Nice', 'Good', 'Great', 'Nice'] } df = pd.DataFrame(data) print("Original Data:") print(df) categorical_columns = df.select_dtypes(include=['object']).columns encoder = OneHotEncoder(sparse_output=False) encoded_data = encoder.fit_transform(df[categorical_columns]) encoded_df = pd.DataFrame( encoded_data, columns=encoder.get_feature_names_out(categorical_columns) ) final_df = pd.concat( [df.drop(columns=categorical_columns), encoded_df], axis=1 ) print("\nOne-Hot Encoded Data:") print(final_df) Output: Output ·
Published: May 29, 2026
🌐
SourceForge
scikit-learn.sourceforge.net › dev › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.17.dev0 documentation
Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
🌐
Codefinity
codefinity.com › courses › v2 › 10db3746-c8ff-4c55-9ac3-4affa0b65c16 › 5c0178cc-e589-40e2-ad4b-6087554af252 › b57d2c2b-7b9f-4fbf-ab76-08b7ba5d8c32
Codefinity: Courses with certificates | Online Learning Platform
It is not c convenient format to work with, so we transform it: We will use OneHotEncoder to create new features for the categorical columns of our dataset. OneHotEncoder cannot process NaNs, so you have to preprocess them first.
🌐
scikit-learn
scikit-learn.org › 0.20 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.20.4 documentation
Given a dataset with two features, ... data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) ......
🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - # creating one hot encoder object onehotencoder = OneHotEncoder() # reshape the 1-D country array to 2-D as fit_transform expects 2-D and fit the encoder X = onehotencoder.fit_transform(df.Country.values.reshape(-1, 1)).toarray()
🌐
W3cubDocs
docs.w3cub.com › scikit_learn › modules › generated › sklearn.preprocessing.onehotencoder
preprocessing.OneHotEncoder() - Scikit-learn - W3cubDocs
Fits transformer to X and y with optional parameters fit_params and returns a transformed version of X. ... Input samples. yarray-like of shape (n_samples,) or (n_samples, n_outputs), default=None · Target values (None for unsupervised transformations).
🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.10.dev0 documentation
Fits transformer to X and y with optional parameters fit_params and returns a transformed version of X. ... Input samples. yarray-like of shape (n_samples,) or (n_samples, n_outputs), default=None · Target values (None for unsupervised ...
🌐
Built In
builtin.com › articles › one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - The usual wisdom is to use sklearn’s sklearn.preprocessing.OneHotEncoder for this purpose, because using its fit/transform paradigm allows you to use the training data set to “teach” categories and apply it to your real-world input data.
🌐
scikit-learn
scikit-learn.org › 0.19 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.19.2 documentation
Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
Top answer
1 of 16
317

Approach 1: You can use pandas' pd.get_dummies.

Example 1:

import pandas as pd
s = pd.Series(list('abca'))
pd.get_dummies(s)
Out[]: 
     a    b    c
0  1.0  0.0  0.0
1  0.0  1.0  0.0
2  0.0  0.0  1.0
3  1.0  0.0  0.0

Example 2:

The following will transform a given column into one hot. Use prefix to have multiple dummies.

import pandas as pd
        
df = pd.DataFrame({
          'A':['a','b','a'],
          'B':['b','a','c']
        })
df
Out[]: 
   A  B
0  a  b
1  b  a
2  a  c

# Get one hot encoding of columns B
one_hot = pd.get_dummies(df['B'])
# Drop column B as it is now encoded
df = df.drop('B',axis = 1)
# Join the encoded df
df = df.join(one_hot)
df  
Out[]: 
       A  a  b  c
    0  a  0  1  0
    1  b  1  0  0
    2  a  0  0  1

Approach 2: Use Scikit-learn

Using a OneHotEncoder has the advantage of being able to fit on some training data and then transform on some other data using the same instance. We also have handle_unknown to further control what the encoder does with unseen data.

Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding.

>>> from sklearn.preprocessing import OneHotEncoder
>>> enc = OneHotEncoder()
>>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]])   
OneHotEncoder(categorical_features='all', dtype=<class 'numpy.float64'>,
   handle_unknown='error', n_values='auto', sparse=True)
>>> enc.n_values_
array([2, 3, 4])
>>> enc.feature_indices_
array([0, 2, 5, 9], dtype=int32)
>>> enc.transform([[0, 1, 1]]).toarray()
array([[ 1.,  0.,  0.,  1.,  0.,  0.,  1.,  0.,  0.]])

Here is the link for this example: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html

2 of 16
149

Much easier to use Pandas for basic one-hot encoding. If you're looking for more options you can use scikit-learn.

For basic one-hot encoding with Pandas you pass your data frame into the get_dummies function.

For example, if I have a dataframe called imdb_movies:

...and I want to one-hot encode the Rated column, I do this:

pd.get_dummies(imdb_movies.Rated)

This returns a new dataframe with a column for every "level" of rating that exists, along with either a 1 or 0 specifying the presence of that rating for a given observation.

Usually, we want this to be part of the original dataframe. In this case, we attach our new dummy coded frame onto the original frame using "column-binding.

We can column-bind by using Pandas concat function:

rated_dummies = pd.get_dummies(imdb_movies.Rated)
pd.concat([imdb_movies, rated_dummies], axis=1)

We can now run an analysis on our full dataframe.

SIMPLE UTILITY FUNCTION

I would recommend making yourself a utility function to do this quickly:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    return(res)

Usage:

encode_and_bind(imdb_movies, 'Rated')

Result:

Also, as per @pmalbu comment, if you would like the function to remove the original feature_to_encode then use this version:

def encode_and_bind(original_dataframe, feature_to_encode):
    dummies = pd.get_dummies(original_dataframe[[feature_to_encode]])
    res = pd.concat([original_dataframe, dummies], axis=1)
    res = res.drop([feature_to_encode], axis=1)
    return(res) 

You can encode multiple features at the same time as follows:

features_to_encode = ['feature_1', 'feature_2', 'feature_3',
                      'feature_4']
for feature in features_to_encode:
    res = encode_and_bind(train_set, feature)
Top answer
1 of 10
49

OneHotEncoder Encodes categorical integer features as a one-hot numeric array. Its Transform method returns a sparse matrix if sparse=True, otherwise it returns a 2-d array.

You can't cast a 2-d array (or sparse matrix) into a Pandas Series. You must create a Pandas Serie (a column in a Pandas dataFrame) for each category.

I would recommend pandas.get_dummies instead:

data = pd.get_dummies(data,prefix=['Profession'], columns = ['Profession'], drop_first=True)

EDIT:

Using Sklearn OneHotEncoder:

transformed = jobs_encoder.transform(data['Profession'].to_numpy().reshape(-1, 1))
#Create a Pandas DataFrame of the hot encoded column
ohe_df = pd.DataFrame(transformed, columns=jobs_encoder.get_feature_names())
#concat with original data
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)

Other Options: If you are doing hyperparameter tuning with GridSearch it's recommanded to use ColumnTransformer and FeatureUnion with Pipeline or directly make_column_transformer

2 of 10
24

So turned out that Scikit-Learns LabelBinarizer gave me better luck in converting the data to one-hot encoded format, with help from Amnie's solution, my final code is as follows

import pandas as pd
from sklearn.preprocessing import LabelBinarizer

jobs_encoder = LabelBinarizer()
jobs_encoder.fit(data['Profession'])
transformed = jobs_encoder.transform(data['Profession'])
ohe_df = pd.DataFrame(transformed)
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)