Try:

oneEncoder.fit_transform(df[["Col2"]]).todense()

Suppose we have:

features = pd.DataFrame({"Col2":["a","b","c"]})

Then:

oneEncoder= OneHotEncoder()
oneEncoder.fit_transform(features[["Col2"]]).todense()
matrix([[1., 0., 0.],
        [0., 1., 0.],
        [0., 0., 1.]])

If you're dealing with a Series object, you may wish to reshape it:

oneEncoder.fit_transform(features.Col2.values.reshape(-1,1)).todense()
matrix([[1., 0., 0.],
        [0., 1., 0.],
        [0., 0., 1.]])

Dropping todense() method will leave your transformation in a sparse matrix.

And finally, you may alway decode what the columns of your matrix mean by:

oneEncoder.categories_
[array(['a', 'b', 'c'], dtype=object)]

Not surpisingly, they are your unique inputs ordered alphabetically.

Answer from Sergey Bushmanov on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Fit OneHotEncoder to X. ... The data to determine the categories of each feature. ... Ignored. This parameter exists only for compatibility with Pipeline. ... Fitted encoder. ... Fit to data, then transform it.
Discussions

python - How to perform OneHotEncoding in Sklearn, getting value error - Stack Overflow
Yes, That's because you are passing rank1 array i.e X[:,0] to onehotencoder.fit_transform which is deprecated. More on stackoverflow.com
🌐 stackoverflow.com
python - sklearn.preprocessing.OneHotEncoder and the way to read it - Stack Overflow
I have been using one-hot encoding for a while now in all pre-processing data pipelines that I have had. But I have run into an issue now that I am trying to pre-process new data automatically with... More on stackoverflow.com
🌐 stackoverflow.com
Onehotencoder.fit_transform - Packages & Environments - Anaconda Forum
Hi onehotencoder = OneHotEncoder(categorical_features = [0]) X = onehotencoder.fit_transform(X).toarray() This method dont work and deprecated OUTPUT integer data will change in version 0.22. Currently, the categories are determined based on the range [0, max(values)], while in the future they ... More on forum.anaconda.com
🌐 forum.anaconda.com
0
August 23, 2023
scikit learn - How to perform one hot encoding on multiple categorical columns - Data Science Stack Exchange
I am trying to perform one-hot encoding on some categorical columns. From the tutorial I am following, I am supposed to do LabelEncoding before One hot encoding. I have successfully performed the More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
April 5, 2020
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
import pandas as pd from sklearn.preprocessing import OneHotEncoder data = { 'Employee_ID': [10, 20, 15, 25, 30], 'Gender': ['M', 'F', 'F', 'M', 'F'], 'Remarks': ['Good', 'Nice', 'Good', 'Great', 'Nice'] } df = pd.DataFrame(data) print("Original Data:") print(df) categorical_columns = df.select_dtypes(include=['object']).columns encoder = OneHotEncoder(sparse_output=False) encoded_data = encoder.fit_transform(df[categorical_columns]) encoded_df = pd.DataFrame( encoded_data, columns=encoder.get_feature_names_out(categorical_columns) ) final_df = pd.concat( [df.drop(columns=categorical_columns), encoded_df], axis=1 ) print("\nOne-Hot Encoded Data:") print(final_df)
Published: May 29, 2026
🌐
Readthedocs
scikit-survival.readthedocs.io › en › stable › api › generated › sksurv.preprocessing.OneHotEncoder.html
sksurv.preprocessing.OneHotEncoder — scikit-survival 0.28.0
Fits the transformer to X by identifying categorical features and then returns a transformed version of X with categorical features one-hot encoded.
🌐
datagy
datagy.io › home › python posts › one-hot encoding in scikit-learn with onehotencoder
One-Hot Encoding in Scikit-Learn with OneHotEncoder • datagy
April 14, 2024 - # One-hot encoding a single column from sklearn.preprocessing import OneHotEncoder from seaborn import load_dataset df = load_dataset('penguins') ohe = OneHotEncoder() transformed = ohe.fit_transform(df[['island']]) print(transformed.toarray()) # Returns: # [[0.
🌐
Medium
datasensei.medium.com › how-to-transform-nominal-data-for-ml-with-onehotencoder-from-scikit-learn-f6febfefb3c6
How to Transform Nominal Data for ML with OneHotEncoder from Scikit-Learn | by Data Seito | Medium
January 18, 2022 - from sklearn.preprocessing import OneHotEncoderX_num = df.select_dtypes(exclude='object') X_cat = df.select_dtypes(include='object')encoder = OneHotEncoder(sparse=False, handle_unknown='error')X_encoded = encoder.fit_transform(X_cat)X_encoded
Find elsewhere
Top answer
1 of 4
7

You can go directly to OneHotEncoding now without using the LabelEncoder, and as we move toward version 0.22 many might want to do things this way to avoid warnings and potential errors (see DOCS and EXAMPLES).


Example code 1 where ALL columns are encoded and where the categories are explicitly specified:

import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]

df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

countries = np.unique(X[:,0])
names = np.unique(X[:,1])

ohe = OneHotEncoder(categories=[countries, names])
X = ohe.fit_transform(X).toarray()

print (X)

Output for code example 1:

[[1. 0. 0. 0. 0. 1. 0.]
 [0. 0. 1. 0. 0. 0. 1.]
 [0. 1. 0. 1. 0. 0. 0.]
 [0. 0. 1. 0. 1. 0. 0.]]

Example code 2 showing the 'auto' option for specification of categories:

The first 3 columns encode the country names, the last four the personal names.

import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]

df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

ohe = OneHotEncoder(categories='auto')
X = ohe.fit_transform(X).toarray()

print (X)

Output for code example 2 (same as for 1):

[[1. 0. 0. 0. 0. 1. 0.]
 [0. 0. 1. 0. 0. 0. 1.]
 [0. 1. 0. 1. 0. 0. 0.]
 [0. 0. 1. 0. 1. 0. 0.]]

Example code 3 where only the first column is one hot encoded:

Now, here's the unique part. What if you only need to One Hot Encode a specific column for your data?

(Note: I've left the last column as strings for easier illustration. In reality it makes more sense to do this WHEN the last column was already numerical).

import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]

df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

countries = np.unique(X[:,0])
names = np.unique(X[:,1])

ohe = OneHotEncoder(categories=[countries]) # specify ONLY unique country names
tmp = ohe.fit_transform(X[:,0].reshape(-1, 1)).toarray()

X = np.append(tmp, names.reshape(-1,1), axis=1)

print (X)

Output for code example 3:

[[1.0 0.0 0.0 'Pechi']
 [0.0 0.0 1.0 'Raj']
 [0.0 1.0 0.0 'Sri']
 [0.0 0.0 1.0 'Vignesh']]
2 of 4
4

Below implementation should work well. Note that the input of onehotencoder fit_transform must not be 1-rank array and also output is sparse and we have used to_array() to expand it.

import pandas as pd
import numpy as np
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder

data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]


df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values

le = LabelEncoder()
X_num = le.fit_transform(X[:,0]).reshape(-1,1)

ohe = OneHotEncoder()
X_num = ohe.fit_transform(X_num)

print (X_num.toarray())

X[:,0] = X_num

print (X)
🌐
Built In
builtin.com › articles › one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - The usual wisdom is to use sklearn’s sklearn.preprocessing.OneHotEncoder for this purpose, because using its fit/transform paradigm allows you to use the training data set to “teach” categories and apply it to your real-world input data.
🌐
Anaconda Forum
forum.anaconda.com › product help › packages & environments
Onehotencoder.fit_transform - Packages & Environments - Anaconda Forum
August 23, 2023 - Hi onehotencoder = OneHotEncoder(categorical_features = [0]) X = onehotencoder.fit_transform(X).toarray() This method dont work and deprecated OUTPUT integer data will change in version 0.22. Currently, the categories are determined based on the range [0, max(values)], while in the future they will be determined based on the unique values.
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-4-data-preprocessing-feature-engineering › applying-encoders
Applying Encoders in Scikit-learn
Instantiate the encoder: Create an instance of OneHotEncoder. Fit and transform: Use the fit_transform method on the selected data.
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - Scikit-learn's OneHotEncoder can handle unknown categories by ignoring them or assigning them to a dedicated column, ensuring the model can still process new data effectively. This example demonstrates how to fit the encoder on the training data and then transform both training and test data, including handling categories that were not present in the training set.
Top answer
1 of 4
21

LabelEncoder is not made to transform the data but the target (also known as labels) as explained here. If you want to encode the data you should use OrdinalEncoder.

If you really need to do it this way:

categorical_cols = ['a', 'b', 'c', 'd'] 

from sklearn.preprocessing import LabelEncoder
# instantiate labelencoder object
le = LabelEncoder()

# apply le on categorical feature columns
data[categorical_cols] = data[categorical_cols].apply(lambda col: le.fit_transform(col))    
from sklearn.preprocessing import OneHotEncoder
ohe = OneHotEncoder()

#One-hot-encode the categorical columns.
#Unfortunately outputs an array instead of dataframe.
array_hot_encoded = ohe.fit_transform(data[categorical_cols])

#Convert it to df
data_hot_encoded = pd.DataFrame(array_hot_encoded, index=data.index)

#Extract only the columns that didnt need to be encoded
data_other_cols = data.drop(columns=categorical_cols)

#Concatenate the two dataframes : 
data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1)

Otherwise:

I suggest you to use pandas.get_dummies if you want to achieve one-hot-encoding from raw data (without having to use OrdinalEncoder before) :

#categorical data
categorical_cols = ['a', 'b', 'c', 'd'] 

#import pandas as pd
df = pd.get_dummies(data, columns = categorical_cols)

You can also use drop_first argument to remove one of the one-hot-encoded columns, as some models require.

2 of 4
8

You can do dummy encoding using Pandas in order to get one-hot encoding as shown below:

import pandas as pd

# Multiple categorical columns
categorical_cols = ['a', 'b', 'c', 'd']

pd.get_dummies(data, columns=categorical_cols)

If you want to do one-hot encoding using sklearn library, you can get it done as shown below:

from sklearn.preprocessing import OneHotEncoder
onehotencoder = OneHotEncoder()

transformed_data = onehotencoder.fit_transform(data[categorical_cols])

# the above transformed_data is an array so convert it to dataframe
encoded_data = pd.DataFrame(transformed_data, index=data.index)

# now concatenate the original data and the encoded data using pandas
concatenated_data = pd.concat([data, encoded_data], axis=1)

If a single column has more than 500 categories, the aforementioned way of one-hot encoding is not a good approach. In this case, we can do one-hot encoding for the top 10 or 20 categories that are occurring most for a particular column. A sample code is shown below:

categorical_cols = ['a', 'b', 'c', 'd']

# Let's say we have a column 'b' which has more than 500 categories.
# Find the top 10 most frequent categories for column 'b'
data.b.value_counts().sort_values(ascending = False).head(20)

# make a list of the most frequent categories of the column
top_10_occurring_cat = [cat for cat in data.b.value_counts().sort_values(ascending = False).head(10).index]

# now make the 10 binary variables
for cat in top_10_occurring_cat:
    data[cat] = np.where(data['b'] == cat, 1, 0) # whenever data['b'] == cat replace it with 1 else 0

# This is done for one categorical column, similarly you can repeat for all categorical columns
🌐
Codefinity
codefinity.com › courses › v2 › 10db3746-c8ff-4c55-9ac3-4affa0b65c16 › 5c0178cc-e589-40e2-ad4b-6087554af252 › b57d2c2b-7b9f-4fbf-ab76-08b7ba5d8c32
Codefinity: Courses with certificates | Online Learning Platform
It is not c convenient format to work with, so we transform it: We will use OneHotEncoder to create new features for the categorical columns of our dataset. OneHotEncoder cannot process NaNs, so you have to preprocess them first. ... 12345678910 from sklearn.preprocessing import OneHotEncoder # data is loaded already # num_cols and cat_cols are created already encoder = OneHotEncoder() new_data = pd.DataFrame(encoder.fit_transform(data[cat_cols]).toarray()) # join new features to the dataset, but remove categorical features data = data[num_cols].join(new_data)
🌐
Readthedocs
scikit-survival.readthedocs.io › en › v0.26.0 › api › generated › sksurv.preprocessing.OneHotEncoder.html
sksurv.preprocessing.OneHotEncoder — scikit-survival 0.26.0
Fits the transformer to X by identifying categorical features and then returns a transformed version of X with categorical features one-hot encoded.
🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - # creating one hot encoder object onehotencoder = OneHotEncoder() # reshape the 1-D country array to 2-D as fit_transform expects 2-D and fit the encoder X = onehotencoder.fit_transform(df.Country.values.reshape(-1, 1)).toarray()