🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
scikit-learn
scikit-learn.org › 0.20 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.20.4 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', ...
🌐
datagy
datagy.io › home › python posts › one-hot encoding in scikit-learn with onehotencoder
One-Hot Encoding in Scikit-Learn with OneHotEncoder • datagy
April 14, 2024 - In this tutorial, you’ll learn ... data in sklearn. One-hot encoding is a process by which categorical data (such as nominal data) are converted into numerical features of a dataset....
🌐
scikit-learn
scikit-learn.org › 0.16 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.16.1 documentation
Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - For more flexibility and control over the encoding process, Scikit-learn offers the OneHotEncoder class. This class provides advanced options, such as handling unknown categories and fitting the encoder to the training data. from sklearn.preprocessing import OneHotEncoder import numpy as np # Creating the encoder enc = OneHotEncoder(handle_unknown='ignore') # Sample data X = [['Red'], ['Green'], ['Blue']] # Fitting the encoder to the data enc.fit(X) # Transforming new data result = enc.transform([['Red']]).toarray() # Displaying the encoded result print(result)
🌐
Built In
builtin.com › articles › one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - For sklearn, the explicit category setting was achieved by passing a parameter to the constructor of the OneHotEncoder class, while for Pandas, we had to set up the categorical data type.
🌐
pythontutorials
pythontutorials.net › blog › one-hot-encode-sklearn
One-Hot Encoding with scikit-learn: A Comprehensive Guide — pythontutorials.net
In this example, we first create a pandas DataFrame and then use pd.get_dummies to one-hot encode the 'color' column. We then split the data into training and testing sets using train_test_split from sklearn.model_selection.
🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.preprocessing.OneHotEncoder.html
One-Hot Encoding in Scikit-learn
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.5.2 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
Find elsewhere
🌐
scikit-learn
scikit-learn.org › 0.19 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.19.2 documentation
Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
🌐
Sklearn
sklearn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.0 documentation - sklearn
The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are encoded using a one-hot (aka ‘one-of-K’ or ‘dummy’) encoding scheme.
🌐
Wordpress
lifewithdatacom.wordpress.com › 2022 › 03 › 09 › onehotencoder-how-to-do-one-hot-encoding-in-sklearn
OneHotEncoder - How to do One Hot Encoding in sklearn.
September 3, 2022 - Build a website. Sell your stuff. Write a blog. And so much more · This site is currently private. Log in to WordPress.com to request access
🌐
scikit-learn
scikit-learn.org › 0.18 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.18.2 documentation
Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, ...
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to one hot encode sequence data in python
How to One Hot Encode Sequence Data in Python - MachineLearningMastery.com
August 14, 2019 - I used the following code to use one hot encode for some categorical variables, but, the model fit throws error after successfully using one hot encoding. There is no error if I use ordinal encoding. Here is the code: ... from sklearn.preprocessing import OneHotEncoder def one_hot_encode_features(df_train,df_test): features = [‘Fare’, ‘Cabin’, ‘Age’, ‘Sex’] #features = [ ‘Cabin’, ‘Sex’] df_combined = pd.concat([df_train[features], df_test[features]]) for feature in features: le = preprocessing.LabelEncoder() onehot_encoder = OneHotEncoder() le = le.fit(df_combined[featu
🌐
scikit-learn
scikit-learn.org › 1.0 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 1.0.2 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
Medium
datasensei.medium.com › how-to-transform-nominal-data-for-ml-with-onehotencoder-from-scikit-learn-f6febfefb3c6
How to Transform Nominal Data for ML with OneHotEncoder from Scikit-Learn | by Data Seito | Medium
January 18, 2022 - The purpose of this article is twofold: first, to give a brief overview of one-hot encoding categorical data, and second to demonstrate how to one-hot encode categorical data using scikit-learn’s OneHotEncoder class.
🌐
scikit-learn
scikit-learn.org › 0.21 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.21.3 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', ...
🌐
Codefinity
codefinity.com › courses › v2 › a65bbc96-309e-4df9-a790-a1eb8c815a1c › 1fce4aa9-710f-4bc9-ad66-16b4b2d30929 › a6d33d0d-3057-4a2f-b8df-4ecd00ffd598
Learn One-Hot Encoder | Preprocessing Data with Scikit-learn
To apply OneHotEncoder, initialize the encoder object and pass the selected columns to .fit_transform(), in the same way as with other transformers. 1234567891011 import pandas as pd from sklearn.preprocessing import OneHotEncoder df = ...
🌐
Ryan Nolan Data
ryanandmattdatascience.com › home › scikit-learn › one hot encoder
How to Use One Hot Encoder in Scikit-Learn (With Examples)
April 12, 2025 - The goal in this article is to One Hot Encode the values for the size column. The next line of code is creating and configuring an instance of the OneHotEncoder class from the sklearn.preprocessing module in scikit-learn
Top answer
1 of 10
49

OneHotEncoder Encodes categorical integer features as a one-hot numeric array. Its Transform method returns a sparse matrix if sparse=True, otherwise it returns a 2-d array.

You can't cast a 2-d array (or sparse matrix) into a Pandas Series. You must create a Pandas Serie (a column in a Pandas dataFrame) for each category.

I would recommend pandas.get_dummies instead:

data = pd.get_dummies(data,prefix=['Profession'], columns = ['Profession'], drop_first=True)

EDIT:

Using Sklearn OneHotEncoder:

transformed = jobs_encoder.transform(data['Profession'].to_numpy().reshape(-1, 1))
#Create a Pandas DataFrame of the hot encoded column
ohe_df = pd.DataFrame(transformed, columns=jobs_encoder.get_feature_names())
#concat with original data
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)

Other Options: If you are doing hyperparameter tuning with GridSearch it's recommanded to use ColumnTransformer and FeatureUnion with Pipeline or directly make_column_transformer

2 of 10
24

So turned out that Scikit-Learns LabelBinarizer gave me better luck in converting the data to one-hot encoded format, with help from Amnie's solution, my final code is as follows

import pandas as pd
from sklearn.preprocessing import LabelBinarizer

jobs_encoder = LabelBinarizer()
jobs_encoder.fit(data['Profession'])
transformed = jobs_encoder.transform(data['Profession'])
ohe_df = pd.DataFrame(transformed)
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)