OneHotEncoder Encodes categorical integer features as a one-hot numeric array. Its Transform method returns a sparse matrix if sparse=True, otherwise it returns a 2-d array.

You can't cast a 2-d array (or sparse matrix) into a Pandas Series. You must create a Pandas Serie (a column in a Pandas dataFrame) for each category.

I would recommend pandas.get_dummies instead:

data = pd.get_dummies(data,prefix=['Profession'], columns = ['Profession'], drop_first=True)

EDIT:

Using Sklearn OneHotEncoder:

transformed = jobs_encoder.transform(data['Profession'].to_numpy().reshape(-1, 1))
#Create a Pandas DataFrame of the hot encoded column
ohe_df = pd.DataFrame(transformed, columns=jobs_encoder.get_feature_names())
#concat with original data
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)

Other Options: If you are doing hyperparameter tuning with GridSearch it's recommanded to use ColumnTransformer and FeatureUnion with Pipeline or directly make_column_transformer

Answer from Amine Benatmane on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
One-Hot Encoding can be implemented in Python using libraries such as Pandas and Scikit-learn, which provide simple and efficient methods for converting categorical data into binary columns.
Published: May 29, 2026
🌐
scikit-learn
scikit-learn.org › 0.19 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.19.2 documentation
class sklearn.preprocessing.OneHotEncoder(n_values=’auto’, categorical_features=’all’, dtype=<class ‘numpy.float64’>, sparse=True, handle_unknown=’error’)[source]¶ · Encode categorical integer features using a one-hot aka one-of-K scheme.
Top answer
1 of 10
49

OneHotEncoder Encodes categorical integer features as a one-hot numeric array. Its Transform method returns a sparse matrix if sparse=True, otherwise it returns a 2-d array.

You can't cast a 2-d array (or sparse matrix) into a Pandas Series. You must create a Pandas Serie (a column in a Pandas dataFrame) for each category.

I would recommend pandas.get_dummies instead:

data = pd.get_dummies(data,prefix=['Profession'], columns = ['Profession'], drop_first=True)

EDIT:

Using Sklearn OneHotEncoder:

transformed = jobs_encoder.transform(data['Profession'].to_numpy().reshape(-1, 1))
#Create a Pandas DataFrame of the hot encoded column
ohe_df = pd.DataFrame(transformed, columns=jobs_encoder.get_feature_names())
#concat with original data
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)

Other Options: If you are doing hyperparameter tuning with GridSearch it's recommanded to use ColumnTransformer and FeatureUnion with Pipeline or directly make_column_transformer

2 of 10
24

So turned out that Scikit-Learns LabelBinarizer gave me better luck in converting the data to one-hot encoded format, with help from Amnie's solution, my final code is as follows

import pandas as pd
from sklearn.preprocessing import LabelBinarizer

jobs_encoder = LabelBinarizer()
jobs_encoder.fit(data['Profession'])
transformed = jobs_encoder.transform(data['Profession'])
ohe_df = pd.DataFrame(transformed)
data = pd.concat([data, ohe_df], axis=1).drop(['Profession'], axis=1)
🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
In Python, we can implement one-hot encoding using the get_dummies() function in the pandas module and the OneHotEncoder class in the sklearn module.
🌐
Built In
builtin.com › articles › one-hot-encoding
One Hot Encoding Explained | Built In
February 15, 2024 - One hot encoding (OHE) is a machine learning technique that encodes categorical data to numerical ones. If you want to perform one hot encoding, both sklearn.preprocessing.OneHotEncoder and pandas.get_dummies are popular choices.
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - For more flexibility and control over the encoding process, Scikit-learn offers the OneHotEncoder class. This class provides advanced options, such as handling unknown categories and fitting the encoder to the training data. from sklearn.preprocessing import OneHotEncoder import numpy as np # Creating the encoder enc = OneHotEncoder(handle_unknown='ignore') # Sample data X = [['Red'], ['Green'], ['Blue']] # Fitting the encoder to the data enc.fit(X) # Transforming new data result = enc.transform([['Red']]).toarray() # Displaying the encoded result print(result)
Find elsewhere
Top answer
1 of 3
6

This library provides several categorical encoders which make sklearn / numpy play nicely with pandas https://github.com/wdm0006/categorical_encoding

However, they do not yet support "handle unknown category"

for now I will use

myEncoder = OneHotEncoder(sparse=False, handle_unknown='ignore')
myEncoder.fit(df[columnsToEncode])

pd.concat([df.drop(columnsToEncode, 1),
          pd.DataFrame(myEncoder.transform(df[columnsToEncode]))], axis=1).reindex()

As this supports unknown datasets. For now, I will stick with half-pandas half-numpy because of the nice pandas labels. for the numeric columns.

2 of 3
3

For One Hot Encoding I recommend using ColumnTransformer and OneHotEncoder instead of get_dummies. That's because OneHotEncoder returns an object which can be used to encode unseen samples using the same mapping that you used on your training data.

The following code encodes all the columns provided in the columns_to_encode variable:

import pandas as pd
import numpy as np

df = pd.DataFrame({'cat_1': ['A1', 'B1', 'C1'], 'num_1': [100, 200, 300], 
                   'cat_2': ['A2', 'B2', 'C2'], 'cat_3': ['A3', 'B3', 'C3'],
                   'label': [1, 0, 0]})

X = df.iloc[:, :-1].values
y = df.iloc[:, -1].values

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
columns_to_encode = [0, 2, 3] # Change here
ct = ColumnTransformer(transformers=[('encoder', OneHotEncoder(), columns_to_encode)], remainder='passthrough')
X = np.array(ct.fit_transform(X))

X:

array([[1.0, 0.0, 0.0, 1.0, 0.0, 0.0, 1.0, 0.0, 0.0, 100],
       [0.0, 1.0, 0.0, 0.0, 1.0, 0.0, 0.0, 1.0, 0.0, 200],
       [0.0, 0.0, 1.0, 0.0, 0.0, 1.0, 0.0, 0.0, 1.0, 300]], dtype=object)

To avoid multicollinearity due to the dummy variable trap, I would also suggest removing one of the columns returned by each column that you encoded. The following code encodes all the columns provided in the columns_to_encode variable AND it removes the last column of each one hot encoded column:

import pandas as pd
import numpy as np

def sum_prev (l_in):
    l_out = []
    l_out.append(l_in[0])
    for i in range(len(l_in)-1):
        l_out.append(l_out[i] + l_in[i+1])
    return [e - 1 for e in l_out]

df = pd.DataFrame({'cat_1': ['A1', 'B1', 'C1'], 'num_1': [100, 200, 300], 
                   'cat_2': ['A2', 'B2', 'C2'], 'cat_3': ['A3', 'B3', 'C3'],
                   'label': [1, 0, 0]})

X = df.iloc[:, :-1].values
y = df.iloc[:, -1].values

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
columns_to_encode = [0, 2, 3] # Change here
ct = ColumnTransformer(transformers=[('encoder', OneHotEncoder(), columns_to_encode)], remainder='passthrough')

columns_to_encode = [df.iloc[:, del_idx].nunique() for del_idx in columns_to_encode]
columns_to_encode = sum_prev(columns_to_encode)
X = np.array(ct.fit_transform(X))
X = np.delete(X, columns_to_encode, 1)

X:

array([[1.0, 0.0, 1.0, 0.0, 1.0, 0.0, 100],
       [0.0, 1.0, 0.0, 1.0, 0.0, 1.0, 200],
       [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 300]], dtype=object)
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-4-data-preprocessing-feature-engineering › encoding-categorical-features
Encoding Categorical Features Scikit-learn
In Scikit-learn, sklearn.preprocessing.OneHotEncoder is the primary tool for this transformation. It's designed to work within Scikit-learn pipelines and offers parameters to handle categories not seen during training (handle_unknown='ignore') or to control the output format (sparse_output=True ...
🌐
scikit-learn
scikit-learn.org › 0.16 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.16.1 documentation
class sklearn.preprocessing.OneHotEncoder(n_values='auto', categorical_features='all', dtype=<type 'float'>, sparse=True, handle_unknown='error')[source]¶ · Encode categorical integer features using a one-hot aka one-of-K scheme.
🌐
Medium
medium.com › data-science › robust-one-hot-encoding-930b5f8943af
Robust One-Hot Encoding. Production grade one-hot encoding… | by Hans Christian Ekne | TDS Archive | Medium
April 26, 2024 - Regarding the issues related to one-hot encoding we described above, they can be mitigated by using the widely available and tested scikit-learn library, and more specifically the sklearn.preprocessing.OneHotEncoder class.
🌐
Towards Data Science
towardsdatascience.com › home › latest › one hot encoding scikit vs pandas
One Hot Encoding scikit vs pandas | Towards Data Science
March 5, 2025 - Both sklearn.preprocessing.OneHotEncoder and pandas.get_dummies are popular choices (well, practically the only choices unless you want want to implement it yourself) to perform One Hot Encoding.
Top answer
1 of 2
1

(My basic answer, post a more informed one and I'll accept it.)

Examples from this article show that the OneHotEncoder will encode for every unique value in a column.

You can check what's being encoded by using the OneHotEncoder().categories_ attribute. This attribute will give you a series of arrays that include all the unique values of each column that have been encoded for. If you feed numerical values into the OneHotEncoder you'll notice that these arrays contain every unique numerical value as a category.

To avoid this, you should pre-select your categorical columns and feed those to the OneHotEncoder. See the following SciKit-Learn Tutorial, and the reference for ColumnTransformer to see how this can be included in a pipeline.

2 of 2
0

Scikit-learn's OneHotEncoder will encode all variables in the dataframe by default.

If you want to encode just a subset, you need to wrap the OneHotEncoder with the ColumnTransformer.

Alternatively, you can use Feature-engine's OneHotEncoder which encodes only variables of type object or categorical by default, leaving the numerical variables the way they are.

See for example this code snippet:

import pandas as pd
from feature_engine.encoding import OneHotEncoder
X = pd.DataFrame(dict(x1 = [1,2,3,4], x2 = ["a", "a", "b", "c"]))
ohe = OneHotEncoder()
ohe.fit(X)
ohe.transform(X)

The result of that transformation is the following dataframe, where the numerical variable remains unchanged and the categorical one was one hot encoded:

   x1  x2_a  x2_b  x2_c
0   1     1     0     0
1   2     1     0     0
2   3     0     1     0
3   4     0     0     1

I leave a link to Feature-engine's OneHotEncoder for more details.

🌐
Train in Data
blog.trainindata.com › one-hot-encoding-categorical-variables
One-hot encoding categorical variables | Train in Data Blog
January 25, 2023 - Let’s import the one hot encoder and a class to transform variable subsets: from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder
Top answer
1 of 1
1

You can't pass a feature containing a merged list directly to the model. You should one-hot encode into separate columns first:

  • If you just want something quick and easy, get_dummies is fine for development, but the following approaches are generally preferred by most sources I've read.
  • If you want to encode your input data, use OneHotEncoder (OHE) to encode one or more columns, then merge with your other features. OHE gives good control over output format, stores intermediate data and has error handling. Good for production.
  • If you need to encode a single column, typically but not limited to labels, use LabelBinarizer to one-hot encode a column with a single value, or use MultiLabelBinarizer to one-hot encode a column with multiple values.

Once you have your one-hot encoded data/labels, you don't need to "tell" the model that certain features are one-hot. You just train the model on the data set using clf.fit(X_train, y_train) and make predictions using clf.predict(X_test).

OHE example

from sklearn.preprocessing import OneHotEncoder
import pandas as pd

X = [['Male', 1], ['Female', 3], ['Female', 2]]
ohe = OneHotEncoder(handle_unknown='ignore')
X_enc = ohe.fit_transform(X).toarray()

# Convert to dataframe if you need to merge this with other features:
df = pd.DataFrame(X_enc, columns=ohe.get_feature_names())

MLB example

from sklearn.preprocessing import MultiLabelBinarizer
import pandas as pd

df = pd.DataFrame({
   'style': ['Folk', 'Rock', 'Classical'],
   'instruments': [['guitar', 'vocals'], ['guitar', 'bass', 'drums', 'vocals'], ['piano']]
})

mlb = MultiLabelBinarizer()
encoded = mlb.fit_transform(df['instruments'])
encoded_df = pd.DataFrame(encoded, columns=mlb.classes_, index=df['instruments'].index)

# Drop old column and merge new encoded columns
df = df.drop('instruments', axis=1)
df = pd.concat([df, encoded_df], axis=1, sort=False)
🌐
Medium
medium.com › @michaeldelsole › what-is-one-hot-encoding-and-how-to-do-it-f0ae272f1179
What is One Hot Encoding and How to Do It | by Michael DelSole | Medium
April 24, 2018 - Now let’s do the actual encoding. Sklearn makes it incredibly easy, but there is a catch. You might have noticed we imported both the labelencoder and the one hot encoder. Sklearn’s one hot encoder doesn’t actually know how to convert categories to numbers, it only knows how to convert numbers to binary.
🌐
SourceForge
scikit-learn.sourceforge.net › 0.15 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.15.2 documentation
class sklearn.preprocessing.OneHotEncoder(n_values='auto', categorical_features='all', dtype=<type 'float'>, sparse=True)¶ · Encode categorical integer features using a one-hot aka one-of-K scheme.