There is actually 2 warnings :

FutureWarning: The handling of integer data will change in version 0.22. Currently, the categories are determined based on the range [0, max(values)], while in the future they will be determined based on the unique values. If you want the future behaviour and silence this warning, you can specify "categories='auto'". In case you used a LabelEncoder before this OneHotEncoder to convert the categories to integers, then you can now use the OneHotEncoder directly.

and the second :

The 'categorical_features' keyword is deprecated in version 0.20 and will be removed in 0.22. You can use the ColumnTransformer instead.
"use the ColumnTransformer instead.", DeprecationWarning)

In the future, you should not define the columns in the OneHotEncoder directly, unless you want to use "categories='auto'". The first message also tells you to use OneHotEncoder directly, without the LabelEncoder first. Finally, the second message tells you to use ColumnTransformer, which is like a Pipe for columns transformations.

Here is the equivalent code for your case :

from sklearn.compose import ColumnTransformer 
ct = ColumnTransformer([("Name_Of_Your_Step", OneHotEncoder(),[0])], remainder="passthrough")) # The last arg ([0]) is the list of columns you want to transform in this step
ct.fit_transform(X)    

See also : ColumnTransformer documentation

For the above example;

Encoding Categorical data (Basically Changing Text to Numerical data i.e, Country Name)

from sklearn.preprocessing import LabelEncoder, OneHotEncoder
from sklearn.compose import ColumnTransformer
#Encode Country Column
labelencoder_X = LabelEncoder()
X[:,0] = labelencoder_X.fit_transform(X[:,0])
ct = ColumnTransformer([("Country", OneHotEncoder(), [0])], remainder = 'passthrough')
X = ct.fit_transform(X)
Answer from CoMartel on Stack Overflow
Top answer
1 of 12
26

There is actually 2 warnings :

FutureWarning: The handling of integer data will change in version 0.22. Currently, the categories are determined based on the range [0, max(values)], while in the future they will be determined based on the unique values. If you want the future behaviour and silence this warning, you can specify "categories='auto'". In case you used a LabelEncoder before this OneHotEncoder to convert the categories to integers, then you can now use the OneHotEncoder directly.

and the second :

The 'categorical_features' keyword is deprecated in version 0.20 and will be removed in 0.22. You can use the ColumnTransformer instead.
"use the ColumnTransformer instead.", DeprecationWarning)

In the future, you should not define the columns in the OneHotEncoder directly, unless you want to use "categories='auto'". The first message also tells you to use OneHotEncoder directly, without the LabelEncoder first. Finally, the second message tells you to use ColumnTransformer, which is like a Pipe for columns transformations.

Here is the equivalent code for your case :

from sklearn.compose import ColumnTransformer 
ct = ColumnTransformer([("Name_Of_Your_Step", OneHotEncoder(),[0])], remainder="passthrough")) # The last arg ([0]) is the list of columns you want to transform in this step
ct.fit_transform(X)    

See also : ColumnTransformer documentation

For the above example;

Encoding Categorical data (Basically Changing Text to Numerical data i.e, Country Name)

from sklearn.preprocessing import LabelEncoder, OneHotEncoder
from sklearn.compose import ColumnTransformer
#Encode Country Column
labelencoder_X = LabelEncoder()
X[:,0] = labelencoder_X.fit_transform(X[:,0])
ct = ColumnTransformer([("Country", OneHotEncoder(), [0])], remainder = 'passthrough')
X = ct.fit_transform(X)
2 of 12
6

As of version 0.22, you can write the same code as below:

from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer
ct = ColumnTransformer([("Country", OneHotEncoder(), [0])], remainder = 'passthrough')
X = ct.fit_transform(X)

As you can see, you don't need to use LabelEncoder anymore.

🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are encoded using a one-hot (aka ‘one-of-K’ or ‘dummy’) encoding scheme.
Discussions

scikit learn - Issue with OneHotEncoder for categorical features - Stack Overflow
I want to encode 3 categorical features out of 10 features in my datasets. I use preprocessing from sklearn.preprocessing to do so as the following: from sklearn import preprocessing cat_features ... More on stackoverflow.com
🌐 stackoverflow.com
python - How to use OneHotEncoder categorical_features - Stack Overflow
I am having trouble encoding only categorical columns using OneHotEncoder and leaving out continuous columns. The encoder encodes all columns no matter what I specify in the categorical_features. ... More on stackoverflow.com
🌐 stackoverflow.com
machine learning - DeprecationWarning: The 'categorical_features' keyword is deprecated in version 0.20 - Data Science Stack Exchange
In case you used a LabelEncoder before this OneHotEncoder to convert the categories to integers, then you can now use the OneHotEncoder directly. More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
python - OneHotEncoder - encoding only some of categorical variable columns - Stack Overflow
The problem is that sklearn's OneHotEncoder needs to have an array of ints as input. But in the array data.values, you still have the string representation of gender. You could, if you wanted, just one hot encode the seniority values, but if you want to know the meaning of those features, it's ... More on stackoverflow.com
🌐 stackoverflow.com
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
import pandas as pd from sklearn.preprocessing import OneHotEncoder data = { 'Employee_ID': [10, 20, 15, 25, 30], 'Gender': ['M', 'F', 'F', 'M', 'F'], 'Remarks': ['Good', 'Nice', 'Good', 'Great', 'Nice'] } df = pd.DataFrame(data) print("Original Data:") print(df) categorical_columns = df.select_dtypes(include=['object']).columns encoder = OneHotEncoder(sparse_output=False) encoded_data = encoder.fit_transform(df[categorical_columns]) encoded_df = pd.DataFrame( encoded_data, columns=encoder.get_feature_names_out(categorical_columns) ) final_df = pd.concat( [df.drop(columns=categorical_columns), encoded_df], axis=1 ) print("\nOne-Hot Encoded Data:") print(final_df) Output: Output ·
Published: May 29, 2026
🌐
scikit-learn
scikit-learn.org › 0.19 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.19.2 documentation
Given a dataset with three features and four samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder() >>> enc.fit([[0, 0, 3], [1, 1, 0], [0, 2, 1], [1, 0, 2]]) OneHotEncoder(categorical_features='all', dtype=<... 'numpy.float64'>, handle_unknown='error', n_values='auto', sparse=True) >>> enc.n_values_ array([2, 3, 4]) >>> enc.feature_indices_ array([0, 2, 5, 9]) >>> enc.transform([[0, 1, 1]]).toarray() array([[ 1., 0., 0., 1., 0., 0., 1., 0., 0.]])
🌐
Medium
medium.com › 0xcode › one-hot-encoding-data-in-python-323b56ea2bfd
One Hot Encoding Data In Python. One Hot Encoding is a technique for… | by VTECH | 0xCODE | Medium
November 25, 2021 - Given a two feature dataset, the encoder can find the unique values per feature and transform the data to a binary one-hot encoding. I have a snippet here with an example of when to use One Hot Encoding.
Top answer
1 of 7
51

If you read the docs for OneHotEncoder you'll see the input for fit is "Input array of type int". So you need to do two steps for your one hot encoded data

from sklearn import preprocessing
cat_features = ['color', 'director_name', 'actor_2_name']
enc = preprocessing.LabelEncoder()
enc.fit(cat_features)
new_cat_features = enc.transform(cat_features)
print new_cat_features # [1 2 0]
new_cat_features = new_cat_features.reshape(-1, 1) # Needs to be the correct shape
ohe = preprocessing.OneHotEncoder(sparse=False) #Easier to read
print ohe.fit_transform(new_cat_features)

Output:

[[ 0.  1.  0.]
 [ 0.  0.  1.]
 [ 1.  0.  0.]]

EDIT

As of 0.20 this became a bit easier, not only because OneHotEncoder now handles strings nicely, but also because we can transform multiple columns easily using ColumnTransformer, see below for an example

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import LabelEncoder, OneHotEncoder
import numpy as np

X = np.array([['apple', 'red', 1, 'round', 0],
              ['orange', 'orange', 2, 'round', 0.1],
              ['bannana', 'yellow', 2, 'long', 0],
              ['apple', 'green', 1, 'round', 0.2]])
ct = ColumnTransformer(
    [('oh_enc', OneHotEncoder(sparse=False), [0, 1, 3]),],  # the column numbers I want to apply this to
    remainder='passthrough'  # This leaves the rest of my columns in place
)
print(ct2.fit_transform(X)) # Notice the output is a string

Output:

[['1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '0.0' '0.0' '1.0' '1' '0']
 ['0.0' '0.0' '1.0' '0.0' '1.0' '0.0' '0.0' '0.0' '1.0' '2' '0.1']
 ['0.0' '1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '1.0' '0.0' '2' '0']
 ['1.0' '0.0' '0.0' '1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '1' '0.2']]
2 of 7
14

You can apply both transformations (from text categories to integer categories, then from integer categories to one-hot vectors) in one shot using the LabelBinarizer class:

cat_features = ['color', 'director_name', 'actor_2_name']
encoder = LabelBinarizer()
new_cat_features = encoder.fit_transform(cat_features)
new_cat_features

Note that this returns a dense NumPy array by default. You can get a sparse matrix instead by passing sparse_output=True to the LabelBinarizer constructor.

Source Hands-On Machine Learning with Scikit-Learn and TensorFlow

🌐
Train in Data
blog.trainindata.com › one-hot-encoding-categorical-variables
One-hot encoding categorical variables | Train in Data Blog
January 25, 2023 - Pandas, Feature-engine and Category Encoders can automatically identify and encode categorical variables, that is, those of type object or categorical. Scikit-learn’s OneHotEncoder(), on the other hand, will encode all variables in the dataset.
Find elsewhere
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to perform one-hot encoding for multi categorical variables
How to Perform One-Hot Encoding For Multi Categorical Variables
February 3, 2025 - # importing pandas import pandas as pd # importing numpy import numpy as np # importing OneHotEncoder from sklearn.preprocessing import OneHotEncoder() Here, we use pandas which are used for data analysis, NumPyused for n-dimensional arrays, and from sklearn, we will use one important class One Hot Encoder for categorical encoding.
🌐
Medium
medium.com › data-science › encoding-categorical-features-21a2651a065c
Encoding Categorical Features. Introduction | by Yang Liu | TDS Archive | Medium
September 20, 2018 - Therefore, for dataframe containing multi class features, a further step of OneHotEncoder is needed. Let’s see the steps to do it. ... # import OneHotEncoder from sklearn.preprocessing import OneHotEncoder# instantiate OneHotEncoder ohe = OneHotEncoder(categorical_features = categorical_feature_mask, sparse=False ) # categorical_features = boolean mask for categorical columns # sparse = False output an array not sparse matrix
🌐
DataCamp
datacamp.com › tutorial › one-hot-encoding-python-tutorial
What Is One Hot Encoding and How to Implement It in Python | DataCamp
June 26, 2024 - We use OneHotEncoder to convert the categorical feature into a one-hot encoded format.
🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
OneHotEncoder converts categorical variables into a numerical format that machine learning models can understand. It converts categorical features into binary vectors where each unique value is represented as a separate feature.
Top answer
1 of 1
12

I think for this case you should stick to pd.get_dummies:

>>> data
   age seniority  gender  salary
0    1    junior    male       5
1    2    senior  female       6
2    3    junior  female       7

# One hot encode with get_dummies
data = pd.concat((data,pd.get_dummies(data.seniority)),1)

>>> data
   age seniority  gender  salary  junior  senior
0    1    junior    male       5       1       0
1    2    senior  female       6       0       1
2    3    junior  female       7       1       0

The problem is that sklearn's OneHotEncoder needs to have an array of ints as input. But in the array data.values, you still have the string representation of gender. You could, if you wanted, just one hot encode the seniority values, but if you want to know the meaning of those features, it's not very nice, you have to pass it the column names manually (which is unfeasible in a lot of cases):

from sklearn.preprocessing import LabelEncoder
label_encoder = LabelEncoder()
data['seniority'] = label_encoder.fit_transform(data['seniority'])

from sklearn.preprocessing import OneHotEncoder
one_hot_encoder = OneHotEncoder(sparse=False)
data[['junior','senior']] = one_hot_encoder.fit_transform(data['seniority'].values.reshape(-1,1))

>>> data
   age  seniority  gender  salary  junior  senior
0    1          0    male       5     1.0     0.0
1    2          1  female       6     0.0     1.0
2    3          0  female       7     1.0     0.0

Or, if the feature names don't matter:

from sklearn.preprocessing import LabelEncoder
label_encoder = LabelEncoder()
data['seniority'] = label_encoder.fit_transform(data['seniority'])

from sklearn.preprocessing import OneHotEncoder
one_hot_encoder = OneHotEncoder(sparse=False)
data = pd.concat((data,pd.DataFrame(one_hot_encoder.fit_transform(data['seniority'].values.reshape(-1,1)))),1)

   age  seniority  gender  salary    0    1
0    1          0    male       5  1.0  0.0
1    2          1  female       6  0.0  1.0
2    3          0  female       7  1.0  0.0

But in the end, pd.get_dummies does the job in a much nicer way (IMO)

🌐
Data School
dataschool.io › encoding-categorical-features-in-python
How to encode categorical features with scikit-learn (video)
June 4, 2024 - In this 28-minute video, you'll learn how to properly encode your categorical features using scikit-learn's OneHotEncoder, ColumnTransformer, and Pipeline.
🌐
scikit-learn
scikit-learn.org › 0.16 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
sklearn.preprocessing.OneHotEncoder — scikit-learn 0.16.1 documentation
This encoding is needed for feeding categorical data to many scikit-learn estimators, notably linear models and SVMs with the standard kernels. ... Given a dataset with three features and two samples, we let the encoder find the maximum value per feature and transform the data to a binary one-hot ...
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › onehot.html
One Hot — Category Encoders 2.11.1 documentation
class category_encoders.one_ho... bool | str | None = None)[source] · Onehot (or dummy) coding for categorical features, produces a binary feature per category....