There are some cases where LabelEncoder or DictVectorizor are useful, but these are quite limited in my opinion due to ordinality.

LabelEncoder can turn [dog,cat,dog,mouse,cat] into [1,2,1,3,2], but then the imposed ordinality means that the average of dog and mouse is cat. Still there are algorithms like decision trees and random forests that can work with categorical variables just fine and LabelEncoder can be used to store values using less disk space.

One-Hot-Encoding has the advantage that the result is binary rather than ordinal and that everything sits in an orthogonal vector space. The disadvantage is that for high cardinality, the feature space can really blow up quickly and you start fighting with the curse of dimensionality. In these cases, I typically employ one-hot-encoding followed by PCA for dimensionality reduction. I find that the judicious combination of one-hot plus PCA can seldom be beat by other encoding schemes. PCA finds the linear overlap, so will naturally tend to group similar features into the same feature.

Answer from AN6U5 on Stack Exchange
🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - Here’s how you can implement one-hot encoding using Scikit-Learn in Python: from sklearn.preprocessing import OneHotEncoder import pandas as pd
🌐
Upgrad
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
July 23, 2025 - You can reverse the transformation in Label Encoding by using the .inverse_transform() method provided by scikit-learn’s LabelEncoder. For One Hot Encoding, you would need to map the binary vector back to its original category using the columns of the encoded matrix.
Discussions

scikit learn - When to use One Hot Encoding vs LabelEncoder vs DictVectorizor? - Data Science Stack Exchange
I have been building models with categorical data for a while now and when in this situation I basically default to using scikit-learn's LabelEncoder function to transform this data prior to buildi... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
December 19, 2015
python - Scikit-learn's LabelBinarizer vs. OneHotEncoder - Stack Overflow
A simple example which encodes an array using LabelEncoder, OneHotEncoder, LabelBinarizer is shown below. I see that OneHotEncoder needs data in integer encoded form first to convert into its respective encoding which is not required in the case of LabelBinarizer. Copyfrom numpy import array from sklearn... More on stackoverflow.com
🌐 stackoverflow.com
machine learning - Difference between One hot encoding and Label Encoding of target/output label - Stack Overflow
I have a problem where there are 20 classes. I have designed a neural network and using the loss as categorical_crossentropy. When dealing with categorical cross entropy the output label must be on... More on stackoverflow.com
🌐 stackoverflow.com
python - LabelEncoding() vs OneHotEncoding() (sklearn,pandas) suggestions - Stack Overflow
71 Scikit-learn's LabelBinarizer vs. OneHotEncoder · 5 LabelEncoder vs. Pandas categorical vs. enumerate? 14 Why shouldn't the sklearn LabelEncoder be used to encode input data? More on stackoverflow.com
🌐 stackoverflow.com
People also ask

How does Label Encoder handle unseen categories in test data?
One of the challenges with Label Encoding is that it cannot handle categories in the test data that were not seen during training. If new categories appear, the model may either assign them an arbitrary value or fail entirely. With Label Encoder vs One Hot Encoder, it’s crucial to ensure that all possible categories are known ahead of time or use models that handle unseen categories.
🌐
upgrad.com
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
Can Label Encoder be used for regression tasks?
Yes, Label Encoder can be used in regression tasks when dealing with ordinal data. If the categorical variables have a natural order, Label Encoder vs One Hot Encoder should be carefully considered. However, for nominal data, it’s better to use One Hot Encoding, as Label Encoding might mislead the model by implying relationships that don’t exist, impacting prediction accuracy.
🌐
upgrad.com
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
When should I use Label Encoder over One Hot Encoder?
You should use Label Encoder when your data has an inherent order, like "Low," "Medium," and "High," as in ordinal data. It assigns an integer to each category, preserving their order. Label Encoder vs One Hot Encoder becomes important when you need to avoid creating unnecessary binary columns, especially in cases where you don’t need to treat categories independently.
🌐
upgrad.com
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
🌐
Mervin Praison
mer.vin › 2022 › 10 › label-encoder-vs-one-hot-encoder
Label Encoder vs. One Hot Encoder - Mervin Praison
October 21, 2022 - from sklearn import preprocessing >>> le = preprocessing.LabelEncoder() >>> le.fit(["paris", "paris", "tokyo", "amsterdam"]) LabelEncoder() >>> list(le.classes_) ['amsterdam', 'paris', 'tokyo'] >>> le.transform(["tokyo", "tokyo", "paris"]) array([2, 2, 1]...) >>> list(le.inverse_transform([2, 2, 1])) ['tokyo', 'tokyo', 'paris'] ... from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) OneHotEncoder(handle_unknown='ignore') >>> enc.categories_ [array(['Female', 'Male'], dtype=object), a
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › one-hot-encoding-vs-label-encoding
One Hot Encoding vs Label Encoding - GeeksforGeeks
January 22, 2026 - Python · import pandas as pd from sklearn.preprocessing import LabelEncoder severity = ['Low', 'Medium', 'High', 'Medium', 'Low'] df = pd.DataFrame({'Severity': severity}) label_encoder = LabelEncoder() df['Severity_encoded'] = label_encoder.fit_transform(df['Severity']) print(df) Output: Label Encoding ·
🌐
DEV Community
dev.to › engrmark › when-to-use-labelencoder-and-onehotencoder-in-machine-learning-7gl
When to Use LabelEncoder and OneHotEncoder in Machine Learning - DEV Community
July 28, 2025 - OneHotEncoder says: I’ll give each category its own column so no one feels more important than the other. If you only have 2 categories, LabelEncoder is fine because it will just give 0 and 1. ... from sklearn.preprocessing import LabelEncoder binary = ["Yes", "No", "Yes", "No"] encoder = ...
Top answer
1 of 4
213

There are some cases where LabelEncoder or DictVectorizor are useful, but these are quite limited in my opinion due to ordinality.

LabelEncoder can turn [dog,cat,dog,mouse,cat] into [1,2,1,3,2], but then the imposed ordinality means that the average of dog and mouse is cat. Still there are algorithms like decision trees and random forests that can work with categorical variables just fine and LabelEncoder can be used to store values using less disk space.

One-Hot-Encoding has the advantage that the result is binary rather than ordinal and that everything sits in an orthogonal vector space. The disadvantage is that for high cardinality, the feature space can really blow up quickly and you start fighting with the curse of dimensionality. In these cases, I typically employ one-hot-encoding followed by PCA for dimensionality reduction. I find that the judicious combination of one-hot plus PCA can seldom be beat by other encoding schemes. PCA finds the linear overlap, so will naturally tend to group similar features into the same feature.

2 of 4
58

While AN6U5 has given a very good answer, I wanted to add a few points for future reference. When considering One Hot Encoding(OHE) and Label Encoding, we must try and understand what model you are trying to build. Namely the two categories of model we will be considering are:

  1. Tree Based Models: Gradient Boosted Decision Trees and Random Forests.
  2. Non-Tree Based Models: Linear, kNN or Neural Network based.

Let's consider when to apply OHE and when to apply Label Encoding while building tree based models.

We apply OHE when:

  1. When the values that are close to each other in the label encoding correspond to target values that aren't close (non-linear data).
  2. When the categorical feature is not ordinal (dog, cat, mouse).

We apply Label encoding when:

  1. The categorical feature is ordinal (Jr. kg, Sr. kg, Primary school, high school, etc).
  2. When we can come up with a label encoder that assigns close labels to similar categories: This leads to less splits in the trees hence reducing the execution time.
  3. When the number of categorical features in the dataset is huge: One-hot encoding a categorical feature with huge number of values can lead to (1) high memory consumption and (2) the case when non-categorical features are rarely used by model. You can deal with the 1st case if you employ sparse matrices. The 2nd case can occur if you build a tree using only a subset of features. For example, if you have 9 numeric features and 1 categorical with 100 unique values and you one-hot-encoded that categorical feature, you will get 109 features. If a tree is built with only a subset of features, initial 9 numeric features will rarely be used. In this case, you can increase the parameter controlling size of this subset. In xgboost it is called colsample_bytree, in sklearn's Random Forest max_features.

In case you want to continue with OHE, as @AN6U5 suggested, you might want to combine PCA with OHE.

Let's consider when to apply OHE and Label Encoding while building non tree based models.

To apply Label encoding, the dependance between feature and target must be linear in order for Label Encoding to be utilised effectively.

Similarly, in case the dependance is non-linear, you might want to use OHE for the same.

Note: Some of the explanation has been referenced from How to Win a Data Science Competition from Coursera.

Find elsewhere
🌐
MLK
machinelearningknowledge.ai › home › categorical data encoding with sklearn labelencoder and onehotencoder
Categorical Data Encoding with Sklearn LabelEncoder and OneHotEncoder - MLK - Machine Learning Knowledge
August 8, 2022 - The objective is to predict the Profit based on the other four independent variables of the dataset. Since one of our variables here, ‘State’ is a categorical variable, we will first be encoding it to the numeric variable by using Sklearn LabelEncoder and OneHotEncoder.
🌐
Contactsunny
blog.contactsunny.com › home › tech
Label Encoder vs. One Hot Encoder in Machine Learning | The ContactSunny Blog
November 6, 2019 - Update: SciKit has a new library called the ColumnTransformer which has replaced LabelEncoding. You can check out this updated post about ColumnTransformer to know more. If you’re new to Machine Learning, you might get confused between these two – Label Encoder and One Hot Encoder. These two encoders are parts of the SciKit Learn library in Python, and they are used to convert categorical data, or text data, into numbers, which our predictive models can better understand.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.9.1 documentation
class sklearn.preprocessing.LabelEncoder[source]# Encode target labels with value between 0 and n_classes-1. This transformer should be used to encode target values, i.e. y, and not the input X. Read more in the User Guide. Added in version 0.12. Attributes: classes_ndarray of shape (n_classes,) Holds the label for each class. See also · OrdinalEncoder · Encode categorical features using an ordinal encoding scheme. OneHotEncoder ·
Top answer
1 of 5
46

A simple example which encodes an array using LabelEncoder, OneHotEncoder, LabelBinarizer is shown below.

I see that OneHotEncoder needs data in integer encoded form first to convert into its respective encoding which is not required in the case of LabelBinarizer.

from numpy import array
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder
from sklearn.preprocessing import LabelBinarizer

# define example
data = ['cold', 'cold', 'warm', 'cold', 'hot', 'hot', 'warm', 'cold', 
'warm', 'hot']
values = array(data)
print "Data: ", values
# integer encode
label_encoder = LabelEncoder()
integer_encoded = label_encoder.fit_transform(values)
print "Label Encoder:" ,integer_encoded

# onehot encode
onehot_encoder = OneHotEncoder(sparse=False)
integer_encoded = integer_encoded.reshape(len(integer_encoded), 1)
onehot_encoded = onehot_encoder.fit_transform(integer_encoded)
print "OneHot Encoder:", onehot_encoded

#Binary encode
lb = LabelBinarizer()
print "Label Binarizer:", lb.fit_transform(values)

Another good link which explains the OneHotEncoder is: Explain onehotencoder using python

There may be other valid differences between the two which experts can probably explain.

2 of 5
31

A difference is that you can use OneHotEncoder for multi column data, while not for LabelBinarizer and LabelEncoder.

from sklearn.preprocessing import LabelBinarizer, LabelEncoder, OneHotEncoder

X = [["US", "M"], ["UK", "M"], ["FR", "F"]]
OneHotEncoder().fit_transform(X).toarray()

# array([[0., 0., 1., 0., 1.],
#        [0., 1., 0., 0., 1.],
#        [1., 0., 0., 1., 0.]])
LabelBinarizer().fit_transform(X)
# ValueError: Multioutput target data is not supported with label binarization

LabelEncoder().fit_transform(X)
# ValueError: bad input shape (3, 2)
🌐
The Neural Base
theneuralbase.com › home › scikit learn › beginner course › labelencoder vs onehotencoder
LabelEncoder vs OneHotEncoder | Scikit Learn Beginner Course | The Neural Base
LabelEncoder is like assigning seat numbers at a concert (1, 2, 3...): the numbers are just IDs with no meaning about distance. OneHotEncoder is like saying 'person A is in section Red, person B is in section Blue': each section is independent, no false proximity.
Top answer
1 of 1
1

Label encoding imposes artificial order: if you label-encode your pet target as 'Dog':0, 'Cat':1, 'Turtle':2, 'Golden Fish':3, then you get the awkward situation where 'Dog' < 'Cat' and 'Turtle is the average of 'Cat' + 'Golden Fish'.

In the case of predictor features (not the target), this is a problem since your Random Forest can be learning something like "if it less than 'Turtle', then...".

Also, you may have categories in the testing set (or even worse, new data during deployment) that were not present in the training, and the transformer doesn't know what to do, so it throws an error. This may be the case or not depending on the particular problem and particular feature you are encoding, obviously not for the target variable.

When hot encoding, if a category absent in the training is present in a prediction, it just get encoded as 0 in each of the encoded features (new columns representing each category), so you don't get an error. Your model still has the other features to make a reasonable guess.

As a general rule, you want to use label encoding for target variables and OHE for predictor features. Note that in general you don't care about artificial order in the target, since the prediction is usually categorical also (A forest will choose a number, not a range of numbers; a network will have one activation unit per category...)

I don't think optimization should be part of the discussion here since they are used for different scenarios demanding different outputs: surely it's more efficient to use the OHE transformer than trying to hack it by performing label encoding and then some pandas trickery to create the same result as with one hot encoding.

Here there are useful comments about the different scenarios (type of model, type of data) and some issues related to efficiency.

Here there's an example on why label encoding is a bad practice for input features.

And let's not forget that the goal of the model is to make predictions, so at the end what's important is not just the output of <transformer>.fit_transform, but also the fitted transformer itself that's going to be applied to the new observations. OHE will deal with new cases differently than label-encoder (e.g. when the value of the feature in the observation was not present in the training set). That's in my opinion enough reason to have different methods, even when they act in a way similar enough so, for some inputs, you may be able to force them to give similar outputs.

🌐
Medium
drlee.io › a-comprehensive-guide-to-categorical-data-encoding-exploring-labelencoder-and-onehotencoder-and-4ecd0f68ac66
A Comprehensive Guide to Categorical Data Encoding: Exploring LabelEncoder and OneHotEncoder and get_dummies with Python | by Dr. Ernesto Lee | Medium
May 25, 2023 - This article elucidates the practice of categorical data encoding using two effective methods available in Sklearn — LabelEncoder and OneHotEncoder. Our discussion will set out by establishing the notion of categorical data, its significance in machine learning, and why it requires a ...
🌐
Medium
abhibvp003.medium.com › label-encoder-vs-one-hot-encoder-in-machine-learning-7b21ed4e08d1
Chapter:1-Label Encoder vs One Hot Encoder in Machine Learning | by ABHISHEK KUMAR | Medium
July 12, 2019 - And to convert this kind of categorical text data into model-understandable numerical data, we use the Label Encoder class. So all we have to do, to label encode the first column, is import the LabelEncoder class from the sklearn library, fit and transform the first column of the data, and then replace the existing text data with the new encoded data.
🌐
Medium
medium.com › p › d7ecf7a46d17
LabelEncoder and OneHotEncoder in sklearn | by Amit Yadav | Medium
February 9, 2023 - LabelEncoder is a preprocessing function in the scikit-learn library in Python that is used to convert categorical labels to numerical values, so that they can be used by machine learning algorithms. For example, suppose you have a categorical feature color with three categories: "Red", "Green", and "Blue". To convert this feature into numerical data, you can use LabelEncoder as follows: from sklearn.preprocessing import LabelEncoder # Create an instance of LabelEncoder le = LabelEncoder() # Fit and transform the data color_labels = le.fit_transform(["Red", "Green", "Blue"]) # Output the transformed data print(color_labels)
🌐
Medium
harshal-soni.medium.com › onehotencoding-vs-labelencoder-vs-pandas-get-dummies-how-and-why-b190dff7a86f
OneHotEncoding vs LabelEncoder vs pandas getdummies — How and Why? | by Harshal Soni | Medium
March 29, 2022 - Even if you have a multi-label multi-class problem, you can use MultiLabelBinarizer for your y labels rather than switching to OneHotEncoder for multi hot encoding · For machine learning, you almost definitely want to use sklearn.OneHotEncoder.
🌐
Medium
contactsunny.medium.com › label-encoder-vs-one-hot-encoder-in-machine-learning-3fc273365621
Label Encoder vs. One Hot Encoder in Machine Learning | by Sunny Srinidhi | Medium
January 9, 2020 - Update: SciKit has a new library called the ColumnTransformer which has replaced LabelEncoding. You can check out this updated post about ColumnTransformer to know more. If you’re new to Machine Learning, you might get confused between these two — Label Encoder and One Hot Encoder. These two encoders are parts of the SciKit Learn library in Python, and they are used to convert categorical data, or text data, into numbers, which our predictive models can better understand.
🌐
GitHub
gist.github.com › CMCDragonkai › 6ed11a9b0c1d77d09f8f227489843eaa
LabelEncoder and OneHotEncoder #python #sklearn · GitHub
LabelEncoder and OneHotEncoder #python #sklearn. GitHub Gist: instantly share code, notes, and snippets.