There are some cases where LabelEncoder or DictVectorizor are useful, but these are quite limited in my opinion due to ordinality.

LabelEncoder can turn [dog,cat,dog,mouse,cat] into [1,2,1,3,2], but then the imposed ordinality means that the average of dog and mouse is cat. Still there are algorithms like decision trees and random forests that can work with categorical variables just fine and LabelEncoder can be used to store values using less disk space.

One-Hot-Encoding has the advantage that the result is binary rather than ordinal and that everything sits in an orthogonal vector space. The disadvantage is that for high cardinality, the feature space can really blow up quickly and you start fighting with the curse of dimensionality. In these cases, I typically employ one-hot-encoding followed by PCA for dimensionality reduction. I find that the judicious combination of one-hot plus PCA can seldom be beat by other encoding schemes. PCA finds the linear overlap, so will naturally tend to group similar features into the same feature.

Answer from AN6U5 on Stack Exchange
Top answer
1 of 4
213

There are some cases where LabelEncoder or DictVectorizor are useful, but these are quite limited in my opinion due to ordinality.

LabelEncoder can turn [dog,cat,dog,mouse,cat] into [1,2,1,3,2], but then the imposed ordinality means that the average of dog and mouse is cat. Still there are algorithms like decision trees and random forests that can work with categorical variables just fine and LabelEncoder can be used to store values using less disk space.

One-Hot-Encoding has the advantage that the result is binary rather than ordinal and that everything sits in an orthogonal vector space. The disadvantage is that for high cardinality, the feature space can really blow up quickly and you start fighting with the curse of dimensionality. In these cases, I typically employ one-hot-encoding followed by PCA for dimensionality reduction. I find that the judicious combination of one-hot plus PCA can seldom be beat by other encoding schemes. PCA finds the linear overlap, so will naturally tend to group similar features into the same feature.

2 of 4
58

While AN6U5 has given a very good answer, I wanted to add a few points for future reference. When considering One Hot Encoding(OHE) and Label Encoding, we must try and understand what model you are trying to build. Namely the two categories of model we will be considering are:

  1. Tree Based Models: Gradient Boosted Decision Trees and Random Forests.
  2. Non-Tree Based Models: Linear, kNN or Neural Network based.

Let's consider when to apply OHE and when to apply Label Encoding while building tree based models.

We apply OHE when:

  1. When the values that are close to each other in the label encoding correspond to target values that aren't close (non-linear data).
  2. When the categorical feature is not ordinal (dog, cat, mouse).

We apply Label encoding when:

  1. The categorical feature is ordinal (Jr. kg, Sr. kg, Primary school, high school, etc).
  2. When we can come up with a label encoder that assigns close labels to similar categories: This leads to less splits in the trees hence reducing the execution time.
  3. When the number of categorical features in the dataset is huge: One-hot encoding a categorical feature with huge number of values can lead to (1) high memory consumption and (2) the case when non-categorical features are rarely used by model. You can deal with the 1st case if you employ sparse matrices. The 2nd case can occur if you build a tree using only a subset of features. For example, if you have 9 numeric features and 1 categorical with 100 unique values and you one-hot-encoded that categorical feature, you will get 109 features. If a tree is built with only a subset of features, initial 9 numeric features will rarely be used. In this case, you can increase the parameter controlling size of this subset. In xgboost it is called colsample_bytree, in sklearn's Random Forest max_features.

In case you want to continue with OHE, as @AN6U5 suggested, you might want to combine PCA with OHE.

Let's consider when to apply OHE and Label Encoding while building non tree based models.

To apply Label encoding, the dependance between feature and target must be linear in order for Label Encoding to be utilised effectively.

Similarly, in case the dependance is non-linear, you might want to use OHE for the same.

Note: Some of the explanation has been referenced from How to Win a Data Science Competition from Coursera.

🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - H20 infact says that they use enum encoding where the categories are given a numerical value , but the numbers themselves are irrelevant(hence not imposing ordinality on nominal variables). But their classification performance doesn't seem to be much different from sklearn's random forest classifier using ordinal encoder)
People also ask

When should I use Label Encoder over One Hot Encoder?
You should use Label Encoder when your data has an inherent order, like "Low," "Medium," and "High," as in ordinal data. It assigns an integer to each category, preserving their order. Label Encoder vs One Hot Encoder becomes important when you need to avoid creating unnecessary binary columns, especially in cases where you don’t need to treat categories independently.
🌐
upgrad.com
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
How can I reverse the transformation done by Label Encoder or One Hot Encoder?
You can reverse the transformation in Label Encoding by using the .inverse_transform() method provided by scikit-learn’s LabelEncoder. For One Hot Encoding, you would need to map the binary vector back to its original category using the columns of the encoded matrix. Understanding Label Encoder vs One Hot Encoder and knowing when to reverse the transformation is crucial for interpreting model outputs.
🌐
upgrad.com
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
How does Label Encoder handle unseen categories in test data?
One of the challenges with Label Encoding is that it cannot handle categories in the test data that were not seen during training. If new categories appear, the model may either assign them an arbitrary value or fail entirely. With Label Encoder vs One Hot Encoder, it’s crucial to ensure that all possible categories are known ahead of time or use models that handle unseen categories.
🌐
upgrad.com
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
🌐
Upgrad
upgrad.com › home › blog › artificial intelligence › label encoder vs one hot encoder in machine learning
Label Encoder vs One Hot Encoder: Is Your Model Ready?
July 23, 2025 - You can reverse the transformation in Label Encoding by using the .inverse_transform() method provided by scikit-learn’s LabelEncoder. For One Hot Encoding, you would need to map the binary vector back to its original category using the columns of the encoded matrix.
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › one-hot-encoding-vs-label-encoding
One Hot Encoding vs Label Encoding - GeeksforGeeks
January 22, 2026 - import pandas as pd from sklearn.preprocessing import LabelEncoder severity = ['Low', 'Medium', 'High', 'Medium', 'Low'] df = pd.DataFrame({'Severity': severity}) label_encoder = LabelEncoder() df['Severity_encoded'] = label_encoder.fit_transform(df['Severity']) print(df)
🌐
DEV Community
dev.to › engrmark › when-to-use-labelencoder-and-onehotencoder-in-machine-learning-7gl
When to Use LabelEncoder and OneHotEncoder in Machine Learning - DEV Community
July 28, 2025 - Use OneHotEncoder when your data has no order (e.g., colors, cities). For 2 categories, LabelEncoder automatically uses 0 and 1.
🌐
Medium
drlee.io › a-comprehensive-guide-to-categorical-data-encoding-exploring-labelencoder-and-onehotencoder-and-4ecd0f68ac66
A Comprehensive Guide to Categorical Data Encoding: Exploring LabelEncoder and OneHotEncoder and get_dummies with Python | by Dr. Ernesto Lee | Medium
May 25, 2023 - For example, LabelEncoder is best suited for ordinal data, while OneHotEncoder works best for nominal data. Additionally, the pandas get_dummies function can serve as a quick and effective alternative for OneHotEncoding.
Find elsewhere
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.9.1 documentation
class sklearn.preprocessing.LabelEncoder[source]# Encode target labels with value between 0 and n_classes-1. This transformer should be used to encode target values, i.e. y, and not the input X. Read more in the User Guide. Added in version 0.12. Attributes: classes_ndarray of shape (n_classes,) Holds the label for each class. See also · OrdinalEncoder · Encode categorical features using an ordinal encoding scheme. OneHotEncoder ·
Top answer
1 of 1
1

Label encoding imposes artificial order: if you label-encode your pet target as 'Dog':0, 'Cat':1, 'Turtle':2, 'Golden Fish':3, then you get the awkward situation where 'Dog' < 'Cat' and 'Turtle is the average of 'Cat' + 'Golden Fish'.

In the case of predictor features (not the target), this is a problem since your Random Forest can be learning something like "if it less than 'Turtle', then...".

Also, you may have categories in the testing set (or even worse, new data during deployment) that were not present in the training, and the transformer doesn't know what to do, so it throws an error. This may be the case or not depending on the particular problem and particular feature you are encoding, obviously not for the target variable.

When hot encoding, if a category absent in the training is present in a prediction, it just get encoded as 0 in each of the encoded features (new columns representing each category), so you don't get an error. Your model still has the other features to make a reasonable guess.

As a general rule, you want to use label encoding for target variables and OHE for predictor features. Note that in general you don't care about artificial order in the target, since the prediction is usually categorical also (A forest will choose a number, not a range of numbers; a network will have one activation unit per category...)

I don't think optimization should be part of the discussion here since they are used for different scenarios demanding different outputs: surely it's more efficient to use the OHE transformer than trying to hack it by performing label encoding and then some pandas trickery to create the same result as with one hot encoding.

Here there are useful comments about the different scenarios (type of model, type of data) and some issues related to efficiency.

Here there's an example on why label encoding is a bad practice for input features.

And let's not forget that the goal of the model is to make predictions, so at the end what's important is not just the output of <transformer>.fit_transform, but also the fitted transformer itself that's going to be applied to the new observations. OHE will deal with new cases differently than label-encoder (e.g. when the value of the feature in the observation was not present in the training set). That's in my opinion enough reason to have different methods, even when they act in a way similar enough so, for some inputs, you may be able to force them to give similar outputs.

🌐
YouTube
youtube.com › watch
Machine learning feature engineering: Label encoding Vs One-Hot encoding (using Scikit-learn) - YouTube
In this tutorial, you will learn how to apply Label encoding & One-hot encoding using Scikit-learn and pandas. Encoding is a method to convert categorical va...
Published: July 12, 2020
🌐
MLK
machinelearningknowledge.ai › home › categorical data encoding with sklearn labelencoder and onehotencoder
Categorical Data Encoding with Sklearn LabelEncoder and OneHotEncoder - MLK - Machine Learning Knowledge
August 8, 2022 - The objective is to predict the Profit based on the other four independent variables of the dataset. Since one of our variables here, ‘State’ is a categorical variable, we will first be encoding it to the numeric variable by using Sklearn LabelEncoder and OneHotEncoder.
🌐
The Neural Base
theneuralbase.com › home › scikit learn › beginner course › labelencoder vs onehotencoder
LabelEncoder vs OneHotEncoder | Scikit Learn Beginner Course | The Neural Base
LabelEncoder is like assigning seat numbers at a concert (1, 2, 3...): the numbers are just IDs with no meaning about distance. OneHotEncoder is like saying 'person A is in section Red, person B is in section Blue': each section is independent, no false proximity. ... from sklearn.preprocessing ...
🌐
Mervin Praison
mer.vin › 2022 › 10 › label-encoder-vs-one-hot-encoder
Label Encoder vs. One Hot Encoder - Mervin Praison
October 21, 2022 - from sklearn import preprocessing >>> le = preprocessing.LabelEncoder() >>> le.fit(["paris", "paris", "tokyo", "amsterdam"]) LabelEncoder() >>> list(le.classes_) ['amsterdam', 'paris', 'tokyo'] >>> le.transform(["tokyo", "tokyo", "paris"]) array([2, 2, 1]...) >>> list(le.inverse_transform([2, 2, 1])) ['tokyo', 'tokyo', 'paris'] ... from sklearn.preprocessing import OneHotEncoder >>> enc = OneHotEncoder(handle_unknown='ignore') >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) OneHotEncoder(handle_unknown='ignore') >>> enc.categories_ [array(['Female', 'Male'], dtype=object), a
🌐
Kaggle
kaggle.com › questions-and-answers › 239564
OneHotEncoder VS OrdinalEncoder VS LabelEncoder in scikit-learn | Kaggle
The third one one, LabelEncoder, is used when you want to transform your dependent variables into classes, e.g., : [1, 1, 2, 6] -> [0, 0, 1, 2]. This is only intended to be used with your LABELS, i.e., your dependent variables, and not your ...
Top answer
1 of 5
46

A simple example which encodes an array using LabelEncoder, OneHotEncoder, LabelBinarizer is shown below.

I see that OneHotEncoder needs data in integer encoded form first to convert into its respective encoding which is not required in the case of LabelBinarizer.

from numpy import array
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder
from sklearn.preprocessing import LabelBinarizer

# define example
data = ['cold', 'cold', 'warm', 'cold', 'hot', 'hot', 'warm', 'cold', 
'warm', 'hot']
values = array(data)
print "Data: ", values
# integer encode
label_encoder = LabelEncoder()
integer_encoded = label_encoder.fit_transform(values)
print "Label Encoder:" ,integer_encoded

# onehot encode
onehot_encoder = OneHotEncoder(sparse=False)
integer_encoded = integer_encoded.reshape(len(integer_encoded), 1)
onehot_encoded = onehot_encoder.fit_transform(integer_encoded)
print "OneHot Encoder:", onehot_encoded

#Binary encode
lb = LabelBinarizer()
print "Label Binarizer:", lb.fit_transform(values)

Another good link which explains the OneHotEncoder is: Explain onehotencoder using python

There may be other valid differences between the two which experts can probably explain.

2 of 5
31

A difference is that you can use OneHotEncoder for multi column data, while not for LabelBinarizer and LabelEncoder.

from sklearn.preprocessing import LabelBinarizer, LabelEncoder, OneHotEncoder

X = [["US", "M"], ["UK", "M"], ["FR", "F"]]
OneHotEncoder().fit_transform(X).toarray()

# array([[0., 0., 1., 0., 1.],
#        [0., 1., 0., 0., 1.],
#        [1., 0., 0., 1., 0.]])
LabelBinarizer().fit_transform(X)
# ValueError: Multioutput target data is not supported with label binarization

LabelEncoder().fit_transform(X)
# ValueError: bad input shape (3, 2)
🌐
Medium
abhibvp003.medium.com › label-encoder-vs-one-hot-encoder-in-machine-learning-7b21ed4e08d1
Chapter:1-Label Encoder vs One Hot Encoder in Machine Learning | by ABHISHEK KUMAR | Medium
July 12, 2019 - Now, as we already discussed, depending on the data we have, we might run into situations where, after label encoding, we might confuse our model into thinking that a column has data with some kind of order or hierarchy when we clearly don’t have it. To avoid this, we ‘OneHotEncode’ that column.
🌐
Contactsunny
blog.contactsunny.com › home › tech
Label Encoder vs. One Hot Encoder in Machine Learning | The ContactSunny Blog
November 6, 2019 - And to convert this kind of categorical text data into model-understandable numerical data, we use the Label Encoder class. So all we have to do, to label encode the first column, is import the LabelEncoder class from the sklearn library, fit and transform the first column of the data, and then replace the existing text data with the new encoded data.
🌐
Medium
harshalisbatman.medium.com › onehotencoding-vs-labelencoder-vs-pandas-get-dummies-how-and-why-b190dff7a86f
OneHotEncoding vs LabelEncoder vs pandas getdummies — How and Why? | by Harshal Soni | Medium
September 7, 2022 - For machine learning, you almost definitely want to use sklearn.OneHotEncoder. For other tasks like simple analyses, you might be able to use pd.get_dummies, which is a bit more convenient. The crux of it is that the sklearn encoder creates a function which persists and can then be applied to new data sets which use the same categorical variables, with consistent results. A quick summary: LabelEncoder — for labels(response variable) coding 1,2,3…
🌐
Medium
medium.com › @amitya_dav › labelencoder-and-onehotencoder-in-sklearn-d7ecf7a46d17
LabelEncoder and OneHotEncoder in sklearn | by Amit Yadav | Medium
February 9, 2023 - LabelEncoder is a preprocessing function in the scikit-learn library in Python that is used to convert categorical labels to numerical values, so that they can be used by machine learning algorithms.