There are some cases where LabelEncoder or DictVectorizor are useful, but these are quite limited in my opinion due to ordinality.

LabelEncoder can turn [dog,cat,dog,mouse,cat] into [1,2,1,3,2], but then the imposed ordinality means that the average of dog and mouse is cat. Still there are algorithms like decision trees and random forests that can work with categorical variables just fine and LabelEncoder can be used to store values using less disk space.

One-Hot-Encoding has the advantage that the result is binary rather than ordinal and that everything sits in an orthogonal vector space. The disadvantage is that for high cardinality, the feature space can really blow up quickly and you start fighting with the curse of dimensionality. In these cases, I typically employ one-hot-encoding followed by PCA for dimensionality reduction. I find that the judicious combination of one-hot plus PCA can seldom be beat by other encoding schemes. PCA finds the linear overlap, so will naturally tend to group similar features into the same feature.

Answer from AN6U5 on Stack Exchange
Top answer
1 of 4
213

There are some cases where LabelEncoder or DictVectorizor are useful, but these are quite limited in my opinion due to ordinality.

LabelEncoder can turn [dog,cat,dog,mouse,cat] into [1,2,1,3,2], but then the imposed ordinality means that the average of dog and mouse is cat. Still there are algorithms like decision trees and random forests that can work with categorical variables just fine and LabelEncoder can be used to store values using less disk space.

One-Hot-Encoding has the advantage that the result is binary rather than ordinal and that everything sits in an orthogonal vector space. The disadvantage is that for high cardinality, the feature space can really blow up quickly and you start fighting with the curse of dimensionality. In these cases, I typically employ one-hot-encoding followed by PCA for dimensionality reduction. I find that the judicious combination of one-hot plus PCA can seldom be beat by other encoding schemes. PCA finds the linear overlap, so will naturally tend to group similar features into the same feature.

2 of 4
58

While AN6U5 has given a very good answer, I wanted to add a few points for future reference. When considering One Hot Encoding(OHE) and Label Encoding, we must try and understand what model you are trying to build. Namely the two categories of model we will be considering are:

  1. Tree Based Models: Gradient Boosted Decision Trees and Random Forests.
  2. Non-Tree Based Models: Linear, kNN or Neural Network based.

Let's consider when to apply OHE and when to apply Label Encoding while building tree based models.

We apply OHE when:

  1. When the values that are close to each other in the label encoding correspond to target values that aren't close (non-linear data).
  2. When the categorical feature is not ordinal (dog, cat, mouse).

We apply Label encoding when:

  1. The categorical feature is ordinal (Jr. kg, Sr. kg, Primary school, high school, etc).
  2. When we can come up with a label encoder that assigns close labels to similar categories: This leads to less splits in the trees hence reducing the execution time.
  3. When the number of categorical features in the dataset is huge: One-hot encoding a categorical feature with huge number of values can lead to (1) high memory consumption and (2) the case when non-categorical features are rarely used by model. You can deal with the 1st case if you employ sparse matrices. The 2nd case can occur if you build a tree using only a subset of features. For example, if you have 9 numeric features and 1 categorical with 100 unique values and you one-hot-encoded that categorical feature, you will get 109 features. If a tree is built with only a subset of features, initial 9 numeric features will rarely be used. In this case, you can increase the parameter controlling size of this subset. In xgboost it is called colsample_bytree, in sklearn's Random Forest max_features.

In case you want to continue with OHE, as @AN6U5 suggested, you might want to combine PCA with OHE.

Let's consider when to apply OHE and Label Encoding while building non tree based models.

To apply Label encoding, the dependance between feature and target must be linear in order for Label Encoding to be utilised effectively.

Similarly, in case the dependance is non-linear, you might want to use OHE for the same.

Note: Some of the explanation has been referenced from How to Win a Data Science Competition from Coursera.

🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › one-hot-encoding-vs-label-encoding
One Hot Encoding vs Label Encoding - GeeksforGeeks
January 22, 2026 - import pandas as pd from sklearn.preprocessing import LabelEncoder severity = ['Low', 'Medium', 'High', 'Medium', 'Low'] df = pd.DataFrame({'Severity': severity}) label_encoder = LabelEncoder() df['Severity_encoded'] = label_encoder.fit_transform(df['Severity']) print(df)
🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - # creating one hot encoder object onehotencoder = OneHotEncoder() # reshape the 1-D country array to 2-D as fit_transform expects 2-D and fit the encoder X = onehotencoder.fit_transform(df.Country.values.reshape(-1, 1)).toarray()
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.9.1 documentation
class sklearn.preprocessing.LabelEncoder[source]# Encode target labels with value between 0 and n_classes-1. This transformer should be used to encode target values, i.e. y, and not the input X. Read more in the User Guide. Added in version 0.12. Attributes: classes_ndarray of shape (n_classes,) Holds the label for each class. See also · OrdinalEncoder · Encode categorical features using an ordinal encoding scheme. OneHotEncoder ·
🌐
DEV Community
dev.to › engrmark › when-to-use-labelencoder-and-onehotencoder-in-machine-learning-7gl
When to Use LabelEncoder and OneHotEncoder in Machine Learning - DEV Community
July 28, 2025 - Use LabelEncoder when your data has a natural order (e.g., Small < Medium < Large). Use OneHotEncoder when your data has no order (e.g., colors, cities).
Find elsewhere
🌐
Medium
medium.com › aimonks › label-encoding-vs-one-hot-encoding-making-sense-of-categorical-data-1181914501f3
Label Encoding vs. One-Hot Encoding: Making Sense of Categorical Data 👨‍💻 | by Pawan Yadav | 𝐀𝐈 𝐦𝐨𝐧𝐤𝐬.𝐢𝐨 | Medium
October 24, 2023 - Label Encoding vs. One-Hot Encoding: Making Sense of Categorical Data 👨‍💻 Categorical data is everywhere in the world around us. It can include things like colors, types of animals, or even …
🌐
SciTePress
scitepress.org › Papers › 2023 › 122594 › 122594.pdf pdf
Encoding Techniques for Handling Categorical Data in Machine
Encoding Techniques for Handling Categorical Data in Machine · Learning-Based Software Development Effort Estimation
🌐
E3S Conferences
e3s-conferences.org › articles › e3sconf › pdf › 2020 › 44 › e3sconf_icmed2020_01011.pdf pdf
A Comparative Study using Feature Selection to Predict the
LabelEncoder and OneHotEncoder of the SciKit python · library. 4. The methodology of our study next required us to · select the best machine learning algorithms that would · generate the highest accuracy in predicting which bank · customer will be staying and who will be exiting.
🌐
University of Toronto
eecg.utoronto.ca › ~jayar › download › a3.pdf pdf
ECE324 Fall 2020 Assignment 3
2. The categorical features in the dataset are represented as strings. Use the LabelEncoder class · of sklearn to turn them into integers and use the OneHotEncoder class to convert the integers · into one-hot vectors. For each categorical feature, call the .fit transform in LabelEncoder ·
🌐
SciTePress
scitepress.org › Papers › 2024 › 129689 › 129689.pdf pdf
The Comparison of Diabetes Risk Prediction Accuracy Across Different Models
The Comparison of Diabetes Risk Prediction Accuracy Across · 1School of Software Engineering, Shan Dong University, Jinan, 250101, China
🌐
Scitevents
kdir.scitevents.org › Abstract.aspx
KDIR 2023 Abstracts
Locally Organized and Hosted by: · INSTICC is Member of:
🌐
scikit-learn
scikit-learn.org › stable › api › index.html
API Reference — scikit-learn 1.9.1 documentation
This is the class and function reference of scikit-learn. Please refer to the full user guide for further details, as the raw specifications of classes and functions may not be enough to give full ...
🌐
Quora
quora.com › How-do-I-re-use-a-label-encoder-and-OneHotEncoder-How-do-I-save-them-on-my-machine-and-then-reuse-them-on-new-data-How-do-I-save-my-machine-learning-model-as-well-and-load-it-again-into-my-environment-using-Python
How to re-use a label encoder and OneHotEncoder? How do I save them on my machine and then reuse them on new data? How do I save my machine learning model as well and load it again into my environment - Quora
Answer: How do I save my machine learning model as well and load it again into my environment? # Import scikit’s joblib pickler… it’s better optimized for model objects than Python’s default pickler from sklearn.externals import joblib # Serialize the model to disk; you can then move/copy ...
🌐
Gteconlinelearning
gteconlinelearning.com › sites › default › files › 2025-06 › CERTIFICATE IN DATA SCIENCE.pdf pdf
CERTIFICATE IN DATA SCIENCE DURATION: 120 Hours TOTAL CREDITS: 4
CERTIFICATE IN DATA SCIENCE · The Certificate Course in Data Science aims to equip learners with the essential skills to analyze,