The main difference between one-hot-encoding and integer-encoding is based on the ordinal relationship between the classes. It means when you apply one-hot-encoding on classes, the distance between class 4 and 5 is equal to the distance of class 4 and 10.

So, both of your approach is totally correct and making the decision of which approach should be taken, is with the nature of your data, but you should note that for classification purpose between multi classes, in many cases, one-hot-encoding can give you the better result(because many ML algorithms are better to work with that such as SVM)

Answer from Mehdi Khademloo on Stack Overflow
🌐
Medium
medium.com › geekculture › machine-learning-one-hot-encoding-vs-integer-encoding-f180eb831cf1
Machine learning: one-hot encoding vs integer encoding | by Stéphanie Crêteur | Geek Culture | Medium
December 16, 2022 - First, one-hot encoding is often considered to be more expressive than integer encoding, because it can more accurately represent the data and its relationships. Indeed, it can better represent the presence or absence of a category.
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › why one-hot encode data in machine learning?
Why One-Hot Encode Data in Machine Learning? - MachineLearningMastery.com
June 30, 2020 - For categorical variables where no such ordinal relationship exists, the integer encoding is not enough. In fact, using this encoding and allowing the model to assume a natural ordering between categories may result in poor performance or unexpected results (predictions halfway between categories). In this case, a one-hot encoding can be applied to the integer representation.
Discussions

Why use one hot encoding instead of integer encoding? - Part 1 (2019) - fast.ai Course Forums
I’m all for the categorization of data such as creating a integer based dictionary of words in any text, where each integer represents a word, but why would you take this integer representation and blow it up into one hot encoding? To me, this just adds n amounts of pointless zeroes. More on forums.fast.ai
🌐 forums.fast.ai
0
April 27, 2019
python - Difference between one-hot-encoded and integer output in Sklearn - Stack Overflow
Consider the case of a multiclass classification problem with 12 classes (classes 0 to 11). These classes are nominal categorical variables (no ranking order). I have trained two models (M1 and M2)... More on stackoverflow.com
🌐 stackoverflow.com
comment:re sklearn -- integer encoding vs 1-hot (py)
(Your post popped up in my twitter feed) I'm not sure why you said you needed to one-hot encode categorical variables for scikit's random forest; I'm fairly certain you do not need to(a... More on github.com
🌐 github.com
12
April 27, 2015
Reasons not to one-hot-encode categorical features - Cross Validated
I could see not one-hot-encoding cat_num maybe being fine in two other cases · 1) Where there is some natural ordering to the categories, e.g. low, medium, and high, and that ordering is reflected in the integer encoding applied, e.g. More on stats.stackexchange.com
🌐 stats.stackexchange.com
🌐
Educative
educative.io › blog › one-hot-encoding
Data Science in 5 Minutes: What is One Hot Encoding?
Some machine learning algorithms ... data must be mapped to integers. One hot encoding is one method of converting data to prepare it for an algorithm and get a better prediction....
🌐
Fast.ai
forums.fast.ai › part 1 (2019)
Why use one hot encoding instead of integer encoding? - Part 1 (2019) - fast.ai Course Forums
April 27, 2019 - I’m all for the categorization of data such as creating a integer based dictionary of words in any text, where each integer represents a word, but why would you take this integer representation and blow it up into one hot encoding? To me, this just adds n amounts of pointless zeroes.
🌐
GitHub
github.com › szilard › benchm-ml › issues › 1
comment:re sklearn -- integer encoding vs 1-hot (py) · Issue #1 · szilard/benchm-ml
April 27, 2015 - It's been awhile since I looked at the source, but I'm pretty sure it handles categorical variables encoded as a single vector of numbers just fine from empirical tests; performance is almost always worse if the features were one-hot encoded.
Author: szilard
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Performs an ordinal (integer) encoding of the categorical features. ... Encodes categorical features using the target. ... Performs a one-hot encoding of dictionary items (also handles string-valued features).
Find elsewhere
🌐
Victorzhou
victorzhou.com › blog › one-hot
One-Hot Encoding, Explained - victorzhou.com
This is known as integer encoding. For Machine Learning, this encoding can be problematic - in this example, we’re essentially saying “green” is the average of “red” and “blue”, which can lead to weird unexpected outcomes. It’s often more useful to use the one-hot encoding instead:
🌐
arXiv
arxiv.org › html › 2312.16930v1
Encoding categorical data: Is there yet anything ‘hotter’ than one-hot encoding?
December 28, 2023 - This might be particularly damaging for linear or deep learning models, especially with high cardinality features. In contrast, the major disadvantage of indicator encoders like one-hot or Helmert contrast coding is that they increase the number of columns in a dataset.
🌐
Kaggle
kaggle.com › questions-and-answers › 396146
One-hot encoded vs. integer encoded, which one do you prefer and why? | Kaggle
In the case of multi-class classification with neural networks, which one of the mentioned encodings do you prefer and why?
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › ordinal and one-hot encodings for categorical data
Ordinal and One-Hot Encodings for Categorical Data - MachineLearningMastery.com
August 17, 2020 - The integer values have a natural ordered relationship between each other and machine learning algorithms may be able to understand and harness this relationship. It is a natural encoding for ordinal variables. For categorical variables, it imposes an ordinal relationship where no such relationship may exist. This can cause problems and a one-hot encoding may be used instead.
🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
In the output, you can see that the original Color column is dropped. Also, the one-hot encoded columns contain boolean values. To get one-hot encoded columns with integers, you can set the dtype parameter to int in the get_dummies() function.
Top answer
1 of 2
4

A note on terminology: As far as I am aware (unfortunately, there are a lot of blogs written by people who overlook the subtle differences and thus mis-information spreads):

One hot encoding is exactly what you described, generating a map from each unique value in a string column to an integer

Dummying is making K new columns (in which K is the number of unique values), of which exactly one column per row must be one.

In the "dog, cat, horse" example, when using a decision tree, consider the following example. Perhaps your target variable is "has it ever meowed?". Clearly what you want your decision tree to do is be able to ask the question "is it a cat? (yes/no)".

If you one-hot encode, such that dog -> 0, cat-> 1, horse->2, the tree can't isolate all of the cats using one question, because decision trees always split using "is feature x greater than or less than X?"

If you're using logistic regression, it also can't assign higher probabilities of meowing to cats.

If you dummy, the tree can explicitly ask the question "the column which signifies cat greater than 0.5?", thus splitting your data into cats and not cats.

If you use logistic regression, your optimiser can learn that the coefficient related to this column should be positive.

Thus in my opinion, whenever you have categorical data which has no implicit ordinality, always dummy, never one-hot encode.

In the case where your data has high cardinality, this could cause problems, especially if the number of examples of each type is tiny, but this is a problem you can't really solve, you simply have too detailed information for the size of your training data and using it would lead to over-fitting.

Nonetheless, one way to mitigate this, is to do some manual clustering (or actual clustering), in which you make a synthetic column, which can take fewer values, and many of the unique values of the original column map to the same value in the new column (e.g. dog, cat, horse-> mammal, pigeon, parrot , chicken -> bird). This makes it easier for the algorithm to learn, and if there's enough data, it can split further within each cluster.

2 of 2
1

I have never been happy with one hot encoding.

See: https://roamanalytics.com/2016/10/28/are-categorical-variables-getting-lost-in-your-random-forests/

Recently, I have tried CatBoost (http://CatBoost.ai), Open Source from Yandex.

CatBoost uses Categorical variables directly. (XGBoost uses one-hot-encoding under the covers.)

It seems to score better and faster than H2O's XGBoost, but training seems to take longer.

You might want to give it a try.

🌐
ScienceDirect
sciencedirect.com › topics › computer-science › one-hot-encoding
One-Hot Encoding - an overview | ScienceDirect Topics
(2019) used a linear interpolation ... to normalize the values of load demand data between 0–1; (IV) one hot encoding is the process of converting categorical features of a dataset to be integer values....
🌐
Data Science Dojo
datasciencedojo.com › home › blog › machine learning › 7 essential encoding techniques for categorical data in machine learning
Categorical Data Encoding: 7 Effective Techniques
January 21, 2026 - Preserves Order: It captures and ... of analyses. Reduces Dimensionality: It reduces the dimensionality of the dataset compared to one-hot encoding, making it more memory-efficient....
🌐
Deepchecks
deepchecks.com › glossary › one-hot encoding
What is One-hot Encoding | Deepchecks
August 5, 2021 - This kind of encoding might be sufficient for some variables. There is an order between integer numbers, which ML algorithms might be capable to grasp and exploit. It creates an ordinal relationship between categorical variables where none previously existed. This can trigger problems, so instead use a one-shot encoding process.
🌐
Statology
statology.org › home › label encoding vs. one hot encoding: what’s the difference?
Label Encoding vs. One Hot Encoding: What's the Difference?
August 8, 2022 - 1. Label Encoding: Assign each categorical value an integer value based on alphabetical order. 2. One Hot Encoding: Create new variables that take on values 0 and 1 to represent the original categorical values.
🌐
LinkedIn
linkedin.com › pulse › title-label-encoding-one-hot-data-preprocessing-shivani-singh
Title: Label Encoding and One-Hot Encoding for Data Preprocessing
July 19, 2023 - A more advanced method for managing categorical values is one-hot encoding. It establishes binary columns for each category instead of allocating integers, and indicates the existence of a category with a 1 and its lack with a 0.
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to one hot encode sequence data in python
How to One Hot Encode Sequence Data in Python - MachineLearningMastery.com
August 14, 2019 - A one hot encoding is a representation of categorical variables as binary vectors. This first requires that the categorical values be mapped to integer values.
🌐
Medium
medium.com › @milanbhadja7932 › one-hot-encoding-and-label-encoding-3a329481984e
ONE HOT ENCODING AND LABEL ENCODING | by milan bhadja | Medium
June 13, 2020 - A one hot encoding is a representation of categorical variables as binary vectors. This first requires that the categorical values be mapped to integer values.