Afaik, both have the same functionality. A bit difference is the idea behind. OrdinalEncoder is for converting features, while LabelEncoder is for converting target variable.

That's why OrdinalEncoder can fit data that has the shape of (n_samples, n_features) while LabelEncoder can only fit data that has the shape of (n_samples,) (though in the past one used LabelEncoder within the loop to handle what has been becoming the job of OrdinalEncoder now)

Answer from ipramusinto on Stack Exchange
Top answer
1 of 4
51

Afaik, both have the same functionality. A bit difference is the idea behind. OrdinalEncoder is for converting features, while LabelEncoder is for converting target variable.

That's why OrdinalEncoder can fit data that has the shape of (n_samples, n_features) while LabelEncoder can only fit data that has the shape of (n_samples,) (though in the past one used LabelEncoder within the loop to handle what has been becoming the job of OrdinalEncoder now)

2 of 4
30

As for differences in OrdinalEncoder and LabelEncoder implementation, the accepted answer mentions the shape of the data:

  • OrdinalEncoder is for 2D data with the shape (n_samples, n_features)
  • LabelEncoder is for 1D data with the shape (n_samples,)

Maybe that's why the top-voted answer suggests OrdinalEncoder is for the "features" (often a 2D array), whereas LabelEncoder is for the "target variable" (often a 1D array).

That's also why a OrdinalEncoder would get an error if trying to fit on 1D data: OrdinalEncoder().fit(['a','b'])

ValueError: Expected 2D array, got 1D array instead:

Another difference between the encoders is the name of their learned parameter;

  • LabelEncoder learns classes_
  • OrdinalEncoder learns categories_

Notice the differences when fitting LabelEncoder vs OrdinalEncoder, and the differences in the values of the learned parameters.

  • LabelEncoder.fit(...) accepts a 1D array; LabelEncoder.classes_ is 1D
  • OrdinalEncoder.fit(...) accepts a 2D array; OrdinalEncoder.categories_ is 2D.
    LabelEncoder().fit(['a','b']).classes_
    # >>> array(['a', 'b'], dtype='<U1')
    
    OrdinalEncoder().fit([['a'], ['b']]).categories_
    # >>> [array(['a', 'b'], dtype=object)]

This is consistent with the idea that

  • OrdinalEncoder is for your X aka your input features (2D)
  • LabelEncoder is for your y aka your target variables (1D) (also mentioned here):

LabelEncoder should be used to encode target values, i.e. y, and not the input X.

Other encoders that work in 2D, including OneHotEncoder, also use the property categories_

More info here about the dtype <U1 (little-endian , Unicode, 1 byte; i.e. a string with length 1)

EDIT

In the comments to my answer, Piotr disagrees with my answer; but Piotr points out the difference between ordinal encoding and label encoding more generally (vs differences in their implementation). Piotr's right about the general definitions/usages:

  • Ordinal encoding should be used for ordinal variables (where order matters, like cold, warm, hot);
  • vs Label encoding should be used for non-ordinal (aka nominal) variables (where order doesn't matter, like blonde, brunette)

This is a good point, but this question asks about the sklearn classes/implementation. If you want ordinal encoding like Piotr describes (i.e. where order is preserved); you must do the ordinal encoding yourself (neither OrdinalEncoder nor LabelEncoder can infer the order! See the OrdinalEncoder constructor parameter called categories).

As for implementation it seems like LabelEncoder and OrdinalEncoder have consistent behavior as far as the chosen integers. They both assign integers based on alphabetical order. For example:

OrdinalEncoder().fit_transform([['cold'],['warm'],['hot']]).reshape((1,3))
# >>> array([[0., 2., 1.]])

LabelEncoder().fit_transform(['cold','warm','hot'])
# >>> array([0, 2, 1], dtype=int64)

Notice how both encoders assigned integers in alphabetical order 'c'<'h'<'w'.

But this part is important: Notice how neither encoder got the "real" order correct (i.e. the real order should reflect the temperature, where order is 'cold'<'warm'<'hot'; 0<1<2). If the encoders used the "real" order, the value 'warm' would have been assigned the integer 1 (instead of the integer 2)

In the blog post referenced by Piotr, the author does not even use OrdinalEncoder(). To achieve ordinal encoding the author does it manually: maps each temperature to a "real" order integer, using a dictionary like {'cold':0, 'warm':1, 'hot':2}:

Refer to this code using Pandas, where first we need to assign the real order of the variable through a dictionary... Though its very straight forward but it requires coding to tell ordinal values and what is the actual mapping from text to integer as per the order.

In other words, if you're wondering whether to use OrdinalEncoder, please note OrdinalEncoder may not actually provide "ordinal encoding" the way you expect!

EDIT @Magnus Persson pointed out that the OrdinalEncoder class accepts an argument called categories, which you can use to determine/assign the resulting order.

OrdinalEncoder(categories=[['cold','warm','hot']])
    .fit_transform([['hot'],['warm'],['warm'],['cold']])
    .reshape((1,-1))[0]

# Output is:
# >>> array([[2., 1., 1., 0.]])    

EDIT @lbcommer pointed out that there is a Python library category_encoders, which has an OrdinalEncoder class. Note how even that class constructor has a mapping argument so you can choose the resulting order:

the value of ‘mapping’ should be a dictionary of ‘original_label’ to ‘encoded_label’.... example mapping: {‘col’: ‘col1’, ‘mapping’: {None: 0, ‘a’: 1, ‘b’: 2}}, {‘col’: ‘col2’, ‘mapping’: {None: 0, ‘x’: 1, ‘y’: 2}}

🌐
Medium
medium.com › @firatozc555 › ordinal-encoding-vs-labelencoding-vs-onehotencoding-6422223a7038
Ordinal Encoding vs LabelEncoding vs OneHotEncoding | by Fırat/Mustafa Özcan | Medium
September 19, 2024 - from sklearn.preprocessing import OrdinalEncoder >>> enc = OrdinalEncoder() >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) OrdinalEncoder() >>> enc.categories_ [array(['Female', 'Male'], dtype=object), array([1, 2, 3], dtype=object)] >>> enc.transform([['Female', 3], ['Male', 1]]) array([[0., 2.], [1., 0.]])
🌐
Medium
medium.com › @vipinnation › unraveling-categorical-variables-understanding-label-ordinal-and-one-hot-encoding-techniques-8f5375151fed
Unraveling Categorical Variables: Understanding Label, Ordinal, and One-Hot Encoding Techniques | by Vipin Singh Inkiya | Medium
April 20, 2024 - In this example, Red is encoded as 2, Green as 1, and Blue as 0. One potential issue with Label Encoding is that it implies ordinal relationships between categories, which might not always be accurate.
🌐
Kaggle
kaggle.com › questions-and-answers › 170936
OrdinalEncoder vs LabelEncoder | Kaggle
Hi everyone, Recently, when I was reading an ML book, I came across sklearn's OrdinalEncoder. I tried it out, and it seems very similar to sklearn's LabelEn...
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.9.1 documentation
OrdinalEncoder · Encode categorical features using an ordinal encoding scheme. OneHotEncoder · Encode categorical features as a one-hot numeric array. Examples · LabelEncoder can be used to normalize labels.
🌐
Towards Data Science
towardsdatascience.com › home › artificial intelligence › machine learning › 3 key encoding techniques for machine learning: a beginner-friendly guide
3 Key Encoding Techniques for Machine Learning: A Beginner-Friendly Guide | Towards Data Science
February 7, 2024 - # Scikit-learnfrom sklearn.preprocessing import LabelEncoderle = LabelEncoder()df_math["Subject_num_scikit"] = le.fit_transform(df_math[['Subject']])print(df_math.iloc[sampled_index])
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › ordinal and one-hot encodings for categorical data
Ordinal and One-Hot Encodings for Categorical Data - MachineLearningMastery.com
August 17, 2020 - This OrdinalEncoder class is intended for input variables that are organized into rows and columns, e.g. a matrix. If a categorical target variable needs to be encoded for a classification predictive modeling problem, then the LabelEncoder ...
🌐
Kaggle
kaggle.com › questions-and-answers › 239564
OneHotEncoder VS OrdinalEncoder VS LabelEncoder in scikit-learn | Kaggle
The third one one, LabelEncoder, is used when you want to transform your dependent variables into classes, e.g., : [1, 1, 2, 6] -> [0, 0, 1, 2]. This is only intended to be used with your LABELS, i.e., your dependent variables, and not your ...
Find elsewhere
🌐
Trainindata
feature-engine.trainindata.com › en › 1.8.x › user_guide › encoding › OrdinalEncoder.html
Ordinal Encoding — 1.8.3
Scikit-learn provides 2 different ... is, categories, with ordinal data. The OrdinalEncoder is designed to transform the predictor variables (those in the training set), while the LabelEncoder is designed to transform the target variable....
🌐
Kaggle
kaggle.com › questions-and-answers › 312940
Ordinal Encoder vs Label encoder | Kaggle
I dont understand the difference ebetween this two (OrdinalEncoder vs Label encoder) can someone explain it ot me ? There both from sklearn.preprocessing
🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - In the given example, the countries have no inherent order, but one hot encoding and label encoding introduces an ordinal relationship based on the encoded integers (e.g., France < Germany < Spain).
Top answer
1 of 5
48

You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

2 of 5
34

I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.

In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)

Then you can do something like:

from sklearn.preprocessing import LabelEncoder

for col in ["Sex","Blood", "Study"]:
    df[col] = LabelEncoder().fit_transform(df[col])

If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).

The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.

More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder

🌐
Medium
medium.com › @vigneshvars2001 › understanding-label-encoding-and-ordinal-encoding-a-deep-dive-into-categorical-data-transformation-fc770a2e4060
Understanding Label Encoding and Ordinal Encoding: A Deep Dive into Categorical Data Transformation | by Vigneshvar Sreekanth | Medium
January 5, 2025 - # Sample data genres = ['Action', 'Comedy', 'Romance', 'Comedy', 'Action']# Initialize and fit the encoder encoder = LabelEncoder() encoded_genres = encoder.fit_transform(genres)print("Original:", genres) print("Encoded:", encoded_genres)
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.9.1 documentation
LabelEncoder · Encodes target labels with values between 0 and n_classes-1. Examples · Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding.
Top answer
1 of 1
28

TL;DR: Using a LabelEncoder to encode ordinal any kind of features is a bad idea!


This is in fact clearly stated in the docs, where it is mentioned that as its name suggests this encoding method is aimed at encoding the label:

This transformer should be used to encode target values, i.e. y, and not the input X.

As you rightly point out in the question, mapping the inherent ordinality of an ordinal feature to a wrong scale will have a very negative impact on the performance of the model (that is, proportional to the relevance of the feature). And the same applies to a categorical feature, just that the original feature has no ordinality.

An intuitive way to think about it, is in the way a decision tree sets its boundaries. During training, a decision tree will learn the optimal features to set at each node, as well as an optimal threshold whereby unseen samples will follow a branch or another depending on these values.

If we encode an ordinal feature using a simple LabelEncoder, that could lead to a feature having say 1 represent warm, 2 which maybe would translate to hot, and a 0 representing boiling. In such case, the result will end up being a tree with an unnecessarily high amount of splits, and hence a much higher complexity for what should be simpler to model.

Instead, the right approach would be to use an OrdinalEncoder, and define the appropriate mapping schemes for the ordinal features. Or in the case of having a categorical feature, we should be looking at OneHotEncoder or the various encoders available in Category Encoders.


Though actually seeing why this is a bad idea will be more intuitive than just words.

Let's use a simple example to illustrate the above, consisting on two ordinal features containing a range with the amount of hours spend by a student preparing for an exam and the average grade of all previous assignments, and a target variable indicating whether the exam was past or not. I've defined the dataframe's columns as pd.Categorical:

df = pd.DataFrame(
        {'Hours of dedication': pd.Categorical(
              values =  ['25-30', '20-25', '5-10', '5-10', '40-45', 
                         '0-5', '15-20', '20-25', '30-35', '5-10',
                         '10-15', '45-50', '20-25'],
              categories=['0-5', '5-10', '10-15', '15-20', 
                          '20-25', '25-30','30-35','40-45', '45-50']),

         'Assignments avg grade': pd.Categorical(
             values =  ['B', 'C', 'F', 'C', 'B', 
                        'D', 'C', 'A', 'B', 'B', 
                        'B', 'A', 'D'],
             categories=['F', 'D', 'C', 'B','A']),

         'Result': pd.Categorical(
             values = ['Pass', 'Pass', 'Fail', 'Fail', 'Pass', 
                       'Fail', 'Fail','Pass','Pass', 'Fail', 
                       'Fail', 'Pass', 'Pass'], 
             categories=['Fail', 'Pass'])
        }
    )

The advantage of defining a categorical column as a pandas' categorical, is that we get to establish an order among its categories, as mentioned earlier. This allows for much faster sorting based on the established order rather than lexical sorting. And it can also be used as a simple way to get codes for the different categories according to their order.

So the dataframe we'll be using looks as follows:

print(df.head())

  Hours_of_dedication   Assignments_avg_grade   Result
0               20-25                       B     Pass
1               20-25                       C     Pass
2                5-10                       F     Fail
3                5-10                       C     Fail
4               40-45                       B     Pass
5                 0-5                       D     Fail
6               15-20                       C     Fail
7               20-25                       A     Pass
8               30-35                       B     Pass
9                5-10                       B     Fail

The corresponding category codes can be obtained with:

X = df.apply(lambda x: x.cat.codes)
X.head()

   Hours_of_dedication   Assignments_avg_grade   Result
0                    4                       3        1
1                    4                       2        1
2                    1                       0        0
3                    1                       2        0
4                    7                       3        1
5                    0                       1        0
6                    3                       2        0
7                    4                       4        1
8                    6                       3        1
9                    1                       3        0

Now let's fit a DecisionTreeClassifier, and see what is how the tree has defined the splits:

from sklearn import tree

dt = tree.DecisionTreeClassifier()
y = X.pop('Result')
dt.fit(X, y)

We can visualise the tree structure using plot_tree:

t = tree.plot_tree(dt, 
                   feature_names = X.columns,
                   class_names=["Fail", "Pass"],
                   filled = True,
                   label='all',
                   rounded=True)

Is that all?? Well… yes! I've actually set the features in such a way that there is this simple and obvious relation between the Hours of dedication feature, and whether the exam is passed or not, making it clear that the problem should be very easy to model.


Now let's try to do the same by directly encoding all features with an encoding scheme we could have obtained for instance through a LabelEncoder, so disregarding the actual ordinality of the features, and just assigning a value at random:

df_wrong = df.copy()
df_wrong['Hours_of_dedication'].cat.set_categories(
             ['0-5','40-45', '25-30', '10-15', '5-10', '45-50','15-20', 
              '20-25','30-35'], inplace=True)
df_wrong['Assignments_avg_grade'].cat.set_categories(
             ['A', 'C', 'F', 'D', 'B'], inplace=True)

rcParams['figure.figsize'] = 14,18
X_wrong = df_wrong.drop(['Result'],1).apply(lambda x: x.cat.codes)
y = df_wrong.Result

dt_wrong = tree.DecisionTreeClassifier()
dt_wrong.fit(X_wrong, y)

t = tree.plot_tree(dt_wrong, 
                   feature_names = X_wrong.columns,
                   class_names=["Fail", "Pass"],
                   filled = True,
                   label='all',
                   rounded=True)

As expected the tree structure is way more complex than necessary for the simple problem we're trying to model. In order for the tree to correctly predict all training samples it has expanded until a depth of 4, when a single node should suffice.

This will imply that the classifier is likely to overfit, since we’re drastically increasing the complexity. And by pruning the tree and tuning the necessary parameters to prevent overfitting we are not solving the problem either, since we’ve added too much noise by wrongly encoding the features.

So to summarize, preserving the ordinality of the features once encoding them is crucial, otherwise as made clear with this example we'll lose all their predictable power and just add noise to our model.

🌐
C# Corner
c-sharpcorner.com › article › ordinal-label-encoding-in-machine-learning
Ordinal & Label Encoding in Machine Learning
May 10, 2024 - #split the data frame into test & train from sklearn.model_selection import train_test_split X_train,X_test,Y_train,Y_test = train_test_split(df.iloc[:,0:2],df.iloc[:,-1],test_size=0.2) # to perform ordinal encoding we will import OrdinalEncoder from sklearn from sklearn.preprocessing import OrdinalEncoder #Lets see our splited dataframe X_train
🌐
Medium
tahera-firdose.medium.com › understanding-categorical-encoding-techniques-ordinal-one-hot-and-label-encoding-4ad209c13e06
Understanding Categorical Encoding Techniques: Ordinal, One-Hot, and Label Encoding | by Tahera Firdose | Medium
December 8, 2023 - Understanding Categorical Encoding Techniques: Ordinal, One-Hot, and Label Encoding Introduction: Categorical variables are an essential part of data analysis, but they cannot be directly processed …
🌐
Skill Certify
skillcertify.org › question › what-is-the-purpose-of-labelencoder-in-scikit-learn-and-when-should-you-use-ordinalencoder-instead
What is the purpose of LabelEncoder in scikit-learn and when should you use OrdinalEncoder instead?
April 29, 2026 - A LabelEncoder is for binary categories; OrdinalEncoder handles more than two categories B They are identical; OrdinalEncoder is simply the updated version of LabelEncoder C LabelEncoder produces one-hot encoded output; OrdinalEncoder produces ...
🌐
Kaggle
kaggle.com › general › 330406
Label encoder Ordinal Encoder and One hot encoder OHE | Kaggle
Encoding is techniques used to convert or transform categorical features into numerical form Different types of encoding techniques are used depends upon typ...