🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.9.1 documentation
class sklearn.preprocessing.OrdinalEncoder(*, categories='auto', dtype=<class 'numpy.float64'>, handle_unknown='error', unknown_value=None, encoded_missing_value=nan, min_frequency=None, max_categories=None)[source]#
🌐
Trainindata
feature-engine.trainindata.com › en › 1.7.x › user_guide › encoding › OrdinalEncoder.html
OrdinalEncoder — 1.7.0
The OrdinalEncoder() replaces the categories by digits, starting from 0 to k-1, where k is the number of different categories. If you select “arbitrary” in the encoding_method, then the encoder will assign numbers as the labels appear in the variable (first come first served).
🌐
Ray
docs.ray.io › en › latest › data › api › doc › ray.data.preprocessors.OrdinalEncoder.html
OrdinalEncoder — Ray 2.58.0
OrdinalEncoder encodes categorical features as integers that range from \(0\) to \(n - 1\), where \(n\) is the number of categories.
🌐
Riverml
riverml.xyz › 0.21.1 › api › preprocessing › OrdinalEncoder
OrdinalEncoder - River
from river import preprocessing ... {"country": "Sweden", "place": "Taco Bell"}, {"country": None, "place": None}, ] encoder = preprocessing.OrdinalEncoder() for x in X: print(encoder.transform_one(x)) encoder.learn_one(x)...
🌐
Snowflake Documentation
docs.snowflake.com › en › developer-guide › snowpark-ml › reference › 1.0.9 › api › modeling › snowflake.ml.modeling.preprocessing.OrdinalEncoder
snowflake.ml.modeling.preprocessing.OrdinalEncoder | Snowflake Documentation
class snowflake.ml.modeling.preprocessing.OrdinalEncoder(*, categories: Union[str, Dict[str, Union[ndarray[Any, dtype[int64]], ndarray[Any, dtype[float64]], ndarray[Any, dtype[str_]], ndarray[Any, dtype[bool_]]]]] = 'auto', handle_unknown: str = 'error', unknown_value: Optional[Union[int, float]] = None, encoded_missing_value: Union[int, float] = nan, input_cols: Optional[Union[str, Iterable[str]]] = None, output_cols: Optional[Union[str, Iterable[str]]] = None, drop_input_cols: Optional[bool] = False)¶
🌐
Medium
medium.com › @bharataameriya › understanding-ordinal-encoding-in-machine-learning-d15ca5c87e5a
Understanding Ordinal Encoding in Machine Learning | by Bharataameriya | Medium
January 30, 2025 - Using OrdinalEncoder from sklearn.preprocessing: from sklearn.preprocessing import OrdinalEncoder import pandas as pd # Sample data data = pd.DataFrame({'Education': ['High School', 'Bachelor\'s', 'Master\'s', 'PhD']} # Define encoder encoder = OrdinalEncoder(categories=[['High School', 'Bachelor\'s', 'Master\'s', 'PhD']]) # Transform data data['Education_Encoded'] = encoder.fit_transform(data[['Education']]) print(data) Output: Education Education_Encoded 0 High School 0.0 1 Bachelor's 1.0 2 Master's 2.0 3 PhD 3.0 ·
Find elsewhere
🌐
Medium
medium.com › @firatozc555 › ordinal-encoding-vs-labelencoding-vs-onehotencoding-6422223a7038
Ordinal Encoding vs LabelEncoding vs OneHotEncoding | by Fırat/Mustafa Özcan | Medium
September 19, 2024 - Example: For the categories “low,” “medium,” and “high,” OrdinalEncoder might assign the values 0, 1, and 2, respectively.
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › ordinal.html
Ordinal — Category Encoders 2.11.1 documentation
class category_encoders.ordinal.OrdinalEncoder(verbose: int = 0, mapping: list[dict[str, str | dict | Series]] | None = None, cols: list[str] = None, drop_invariant: bool = False, return_df: bool = True, handle_unknown: str = 'value', handle_missing: str = 'value', index_start: int = 1, min_group_size: int | float | None = None, min_group_name: str | None = None, combine_min_nan_groups: bool | str | None = None)[source]
Top answer
1 of 5
48

You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

2 of 5
34

I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.

In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)

Then you can do something like:

from sklearn.preprocessing import LabelEncoder

for col in ["Sex","Blood", "Study"]:
    df[col] = LabelEncoder().fit_transform(df[col])

If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).

The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.

More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder

🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › how-to-perform-ordinal-encoding-using-sklearn
How to Perform Ordinal Encoding Using Sklearn - GeeksforGeeks
August 5, 2025 - By using scikit-learn's OrdinalEncoder, we can easily encode features that have a natural hierarchy, ensuring our models interpret the underlying order correctly.
Top answer
1 of 4
51

Afaik, both have the same functionality. A bit difference is the idea behind. OrdinalEncoder is for converting features, while LabelEncoder is for converting target variable.

That's why OrdinalEncoder can fit data that has the shape of (n_samples, n_features) while LabelEncoder can only fit data that has the shape of (n_samples,) (though in the past one used LabelEncoder within the loop to handle what has been becoming the job of OrdinalEncoder now)

2 of 4
30

As for differences in OrdinalEncoder and LabelEncoder implementation, the accepted answer mentions the shape of the data:

  • OrdinalEncoder is for 2D data with the shape (n_samples, n_features)
  • LabelEncoder is for 1D data with the shape (n_samples,)

Maybe that's why the top-voted answer suggests OrdinalEncoder is for the "features" (often a 2D array), whereas LabelEncoder is for the "target variable" (often a 1D array).

That's also why a OrdinalEncoder would get an error if trying to fit on 1D data: OrdinalEncoder().fit(['a','b'])

ValueError: Expected 2D array, got 1D array instead:

Another difference between the encoders is the name of their learned parameter;

  • LabelEncoder learns classes_
  • OrdinalEncoder learns categories_

Notice the differences when fitting LabelEncoder vs OrdinalEncoder, and the differences in the values of the learned parameters.

  • LabelEncoder.fit(...) accepts a 1D array; LabelEncoder.classes_ is 1D
  • OrdinalEncoder.fit(...) accepts a 2D array; OrdinalEncoder.categories_ is 2D.
    LabelEncoder().fit(['a','b']).classes_
    # >>> array(['a', 'b'], dtype='<U1')
    
    OrdinalEncoder().fit([['a'], ['b']]).categories_
    # >>> [array(['a', 'b'], dtype=object)]

This is consistent with the idea that

  • OrdinalEncoder is for your X aka your input features (2D)
  • LabelEncoder is for your y aka your target variables (1D) (also mentioned here):

LabelEncoder should be used to encode target values, i.e. y, and not the input X.

Other encoders that work in 2D, including OneHotEncoder, also use the property categories_

More info here about the dtype <U1 (little-endian , Unicode, 1 byte; i.e. a string with length 1)

EDIT

In the comments to my answer, Piotr disagrees with my answer; but Piotr points out the difference between ordinal encoding and label encoding more generally (vs differences in their implementation). Piotr's right about the general definitions/usages:

  • Ordinal encoding should be used for ordinal variables (where order matters, like cold, warm, hot);
  • vs Label encoding should be used for non-ordinal (aka nominal) variables (where order doesn't matter, like blonde, brunette)

This is a good point, but this question asks about the sklearn classes/implementation. If you want ordinal encoding like Piotr describes (i.e. where order is preserved); you must do the ordinal encoding yourself (neither OrdinalEncoder nor LabelEncoder can infer the order! See the OrdinalEncoder constructor parameter called categories).

As for implementation it seems like LabelEncoder and OrdinalEncoder have consistent behavior as far as the chosen integers. They both assign integers based on alphabetical order. For example:

OrdinalEncoder().fit_transform([['cold'],['warm'],['hot']]).reshape((1,3))
# >>> array([[0., 2., 1.]])

LabelEncoder().fit_transform(['cold','warm','hot'])
# >>> array([0, 2, 1], dtype=int64)

Notice how both encoders assigned integers in alphabetical order 'c'<'h'<'w'.

But this part is important: Notice how neither encoder got the "real" order correct (i.e. the real order should reflect the temperature, where order is 'cold'<'warm'<'hot'; 0<1<2). If the encoders used the "real" order, the value 'warm' would have been assigned the integer 1 (instead of the integer 2)

In the blog post referenced by Piotr, the author does not even use OrdinalEncoder(). To achieve ordinal encoding the author does it manually: maps each temperature to a "real" order integer, using a dictionary like {'cold':0, 'warm':1, 'hot':2}:

Refer to this code using Pandas, where first we need to assign the real order of the variable through a dictionary... Though its very straight forward but it requires coding to tell ordinal values and what is the actual mapping from text to integer as per the order.

In other words, if you're wondering whether to use OrdinalEncoder, please note OrdinalEncoder may not actually provide "ordinal encoding" the way you expect!

EDIT @Magnus Persson pointed out that the OrdinalEncoder class accepts an argument called categories, which you can use to determine/assign the resulting order.

OrdinalEncoder(categories=[['cold','warm','hot']])
    .fit_transform([['hot'],['warm'],['warm'],['cold']])
    .reshape((1,-1))[0]

# Output is:
# >>> array([[2., 1., 1., 0.]])    

EDIT @lbcommer pointed out that there is a Python library category_encoders, which has an OrdinalEncoder class. Note how even that class constructor has a mapping argument so you can choose the resulting order:

the value of ‘mapping’ should be a dictionary of ‘original_label’ to ‘encoded_label’.... example mapping: {‘col’: ‘col1’, ‘mapping’: {None: 0, ‘a’: 1, ‘b’: 2}}, {‘col’: ‘col2’, ‘mapping’: {None: 0, ‘x’: 1, ‘y’: 2}}

🌐
H1ros
h1ros.github.io › posts › ordinal-encoding-using-scikit-learn
Ordinal Encoding using Scikit-learn | Step-by-step Data Science
May 1, 2019 - # Instanciate ordinal encoder class oe = sklearn.preprocessing.OrdinalEncoder() # Learn the mapping from categories to the numbers oe.fit(df.loc[:, ['type']]) Out[3]: OrdinalEncoder(categories='auto', dtype=<class 'numpy.float64'>) In [4]: # Apply this ordinal encoder to new data oe.transform(pd.DataFrame(['cat'] * 3 + ['dog'] * 2 + ['sheep'] * 5)) Out[4]: array([[0.], [0.], [0.], [1.], [1.], [2.], [2.], [2.], [2.], [2.]]) Comments powered by Disqus
🌐
Trainindata
feature-engine.trainindata.com › en › 1.8.x › user_guide › encoding › OrdinalEncoder.html
Ordinal Encoding — 1.8.3
If the encoding_method is defined as “ordered”, then OrdinalEncoder() will assign numeric values according to the mean of the target variable for each category. The categories with the highest target mean value will be replaced by an integer value k-1, while the category with the lowest target mean value will be replaced by 0.
🌐
Medium
medium.com › @paghadalsneh › encoding-categorical-data-ordinal-encoding-86f91fd07fcf
Encoding Categorical Data | Ordinal Encoding | by Sneh Paghdal | Medium
January 23, 2025 - To apply ordinal encoding in Python, you can use the OrdinalEncoder from the scikit-learn library.
🌐
Kaggle
kaggle.com › code › rhythmcam › preprocess-basic-ordinalencoder-usage
[Preprocess Basic]OrdinalEncoder usage | Kaggle
March 23, 2022 - Explore and run AI code with Kaggle Notebooks | Using data from Titanic - Machine Learning from Disaster
🌐
PythonProg
pythonprog.com › home › scikit-learn’s preprocessing.ordinalencoder in python (with examples)
Scikit-Learn's preprocessing.OrdinalEncoder in Python (with Examples) | PythonProg
February 8, 2024 - The OrdinalEncoder is designed to transform ordinal categorical variables into numerical values while preserving the order information.