You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

Answer from abcdaire on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.9.1 documentation
Encode categorical features as an integer array. The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are converted to ordinal integers.
🌐
scikit-learn
scikit-learn.org › 0.20 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder — scikit-learn 0.20.4 documentation
The features are converted to ordinal integers. This results in a single column of integers (0 to n_categories - 1) per feature. Read more in the User Guide. See also · sklearn.preprocessing.OneHotEncoder · performs a one-hot encoding of categorical features.
🌐
scikit-learn
scikit-learn.org › 1.0 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder — scikit-learn 1.0.2 documentation
class sklearn.preprocessing.OrdinalEncoder(*, categories='auto', dtype=<class 'numpy.float64'>, handle_unknown='error', unknown_value=None)[source]¶ · Encode categorical features as an integer array.
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › how-to-perform-ordinal-encoding-using-sklearn
How to Perform Ordinal Encoding Using Sklearn - GeeksforGeeks
August 5, 2025 - Initializes OrdinalEncoder and explicitly sets the order: 'A' < 'B' < 'C'. Transforms the 'Grade' column into numeric codes (0, 1, 2). Stores the result in a new column Grade_encoded.
🌐
scikit-learn
scikit-learn.org › 0.22 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder — scikit-learn 0.22.2 documentation
class sklearn.preprocessing.OrdinalEncoder(categories='auto', dtype=<class 'numpy.float64'>)[source]¶ · Encode categorical features as an integer array.
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › ordinal.html
Ordinal — Category Encoders 2.11.1 documentation
Perform the inverse transformation to encoded data. Will attempt best case reconstruction, which means it will return nan for handle_missing and handle_unknown settings that break the bijection. We issue warnings when some of those cases occur. ... static ordinal_encoding(X_in: DataFrame, mapping: list[dict[str, str | dict | Series]] | None = None, cols: list[str] = None, handle_unknown: str = 'value', handle_missing: str = 'value', index_start: int = 1) → tuple[DataFrame, list[dict]][source]
🌐
scikit-learn
scikit-learn.org › 0.24 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder — scikit-learn 0.24.2 documentation
class sklearn.preprocessing.OrdinalEncoder(*, categories='auto', dtype=<class 'numpy.float64'>, handle_unknown='error', unknown_value=None)[source]¶ · Encode categorical features as an integer array.
🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.10.dev0 documentation
Encode categorical features as an integer array. The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are converted to ordinal integers.
Find elsewhere
Top answer
1 of 5
48

You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

2 of 5
34

I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.

In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)

Then you can do something like:

from sklearn.preprocessing import LabelEncoder

for col in ["Sex","Blood", "Study"]:
    df[col] = LabelEncoder().fit_transform(df[col])

If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).

The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.

More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder

🌐
Sklearn
sklearn.org › 1.6 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.6.0 documentation - sklearn
Encode categorical features as an integer array. The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are converted to ordinal integers.
🌐
scikit-learn
scikit-learn.org › 0.23 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder — scikit-learn 0.23.2 documentation
class sklearn.preprocessing.OrdinalEncoder(*, categories='auto', dtype=<class 'numpy.float64'>)[source]¶ · Encode categorical features as an integer array.
🌐
Thomasjpfan
thomasjpfan.github.io › scikit-learn-website › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder — scikit-learn 0.22.dev0 documentation
class sklearn.preprocessing.OrdinalEncoder(categories='auto', dtype=<class 'numpy.float64'>)[source]¶ · Encode categorical features as an integer array.
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.5.2 documentation
Encode categorical features as an integer array. The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are converted to ordinal integers.
🌐
Trainindata
feature-engine.trainindata.com › en › 1.8.x › user_guide › encoding › OrdinalEncoder.html
Ordinal Encoding — 1.8.3
We’ll show how ordinal encoding is implemented by Feature-engine’s OrdinalEncoder() using the Titanic Dataset. Let’s load the dataset and split it into train and test sets: import pandas as pd from sklearn.model_selection import train_test_split from feature_engine.datasets import load_titanic from feature_engine.encoding import OrdinalEncoder X, y = load_titanic( return_X_y_frame=True, handle_missing=True, predictors_only=True, cabin="letter_only", ) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=0, ) print(X_train.head())
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › ordinal and one-hot encodings for categorical data
Ordinal and One-Hot Encodings for Categorical Data - MachineLearningMastery.com
August 17, 2020 - I checked the syntax documentation here: https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html?highlight=onehotencoder#sklearn.preprocessing.OneHotEncoder · The code seems to be right, but I still get the error. Can you please help? Thanks so much for your tutorials – they are great! ... Ugh. SO sorry. I thought I was on the latest version. I updated and the dummy encoding code works now. Thanks for taking the time to get back to me! ... No problem at all. ... Please, i am trying to fit the OrdinalEncoder on the training dataset and use it to transform the train and test datasets as follows; ordinal_encoder = OrdinalEncoder() ordinal_encoder.fit(X_train) X_train = ordinal_encoder.transform(X_train) X_test = ordinal_encoder.transform(X_test)
🌐
The Security Buddy
thesecuritybuddy.com › home › data preprocessing
How to perform ordinal encoding using sklearn? - The Security Buddy
November 16, 2022 - When a column in a dataset contains ordinal values, we use ordinal encoding to encode the ordinal data. We can use the OrdinalEncoder class from the sklearn.preprocessing module to perform ordinal encoding.
🌐
Trainindata
feature-engine.trainindata.com › en › 1.7.x › user_guide › encoding › OrdinalEncoder.html
OrdinalEncoder — 1.7.0
Arbitrary ordinal encoding: the numbers will be assigned arbitrarily to the categories, on a first seen first served basis. Let’s look at an example using the Titanic Dataset. First, let’s load the data and separate it into train and test: from sklearn.model_selection import train_test_split from feature_engine.datasets import load_titanic from feature_engine.encoding import OrdinalEncoder X, y = load_titanic( return_X_y_frame=True, handle_missing=True, predictors_only=True, cabin="letter_only", ) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=0, ) print(X_train.head())
🌐
Trainindata
feature-engine.trainindata.com › en › latest › user_guide › encoding › OrdinalEncoder.html
Ordinal Encoding — 1.9.4 - Feature-engine
We’ll show how ordinal encoding is implemented by Feature-engine’s OrdinalEncoder() using the Titanic Dataset. Let’s load the dataset and split it into train and test sets: import pandas as pd from sklearn.model_selection import train_test_split from feature_engine.datasets import load_titanic from feature_engine.encoding import OrdinalEncoder X, y = load_titanic( return_X_y_frame=True, handle_missing=True, predictors_only=True, cabin="letter_only", ) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=0, ) print(X_train.head())
🌐
Towards Data Science
towardsdatascience.com › home › latest › feature engineering ordinal variables
Feature Engineering Ordinal Variables | Towards Data Science
January 16, 2025 - Variables with an ordered sequence are Ordinal variables; group labels are in ascending/ descending order. ... 📔 For many algorithms (machine learning models), the input must be numerical.