You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

Answer from abcdaire on Stack Overflow
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ stable โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder โ€” scikit-learn 1.9.1 documentation
This encoding is typically suitable for high cardinality categorical variables. ... Encodes target labels with values between 0 and n_classes-1. ... Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding. >>> from sklearn.preprocessing import OrdinalEncoder >>> enc = OrdinalEncoder() >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) OrdinalEncoder() >>> enc.categories_ [array(['Female', 'Male'], dtype=object), array([1, 2, 3], dtype=object)] >>> enc.transform([['Female', 3], ['Male', 1]]) array([[0., 2.], [1., 0.]])
๐ŸŒ
MachineLearningMastery
machinelearningmastery.com โ€บ home โ€บ blog โ€บ ordinal and one-hot encodings for categorical data
Ordinal and One-Hot Encodings for Categorical Data - MachineLearningMastery.com
August 17, 2020 - For categorical variables, it imposes an ordinal relationship where no such relationship may exist. This can cause problems and a one-hot encoding may be used instead. This ordinal encoding transform is available in the scikit-learn Python machine learning library via the OrdinalEncoder class.
Top answer
1 of 5
48

You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

2 of 5
34

I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.

In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)

Then you can do something like:

from sklearn.preprocessing import LabelEncoder

for col in ["Sex","Blood", "Study"]:
    df[col] = LabelEncoder().fit_transform(df[col])

If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).

The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.

More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder

๐ŸŒ
APXML
apxml.com โ€บ courses โ€บ intro-feature-engineering โ€บ chapter-3-encoding-categorical-features โ€บ ordinal-encoding
Ordinal Encoding for Ordered Features
Crucially, for ordinal data, you ... rather than an arbitrary one based on the order the categories appear in the data or alphabetically. import pandas as pd from sklearn.preprocessing import OrdinalEncoder # Sample data data = {'ID': [1, 2, 3, 4, 5], 'Satisfaction': ...
๐ŸŒ
Thomasjpfan
thomasjpfan.github.io โ€บ scikit-learn-website โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder โ€” scikit-learn 0.22.dev0 documentation
The categories of each feature determined during fitting (in order of the features in X and corresponding with the output of transform). ... Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding.
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ machine learning โ€บ how-to-perform-ordinal-encoding-using-sklearn
How to Perform Ordinal Encoding Using Sklearn - GeeksforGeeks
August 5, 2025 - encoder = OrdinalEncoder(categories=[['A', 'B', 'C']]) df['Grade_encoded'] = encoder.fit_transform(df[['Grade']]) print(df)
Find elsewhere
๐ŸŒ
Medium
medium.com โ€บ @bharataameriya โ€บ understanding-ordinal-encoding-in-machine-learning-d15ca5c87e5a
Understanding Ordinal Encoding in Machine Learning | by Bharataameriya | Medium
January 30, 2025 - โŒ When categorical variables have no inherent ranking (e.g., city names, colors, or country names). In such cases, one-hot encoding is a better choice. ... from sklearn.preprocessing import OrdinalEncoder import pandas as pd # Sample data data = pd.DataFrame({'Education': ['High School', 'Bachelor\'s', 'Master\'s', 'PhD']} # Define encoder encoder = OrdinalEncoder(categories=[['High School', 'Bachelor\'s', 'Master\'s', 'PhD']]) # Transform data data['Education_Encoded'] = encoder.fit_transform(data[['Education']]) print(data)
Top answer
1 of 1
16

I'm not sure if you ever figured this out but I was trying to find answers on this exact same question and there aren't really any good answers in my opinion. I finally figured it out though. OrdinalEncoder is capable of encoding multiple columns in a dataframe. So, when you instantiate OrdinalEncoder(), you give the categories parameter a list of lists:

enc = OrdinalEncoder(categories=[list_of_values_cat1, list_of_values_cat2, etc])

Specifically, in your example above, you would just put ['low', 'med', 'high'] inside another list:

end = OrdinalEncoder(categories=[['low', 'med', 'high']])
enc.fit_transform(X.loc[:,['animals']])
>>array([[0.],
         [1.],
         [0.],
         [2.],
         [0.],
         [2.]])
# Now 'low' is correctly mapped to 0, 'med' to 1, and 'high' to 2

To see how you can encode multiple columns with their own individual ordinal values, try this:

# Sample dataframe with 2 ordinal categorical columns: 'temp' and 'place'
categorical_df = pd.DataFrame({'my_id': ['101', '102', '103', '104'],
                               'temp': ['hot', 'warm', 'cool', 'cold'], 
                               'place': ['third', 'second', 'first', 'second']})

# In the 'temp' column, I want 'cold' to be 0, 'cool' to be 1, 'warm' to be 2, and 'hot' to be 3
# In the 'place' column, I want 'first' to be 0, 'second' to be 1, and 'third' to be 2
temp_categories = ['cold', 'cool', 'warm', 'hot']
place_categories = ['first', 'second', 'third']

# Now, when you instantiate the encoder, both of these lists go in one big categories list:
encoder = OrdinalEncoder(categories=[temp_categories, place_categories])

encoder.fit_transform(categorical_df[['temp', 'place']])
>>array([[3., 2.],
         [2., 1.],
         [1., 0.],
         [0., 1.]])
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 1.0 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder โ€” scikit-learn 1.0.2 documentation
Performs a one-hot encoding of categorical features. ... Encodes target labels with values between 0 and n_classes-1. ... Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal ...
๐ŸŒ
Scikit-learn
contrib.scikit-learn.org โ€บ category_encoders โ€บ ordinal.html
Ordinal โ€” Category Encoders 2.11.1 documentation
class category_encoders.ordina... str | None = None, combine_min_nan_groups: bool | str | None = None)[source]๏ƒ ยท Encodes categorical features as ordinal, in one ordered feature....
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.24 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder โ€” scikit-learn 0.24.2 documentation
Performs a one-hot encoding of categorical features. ... Encodes target labels with values between 0 and n_classes-1. ... Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding. >>> from sklearn.preprocessing import ...
๐ŸŒ
Trainindata
feature-engine.trainindata.com โ€บ en โ€บ latest โ€บ user_guide โ€บ encoding โ€บ OrdinalEncoder.html
Ordinal Encoding โ€” 1.9.4 - Feature-engine
That is, it encodes categorical features by replacing each category with a unique number ranging from 0 to k-1, where โ€˜kโ€™ is the distinct number of categories in the dataset. OrdinalEncoder() supports both arbitrary and ordered encoding methods.
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.20 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder โ€” scikit-learn 0.20.4 documentation
OrdinalEncoder(categories='auto', dtype=<... 'numpy.float64'>) >>> enc.categories_ [array(['Female', 'Male'], dtype=object), array([1, 2, 3], dtype=object)] >>> enc.transform([['Female', 3], ['Male', 1]]) array([[0., 2.], [1., 0.]])
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 1.5 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder โ€” scikit-learn 1.5.2 documentation
With a high proportion of nan values, inferring categories becomes slow with Python versions before 3.10. The handling of nan values was improved from Python 3.10 onwards, (c.f. bpo-43475). ... Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding. >>> from sklearn.preprocessing import OrdinalEncoder >>> enc = OrdinalEncoder() >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) OrdinalEncoder() >>> enc.categories_ [array(['Female', 'Male'], dtype=object), array([1, 2, 3], dtype=object)] >>> enc.transform([['Female', 3], ['Male', 1]]) array([[0., 2.], [1., 0.]])
๐ŸŒ
Applied AI Blog
appliedaicourse.com โ€บ home โ€บ data science โ€บ ordinal encoding โ€” a brief guide
Ordinal Encoding โ€” A Brief Guide
April 28, 2025 - # Define explicit orderings education_order = ['High School', 'Bachelor', 'Master', 'PhD'] satisfaction_order = ['Poor', 'Average', 'Good', 'Excellent'] # Initialize the encoder with category order encoder = OrdinalEncoder(categories=[education_order, satisfaction_order], handle_unknown='use_encoded_value', unknown_value=-1) # Apply encoding df_encoded = encoder.fit_transform(df) df_encoded = pd.DataFrame(df_encoded, columns=['Education', 'Satisfaction']) print(df_encoded)
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.22 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.OrdinalEncoder.html
sklearn.preprocessing.OrdinalEncoder โ€” scikit-learn 0.22.2 documentation
Performs a one-hot encoding of categorical features. ... Encodes target labels with values between 0 and n_classes-1. ... Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding. >>> from sklearn.preprocessing import ...