You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

Answer from abcdaire on Stack Overflow
Top answer
1 of 5
48

You were almost there !

Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.

You have to call transform to transform the data , or use fit_transform which fit and transform the same data.

enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])

or directly

enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])

Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.

As you can see from enc.categories_

[array(['F', 'M'], dtype=object),
 array(['A', 'AB', 'B', 'O'], dtype=object),
 array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```

Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)

2 of 5
34

I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.

In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)

Then you can do something like:

from sklearn.preprocessing import LabelEncoder

for col in ["Sex","Blood", "Study"]:
    df[col] = LabelEncoder().fit_transform(df[col])

If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).

The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.

More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder

🌐
Medium
medium.com › bycodegarage › encoding-categorical-data-in-machine-learning-def03ccfbf40
Encoding Categorical data in Machine Learning | by Akhil Reddy Mallidi | #ByCodeGarage | Medium
September 6, 2019 - # Creating an Pandas dataframe for ordinal datadata = {'Employee Id' : [112, 113, 114, 115], 'Income Range' : ['Low', 'High', 'Medium', 'High']}df_ordinal = pd.DataFrame(data)# Viewing few rows of created dataframedf_ordinal.head() Few rows of sample dataframe · # Encoding above ordinal data using OrdinalEncoderfrom sklearn.preprocessing import OrdinalEncoderordinalencoder = OrdinalEncoder()ordinalencoder.fit_transform(df_ordinal[['Income Range']]) Output returned after encoding using Ordinal Encoder ·
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › how-to-perform-ordinal-encoding-using-sklearn
How to Perform Ordinal Encoding Using Sklearn - GeeksforGeeks
August 5, 2025 - import pandas as pd from sklearn.preprocessing import OrdinalEncoder import matplotlib.pyplot as plt import seaborn as sns
🌐
Towards Data Science
towardsdatascience.com › home › latest › feature engineering ordinal variables
Feature Engineering Ordinal Variables | Towards Data Science
January 16, 2025 - Ideally, we would want something per High(2) > Med(1) > Low(0). We can do this via Pandas. # Define a dictionary for encoding target variable enc_dict = {'Low':0, 'Med':1, 'High':2} # Create the mapped values in a new column Ex['target_ordinal'] = Ex['retnRisks'].map(enc_dict) ... Next, the predictor variables. # Instantiate ordinal encoder ordinal_encoder = OrdinalEncoder()
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.9.1 documentation
class sklearn.preprocessing.OrdinalEncoder(*, categories='auto', dtype=<class 'numpy.float64'>, handle_unknown='error', unknown_value=None, encoded_missing_value=nan, min_frequency=None, max_categories=None)[source]#
🌐
Trainindata
feature-engine.trainindata.com › en › 1.8.x › api_doc › encoding › OrdinalEncoder.html
OrdinalEncoder — 1.8.3
>>> import pandas as pd >>> from feature_engine.encoding import OrdinalEncoder >>> X = pd.DataFrame(dict(x1 = [1,2,3,4], x2 = ["c", "a", "b", "c"])) >>> y = pd.Series([0,1,1,0]) >>> od = OrdinalEncoder(encoding_method='arbitrary') >>> od.fit(X) >>> od.transform(X) x1 x2 0 1 0 1 2 1 2 3 2 3 4 0 ·
Top answer
1 of 1
16

I'm not sure if you ever figured this out but I was trying to find answers on this exact same question and there aren't really any good answers in my opinion. I finally figured it out though. OrdinalEncoder is capable of encoding multiple columns in a dataframe. So, when you instantiate OrdinalEncoder(), you give the categories parameter a list of lists:

enc = OrdinalEncoder(categories=[list_of_values_cat1, list_of_values_cat2, etc])

Specifically, in your example above, you would just put ['low', 'med', 'high'] inside another list:

end = OrdinalEncoder(categories=[['low', 'med', 'high']])
enc.fit_transform(X.loc[:,['animals']])
>>array([[0.],
         [1.],
         [0.],
         [2.],
         [0.],
         [2.]])
# Now 'low' is correctly mapped to 0, 'med' to 1, and 'high' to 2

To see how you can encode multiple columns with their own individual ordinal values, try this:

# Sample dataframe with 2 ordinal categorical columns: 'temp' and 'place'
categorical_df = pd.DataFrame({'my_id': ['101', '102', '103', '104'],
                               'temp': ['hot', 'warm', 'cool', 'cold'], 
                               'place': ['third', 'second', 'first', 'second']})

# In the 'temp' column, I want 'cold' to be 0, 'cool' to be 1, 'warm' to be 2, and 'hot' to be 3
# In the 'place' column, I want 'first' to be 0, 'second' to be 1, and 'third' to be 2
temp_categories = ['cold', 'cool', 'warm', 'hot']
place_categories = ['first', 'second', 'third']

# Now, when you instantiate the encoder, both of these lists go in one big categories list:
encoder = OrdinalEncoder(categories=[temp_categories, place_categories])

encoder.fit_transform(categorical_df[['temp', 'place']])
>>array([[3., 2.],
         [2., 1.],
         [1., 0.],
         [0., 1.]])
Find elsewhere
🌐
APXML
apxml.com › courses › intro-feature-engineering › chapter-3-encoding-categorical-features › ordinal-encoding
Ordinal Encoding for Ordered Features
You can implement Ordinal Encoding manually using Pandas or leverage Scikit-learn's dedicated transformer.
Top answer
1 of 2
5

You are using the 'mapping' param wrong.

The format should be:

'mapping' param should be a list of dicts where internal dicts should contain the keys 'col' and 'mapping' and in that the 'mapping' key should have a list of tuples of format (original_label, encoded_label) as value.

Something like this:

ordinal_cols_mapping = [{
    "col":"ExterQual",    
    "mapping": [
        ('Ex',5), 
        ('Gd',4), 
        ('TA',3), 
        ('Fa',2), 
        ('Po',1), 
        ('NA',np.nan)
    ]},
]

Then no need to set the 'cols' param separately. Column names will be used from the 'mapping' param.

Just do this:

encoder = OrdinalEncoder(mapping = ordinal_cols_mapping, 
                         return_df = True)  
df_train = encoder.fit_transform(train_data)

Hope that this makes it clear.

2 of 2
2

In case all of your columns to encode are already pandas categoricals, you can construct a mapping like this.

In [82]:
from category_encoders.ordinal import OrdinalEncoder
import pandas as pd
from pandas.api.types import CategoricalDtype

# define a categorical dtype
platforms = ['android', 'ios', 'amazon']
platform_category = CategoricalDtype(categories=platforms, ordered=False)

# create a dataframe
df = pd.DataFrame([
    {'id': 1, 'platform': 'android'},
    {'id': 2, 'platform': 'ios'},
    {'id': 3, 'platform': 'amazon'},
])
# apply the categorical dtype
df['platform'] = df['platform'].astype(platform_category)

# create a mapping from all categorical columns that can be used with OrdinalEncoder
categorical_columns = list(df.select_dtypes(['category']).columns)
category_mapping = [
    {'col': column_name, 'mapping': list(zip(df[column_name].cat.categories, df[column_name].cat.codes))} 
    for column_name in categorical_columns
]

# pass this as the mapping
cat_encoder = OrdinalEncoder(cols=categorical_columns, mapping=category_mapping)
cat_encoder
Out[82]:
OrdinalEncoder(cols=['platform'], drop_invariant=False,
        handle_unknown='impute', impute_missing=True,
        mapping=[{'col': 'platform', 'mapping': [('android', 0), ('ios', 1), ('amazon', 2)]}],
        return_df=True, verbose=0)
🌐
Trainindata
feature-engine.trainindata.com › en › latest › api_doc › encoding › OrdinalEncoder.html
OrdinalEncoder — 1.9.4
>>> import pandas as pd >>> from feature_engine.encoding import OrdinalEncoder >>> X = pd.DataFrame(dict(x1 = [1,2,3,4], x2 = ["c", "a", "b", "c"])) >>> y = pd.Series([0,1,1,0]) >>> od = OrdinalEncoder(encoding_method='arbitrary') >>> od.fit(X) >>> od.transform(X) x1 x2 0 1 0 1 2 1 2 3 2 3 4 0 ·
🌐
Seaborn
deeplearningnerds.com › pandas-encode-ordinal-categorical-features
Pandas - Ordinal Encoding
November 16, 2023 - To encode the categorical values, we use the replace() method of Pandas and pass a dictionary with the mapping between categorical and numerical values:
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › ordinal and one-hot encodings for categorical data
Ordinal and One-Hot Encodings for Categorical Data - MachineLearningMastery.com
August 17, 2020 - An interesting discussion on “Why OneHotEncoder not get_dummies?” in sklearn can be found here: https://stackoverflow.com/questions/36631163/what-are-the-pros-and-cons-between-get-dummies-pandas-and-onehotencoder-sciki ... Thanks for sharing. ... In case of the OrdinalEncoding, I have observed shift in the encoded value after introducing new value E.g.
🌐
Medium
medium.com › @bharataameriya › understanding-ordinal-encoding-in-machine-learning-d15ca5c87e5a
Understanding Ordinal Encoding in Machine Learning | by Bharataameriya | Medium
January 30, 2025 - from sklearn.preprocessing import OrdinalEncoder import pandas as pd # Sample data data = pd.DataFrame({'Education': ['High School', 'Bachelor\'s', 'Master\'s', 'PhD']} # Define encoder encoder = OrdinalEncoder(categories=[['High School', 'Bachelor\'s', 'Master\'s', 'PhD']]) # Transform data data['Education_Encoded'] = encoder.fit_transform(data[['Education']]) print(data) Output: Education Education_Encoded 0 High School 0.0 1 Bachelor's 1.0 2 Master's 2.0 3 PhD 3.0 ·
🌐
Dask
ml.dask.org › modules › generated › dask_ml.preprocessing.OrdinalEncoder.html
dask_ml.preprocessing.OrdinalEncoder — dask-ml 2025.1.1 documentation
This transformer only applies to dask and pandas DataFrames. For dask DataFrames, all of your categoricals should be known. The inverse transformation can be used on a dataframe or array. Examples · >>> data = pd.DataFrame({"A": [1, 2, 3, 4], ... "B": pd.Categorical(['a', 'a', 'a', 'b'])}) >>> enc = OrdinalEncoder() >>> trn = enc.fit_transform(data) >>> trn A B 0 1 0 1 2 0 2 3 0 3 4 1 ·
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › ordinal.html
Ordinal — Category Encoders 2.11.1 documentation
class category_encoders.ordinal.OrdinalEncoder(verbose: int = 0, mapping: list[dict[str, str | dict | Series]] | None = None, cols: list[str] = None, drop_invariant: bool = False, return_df: bool = True, handle_unknown: str = 'value', handle_missing: str = 'value', index_start: int = 1, min_group_size: int | float | None = None, min_group_name: str | None = None, combine_min_nan_groups: bool | str | None = None)[source]
🌐
Trainindata
feature-engine.trainindata.com › en › 1.8.x › user_guide › encoding › OrdinalEncoder.html
Ordinal Encoding — 1.8.3
We’ll show how ordinal encoding is implemented by Feature-engine’s OrdinalEncoder() using the Titanic Dataset. Let’s load the dataset and split it into train and test sets: import pandas as pd from sklearn.model_selection import train_test_split from feature_engine.datasets import load_titanic from feature_engine.encoding import OrdinalEncoder X, y = load_titanic( return_X_y_frame=True, handle_missing=True, predictors_only=True, cabin="letter_only", ) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=0, ) print(X_train.head())