I'm not sure if you ever figured this out but I was trying to find answers on this exact same question and there aren't really any good answers in my opinion. I finally figured it out though. OrdinalEncoder is capable of encoding multiple columns in a dataframe. So, when you instantiate OrdinalEncoder(), you give the categories parameter a list of lists:
enc = OrdinalEncoder(categories=[list_of_values_cat1, list_of_values_cat2, etc])
Specifically, in your example above, you would just put ['low', 'med', 'high'] inside another list:
end = OrdinalEncoder(categories=[['low', 'med', 'high']])
enc.fit_transform(X.loc[:,['animals']])
>>array([[0.],
[1.],
[0.],
[2.],
[0.],
[2.]])
# Now 'low' is correctly mapped to 0, 'med' to 1, and 'high' to 2
To see how you can encode multiple columns with their own individual ordinal values, try this:
# Sample dataframe with 2 ordinal categorical columns: 'temp' and 'place'
categorical_df = pd.DataFrame({'my_id': ['101', '102', '103', '104'],
'temp': ['hot', 'warm', 'cool', 'cold'],
'place': ['third', 'second', 'first', 'second']})
# In the 'temp' column, I want 'cold' to be 0, 'cool' to be 1, 'warm' to be 2, and 'hot' to be 3
# In the 'place' column, I want 'first' to be 0, 'second' to be 1, and 'third' to be 2
temp_categories = ['cold', 'cool', 'warm', 'hot']
place_categories = ['first', 'second', 'third']
# Now, when you instantiate the encoder, both of these lists go in one big categories list:
encoder = OrdinalEncoder(categories=[temp_categories, place_categories])
encoder.fit_transform(categorical_df[['temp', 'place']])
>>array([[3., 2.],
[2., 1.],
[1., 0.],
[0., 1.]])
Answer from fugumagu on Stack ExchangeYou were almost there !
Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.
You have to call transform to transform the data , or use fit_transform which fit and transform the same data.
enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])
or directly
enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])
Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.
As you can see from enc.categories_
[array(['F', 'M'], dtype=object),
array(['A', 'AB', 'B', 'O'], dtype=object),
array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```
Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)
I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.
In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)
Then you can do something like:
from sklearn.preprocessing import LabelEncoder
for col in ["Sex","Blood", "Study"]:
df[col] = LabelEncoder().fit_transform(df[col])
If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).
The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.
More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder