You were almost there !
Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.
You have to call transform to transform the data , or use fit_transform which fit and transform the same data.
enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])
or directly
enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])
Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.
As you can see from enc.categories_
[array(['F', 'M'], dtype=object),
array(['A', 'AB', 'B', 'O'], dtype=object),
array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```
Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)
Answer from abcdaire on Stack OverflowYou were almost there !
Basically the fit method, prepare the encoder (fit on your data i.e. prepare the mapping) but don't transform the data.
You have to call transform to transform the data , or use fit_transform which fit and transform the same data.
enc = OrdinalEncoder()
enc.fit(df[["Sex","Blood", "Study"]])
df[["Sex","Blood", "Study"]] = enc.transform(df[["Sex","Blood", "Study"]])
or directly
enc = OrdinalEncoder()
df[["Sex","Blood", "Study"]] = enc.fit_transform(df[["Sex","Blood", "Study"]])
Note: The values won't be the one that you provided, since internally the fit method use numpy.unique which gives result sorted in alphabetic order and not by order of appearance.
As you can see from enc.categories_
[array(['F', 'M'], dtype=object),
array(['A', 'AB', 'B', 'O'], dtype=object),
array(['Biology', 'English', 'Math', 'Science'], dtype=object)]```
Each value in the array is encoded by it's position. (F will be encoded as 0 , M as 1)
I think it is important to point out that this is not an example for an ordinal encoding of variables. Sex, Blood and Study should all not have an ordinal scale (and was also not suggested by the person, who asked the question). Ordinal data has a ranking (see e.g. https://en.wikipedia.org/wiki/Ordinal_data) Those examples here do not have a ranking.
In the case that your variable is a target variable you can use the LabelEncoder.(https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html)
Then you can do something like:
from sklearn.preprocessing import LabelEncoder
for col in ["Sex","Blood", "Study"]:
df[col] = LabelEncoder().fit_transform(df[col])
If your variables are features you should use the Ordinalencoder for accomplishing this. (See comments to my answer).
The naming for the Ordinalencoder is quite unfortunate as "ordinal" is seen from a mathematical and not a statistical naming perspective.
More on the difference between ordinal- and labelencoder in sklearn: https://datascience.stackexchange.com/questions/39317/difference-between-ordinalencoder-and-labelencoder
You are using the 'mapping' param wrong.
The format should be:
'mapping'param should be alistofdictswhere internaldictsshould contain the keys'col'and'mapping'and in that the'mapping'key should have a list of tuples of format(original_label, encoded_label)as value.
Something like this:
ordinal_cols_mapping = [{
"col":"ExterQual",
"mapping": [
('Ex',5),
('Gd',4),
('TA',3),
('Fa',2),
('Po',1),
('NA',np.nan)
]},
]
Then no need to set the 'cols' param separately. Column names will be used from the 'mapping' param.
Just do this:
encoder = OrdinalEncoder(mapping = ordinal_cols_mapping,
return_df = True)
df_train = encoder.fit_transform(train_data)
Hope that this makes it clear.
In case all of your columns to encode are already pandas categoricals, you can construct a mapping like this.
In [82]:
from category_encoders.ordinal import OrdinalEncoder
import pandas as pd
from pandas.api.types import CategoricalDtype
# define a categorical dtype
platforms = ['android', 'ios', 'amazon']
platform_category = CategoricalDtype(categories=platforms, ordered=False)
# create a dataframe
df = pd.DataFrame([
{'id': 1, 'platform': 'android'},
{'id': 2, 'platform': 'ios'},
{'id': 3, 'platform': 'amazon'},
])
# apply the categorical dtype
df['platform'] = df['platform'].astype(platform_category)
# create a mapping from all categorical columns that can be used with OrdinalEncoder
categorical_columns = list(df.select_dtypes(['category']).columns)
category_mapping = [
{'col': column_name, 'mapping': list(zip(df[column_name].cat.categories, df[column_name].cat.codes))}
for column_name in categorical_columns
]
# pass this as the mapping
cat_encoder = OrdinalEncoder(cols=categorical_columns, mapping=category_mapping)
cat_encoder
Out[82]:
OrdinalEncoder(cols=['platform'], drop_invariant=False,
handle_unknown='impute', impute_missing=True,
mapping=[{'col': 'platform', 'mapping': [('android', 0), ('ios', 1), ('amazon', 2)]}],
return_df=True, verbose=0)
You don't want dummies, you want factors/categories.
Use pandas.factorize:
df['Size_Numerical'] = pd.factorize(df['Size'])[0] + 1
output:
Size Size_Numerical
0 Big 1
1 Medium 2
2 Small 3
I think OneHotEncoding has a similar issue that it expands and creates n-dimensions as labels. You need to use LabelEncoder so that:
from sklearn import preprocessing
le = preprocessing.LabelEncoder()
le.fit(df['Sizes'])
df['Category'] = le.transform(df['Sizes']) + 1
Outputs:
Sizes Category
0 Small 3
1 Medium 2
2 Large 1
If you want to perform this operation using OrdinalEncoder, you can use the categories parameter to specify the ordering.
As follows:
OrdinalEncoder(categories=[['low', 'medium', 'high']]).fit_transform(df[['salary']])
Output:
array([[0.],
[1.],
[1.],
[2.],
[2.]])
As @BENY says, you can stay in pandas and do what you want. factorize is great if "low" appears first, "medium" second and "high" third in the data (as shown in your example). If that's not the case, factorize may not produce what you want.
A possible solution is to create a dictionary that maps salary levels to numbers and use map:
mapper = dict([['low', 1], ['medium', 2], ['high', 3]])
df['salary'] = df['salary'].map(mapper)
Output:
department salary tenure
0 operations 1 5
1 operations 2 6
2 support 2 6
3 logics 3 8
4 sales 3 5