Try:
oneEncoder.fit_transform(df[["Col2"]]).todense()
Suppose we have:
features = pd.DataFrame({"Col2":["a","b","c"]})
Then:
oneEncoder= OneHotEncoder()
oneEncoder.fit_transform(features[["Col2"]]).todense()
matrix([[1., 0., 0.],
[0., 1., 0.],
[0., 0., 1.]])
If you're dealing with a Series object, you may wish to reshape it:
oneEncoder.fit_transform(features.Col2.values.reshape(-1,1)).todense()
matrix([[1., 0., 0.],
[0., 1., 0.],
[0., 0., 1.]])
Dropping todense() method will leave your transformation in a sparse matrix.
And finally, you may alway decode what the columns of your matrix mean by:
oneEncoder.categories_
[array(['a', 'b', 'c'], dtype=object)]
Not surpisingly, they are your unique inputs ordered alphabetically.
Answer from Sergey Bushmanov on Stack OverflowTry:
oneEncoder.fit_transform(df[["Col2"]]).todense()
Suppose we have:
features = pd.DataFrame({"Col2":["a","b","c"]})
Then:
oneEncoder= OneHotEncoder()
oneEncoder.fit_transform(features[["Col2"]]).todense()
matrix([[1., 0., 0.],
[0., 1., 0.],
[0., 0., 1.]])
If you're dealing with a Series object, you may wish to reshape it:
oneEncoder.fit_transform(features.Col2.values.reshape(-1,1)).todense()
matrix([[1., 0., 0.],
[0., 1., 0.],
[0., 0., 1.]])
Dropping todense() method will leave your transformation in a sparse matrix.
And finally, you may alway decode what the columns of your matrix mean by:
oneEncoder.categories_
[array(['a', 'b', 'c'], dtype=object)]
Not surpisingly, they are your unique inputs ordered alphabetically.
you could do directly
categories = pd.get_dummies(features['COL2'])
otherwise you could pass a 2d array instead
oneEncoder.fit_transform( features[['COL2']].values)
python - How to perform OneHotEncoding in Sklearn, getting value error - Stack Overflow
python - sklearn.preprocessing.OneHotEncoder and the way to read it - Stack Overflow
Onehotencoder.fit_transform - Packages & Environments - Anaconda Forum
scikit learn - How to perform one hot encoding on multiple categorical columns - Data Science Stack Exchange
You can go directly to OneHotEncoding now without using the LabelEncoder, and as we move toward version 0.22 many might want to do things this way to avoid warnings and potential errors (see DOCS and EXAMPLES).
Example code 1 where ALL columns are encoded and where the categories are explicitly specified:
import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder
data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]
df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values
countries = np.unique(X[:,0])
names = np.unique(X[:,1])
ohe = OneHotEncoder(categories=[countries, names])
X = ohe.fit_transform(X).toarray()
print (X)
Output for code example 1:
[[1. 0. 0. 0. 0. 1. 0.]
[0. 0. 1. 0. 0. 0. 1.]
[0. 1. 0. 1. 0. 0. 0.]
[0. 0. 1. 0. 1. 0. 0.]]
Example code 2 showing the 'auto' option for specification of categories:
The first 3 columns encode the country names, the last four the personal names.
import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder
data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]
df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values
ohe = OneHotEncoder(categories='auto')
X = ohe.fit_transform(X).toarray()
print (X)
Output for code example 2 (same as for 1):
[[1. 0. 0. 0. 0. 1. 0.]
[0. 0. 1. 0. 0. 0. 1.]
[0. 1. 0. 1. 0. 0. 0.]
[0. 0. 1. 0. 1. 0. 0.]]
Example code 3 where only the first column is one hot encoded:
Now, here's the unique part. What if you only need to One Hot Encode a specific column for your data?
(Note: I've left the last column as strings for easier illustration. In reality it makes more sense to do this WHEN the last column was already numerical).
import pandas as pd
import numpy as np
from sklearn.preprocessing import OneHotEncoder
data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]
df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values
countries = np.unique(X[:,0])
names = np.unique(X[:,1])
ohe = OneHotEncoder(categories=[countries]) # specify ONLY unique country names
tmp = ohe.fit_transform(X[:,0].reshape(-1, 1)).toarray()
X = np.append(tmp, names.reshape(-1,1), axis=1)
print (X)
Output for code example 3:
[[1.0 0.0 0.0 'Pechi']
[0.0 0.0 1.0 'Raj']
[0.0 1.0 0.0 'Sri']
[0.0 0.0 1.0 'Vignesh']]
Below implementation should work well. Note that the input of onehotencoder
fit_transform must not be 1-rank array and also output is sparse and we have used to_array() to expand it.
import pandas as pd
import numpy as np
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder
data= [["AUS", "Sri"],["USA","Vignesh"],["IND", "Pechi"],["USA","Raj"]]
df = pd.DataFrame(data, columns=['Country', 'Name'])
X = df.values
le = LabelEncoder()
X_num = le.fit_transform(X[:,0]).reshape(-1,1)
ohe = OneHotEncoder()
X_num = ohe.fit_transform(X_num)
print (X_num.toarray())
X[:,0] = X_num
print (X)
IIUC, use get_feature_names_out():
import pandas as pd
from sklearn.preprocessing import OneHotEncoder
df = pd.DataFrame({'A': [0, 1, 2], 'B': [3, 1, 0],
'C': [0, 2, 2], 'D': [0, 1, 1]})
ohe = OneHotEncoder()
data = ohe.fit_transform(df)
df1 = pd.DataFrame(data.toarray(), columns=ohe.get_feature_names_out(), dtype=int)
Output:
>>> df
A B C D
0 0 3 0 0
1 1 1 2 1
2 2 0 2 1
>>> df1
A_0 A_1 A_2 B_0 B_1 B_3 C_0 C_2 D_0 D_1
0 1 0 0 0 0 1 1 0 1 0
1 0 1 0 0 1 0 0 1 0 1
2 0 0 1 1 0 0 0 1 0 1
>>> pd.Series(ohe.get_feature_names_out()).str.rsplit('_', 1).str[0]
0 A
1 A
2 A
3 B
4 B
5 B
6 C
7 C
8 D
9 D
dtype: object
An alternative to sklearn's OneHotEncoder is Feature-engine's OneHotEncoder, which returns clearly named dummy variables:
import pandas as pd
from feature_engine.encoding import OneHotEncoder
X = pd.DataFrame(dict(x1 = [1,2,3,4], x2 = ["a", "a", "b", "c"]))
ohe = OneHotEncoder()
ohe.fit(X)
ohe.transform(X)
The previous code returns the following dataframe:
x1 x2_a x2_b x2_c
0 1 1 0 0
1 2 1 0 0
2 3 0 1 0
3 4 0 0 1
The encoded variables are named with the variable name, then underscore, and then the category, so they are very easy to identify.
I leave the link to Feature-engine's OneHotEncoder for more details.
LabelEncoder is not made to transform the data but the target (also known as labels) as explained here. If you want to encode the data you should use OrdinalEncoder.
If you really need to do it this way:
categorical_cols = ['a', 'b', 'c', 'd']
from sklearn.preprocessing import LabelEncoder
# instantiate labelencoder object
le = LabelEncoder()
# apply le on categorical feature columns
data[categorical_cols] = data[categorical_cols].apply(lambda col: le.fit_transform(col))
from sklearn.preprocessing import OneHotEncoder
ohe = OneHotEncoder()
#One-hot-encode the categorical columns.
#Unfortunately outputs an array instead of dataframe.
array_hot_encoded = ohe.fit_transform(data[categorical_cols])
#Convert it to df
data_hot_encoded = pd.DataFrame(array_hot_encoded, index=data.index)
#Extract only the columns that didnt need to be encoded
data_other_cols = data.drop(columns=categorical_cols)
#Concatenate the two dataframes :
data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1)
Otherwise:
I suggest you to use pandas.get_dummies if you want to achieve one-hot-encoding from raw data (without having to use OrdinalEncoder before) :
#categorical data
categorical_cols = ['a', 'b', 'c', 'd']
#import pandas as pd
df = pd.get_dummies(data, columns = categorical_cols)
You can also use drop_first argument to remove one of the one-hot-encoded columns, as some models require.
You can do dummy encoding using Pandas in order to get one-hot encoding as shown below:
import pandas as pd
# Multiple categorical columns
categorical_cols = ['a', 'b', 'c', 'd']
pd.get_dummies(data, columns=categorical_cols)
If you want to do one-hot encoding using sklearn library, you can get it done as shown below:
from sklearn.preprocessing import OneHotEncoder
onehotencoder = OneHotEncoder()
transformed_data = onehotencoder.fit_transform(data[categorical_cols])
# the above transformed_data is an array so convert it to dataframe
encoded_data = pd.DataFrame(transformed_data, index=data.index)
# now concatenate the original data and the encoded data using pandas
concatenated_data = pd.concat([data, encoded_data], axis=1)
If a single column has more than 500 categories, the aforementioned way of one-hot encoding is not a good approach. In this case, we can do one-hot encoding for the top 10 or 20 categories that are occurring most for a particular column. A sample code is shown below:
categorical_cols = ['a', 'b', 'c', 'd']
# Let's say we have a column 'b' which has more than 500 categories.
# Find the top 10 most frequent categories for column 'b'
data.b.value_counts().sort_values(ascending = False).head(20)
# make a list of the most frequent categories of the column
top_10_occurring_cat = [cat for cat in data.b.value_counts().sort_values(ascending = False).head(10).index]
# now make the 10 binary variables
for cat in top_10_occurring_cat:
data[cat] = np.where(data['b'] == cat, 1, 0) # whenever data['b'] == cat replace it with 1 else 0
# This is done for one categorical column, similarly you can repeat for all categorical columns