You you are almost there... Like you said you can add all the columns you want to encode in fit_transform directly.

ohe = OneHotEncoder(categories='auto')
feature_arr = ohe.fit_transform(df[['phone','city']]).toarray()
feature_labels = ohe.categories_

And then you just need to do the following:

feature_labels = np.array(feature_labels).ravel()

Which enables you to name your columns like you wanted:

features = pd.DataFrame(feature_arr, columns=feature_labels)
Answer from MaximeKan on Stack Overflow
Top answer
1 of 4
21

LabelEncoder is not made to transform the data but the target (also known as labels) as explained here. If you want to encode the data you should use OrdinalEncoder.

If you really need to do it this way:

categorical_cols = ['a', 'b', 'c', 'd'] 

from sklearn.preprocessing import LabelEncoder
# instantiate labelencoder object
le = LabelEncoder()

# apply le on categorical feature columns
data[categorical_cols] = data[categorical_cols].apply(lambda col: le.fit_transform(col))    
from sklearn.preprocessing import OneHotEncoder
ohe = OneHotEncoder()

#One-hot-encode the categorical columns.
#Unfortunately outputs an array instead of dataframe.
array_hot_encoded = ohe.fit_transform(data[categorical_cols])

#Convert it to df
data_hot_encoded = pd.DataFrame(array_hot_encoded, index=data.index)

#Extract only the columns that didnt need to be encoded
data_other_cols = data.drop(columns=categorical_cols)

#Concatenate the two dataframes : 
data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1)

Otherwise:

I suggest you to use pandas.get_dummies if you want to achieve one-hot-encoding from raw data (without having to use OrdinalEncoder before) :

#categorical data
categorical_cols = ['a', 'b', 'c', 'd'] 

#import pandas as pd
df = pd.get_dummies(data, columns = categorical_cols)

You can also use drop_first argument to remove one of the one-hot-encoded columns, as some models require.

2 of 4
8

You can do dummy encoding using Pandas in order to get one-hot encoding as shown below:

import pandas as pd

# Multiple categorical columns
categorical_cols = ['a', 'b', 'c', 'd']

pd.get_dummies(data, columns=categorical_cols)

If you want to do one-hot encoding using sklearn library, you can get it done as shown below:

from sklearn.preprocessing import OneHotEncoder
onehotencoder = OneHotEncoder()

transformed_data = onehotencoder.fit_transform(data[categorical_cols])

# the above transformed_data is an array so convert it to dataframe
encoded_data = pd.DataFrame(transformed_data, index=data.index)

# now concatenate the original data and the encoded data using pandas
concatenated_data = pd.concat([data, encoded_data], axis=1)

If a single column has more than 500 categories, the aforementioned way of one-hot encoding is not a good approach. In this case, we can do one-hot encoding for the top 10 or 20 categories that are occurring most for a particular column. A sample code is shown below:

categorical_cols = ['a', 'b', 'c', 'd']

# Let's say we have a column 'b' which has more than 500 categories.
# Find the top 10 most frequent categories for column 'b'
data.b.value_counts().sort_values(ascending = False).head(20)

# make a list of the most frequent categories of the column
top_10_occurring_cat = [cat for cat in data.b.value_counts().sort_values(ascending = False).head(10).index]

# now make the 10 binary variables
for cat in top_10_occurring_cat:
    data[cat] = np.where(data['b'] == cat, 1, 0) # whenever data['b'] == cat replace it with 1 else 0

# This is done for one categorical column, similarly you can repeat for all categorical columns
Discussions

python - OneHotEncoding multiple columns in dataset at once - Stack Overflow
Consider the following simple snipped dataset: 2 Columns: X, Y. Both X and Y columns have only 3 optional category values I want to onehotencode these columns. My try: import numpy as np import p... More on stackoverflow.com
🌐 stackoverflow.com
July 20, 2020
pandas - OneHotEncoder Multiple Columns - Stack Overflow
3 One hot encoding - encode multiple columns as one · 4 How do I use OneHotEncoder on a pandas series of lists? More on stackoverflow.com
🌐 stackoverflow.com
March 6, 2019
python - Using OneHotEncoder in multiple columns with repetead categories amongst columns? - Stack Overflow
For column X, there are three different values [a, b, c] so you need 3 columns to encode them. For instance: ... Note the identity matrix. Let's use pd.get_dummies rather than OneHotEncoder to a better understanding: More on stackoverflow.com
🌐 stackoverflow.com
python - OneHotEncoder - multiple columns at once - Stack Overflow
I have a dataset (200, 600) where 180 - 199 are strings, that I need to convert. I am using OneHotEncoder, but I am stuck on how to convert all string columns at once. ohe = OneHotEncoder() aa = np. More on stackoverflow.com
🌐 stackoverflow.com
Top answer
1 of 7
25
import pandas as pd
df = pd.DataFrame({'name': ['Manie', 'Joyce', 'Ami'],
                   'Org':  ['ABC2', 'ABC1', 'NSV2'],
                   'Dept': ['Finance', 'HR', 'HR']        
        })


df_2 = pd.get_dummies(df,drop_first=True)

test:

print(df_2)
   Dept_HR  Org_ABC2  Org_NSV2  name_Joyce  name_Manie
0        0         1         0           0           1
1        1         0         0           1           0
2        1         0         1           0           0 

UPDATE regarding your error with pd.get_dummies(X, columns =[1:]:

Per the documentation page, the columns parameter takes "Column Names". So the following code would work:

df_2 = pd.get_dummies(df, columns=['Org', 'Dept'], drop_first=True)

output:

    name  Org_ABC2  Org_NSV2  Dept_HR
0  Manie         1         0        0
1  Joyce         0         0        1
2    Ami         0         1        1

If you really want to define your columns positionally, you could do it this way:

column_names_for_onehot = df.columns[1:]
df_2 = pd.get_dummies(df, columns=column_names_for_onehot, drop_first=True)
2 of 7
5

I use my own template for doing that:

from sklearn.base import TransformerMixin
import pandas as pd
import numpy as np
class DataFrameEncoder(TransformerMixin):

    def __init__(self):
        """Encode the data.

        Columns of data type object are appended in the list. After 
        appending Each Column of type object are taken dummies and 
        successively removed and two Dataframes are concated again.

        """
    def fit(self, X, y=None):
        self.object_col = []
        for col in X.columns:
            if(X[col].dtype == np.dtype('O')):
                self.object_col.append(col)
        return self

    def transform(self, X, y=None):
        dummy_df = pd.get_dummies(X[self.object_col],drop_first=True)
        X = X.drop(X[self.object_col],axis=1)
        X = pd.concat([dummy_df,X],axis=1)
        return X

And for using this code just put this template in current directory with filename let's suppose CustomeEncoder.py and type in your code:

from customEncoder import DataFrameEncoder
data = DataFrameEncoder().fit_transormer(data)

And all the object type data removed, Encoded, removed first and joined together to give the final desired output.
PS: That the input file to this template is Pandas Dataframe.

🌐
datagy
datagy.io › home › python posts › one-hot encoding in scikit-learn with onehotencoder
One-Hot Encoding in Scikit-Learn with OneHotEncoder • datagy
April 14, 2024 - We used the remainder='passthrough' parameter to specify that all other columns should be left untouched. We then applied the .fit_transform() method to our DataFrame. ... In the next section, you’ll learn how to use the make_column_transformer() function to one-hot encode multiple columns ...
🌐
Medium
medium.com › @sami.yousuf.azad › one-hot-encoding-with-pandas-dataframe-49a304e8507a
One hot encoding with Pandas dataframe | by Yousuf Azad Sami | Medium
April 27, 2022 - Way 3: One hot encoding with sklearn OneHotEncoding class with Pandas dataframe (nicer apporach) # One-hot encoding multiple columns from sklearn.preprocessing import OneHotEncoder from sklearn.compose import make_column_transformer from seaborn import load_dataset import pandas as pddf = load_dataset('penguins') # taking a smaller subset of the df, so it is easy to work with df = df[['island', 'sex', 'body_mass_g']] df = df.dropna() display(df)transformer = make_column_transformer( (OneHotEncoder(), ['island', 'sex']), remainder='passthrough')transformed = transformer.fit_transform(df) transformed_df = pd.DataFrame( transformed, columns=transformer.get_feature_names()) display(transformed_df.head())
🌐
Stack Overflow
stackoverflow.com › questions › 63001203 › onehotencoding-multiple-columns-in-dataset-at-once
python - OneHotEncoding multiple columns in dataset at once - Stack Overflow
July 20, 2020 - import numpy as np import pandas as pd from sklearn.impute import SimpleImputer from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder dataset = pd.read_csv('./forestfires.csv') X = dataset.iloc[:, :-1].values y = dataset.iloc[:, -1].values imputer = SimpleImputer(missing_values=np.nan, strategy='mean') imputer.fit(X[:, 4:12]) X[:, 4:12] = imputer.transform(X[:, 4:12]) ct = ColumnTransformer(transformers=[('encoder', OneHotEncoder(), [0])], remainder='passthrough') X = np.array(ct.fit_transform(X)) At the current state, it only encodes the first column. I can't really understand sklearn documentation for this function ColumnTransformer. How would I select multiple columns to encode all at once?
Find elsewhere
🌐
Coding Infinite
codinginfinite.com › home › one hot encoding in python
One Hot Encoding in Python - Coding Infinite
June 24, 2023 - You can also perform one hot encoding on multiple features using a single OneHotEncoder object. For this, you can simply pass the 2-D list containing all the rows and columns as input to the fit() method as shown below.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
The input to this transformer should be an array-like of integers or strings, denoting the values taken on by categorical (discrete) features. The features are encoded using a one-hot (aka ‘one-of-K’ or ‘dummy’) encoding scheme. This creates a binary column for each category and returns a sparse matrix or dense array (depending on the sparse_output parameter).
🌐
Stack Overflow
stackoverflow.com › questions › 55008094 › onehotencoder-multiple-columns
pandas - OneHotEncoder Multiple Columns - Stack Overflow
March 6, 2019 - for b in batches: batch_ohe = pd.get_dummies(b, columns=['label_col']) ohe = pd.concat([ohe, batch_ohe], axis=0) ohe = ohe.fillna(0)
🌐
Apache
nightlies.apache.org › flink › flink-ml-docs-release-2.0 › docs › operators › feature › onehotencoder
One Hot Encoder | Apache Flink Machine Learning Library
One Hot Encoder # One-hot encoding ... which expect continuous features, such as Logistic Regression, to use categorical features. OneHotEncoder can transform multiple columns, returning an one-hot-encoded output vector column for each input column....
🌐
Stack Overflow
stackoverflow.com › questions › 69963760 › onehotencoder-multiple-columns-at-once
python - OneHotEncoder - multiple columns at once - Stack Overflow
ohe = OneHotEncoder() aa = np.array( X[:,199]).reshape(-1, 1) one_hot_encoder = ohe.fit(aa) ohe_coded = ohe.transform(aa).toarray() This is what I have for now. I can convert one column with X[:,199], but I have no idea how to convert multiple at once, or if it is even possible in some simple way.
🌐
Sdv
docs.sdv.dev › rdt › transformers-glossary › categorical › onehotencoder
OneHotEncoder | RDT
The OneHotEncoder transforms data that represents unordered, categorical values into multiple columns -- one for each category. These new columns have 0 and 1 values.
🌐
GitHub
github.com › ksator › Machine_Learning_with_Python › wiki › split-a-column-into-multiple-columns-(One-Hot-Encode)
split a column into multiple columns (One Hot Encode)
June 30, 2019 - OneHotEncoder takes a column and split it into multiple columns OneHotEncoder splits the feature country into 3 features (Germany, France, and Spain) which are all binary (0 or 1)
Author: ksator
🌐
KNIME Community
forum.knime.com › knime analytics platform
One Hot Encoding with multiple columns - KNIME Analytics Platform - KNIME Community Forum
September 1, 2023 - How can I update a dataset that was already been hot encoded once? For example, there are two columns A and B both with similar data that needs to be hot encoded. I understand using the One to Many Node to hot encode and execute hot encoding on Column A. But how can I then take the data that is in Column B to update the newly formed hot encoded columns?
🌐
KDnuggets
kdnuggets.com › 2023 › 07 › pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - For the categorical column, we can break it down into multiple columns. For this, we use pandas.get_dummies() method.