From the docs, OneHotEncoder can take a dataframe and convert the categorical columns into the vectors you see. LabelEncoder takes a Series(your y / dependent variable) and generates new labels.

OnHotEncoder's usage: fit_transform(X,[y])

LabelEncoder's usage: fit_transform(y)

That's why it'll tell you: "fit_transform() takes 2 positional arguments but 3 were given"

Just call LabelEncoder fit_transform on the y directly if you really want to use it. Here is a similar question: How to use sklearn Column Transformer?

Here are the docs:

  1. https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html
  2. https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html
Answer from joesph nguyen on Stack Overflow
🌐
Towards Data Science
towardsdatascience.com › home › latest › columntransformer in scikit for labelencoding and onehotencoding in machine learning
ColumnTransformer in SciKit for LabelEncoding and OneHotEncoding in Machine Learning | Towards Data Science
January 23, 2025 - The developers of the library might have realised that people use LabelEncoding and OneHotEncoding very frequently. So they decided to come up with a new library called the ColumnTransformer, which will basically combine LabelEncoding and OneHotEncoding into just one line of code.
🌐
Kaggle
kaggle.com › code › pritampatil004 › label-encoder-one-hot-encoder-columntransformer
Label Encoder, One Hot Encoder, ColumnTransformer
Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
Neuraxle
neuraxle.org › stable › examples › getting_started › plot_label_encoder_across_multiple_columns.html
Create label encoder across multiple columns — Neuraxle 0.8.0 documentation
""" p = FlattenForEach(LabelEncoder(), then_unflatten=True) p, predicted_output = p.fit_transform(df.values) expected_output = np.array([ [6, 7, 6, 8, 7, 7], [1, 3, 0, 1, 5, 3], [4, 2, 2, 4, 4, 2] ]).transpose() assert np.array_equal(predicted_output, expected_output) def _apply_different_encoders_to_columns(): """ One standalone LabelEncoder will be applied on the pets, and another one will be shared for the columns owner and location. """ p = ColumnTransformer([ # A different encoder will be used for column 0 with name "pets": (0, FlattenForEach(LabelEncoder(), then_unflatten=True)), # A sha
Top answer
1 of 9
32

It is a bit strange to encode continuous data as Salary. It makes no sense unless you have binned your salary to certain ranges/categories. If I were you I would do:

import pandas as pd
import numpy as np

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder



numeric_features = ['Salary']
numeric_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler())])

categorical_features = ['Age','Country']
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='constant', fill_value='missing')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))])

preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, numeric_features),
        ('cat', categorical_transformer, categorical_features)])

from here you can pipe it with a classifier e.g.

clf = Pipeline(steps=[('preprocessor', preprocessor),
                  ('classifier', LogisticRegression(solver='lbfgs'))])  
                  

Use it as so:

clf.fit(X_train,y_train)

this will apply the preprocessor and then pass transformed data to the predictor.

Updates:

If we want to select data types on the fly, we can modify our preprocessor to use column selector by data dtypes:

from sklearn.compose import make_column_selector as selector

preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, selector(dtype_include="numeric")),
        ('cat', categorical_transformer, selector(dtype_include="category"))])

Using GridSearch

param_grid = {
    'preprocessor__num__imputer__strategy': ['mean', 'median'],
    'classifier__C': [0.1, 1.0, 10, 100],
    'classifier__solver': ['lbfgs', 'sag'],
}

grid_search = GridSearchCV(clf, param_grid, cv=10)
grid_search.fit(X_train,y_train)

Getting names of features


preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, selector(dtype_include="numeric")),
        ('cat', categorical_transformer, selector(dtype_include="category"))],
    verbose_feature_names_out=False, # added this line
)

# now we can access feature names with

clf[:-1]. get_feature_names_out() # step before estimator

2 of 9
13

I think the poster is not trying to transform the Age and Salary. From the documentation (https://scikit-learn.org/stable/modules/generated/sklearn.compose.make_column_transformer.html), you ColumnTransformer (and make_column_transformer) only columns specified in the transformer (i.e., [0] in your example). You should set remainder="passthrough" to get the rest of the columns. In other words:

preprocessor = make_column_transformer( (OneHotEncoder(),[0]),remainder="passthrough")
x = preprocessor.fit_transform(x)
🌐
Contactsunny
blog.contactsunny.com › home › data science
ColumnTransformer in SciKit for LabelEncoding and OneHotEncoding in Machine Learning | The ContactSunny Blog
November 7, 2019 - The developers of the library might have realised that people use LabelEncoding and OneHotEncoding very frequently. So they decided to come up with a new library called the ColumnTransformer, which will basically combine LabelEncoding and OneHotEncoding into just one line of code.
Top answer
1 of 16
609

You can easily do this though,

df.apply(LabelEncoder().fit_transform)

EDIT2:

In scikit-learn 0.20, the recommended way is

OneHotEncoder().fit_transform(df)

as the OneHotEncoder now supports string input. Applying OneHotEncoder only to certain columns is possible with the ColumnTransformer.

EDIT:

Since this original answer is over a year ago, and generated many upvotes (including a bounty), I should probably extend this further.

For inverse_transform and transform, you have to do a little bit of hack.

from collections import defaultdict
d = defaultdict(LabelEncoder)

With this, you now retain all columns LabelEncoder as dictionary.

# Encoding the variable
fit = df.apply(lambda x: d[x.name].fit_transform(x))

# Inverse the encoded
fit.apply(lambda x: d[x.name].inverse_transform(x))

# Using the dictionary to label future data
df.apply(lambda x: d[x.name].transform(x))

MOAR EDIT:

Using Neuraxle's FlattenForEach step, it's possible to do this as well to use the same LabelEncoder on all the flattened data at once:

FlattenForEach(LabelEncoder(), then_unflatten=True).fit_transform(df)

For using separate LabelEncoders depending for your columns of data, or if only some of your columns of data needs to be label-encoded and not others, then using a ColumnTransformer is a solution that allows for more control on your column selection and your LabelEncoder instances.

2 of 16
132

As mentioned by larsmans, LabelEncoder() only takes a 1-d array as an argument. That said, it is quite easy to roll your own label encoder that operates on multiple columns of your choosing, and returns a transformed dataframe. My code here is based in part on Zac Stewart's excellent blog post found here.

Creating a custom encoder involves simply creating a class that responds to the fit(), transform(), and fit_transform() methods. In your case, a good start might be something like this:

import pandas as pd
from sklearn.preprocessing import LabelEncoder
from sklearn.pipeline import Pipeline

# Create some toy data in a Pandas dataframe
fruit_data = pd.DataFrame({
    'fruit':  ['apple','orange','pear','orange'],
    'color':  ['red','orange','green','green'],
    'weight': [5,6,3,4]
})

class MultiColumnLabelEncoder:
    def __init__(self,columns = None):
        self.columns = columns # array of column names to encode

    def fit(self,X,y=None):
        return self # not relevant here

    def transform(self,X):
        '''
        Transforms columns of X specified in self.columns using
        LabelEncoder(). If no columns specified, transforms all
        columns in X.
        '''
        output = X.copy()
        if self.columns is not None:
            for col in self.columns:
                output[col] = LabelEncoder().fit_transform(output[col])
        else:
            for colname,col in output.iteritems():
                output[colname] = LabelEncoder().fit_transform(col)
        return output

    def fit_transform(self,X,y=None):
        return self.fit(X,y).transform(X)

Suppose we want to encode our two categorical attributes (fruit and color), while leaving the numeric attribute weight alone. We could do this as follows:

MultiColumnLabelEncoder(columns = ['fruit','color']).fit_transform(fruit_data)

Which transforms our fruit_data dataset from

to

Passing it a dataframe consisting entirely of categorical variables and omitting the columns parameter will result in every column being encoded (which I believe is what you were originally looking for):

MultiColumnLabelEncoder().fit_transform(fruit_data.drop('weight',axis=1))

This transforms

to

.

Note that it'll probably choke when it tries to encode attributes that are already numeric (add some code to handle this if you like).

Another nice feature about this is that we can use this custom transformer in a pipeline:

encoding_pipeline = Pipeline([
    ('encoding',MultiColumnLabelEncoder(columns=['fruit','color']))
    # add more pipeline steps as needed
])
encoding_pipeline.fit_transform(fruit_data)
🌐
GeeksforGeeks
geeksforgeeks.org › label-encoding-across-multiple-columns-in-scikit-learn
Label Encoding Across Multiple Columns in Scikit-Learn - GeeksforGeeks
August 22, 2024 - Scikit-Learn's ColumnTransformer allows for more complex transformations, applying different preprocessing steps to different columns.
Find elsewhere
Top answer
1 of 4
6

LabelEncoding your features is a bad practice

You should avoid using LabelEncoder to encode your input features! Don't believe me? Here's what scikit-learn's official documentation for LabelEncoder says:

This transformer should be used to encode target values, i.e. y, and not the input X.

That's why it's called LabelEncoding.

Why you shouldn't use LabelEncoder to encode features.

This encoder simply makes a mapping of a feature's unique values to integers. For example, let's say we want to encode a feature called shirt color, which represents the color of the shirt someone's wearing. This feature has values ['red', 'green', 'blue', ...]. If you encode these into integers, i.e. [1, 2, 3, ...], you might confuse your model by because you have now given relationships to these values that don't exist in the real world, e.g. red < greed < blue or red + green = blue. This type of feature is called nominal and preferably should be one-hot encoded.

There are features however, where you might want to map their values to integers. These are called ordinal. For example, the feature rating, which has values ['bad', 'good', 'excellent', ...]. By mapping these to integers you actually preserve the relationsips these values hold in the real world, e.g. bad < good < excellent. There is a catch to this however, in order to do the above, you need to map each value with a specific integer (e.g. we can't map 'good' -> 1, 'bad' -> 2, 'excellent' -> 3, because that doesn't preserve the real-world relationship of these values). The computer doesn't know which number to map to each value, though, so if you use LabelEncoder even on ordinal variables, it most likely won't generate the correct encoding.

How to properly encode ordinal features

A more proper way of encoding ordinal variables is manually choosing the mapping. This requires more work and isn't as elegant as a one-liner that encodes all values, but is the only correct way. Let's see how we can do this in pandas.

custom_mapping = {'bad': 1, 'good': 2, 'excellent': 3}


df['rating'] = df['rating'].map(custom_mapping)

Obviously this needs to be done for each ordinal feature.


At this point I think it's clear that I strongly recommend against using LabelEncoder, but if you still want to do it at least do it correctly.

If you still want to use LabelEncoding

While both answers by @ggordon and @Anan Srivastava will do what you want, they don't have much value in practice. The problem isthat by not bounding the fitted LabelEncoder to a variable, you are loosing the mapping from categories to numbers. If you want to predict on future data, you won't know which number to encode each category with.

Expanding upon @ggordon's answer

columns_to_be_encoded = [...]  # list of column names you want encoded

# Instantiate the encoders
encoders = {column: LabelEncoder() for column in columns_to_be_encoded}

for column in columns_to_be_encoded:
    df[column] = encoders[column].fit_transform(df[column])

This way you have a dictionary of fitted encoders so that you can reuse the same encoding if you wish.

2 of 4
1

+1 to @Djib2011: LabelEncoder is for the targets/labels, not for other data columns. Also, I agree that generally you don't want an ordinal encoding, when one-hot is more faithful to the original data.

But, if you do want to ordinal encode, there's a better way: OrdinalEncoder. And if you want it to only apply to certain columns, you can use ColumnTransformer, e.g.

encoder = ColumnTransformer(
              transformers=[('ord_enc', OrdinalEncoder(), object_columns)],
              remainder='passthrough'
              )
df_enc = encoder.fit_transform(df)
🌐
Stack Exchange
datascience.stackexchange.com › questions › 110241 › not-able-to-encode-multiple-categorical-columns-at-once
machine learning - Not able to encode multiple categorical columns at once - Data Science Stack Exchange
Difference between OrdinalEncoder and LabelEncoder (4 answers) Closed 11 months ago. I have written the following code for encoding categorical features of the dataframe( named 't') - from sklearn.compose import ColumnTransformer categorical_columns = ['warehouse_ID', 'Product_Type','month', 'is_weekend', 'is_warehouse_closed'] transformer = ColumnTransformer(transformers= [('le',preprocessing.LabelEncoder() ,categorical_columns)],remainder= 'passthrough') Train_transform = transformer.fit_transform(t) But it is showing this error - TypeError: fit_transform() takes 2 positional arguments but 3 were given ·
🌐
Kaggle
kaggle.com › getting-started › 146568
LabelEncoder is not working with pipelines | Kaggle
Hello all, I have an issue when I use LabelEncoder with sklearn pipelines. It throw me this error: TypeError: fit_transform() takes 2 positional arguments ...
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - These cannot directly go into the column transformer? 2. I assume we can do multiple splits of the dataframe too? Not just 2? Like if we want to LabelEncode some cat columns, ‘most_frequent’ impute & OneHotEncode other cat columns, impute ‘median’ for a few num columns and impute ‘mean’ for others followed by scaling them?
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 21604
`LabelEncoder` uses `y`. Therefore breaking transformers. · Issue #21604 · scikit-learn/scikit-learn
November 9, 2021 - If you try to pipeline a the LabelEncoder in like in the example below it will break because LabelEncoder uses fit_transform(y) rather than fit_transform(X, y=None). preprocessor = ColumnTransformer( transformers=[ ('categorical', LabelEncoder(), [c for c in x.columns if x[c].dtype == 'object']) ], remainder='passthrough', ) embedder = RandomTreesEmbedding() self.embedding_model_ = Pipeline( steps=[ ('cat-preprocessing', preprocessor), ('embedder', embedder) ] ) x_prime = self.embedding_model_.fit_transform(x, y=None) It's a trivial example to fix and I can do it myself if you'd like?
Author: scikit-learn
🌐
Stack Overflow
stackoverflow.com › questions › 63187686 › apply-label-encoder-for-multiple-columns-in-train-and-test-dataset
python - apply label encoder for multiple columns in train and test dataset - Stack Overflow
July 31, 2020 - You can use ColumnTransformer and Pipeline to encode all categorical columns. After you can also add transformation for the numerical columns. categorical_features = ['A0', 'A1', 'A2', 'A3', 'A4', 'A5', 'A6', 'A8'] categorical_transformer = Pipeline(steps=[('le', LabelEncoder())]) preprocessor = ColumnTransformer(transformers=[('cat', categorical_transformer, categorical_features)]) pipeline = Pipeline(steps=[('preprocessor', preprocessor)]) pipeline.fit(X_train) Share ·
Top answer
1 of 1
1

As you can see in your output, when you are trying to inverse_transform, it seems that the code is only using the information he obtained for the last column "name". You can see that because now, all the rows of your columns have values related to names. You should have one LabelEncoder() for each column.

The key here is to have one LabelEncoder fitted for each different column. To do this, I recommend you save them in a dictionary:

to_encode = ["animal", "color", "sex", "name"]
d={}
for col in to_encode:
    d[col]=preprocessing.LabelEncoder().fit(df[col]) #For each column, we create one instance in the dictionary. Take care we are only fitting now.

If we print the dictionary now, we will obtain something like this:

{'animal': LabelEncoder(),
 'color': LabelEncoder(),
 'sex': LabelEncoder(),
 'name': LabelEncoder()}

As we can see, for each column we want to transform, we have his LabelEncoder() information. This means, for example, that for the animal LabelEncoder it saves that 0 is equal to bird, 1 equal to cat, ... And the same for each column.

Once we have every column fitted, we can proceed to transform, and then, if we want to inverse_transform. The only thing to be aware is that every transform/inverse_transform have to use the corresponding LabelEncoder of this column.

Here we transform:

for col in to_encode:
    df[col] = d[col].transform(df[col]) #Be aware we are using the dictionary

df

animal  color   age pet sex name
0   2   0   1   1   1   2
1   0   1   10  0   1   1
2   2   2   3   1   0   3
3   1   0   6   1   0   0

And, once the df is transformed, we can inverse_transform:

for col in to_encode:
    df[col] = d[col].inverse_transform(df[col])

df

animal  color   age pet sex name
0   Dog Black   1   1   m   Rex
1   Bird Blue   10  0   m   Gizmo
2   Dog Brown   3   1   f   Suzy
3   Cat Black   6   1   f   Boo

One interesting idea could be using ColumnTransformer, but unfortunately, it doesn't suppport inverse_transform().