You can easily do this though,

df.apply(LabelEncoder().fit_transform)

EDIT2:

In scikit-learn 0.20, the recommended way is

OneHotEncoder().fit_transform(df)

as the OneHotEncoder now supports string input. Applying OneHotEncoder only to certain columns is possible with the ColumnTransformer.

EDIT:

Since this original answer is over a year ago, and generated many upvotes (including a bounty), I should probably extend this further.

For inverse_transform and transform, you have to do a little bit of hack.

from collections import defaultdict
d = defaultdict(LabelEncoder)

With this, you now retain all columns LabelEncoder as dictionary.

# Encoding the variable
fit = df.apply(lambda x: d[x.name].fit_transform(x))

# Inverse the encoded
fit.apply(lambda x: d[x.name].inverse_transform(x))

# Using the dictionary to label future data
df.apply(lambda x: d[x.name].transform(x))

MOAR EDIT:

Using Neuraxle's FlattenForEach step, it's possible to do this as well to use the same LabelEncoder on all the flattened data at once:

FlattenForEach(LabelEncoder(), then_unflatten=True).fit_transform(df)

For using separate LabelEncoders depending for your columns of data, or if only some of your columns of data needs to be label-encoded and not others, then using a ColumnTransformer is a solution that allows for more control on your column selection and your LabelEncoder instances.

Answer from Napitupulu Jon on Stack Overflow
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ stable โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.LabelEncoder.html
LabelEncoder โ€” scikit-learn 1.9.1 documentation
This transformer should be used to encode target values, i.e. y, and not the input X. Read more in the User Guide. Added in version 0.12. ... Holds the label for each class.
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ machine learning โ€บ ml-label-encoding-of-datasets-in-python
Label Encoding in Python - GeeksforGeeks
... from sklearn.preprocessing import LabelEncoder import pandas as pd data = pd.DataFrame({ 'Fruit': ['Apple', 'Banana', 'Orange', 'Apple', 'Orange', 'Banana'], 'Price': [1.2, 0.5, 0.8, 1.3, 0.9, 0.6] }) le = LabelEncoder() data['Fruit_Encoded'] ...
Published: June 11, 2026
๐ŸŒ
Medium
medium.com โ€บ @prathik.codes โ€บ labelencoder-in-scikit-learn-c1b7bccec412
LabelEncoder in scikit-learn. ML Quickies #24 | by Prathik C | Medium
October 9, 2025 - from sklearn.preprocessing import LabelEncoder # Example categorical labels y = ["cat", "dog", "cat", "bird"] # Create and fit the encoder le = LabelEncoder() y_encoded = le.fit_transform(y) print("Encoded labels:", y_encoded) Output: Encoded ...
๐ŸŒ
Medium
medium.com โ€บ @kattilaxman4 โ€บ a-practical-guide-for-python-label-encoding-with-python-fb0b0e7079c5
A Practical Guide for Python: Label Encoding with Python | by Kattilaxman | Medium
October 25, 2023 - Hereโ€™s an example of label encoding with a dataset loaded from a CSV file: import pandas as pd ยท from sklearn.preprocessing import LabelEncoder ยท # Load the dataset ยท data = pd.read_csv(โ€˜your_data.csvโ€™) # Initialize the label encoder ยท label_encoder = LabelEncoder() # Apply label encoding to a specific column ยท
๐ŸŒ
VitalFlux
vitalflux.com โ€บ home โ€บ data science โ€บ sklearn labelencoder example โ€“ single & multiple columns
Sklearn LabelEncoder Example - Single & Multiple Columns
September 13, 2024 - Label encoding technique is implemented using sklearn LabelEncoder. You would learn the concept and usage of sklearn LabelEncoder using code examples, for handling encoding labels related to categorical features of single and multiple columns in Python Pandas Dataframe.
๐ŸŒ
Great Learning
mygreatlearning.com โ€บ blog โ€บ ai and machine learning โ€บ label encoding in python
What is Label Encoding in Python | Great Learning
December 18, 2024 - Getting back to our example, in Python, this process can be implemented using 2 approaches as follows: ... As one-hot encoding is also part of data preprocessing, hence we will take an help of preprocessing module from sklearn package and them import OneHotEncoder class as below
Find elsewhere
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.21 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.LabelEncoder.html
sklearn.preprocessing.LabelEncoder โ€” scikit-learn 0.21.3 documentation
Encode labels with value between 0 and n_classes-1. Read more in the User Guide. ... LabelEncoder can be used to normalize labels. >>> from sklearn import preprocessing >>> le = preprocessing.LabelEncoder() >>> le.fit([1, 2, 2, 6]) LabelEncoder() >>> le.classes_ array([1, 2, 6]) >>> le.transform([1, 1, 2, 6]) array([0, 0, 1, 2]...) >>> le.inverse_transform([0, 0, 1, 2]) array([1, 1, 2, 6])
Top answer
1 of 16
609

You can easily do this though,

df.apply(LabelEncoder().fit_transform)

EDIT2:

In scikit-learn 0.20, the recommended way is

OneHotEncoder().fit_transform(df)

as the OneHotEncoder now supports string input. Applying OneHotEncoder only to certain columns is possible with the ColumnTransformer.

EDIT:

Since this original answer is over a year ago, and generated many upvotes (including a bounty), I should probably extend this further.

For inverse_transform and transform, you have to do a little bit of hack.

from collections import defaultdict
d = defaultdict(LabelEncoder)

With this, you now retain all columns LabelEncoder as dictionary.

# Encoding the variable
fit = df.apply(lambda x: d[x.name].fit_transform(x))

# Inverse the encoded
fit.apply(lambda x: d[x.name].inverse_transform(x))

# Using the dictionary to label future data
df.apply(lambda x: d[x.name].transform(x))

MOAR EDIT:

Using Neuraxle's FlattenForEach step, it's possible to do this as well to use the same LabelEncoder on all the flattened data at once:

FlattenForEach(LabelEncoder(), then_unflatten=True).fit_transform(df)

For using separate LabelEncoders depending for your columns of data, or if only some of your columns of data needs to be label-encoded and not others, then using a ColumnTransformer is a solution that allows for more control on your column selection and your LabelEncoder instances.

2 of 16
132

As mentioned by larsmans, LabelEncoder() only takes a 1-d array as an argument. That said, it is quite easy to roll your own label encoder that operates on multiple columns of your choosing, and returns a transformed dataframe. My code here is based in part on Zac Stewart's excellent blog post found here.

Creating a custom encoder involves simply creating a class that responds to the fit(), transform(), and fit_transform() methods. In your case, a good start might be something like this:

import pandas as pd
from sklearn.preprocessing import LabelEncoder
from sklearn.pipeline import Pipeline

# Create some toy data in a Pandas dataframe
fruit_data = pd.DataFrame({
    'fruit':  ['apple','orange','pear','orange'],
    'color':  ['red','orange','green','green'],
    'weight': [5,6,3,4]
})

class MultiColumnLabelEncoder:
    def __init__(self,columns = None):
        self.columns = columns # array of column names to encode

    def fit(self,X,y=None):
        return self # not relevant here

    def transform(self,X):
        '''
        Transforms columns of X specified in self.columns using
        LabelEncoder(). If no columns specified, transforms all
        columns in X.
        '''
        output = X.copy()
        if self.columns is not None:
            for col in self.columns:
                output[col] = LabelEncoder().fit_transform(output[col])
        else:
            for colname,col in output.iteritems():
                output[colname] = LabelEncoder().fit_transform(col)
        return output

    def fit_transform(self,X,y=None):
        return self.fit(X,y).transform(X)

Suppose we want to encode our two categorical attributes (fruit and color), while leaving the numeric attribute weight alone. We could do this as follows:

MultiColumnLabelEncoder(columns = ['fruit','color']).fit_transform(fruit_data)

Which transforms our fruit_data dataset from

to

Passing it a dataframe consisting entirely of categorical variables and omitting the columns parameter will result in every column being encoded (which I believe is what you were originally looking for):

MultiColumnLabelEncoder().fit_transform(fruit_data.drop('weight',axis=1))

This transforms

to

.

Note that it'll probably choke when it tries to encode attributes that are already numeric (add some code to handle this if you like).

Another nice feature about this is that we can use this custom transformer in a pipeline:

encoding_pipeline = Pipeline([
    ('encoding',MultiColumnLabelEncoder(columns=['fruit','color']))
    # add more pipeline steps as needed
])
encoding_pipeline.fit_transform(fruit_data)
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ label-encoding-across-multiple-columns-in-scikit-learn
Label Encoding Across Multiple Columns in Scikit-Learn - GeeksforGeeks
August 22, 2024 - To label encode multiple columns, you need to apply the LabelEncoder to each column individually. Here is an example of how you can do this: ... import pandas as pd from sklearn.preprocessing import LabelEncoder df = pd.DataFrame({ 'team': ['A', 'A', 'B', 'B', 'B', 'C', 'C', 'D'], 'position': ['G', 'F', 'G', 'F', 'F', 'G', 'G', 'F'], 'all_star': ['Y', 'N', 'Y', 'Y', 'Y', 'N', 'Y', 'N'], 'points': [11, 8, 10, 6, 6, 5, 9, 12] }) le_team = LabelEncoder() le_position = LabelEncoder() le_all_star = LabelEncoder() df['team'] = le_team.fit_transform(df['team']) df['position'] = le_position.fit_transform(df['position']) df['all_star'] = le_all_star.fit_transform(df['all_star']) print(df)
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.18 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.LabelEncoder.html
sklearn.preprocessing.LabelEncoder โ€” scikit-learn 0.18.2 documentation
Encode labels with value between 0 and n_classes-1. Read more in the User Guide. ... LabelEncoder can be used to normalize labels. >>> from sklearn import preprocessing >>> le = preprocessing.LabelEncoder() >>> le.fit([1, 2, 2, 6]) LabelEncoder() >>> le.classes_ array([1, 2, 6]) >>> le.transform([1, 1, 2, 6]) array([0, 0, 1, 2]...) >>> le.inverse_transform([0, 0, 1, 2]) array([1, 1, 2, 6])
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ 0.17 โ€บ modules โ€บ generated โ€บ sklearn.preprocessing.LabelEncoder.html
sklearn.preprocessing.LabelEncoder โ€” scikit-learn 0.17.1 documentation
>>> le = preprocessing.LabelEncoder() >>> le.fit(["paris", "paris", "tokyo", "amsterdam"]) LabelEncoder() >>> list(le.classes_) ['amsterdam', 'paris', 'tokyo'] >>> le.transform(["tokyo", "tokyo", "paris"]) array([2, 2, 1]...) >>> list(le.inverse_transform([2, 2, 1])) ['tokyo', 'tokyo', 'paris']
๐ŸŒ
Medium
medium.com โ€บ @vtalladin06 โ€บ label-encoding-in-python-ec0bbe6f0e0f
Label Encoding in Python. Introduction: | by Tahseen Alladin | Medium
February 1, 2024 - from sklearn.preprocessing import LabelEncoder # Example data data = {'ordinal_category': ['low', 'medium', 'high', 'medium', 'low']} # Create a DataFrame df = pd.DataFrame(data) # Create a LabelEncoder le = LabelEncoder() # Fit the encoder ...
๐ŸŒ
Analytics Vidhya
analyticsvidhya.com โ€บ home โ€บ one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - For decision tree algorithms like random forest, even if the categorical variable is nominal, it doesn't seem to have a problem with being represented as ordinal using label or ordinal encoder. Seems unintuitive. Can someone please explain? H20 infact says that they use enum encoding where the categories are given a numerical value , but the numbers themselves are irrelevant(hence not imposing ordinality on nominal variables). But their classification performance doesn't seem to be much different from sklearn's random forest classifier using ordinal encoder)
๐ŸŒ
Snyk
snyk.io โ€บ advisor โ€บ sklearn โ€บ functions โ€บ sklearn.preprocessing.labelencoder
How to use the sklearn.preprocessing.LabelEncoder function in sklearn | Snyk
a = alg.split('+') alg_list = [alg_list[alg_abbreviation.index(i)] for i in a] num_cores = multiprocessing.cpu_count() os.chdir(path + 'data/') print "--- Data preprocessing ---" df = pd.read_csv(data, sep='\t') df = df.sort('fname', ascending=True).reset_index(drop=True) df['label'] = pd.read_csv(label, sep='\t', header=None) #for i in xrange(len(df.fname.tolist())): # df.label[i] = re.sub(r"(.*)(.*)( \([0-9]+\).xml)", r"\1", df.fname.tolist()[i]) y_unencoded = df.label print "Label encoding" encoder = LabelEncoder() encoder.fit(y_unencoded) y = encoder.transform(y_unencoded) pred_table = pd.
๐ŸŒ
Pythonclass
pythonclass.in โ€บ labelencoder-sklearn.php
labelencoder sklearn | labelencoder scikit
Then applying label encoding, the Height column is converted into: Where 0 is the label for tall and 1 is the label for medium, where 2 is label for short height. Example:- import numpy as np import pandas as pd df=pd.read_csv (โ€˜.../../data/Iris.csvโ€™) df[โ€˜speciesโ€™].unique () Output:- ...
๐ŸŒ
Spot Intelligence
spotintelligence.com โ€บ home โ€บ practical guide and tutorial to label encoding in python
Practical Guide And Tutorial To Label Encoding In Python
October 11, 2024 - Hereโ€™s a Python example using the LabelEncoder class from the scikit-learn library to perform label encoding: from sklearn.preprocessing import LabelEncoder # Sample categorical data colors = ["red", "green", "blue", "red", "green"] # Initialize the LabelEncoder label_encoder = LabelEncoder() # Fit the encoder to the data and transform the categories into integers encoded_colors = label_encoder.fit_transform(colors) print(encoded_colors)
๐ŸŒ
AskPython
askpython.com โ€บ python โ€บ examples โ€บ label-encoding
Label Encoding in Python - A Quick Guide! - AskPython
February 16, 2023 - For example, if a dataset contains a variable โ€˜Genderโ€™ with labels โ€˜Maleโ€™ and โ€˜Femaleโ€™, then the label encoder would convert these labels into a number format and the resultant outcome would be [0,1]. Thus, by converting the labels into the integer format, the machine learning model can have a better understanding in terms of operating the dataset. Python sklearn library provides us with a pre-defined function to carry out Label Encoding on the dataset.
๐ŸŒ
Analytics Vidhya
analyticsvidhya.com โ€บ home โ€บ how to perform label encoding in python?
Label Encoding in Python Explained with Examples
December 19, 2023 - By transforming category data into numerical labels, label encoding enables us to use them in various algorithms. This post will explain label encoding, show where it may be applied in Python, and give examples of how to apply it with the well-liked sci-kit-learn module.