Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

Answer from Tinu on Stack Exchange
Top answer
1 of 3
40

Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

2 of 3
30

The accepted answer seems rather eccentric to me. I think that is rarely done, if ever, and will usually yield bad results.

There's a much more common, sensible use case for this. "Multi-hot encoding" doesn't seem to be a standard term, but I'm not sure there's any standard term. scikit-learn refers to a multi label binarizer.

This is simply used for multi label problems. That is, problems where more than one label can be associated with each example.

For example, say you are trying to detect whether certain types of animal are in a photo. Note that multiple types of animal can be in a single photo. Say the possible types of animal are ['cat', 'dog', 'fish', 'bird', 'ant']. A photo containing cats and dogs would be represented as [1, 1, 0, 0, 0].

🌐
GitHub
github.com › scikit-learn-contrib › categorical-encoding › issues › 161
Multi-hot encoding for ambiguous input · Issue #161 · scikit-learn-contrib/category_encoders
January 2, 2019 - Efficiency: If we prepare a mapping which represents relationships between ambiguous|dirty categories and feature without ambiguity, we will need a lot of memory capacity for j-th feature (O(2^C_j), where C_j is the cardinality of j-th feature). Multi-hot encoding needs O(C_j) memory.
Author: scikit-learn-contrib
Discussions

keras - How to create multi-hot encoding from a list column in dataframe? - Data Science Stack Exchange
I have dataframe like this Label IDs 0 [10, 1] 1 [15] 0 [14] I want to create a multihot encoding of the feature IDs. It should look like this Label ID_10 ID_1 ID_15 ID_14 0 ... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
scikit learn - How to perform one hot encoding on multiple categorical columns - Data Science Stack Exchange
From the tutorial I am following, I am supposed to do LabelEncoding before One hot encoding. I have successfully performed the labelencoding as shown below · #categorical data categorical_cols = ['a', 'b', 'c', 'd'] from sklearn.preprocessing import LabelEncoder # instantiate labelencoder ... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
April 5, 2020
python - One-hot-encoding multiple columns in sklearn and naming columns - Stack Overflow
I have the following code to one-hot-encode 2 columns I have. # encode city labels using one-hot encoding scheme city_ohe = OneHotEncoder(categories='auto') city_feature_arr = city_ohe.fit_transfo... More on stackoverflow.com
🌐 stackoverflow.com
scikit learn - Multi-Feature One-Hot-Encoder with varying amount of feature instances - Data Science Stack Exchange
I am wondering how I can encode the third element of these datapoints. For multiple features values we could use sklearn's OneHotEncoder, but as far as I could find out, it cannot handle inputs of different length. More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
January 29, 2021
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Transforms between iterable of iterables and a multilabel format, e.g. a (samples x classes) binary matrix indicating the presence of a class label. ... Given a dataset with two features, we let the encoder find the unique values per feature ...
Top answer
1 of 4
21

LabelEncoder is not made to transform the data but the target (also known as labels) as explained here. If you want to encode the data you should use OrdinalEncoder.

If you really need to do it this way:

categorical_cols = ['a', 'b', 'c', 'd'] 

from sklearn.preprocessing import LabelEncoder
# instantiate labelencoder object
le = LabelEncoder()

# apply le on categorical feature columns
data[categorical_cols] = data[categorical_cols].apply(lambda col: le.fit_transform(col))    
from sklearn.preprocessing import OneHotEncoder
ohe = OneHotEncoder()

#One-hot-encode the categorical columns.
#Unfortunately outputs an array instead of dataframe.
array_hot_encoded = ohe.fit_transform(data[categorical_cols])

#Convert it to df
data_hot_encoded = pd.DataFrame(array_hot_encoded, index=data.index)

#Extract only the columns that didnt need to be encoded
data_other_cols = data.drop(columns=categorical_cols)

#Concatenate the two dataframes : 
data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1)

Otherwise:

I suggest you to use pandas.get_dummies if you want to achieve one-hot-encoding from raw data (without having to use OrdinalEncoder before) :

#categorical data
categorical_cols = ['a', 'b', 'c', 'd'] 

#import pandas as pd
df = pd.get_dummies(data, columns = categorical_cols)

You can also use drop_first argument to remove one of the one-hot-encoded columns, as some models require.

2 of 4
8

You can do dummy encoding using Pandas in order to get one-hot encoding as shown below:

import pandas as pd

# Multiple categorical columns
categorical_cols = ['a', 'b', 'c', 'd']

pd.get_dummies(data, columns=categorical_cols)

If you want to do one-hot encoding using sklearn library, you can get it done as shown below:

from sklearn.preprocessing import OneHotEncoder
onehotencoder = OneHotEncoder()

transformed_data = onehotencoder.fit_transform(data[categorical_cols])

# the above transformed_data is an array so convert it to dataframe
encoded_data = pd.DataFrame(transformed_data, index=data.index)

# now concatenate the original data and the encoded data using pandas
concatenated_data = pd.concat([data, encoded_data], axis=1)

If a single column has more than 500 categories, the aforementioned way of one-hot encoding is not a good approach. In this case, we can do one-hot encoding for the top 10 or 20 categories that are occurring most for a particular column. A sample code is shown below:

categorical_cols = ['a', 'b', 'c', 'd']

# Let's say we have a column 'b' which has more than 500 categories.
# Find the top 10 most frequent categories for column 'b'
data.b.value_counts().sort_values(ascending = False).head(20)

# make a list of the most frequent categories of the column
top_10_occurring_cat = [cat for cat in data.b.value_counts().sort_values(ascending = False).head(10).index]

# now make the 10 binary variables
for cat in top_10_occurring_cat:
    data[cat] = np.where(data['b'] == cat, 1, 0) # whenever data['b'] == cat replace it with 1 else 0

# This is done for one categorical column, similarly you can repeat for all categorical columns
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to perform one-hot encoding for multi categorical variables
How to Perform One-Hot Encoding For Multi Categorical Variables
February 3, 2025 - Now, we have to import important modules from python that will use for the one-hot encoding · # importing pandas import pandas as pd # importing numpy import numpy as np # importing OneHotEncoder from sklearn.preprocessing import OneHotEncoder()
Find elsewhere
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.MultiLabelBinarizer.html
MultiLabelBinarizer — scikit-learn 1.9.1 documentation
Encode categorical features using a one-hot aka one-of-K scheme. Examples · >>> from sklearn.preprocessing import MultiLabelBinarizer >>> mlb = MultiLabelBinarizer() >>> mlb.fit_transform([(1, 2), (3,)]) array([[1, 1, 0], [0, 0, 1]]) >>> ...
🌐
Towards Data Science
towardsdatascience.com › home › data science › from encodings to embeddings
From Encodings to Embeddings | Towards Data Science
September 7, 2023 - 💻 We can use scikit-learn or pandas to achieve multi-hot encoding in Python. Here's how to do it using scikit-learn's MultiLabelBinarizer: ... from sklearn.preprocessing import MultiLabelBinarizer# Sample data: List of sets representing ...
🌐
ProjectPro
projectpro.io › recipes › one-hot-encoding-with-multiple-labels-in-python
One hot encoding for multi label classification - Projectpro
December 20, 2022 - So what we can do is we can make different columns acconding to the labels and assign bool values in it. This python source code does the following: 1. Converts categorical into numerical types. 2. Loads the important libraries and modules. 3. Implements multi label binarizer.
Top answer
1 of 2
1

I believe @David Masip's answer can now be further improved upon by providing a custom analyzer to CountVectorizer.

Joining all the features into a single string has a few issues:

  • Wouldn't work well with features containing whitespace
  • Wouldn't work well with features containing accents, stop words, etc. due to CountVectorizer's tokenization

Luckily, CountVectorizer now offers a way to provide a custom tokenizer via the analyzer argument. So, we could just keep passing lists of values as input, and have a custom analyzer pass those lists through without tokenization:

from sklearn.feature_extraction.text import CountVectorizer

X = [["banana","apple","cucumber"], ["orange","banana", "cucumber"]]
enc = CountVectorizer(analyzer=lambda lst: lst)
print(enc.fit_transform(X).toarray())

This produces the same result

[[1 1 1 0]
 [0 1 1 1]]

while avoiding the issues listed above, and also saving CPU cycles/memory by not joining and re-tokenizing features.

2 of 2
3

I think you can transform this into a text preprocessing problem and then use CountVectorizer. You basically build "documents" by putting together all the words in your raw data and then use CountVectorizer on those documents.

from sklearn.feature_extraction.text import CountVectorizer

X = [["banana","apple","cucumber"], ["orange","banana", "cucumber"]]
# Create documents
X_ = [' '.join(x) for x in X]
enc = CountVectorizer()
print(enc.fit_transform(X_).toarray())

Returns

[[1 1 1 0]
 [0 1 1 1]]

which has 4 different values as you expected.

🌐
Analytics Vidhya
analyticsvidhya.com › home › one hot encoding vs label encoding in machine learning
One Hot Encoding vs Label Encoding in Machine Learning
April 23, 2025 - Here’s how you can implement one-hot encoding using Scikit-Learn in Python: from sklearn.preprocessing import OneHotEncoder import pandas as pd
Top answer
1 of 5
46

A simple example which encodes an array using LabelEncoder, OneHotEncoder, LabelBinarizer is shown below.

I see that OneHotEncoder needs data in integer encoded form first to convert into its respective encoding which is not required in the case of LabelBinarizer.

from numpy import array
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder
from sklearn.preprocessing import LabelBinarizer

# define example
data = ['cold', 'cold', 'warm', 'cold', 'hot', 'hot', 'warm', 'cold', 
'warm', 'hot']
values = array(data)
print "Data: ", values
# integer encode
label_encoder = LabelEncoder()
integer_encoded = label_encoder.fit_transform(values)
print "Label Encoder:" ,integer_encoded

# onehot encode
onehot_encoder = OneHotEncoder(sparse=False)
integer_encoded = integer_encoded.reshape(len(integer_encoded), 1)
onehot_encoded = onehot_encoder.fit_transform(integer_encoded)
print "OneHot Encoder:", onehot_encoded

#Binary encode
lb = LabelBinarizer()
print "Label Binarizer:", lb.fit_transform(values)

Another good link which explains the OneHotEncoder is: Explain onehotencoder using python

There may be other valid differences between the two which experts can probably explain.

2 of 5
31

A difference is that you can use OneHotEncoder for multi column data, while not for LabelBinarizer and LabelEncoder.

from sklearn.preprocessing import LabelBinarizer, LabelEncoder, OneHotEncoder

X = [["US", "M"], ["UK", "M"], ["FR", "F"]]
OneHotEncoder().fit_transform(X).toarray()

# array([[0., 0., 1., 0., 1.],
#        [0., 1., 0., 0., 1.],
#        [1., 0., 0., 1., 0.]])
LabelBinarizer().fit_transform(X)
# ValueError: Multioutput target data is not supported with label binarization

LabelEncoder().fit_transform(X)
# ValueError: bad input shape (3, 2)
🌐
Reddit
reddit.com › r/learnpython › multi-hot encoding strings without taking forever?
r/learnpython on Reddit: Multi-hot encoding strings without taking forever?
February 21, 2020 -

I have a pandas data series of strings that each have a bunch of text symbols in them (I'll call them words for discussion's sake, but in my use case they aren't actually words). The series of strings is already parsed to give me a vocabulary of all the words found anywhere in the series.

I've taken that vocabulary (vocab1, a list of all the words in the vocabulary) and made a dict (vocab) to assign an index to each word:

vocab = {c:i for i,c in enumerate(vocab1)}

I need to change these strings into a multi-hot encoded numpy array (for machine learning blah blah). The arrays are the size of the vocabulary. 1 at a given position in the array indicates the absence of the corresponding word in the string, while 0 indicates its absence. The strings typically only have a few words in them (which sometimes are redundantly repeated in the original data) but there are 30k different words in the vocabulary, and in the broader use case this goes up to 260k or so.

Anyway, vocablength is the number of words in the vocabulary. I'm using pandas.Series.transform to apply my function to the entire series, which is over 200k strings (and will be millions in the broader use case).

def multihot_codes(codes):
    o = np.zeros((vocablength,),dtype=np.int32)
    for k in codes.split():
        o[vocab[k]] = 1
    return o

df["codesmh"] = df["codess"].transform(multihot_codes)

This takes close to forever and has significant disk usage. There's probably a faster way to do this. For example, Tensorflow generates binary one-hot vectors for each word and then does a row reduce on them, but if you dig into the code, they have a compiled C++ module that does the heavy lifting. Not exactly what I'm prepared to do here myself....

Any ideas on how to improve this, or if there's something already in numpy/scipy or another package that's suited for use here? Thanks in advance!

🌐
KDnuggets
kdnuggets.com › 2023 › 01 › encoding-categorical-features-multilabelbinarizer.html
Encoding Categorical Features with MultiLabelBinarizer - KDnuggets
January 20, 2023 - We will now use the Scikit-learn MultiLabelBinarizer to convert iterable of iterables and multilabel targets into binary encoding.
🌐
GeeksforGeeks
geeksforgeeks.org › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
Pandas offers the get_dummies function which is a simple and effective way to perform one-hot encoding. This method converts categorical variables into multiple binary columns. For example the Gender column with values 'M' and 'F' becomes two binary columns: Gender_F and Gender_M. drop_first=True in pandas drops one redundant column e.g., keeps only Gender_F to avoid multicollinearity. ... import pandas as pd from sklearn.preprocessing import OneHotEncoder data = { 'Employee id': [10, 20, 15, 25, 30], 'Gender': ['M', 'F', 'F', 'M', 'F'], 'Remarks': ['Good', 'Nice', 'Good', 'Great', 'Nice'] } d
Published: February 7, 2025