Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

Answer from Tinu on Stack Exchange
Top answer
1 of 3
40

Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

2 of 3
30

The accepted answer seems rather eccentric to me. I think that is rarely done, if ever, and will usually yield bad results.

There's a much more common, sensible use case for this. "Multi-hot encoding" doesn't seem to be a standard term, but I'm not sure there's any standard term. scikit-learn refers to a multi label binarizer.

This is simply used for multi label problems. That is, problems where more than one label can be associated with each example.

For example, say you are trying to detect whether certain types of animal are in a photo. Note that multiple types of animal can be in a single photo. Say the possible types of animal are ['cat', 'dog', 'fish', 'bird', 'ant']. A photo containing cats and dogs would be represented as [1, 1, 0, 0, 0].

🌐
Analytics Vidhya
analyticsvidhya.com › home › how to perform one-hot encoding for multi categorical variables
How to Perform One-Hot Encoding For Multi Categorical Variables
February 3, 2025 - Here, we use pandas which are used for data analysis, NumPyused for n-dimensional arrays, and from sklearn, we will use one important class One Hot Encoder for categorical encoding. Now we have to read this data using Python.
🌐
Reddit
reddit.com › r/learnpython › multi-hot encoding in pandas
r/learnpython on Reddit: Multi-Hot Encoding in Pandas
September 25, 2019 -

I’m training a sentiment classification model of sorts that takes as input a sentence and then outputs however many categories it thinks that sentence belongs to. Often times it will only belong to one category, maybe “happy” or “sad,” but sometimes it could belong to several, like “happy, excited, nervous, etc.”

Right now I have the data in a pandas dataframe with the sentences in one column and then the categories that those sentences belong to in another column separated by commas:

SENTENCES SENTIMENTS

sentence 1 || sentiment 1, sentiment 2

sentence 2 || sentiment 1

sentence 3 || sentiment 2, sentiment 3, sentiment 6, sentiment 7

sentence 4 || sentiment 4, sentiment 5

I was able to find info on one-hot encoding, but I cannot for the life of me figure out how to multi-hot encode the sentiments column.

Does anyone know of a way to multi-hot encode a column in a Pandas dataframe?

Surely this is a common thing?

(Sorry for the bad representations of the data. I’m on my phone and haven’t yet figured out how to insert code into a post using the app.)

Edit: formatting

🌐
Reddit
reddit.com › r/learnmachinelearning › multi hot encoding question
r/learnmachinelearning on Reddit: Multi Hot Encoding question
October 11, 2023 -

If I understand Multi Hot encoding correctly basically we take arrays of different lengths, and turn them into arrays/lists of all the same length with each index of the new array/list being on(0) or off(1) for the various entries in the original array.

so

[1,3] would become [0, 1, 0, 1]

My question is, how do we account for if a number appears twice in the original array/list,

for example [1,3,3]

🌐
ProjectPro
projectpro.io › recipes › one-hot-encoding-with-multiple-labels-in-python
One hot encoding for multi label classification - Projectpro
December 20, 2022 - In many datasets we find that there are multiple labels and machine learning model can not be trained on the labels. To solve this problem we may assign numbers to this labels but machine learning models can compare numbers and will give different weightage to different labels and as a result it will be bias towards a label. So what we can do is we can make different columns acconding to the labels and assign bool values in it. This python source code does the following: 1.
🌐
Reddit
reddit.com › r/learnpython › multi-hot encoding strings without taking forever?
r/learnpython on Reddit: Multi-hot encoding strings without taking forever?
February 21, 2020 -

I have a pandas data series of strings that each have a bunch of text symbols in them (I'll call them words for discussion's sake, but in my use case they aren't actually words). The series of strings is already parsed to give me a vocabulary of all the words found anywhere in the series.

I've taken that vocabulary (vocab1, a list of all the words in the vocabulary) and made a dict (vocab) to assign an index to each word:

vocab = {c:i for i,c in enumerate(vocab1)}

I need to change these strings into a multi-hot encoded numpy array (for machine learning blah blah). The arrays are the size of the vocabulary. 1 at a given position in the array indicates the absence of the corresponding word in the string, while 0 indicates its absence. The strings typically only have a few words in them (which sometimes are redundantly repeated in the original data) but there are 30k different words in the vocabulary, and in the broader use case this goes up to 260k or so.

Anyway, vocablength is the number of words in the vocabulary. I'm using pandas.Series.transform to apply my function to the entire series, which is over 200k strings (and will be millions in the broader use case).

def multihot_codes(codes):
    o = np.zeros((vocablength,),dtype=np.int32)
    for k in codes.split():
        o[vocab[k]] = 1
    return o

df["codesmh"] = df["codess"].transform(multihot_codes)

This takes close to forever and has significant disk usage. There's probably a faster way to do this. For example, Tensorflow generates binary one-hot vectors for each word and then does a row reduce on them, but if you dig into the code, they have a compiled C++ module that does the heavy lifting. Not exactly what I'm prepared to do here myself....

Any ideas on how to improve this, or if there's something already in numpy/scipy or another package that's suited for use here? Thanks in advance!

Find elsewhere
Top answer
1 of 2
5

Considering the input data as:

import pandas as pd
data = {'Title': ['Movie 1', 'Movie 2', 'Movie 3', 'Movie 4', 'Movie 5'], 
        'column': ['Action, Fantasy', 'Fantasy, Drama', 'Action', 'Sci-Fi, Romance, Comedy', np.nan]}
df = pd.DataFrame(data)
df
    Title   column
0   Movie 1 Action, Fantasy
1   Movie 2 Fantasy, Drama
2   Movie 3 Action
3   Movie 4 Sci-Fi, Romance, Comedy
4   Movie 5 NaN

This code produces the desired output:

# treat null values
df['column'].fillna('NA', inplace = True)

# separate all genres into one list, considering comma + space as separators
genre = df['column'].str.split(', ').tolist()

# flatten the list
flat_genre = [item for sublist in genre for item in sublist]

# convert to a set to make unique
set_genre = set(flat_genre)

# back to list
unique_genre = list(set_genre)

# remove NA
unique_genre.remove('NA')

# create columns by each unique genre
df = df.reindex(df.columns.tolist() + unique_genre, axis=1, fill_value=0)

# for each value inside column, update the dummy
for index, row in df.iterrows():
    for val in row.column.split(', '):
        if val != 'NA':
            df.loc[index, val] = 1

df.drop('column', axis = 1, inplace = True)    
df
    Title   Action  Fantasy Comedy  Sci-Fi  Drama   Romance
0   Movie 1 1       1       0       0       0       0
1   Movie 2 0       1       0       0       1       0
2   Movie 3 1       0       0       0       0       0
3   Movie 4 0       0       1       1       0       1
4   Movie 5 0       0       0       0       0       0

UPDATE: I've added a null value into the test data, and treat it appropriately in the first line of the solution.

2 of 2
1
### Import libraries and load sample data

import numpy as np
import pandas as pd

data = {
    'Movie 1': ['Action, Fantasy'],
    'Movie 2': ['Fantasy, Drama'],
    'Movie 3': ['Action'],
    'Movie 4': ['Sci-Fi, Romance, Comedy'],
    'Movie 5': ['NA'],
}

df = pd.DataFrame.from_dict(data, orient='index')
df.rename(columns={0:'column'}, inplace=True)

At this stage our DataFrame looks like this:

           column
Movie 1    Action, Fantasy
Movie 2    Fantasy, Drama
Movie 3    Action
Movie 4    Sci-Fi, Romance, Comedy
Movie 5    NA

Now, the question we're asking is - does a given genre word ("sub-string") occur in 'column' for a given movie?

To do this we'll first need a list of genre words:

### Join every string in every row, split the result, pull out the unique values.
genres = np.unique(', '.join(df['column']).split(', '))
### Drop 'NA'
genres = np.delete(genres, np.where(genres == 'NA'))

Depending on how large your dataset is, this could be computationally costly. You mentioned that you know the unique values already. So you could just define the iterable 'genres' manually.

Getting the OneHotVectors:

for genre in genres:
    df[genre] = df['column'].str.contains(genre).astype('int')

df.drop('column', axis=1, inplace=True)

We loop through each genre, we ask whether the genre exists in 'column', this returns a True or False, which is converted to 1 or 0 respectively - when we cast to type('int').

We end up with:

          Action    Comedy  Drama   Fantasy Romance Sci-Fi
Movie 1        1         0      0         1       0      0
Movie 2        0         0      1         1       0      0
Movie 3        1         0      0         0       0      0
Movie 4        0         1      0         0       1      1
Movie 5        0         0      0         0       0      0

🌐
Kaggle
kaggle.com › code › adityasingh3519 › one-hot-encoding-for-multi-categorical-variables
One Hot Encoding for Multi Categorical Variables
Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
GitHub
github.com › binarymax › multilabel
GitHub - binarymax/multilabel: Multi-hot label encoding for NumPy
Multi-hot label encoding for NumPy · Add labels to set, and combine them to form multihot encoded numpy arrays. import numpy as np import multilabel as ml labels = ml.MultiLabel() labels.add("news") labels.add("tech") labels.add("law") labels.add("culture") labels.add("politics") labels.add("linux") labels.add("python") print(labels.combine(["news","python"])) #> [ 1.
Author: binarymax
🌐
Shiksha
shiksha.com › home › it & software › it & software articles › software tools articles › one hot encoding for multi categorical variables
One hot encoding for multi categorical variables - Shiksha Online
September 20, 2022 - One hot encoding can be used to handle multiple categorical categories also. In this blog we will learn this theoretical as well will implement python code with a practical example.
🌐
Stack Overflow
stackoverflow.com › questions › 45380454 › one-hot-encoding-multi-dimensional-data
python - One hot encoding multi dimensional data - Stack Overflow
Using code below I'm attempting to one hot code multi dimensional data. In this case the data is 2d. The code works as expected for 1d data but for 2d data each column is one hot encoded instead of the entire row.
Top answer
1 of 1
1

TL;DR: you can use Numba to optimize np.dot to only operate only on binary values. More specifically, you can perform SIMD-like operations on 8 bytes at once using 64-bit views.




Converting lists to arrays

First of all, the lists can be efficiently converted to relatively-compact arrays using this approach:

vector = np.fromiter(vector, np.uint8)
conflicts = np.array([np.fromiter(conflicts[i], np.uint8) for i in range(len(conflicts))])

This is faster than using the automatic Numpy conversion or np.array (there is less check to perform in the Numpy code internally and Numpy, Numpy know what type of array to build and the resulting one is smaller in memory and thus faster to fill). This step can be used to speed up your np.dot-based solution.

If the input are already a Numpy array, then check they are of type np.uint8 or np.int8. Otherwise, please cast them to such type using conflits = conflits.astype(np.uint8) for example.


First try

Then, one solution could be to use np.packbits to pack the input binary values much as possible in an array of bits in memory, and then perform logical ANDs. But it turns out that np.packbits is pretty slow. Thus, this solution is not a good idea in the end. In fact, any solution creating temporary arrays with a shape similar to conflicts will be slow since writing such an array in memory is generally slower than np.dot (which read conflicts from memory once).


Using Numba

Since np.dot is pretty well optimized, the only solution to defeat it is to use an optimized native code. Numba can be used to generate a native executable code at runtime from a Numpy-based Python code thanks to a just-in-time compiler. The idea is to perform a logical ANDs between vector and rows of conflicts per block. Conflict are check for each block so to stop the computation as early as possible. Blocks can be efficiently compared by groups of 8 octets by comparing the uint64 views of the two arrays (in a SIMD-friendly way).

import numba as nb

@nb.njit('bool_(uint8[::1], uint8[:,::1])')
def check_valid(vector, conflicts):
    n, m = conflicts.shape
    assert vector.size == m

    for i in range(n):
        block_size = 128 # In the range: 8,16,...,248
        conflicts_row = conflicts[i,:]
        gsum = 0 # Global sum of conflicts
        m_limit = m // block_size * block_size

        for j in range(0, m_limit, block_size):
            vector_block = vector[j:j+block_size].view(np.uint64)
            conflicts_block = conflicts_row[j:j+block_size].view(np.uint64)

            # Matching
            lsum = np.uint64(0) # 8 local sums of conflicts
            for k in range(block_size//8):
                lsum += vector_block[k] & conflicts_block[k]

            # Trick to perform the reduction of all the bytes in lsum
            lsum += lsum >> 32
            lsum += lsum >> 16
            lsum += lsum >> 8
            gsum += lsum & 0xFF

            # Check if there is a conflict
            if gsum >= 2:
                return False

        # Remaining part
        for j in range(m_limit, m):
            gsum += vector[j] & conflicts_row[j]

        if gsum >= 2:
            return False

    return True

Results

This is about 9 times faster than np.dot on my machine for a large conflicts array of shape (16, 65536) (without conflicts). The time to convert lists is not included in both cases. When there are conflicts, the provided solution is much faster since it can early stop the computation.

Theoretically, the computation should be even faster, but the Numba JIT do not succeed to vectorize the loop using SIMD instructions. That being said, it seems the same issue appears for np.dot. If the arrays are even bigger, you can parallelize the computation of the blocks (at the expense of a slower computation if the function return False).

🌐
Codecademy
codecademy.com › article › what-is-one-hot-encoding-and-how-to-implement-it-in-python
What is One Hot Encoding and How to Implement it in Python? | Codecademy
We can also perform one-hot encoding on multiple columns at once using the get_dummies() function in Python. For this, we need to pass all the column names we want to encode in a list to the columns parameter.
🌐
Quora
quora.com › What-is-a-multi-hot-vector-in-machine-learning
What is a multi-hot vector in machine learning? - Quora
Answer (1 of 2): Multi hot vector is an artificial vector created on machine learning processes in order to represent categorical variables in a multidimensional space by encoding them into numerical values. For example, consider that we have a machine learning problem to predict whether a perso...