Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

Answer from Tinu on Stack Exchange
Top answer
1 of 3
40

Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

2 of 3
30

The accepted answer seems rather eccentric to me. I think that is rarely done, if ever, and will usually yield bad results.

There's a much more common, sensible use case for this. "Multi-hot encoding" doesn't seem to be a standard term, but I'm not sure there's any standard term. scikit-learn refers to a multi label binarizer.

This is simply used for multi label problems. That is, problems where more than one label can be associated with each example.

For example, say you are trying to detect whether certain types of animal are in a photo. Note that multiple types of animal can be in a single photo. Say the possible types of animal are ['cat', 'dog', 'fish', 'bird', 'ant']. A photo containing cats and dogs would be represented as [1, 1, 0, 0, 0].

🌐
Reddit
reddit.com › r/learnmachinelearning › multi hot encoding question
r/learnmachinelearning on Reddit: Multi Hot Encoding question
October 11, 2023 -

If I understand Multi Hot encoding correctly basically we take arrays of different lengths, and turn them into arrays/lists of all the same length with each index of the new array/list being on(0) or off(1) for the various entries in the original array.

so

[1,3] would become [0, 1, 0, 1]

My question is, how do we account for if a number appears twice in the original array/list,

for example [1,3,3]

Discussions

python - How to do Multi-hot Encoding but with actual values instead of ones - Stack Overflow
I am able to perform a Multi-hot encoding of ratings to movies by: from sklearn.preprocessing import MultiLabelBinarizer def multihot_encode(actual_values, ordered_possible_values) -> np.array: ... More on stackoverflow.com
🌐 stackoverflow.com
Muti-hot encoding vs Label-Encoding - Data Science Stack Exchange
I am learning about different input-vector representations for Neural Networks One of the alternatives to sparse One-Hot encoded vector is the Multi-Hot encoding. Do I understand correctly that a More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
Multi-hot encoding for ambiguous input
I propose to implement simple multi-hot encoding which allows ambiguous input and outputs non-negative value. Let x_j be a realization of department of a student. Usually, we assume that x_j is def... More on github.com
🌐 github.com
2
January 2, 2019
python - Determining the validity of a multi-hot encoding - Stack Overflow
Suppose I have N items and a multi-hot vector of values {0, 1} that represents inclusion of these items in a result: N = 4 # items 1 and 3 will be included in the result vector = [0, 1, 0, 1] # i... More on stackoverflow.com
🌐 stackoverflow.com
🌐
Google
developers.google.com › machine learning › categorical data: vocabulary and one-hot encoding
Categorical data: Vocabulary and one-hot encoding | Machine Learning | Google for Developers
Note: In a true one-hot encoding, only one element has the value 1.0. In a variant known as multi-hot encoding, multiple values can be 1.0.
🌐
Towards Data Science
towardsdatascience.com › home › data science › from encodings to embeddings
From Encodings to Embeddings | Towards Data Science
September 7, 2023 - short of semantic: encoding of two words 'good' and 'great' are as different as encoding of 'good' and 'bad'! ✏️ In a nutshell, use one-hot/multi-hot encoding when the number of categories is small; usually lesser than 15 or so.
Top answer
1 of 2
5

You can think of binary encoding as a compromise between label encoding and one-hot encoding. For distinct categories, label encoding introduces a false linear order that brings a lot of noise into the model (category 1 < category 2 < category 3....) . Binary encoding introduces false additive relationships between the categories (e.g. category 4 + category 1 = category 5 or 100 + 001 = 101) but fewer of them.

Therefore, binary will usually work better than label encoding, however only one-hot encoding will usually preserve the full information in the data.

Unless your algorithm (or computing power) is limited in the number of categories it can handle, one-hot encoding will be preferred over other encoding schemes. If you are limited, mean encoding is a powerful alternative because it transform a categorical feature into a numeric one (giving you the minimal number of inputs) while preserving the most important information in the data.

[Mean encoding replaces every category with its target mean. The mean encodings need to be constructed carefully on a separate dataset to avoid data leakage. If you want to reduce your input dimensions as much as possible, this can be helpful. Mean encoding can also be helpful as an additional feature and is very popular on Kaggle to squeeze out some extra performance. This is also known as target encoding or likelihood encoding.]

2 of 2
0

Do I understand correctly that a traditional binary approach to counting numbers is exactly what the Multi-Hot is? We can imagine a byte as a vector of 8 components, and each entry is either 0 or 1.0

I strongly believe so.

This means I can't use it for 255 distinct categories (or for relatively unrelated categories)

Is the effect as bad in Multi-Hot encoding, in particular in a binary approach?

I found a pretty interesting answer. It seems that binary can actually be used with classification tasks really well!
Have a look at the table on the last page of paper "A Comparative Study of Categorical Variable Encoding Techniques for Neural Network Classifiers" 2017

This would mean we can save the number of input neurons BY AN INCREDIBLE AMOUNT, instead of using traditional one-hot encoding.

Personally, I can get it intuitively: Label-encoding tells us to set a neuron with different values: 1,2,3,4... It's really easy for the network to linearly-interpolate from 1 to 2 and from 2 to 3, by using fractions. Thus, there is a really strong precidence between such input values, and the network will easily pick up on that. So we can't use Label-encoding for categories.

Contrary to that, Binary encoding exhibits a more "integer-like" behavior. In other words, it's not as blatantly evident how to linearly interpolate from 1 to 2 in binary form (from 0001 to 0010), which resembles one-hot approach too :)

🌐
Google
docs.cloud.google.com › bigquery › the ml.multi_hot_encoder function
The ML.MULTI_HOT_ENCODER function | BigQuery | Google Cloud Documentation
This document describes the ML.MULTI_HOT_ENCODER function, which lets you encode a string array expression by using a multi-hot encoding scheme.
Find elsewhere
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to perform one-hot encoding for multi categorical variables
How to Perform One-Hot Encoding For Multi Categorical Variables
February 3, 2025 - Learn multiple categorical variables using One-Hot Encoding in machine learning, including techniques for top-n frequent categories.
🌐
GitHub
github.com › scikit-learn-contrib › categorical-encoding › issues › 161
Multi-hot encoding for ambiguous input · Issue #161 · scikit-learn-contrib/category_encoders
January 2, 2019 - Efficiency: If we prepare a mapping which represents relationships between ambiguous|dirty categories and feature without ambiguity, we will need a lot of memory capacity for j-th feature (O(2^C_j), where C_j is the cardinality of j-th feature). Multi-hot encoding needs O(C_j) memory.
Author: scikit-learn-contrib
🌐
TheCVF
openaccess.thecvf.com › content_CVPRW_2020 › papers › w45 › Jaiswal_MUTE_Inter-Class_Ambiguity_Driven_Multi-Hot_Target_Encoding_for_Deep_Neural_CVPRW_2020_paper.pdf pdf
Inter-Class Ambiguity Driven Multi-Hot Target Encoding for ...
in models trained with one-hot encodings (c), get well-separated in ... Figure 2. MUTE generation. The inter-class ambiguity for a given data is used to compute weights between classes. These weights are used · in Algorithm 1 to optimally assign target codes to classes. The chosen set of MUTE are highlighted in green color. Details are in Section 3. obtain a multi-hot encoding such that semantically closer
🌐
ResearchGate
researchgate.net › publication › 358098099_Effective_Multi-Hot_Encoding_and_Classifier_for_Lightweight_Scene_Text_Recognition_with_a_Large_Character_Set
(PDF) Effective Multi-Hot Encoding and Classifier for Lightweight Scene Text Recognition with a Large Character Set
August 1, 2022 - However, the conventional softmax-based one-hot classification module becomes a cumbersome obstacle when handling multi-languages or languages with large character set ( e.g. , Chinese) due to the rapid expansion of model parameters with the number of classes. To this end, we propose an Effective Multi-hot encoding and classification modUle (EMU) for scene text recognition in the scenario of multi-languages or languages with large character set.
🌐
IEEE Xplore
ieeexplore.ieee.org › document › 9691322
EMU: Effective Multi-Hot Encoding Net for Lightweight Scene Text Recognition With a Large Character Set | IEEE Journals & Magazine | IEEE Xplore
To this end, we propose an Effective Multi-hot encoding and classification modUle (EMU) for scene text recognition in the scenario of multi-languages or languages with large character set. Specifically, EMU generates a binary multi-hot label for each class with a real-valued sub-network in ...
🌐
SSRN
papers.ssrn.com › sol3 › papers.cfm
Multi-Hot Encoding of Categorical Dataset for K-Means Clustering Cropping Based Clustering of Districts in Bangladesh by Mozammel H. A. Khan :: SSRN
March 25, 2022 - For k-Means clustering of categorical datasets, we propose a multi-hot encoding of categorical values followed by integer conversion of the encoded binary strin
🌐
Ray
docs.ray.io › en › latest › data › api › doc › ray.data.preprocessors.MultiHotEncoder.html
ray.data.preprocessors.MultiHotEncoder — Ray 2.55.1
If you specify max_categories, then MultiHotEncoder creates features for only the most frequent categories. >>> encoder = MultiHotEncoder(columns=["genre"], max_categories={"genre": 3}) >>> encoder.fit_transform(ds).to_pandas() name genre 0 Shaolin Soccer [1, 1, 1] 1 Moana [1, 1, 0] 2 The Smartest Guys in the Room [0, 0, 0] >>> encoder.stats_ OrderedDict([('unique_values(genre)', {'comedy': 0, 'action': 1, 'sports': 2})])
🌐
Quora
quora.com › What-is-a-multi-hot-vector-in-machine-learning
What is a multi-hot vector in machine learning? - Quora
Answer (1 of 2): Multi hot vector is an artificial vector created on machine learning processes in order to represent categorical variables in a multidimensional space by encoding them into numerical values. For example, consider that we have a machine learning problem to predict whether a perso...
🌐
Justrocketscience
justrocketscience.com › post › jax_multi_hot
Multi-hot encoding in JAX :: Blog
August 19, 2022 - JAX offers a range of practical functions, also for data preparation. One of these is jax.nn.one_hot, which performs a classic one-hot encoding. Unfortunately, I was not able find a suitable multi-hot equivalent for multi-label applications. However, it is also quite easy to implement the functionality directly in JAX: import jax.numpy as jnp from functools import partial @partial(jax.jit, static_argnames=("num_classes")) def multi_hot(labels, num_classes: int): return jnp.take(jnp.eye(num_classes), jnp.array(labels), axis=0).sum(axis=0)
Top answer
1 of 1
1

TL;DR: you can use Numba to optimize np.dot to only operate only on binary values. More specifically, you can perform SIMD-like operations on 8 bytes at once using 64-bit views.




Converting lists to arrays

First of all, the lists can be efficiently converted to relatively-compact arrays using this approach:

vector = np.fromiter(vector, np.uint8)
conflicts = np.array([np.fromiter(conflicts[i], np.uint8) for i in range(len(conflicts))])

This is faster than using the automatic Numpy conversion or np.array (there is less check to perform in the Numpy code internally and Numpy, Numpy know what type of array to build and the resulting one is smaller in memory and thus faster to fill). This step can be used to speed up your np.dot-based solution.

If the input are already a Numpy array, then check they are of type np.uint8 or np.int8. Otherwise, please cast them to such type using conflits = conflits.astype(np.uint8) for example.


First try

Then, one solution could be to use np.packbits to pack the input binary values much as possible in an array of bits in memory, and then perform logical ANDs. But it turns out that np.packbits is pretty slow. Thus, this solution is not a good idea in the end. In fact, any solution creating temporary arrays with a shape similar to conflicts will be slow since writing such an array in memory is generally slower than np.dot (which read conflicts from memory once).


Using Numba

Since np.dot is pretty well optimized, the only solution to defeat it is to use an optimized native code. Numba can be used to generate a native executable code at runtime from a Numpy-based Python code thanks to a just-in-time compiler. The idea is to perform a logical ANDs between vector and rows of conflicts per block. Conflict are check for each block so to stop the computation as early as possible. Blocks can be efficiently compared by groups of 8 octets by comparing the uint64 views of the two arrays (in a SIMD-friendly way).

import numba as nb

@nb.njit('bool_(uint8[::1], uint8[:,::1])')
def check_valid(vector, conflicts):
    n, m = conflicts.shape
    assert vector.size == m

    for i in range(n):
        block_size = 128 # In the range: 8,16,...,248
        conflicts_row = conflicts[i,:]
        gsum = 0 # Global sum of conflicts
        m_limit = m // block_size * block_size

        for j in range(0, m_limit, block_size):
            vector_block = vector[j:j+block_size].view(np.uint64)
            conflicts_block = conflicts_row[j:j+block_size].view(np.uint64)

            # Matching
            lsum = np.uint64(0) # 8 local sums of conflicts
            for k in range(block_size//8):
                lsum += vector_block[k] & conflicts_block[k]

            # Trick to perform the reduction of all the bytes in lsum
            lsum += lsum >> 32
            lsum += lsum >> 16
            lsum += lsum >> 8
            gsum += lsum & 0xFF

            # Check if there is a conflict
            if gsum >= 2:
                return False

        # Remaining part
        for j in range(m_limit, m):
            gsum += vector[j] & conflicts_row[j]

        if gsum >= 2:
            return False

    return True

Results

This is about 9 times faster than np.dot on my machine for a large conflicts array of shape (16, 65536) (without conflicts). The time to convert lists is not included in both cases. When there are conflicts, the provided solution is much faster since it can early stop the computation.

Theoretically, the computation should be even faster, but the Numba JIT do not succeed to vectorize the loop using SIMD instructions. That being said, it seems the same issue appears for np.dot. If the arrays are even bigger, you can parallelize the computation of the blocks (at the expense of a slower computation if the function return False).

🌐
Keras
keras.io › api › layers › preprocessing_layers › categorical › category_encoding
Keras documentation: CategoryEncoding layer
If the last dimension is not size 1, will append a new dimension for the encoded output. - "multi_hot": Encodes each sample in the input into a single array of num_tokens size, containing a 1 for each vocabulary term present in the sample.