Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

Answer from Tinu on Stack Exchange
Top answer
1 of 3
40

Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .

Examples:

'cat'  = [0,0,0]  
'dog'  = [0,0,1]  
'fish' = [0,1,0]  
'bird' = [0,1,1]  
'ant'  = [1,0,0]   

This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.

Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.

2 of 3
30

The accepted answer seems rather eccentric to me. I think that is rarely done, if ever, and will usually yield bad results.

There's a much more common, sensible use case for this. "Multi-hot encoding" doesn't seem to be a standard term, but I'm not sure there's any standard term. scikit-learn refers to a multi label binarizer.

This is simply used for multi label problems. That is, problems where more than one label can be associated with each example.

For example, say you are trying to detect whether certain types of animal are in a photo. Note that multiple types of animal can be in a single photo. Say the possible types of animal are ['cat', 'dog', 'fish', 'bird', 'ant']. A photo containing cats and dogs would be represented as [1, 1, 0, 0, 0].

🌐
Towards Data Science
towardsdatascience.com › home › data science › from encodings to embeddings
From Encodings to Embeddings | Towards Data Science
September 7, 2023 - Multi-hot encoding is an extension of one-hot encoding when a categorical variable can take multiple values at the same time. For example, there are 28 distinct IMDB genres, and a movie can take multiple genres, e.g.
Discussions

Multi Hot Encoding question
You're doing it wrong, but this will help you: https://stats.stackexchange.com/questions/467633/what-exactly-is-multi-hot-encoding-and-how-is-it-different-from-one-hot More on reddit.com
🌐 r/learnmachinelearning
3
2
October 11, 2023
python - How to do Multi-hot Encoding but with actual values instead of ones - Stack Overflow
I am able to perform a Multi-hot encoding of ratings to movies by: from sklearn.preprocessing import MultiLabelBinarizer def multihot_encode(actual_values, ordered_possible_values) -> np.array: ... More on stackoverflow.com
🌐 stackoverflow.com
Muti-hot encoding vs Label-Encoding - Data Science Stack Exchange
I would like to use Multi-Hot encoding to express up to 255 possible input values. If I use the binary-approach, will this be identical to Label-encoding? In a sense that network will figure out that 00000010 is "superior" to 00000001, that there is a "strong correlation and precedence"? Or is it less exaggerated? For example... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
Multi-hot encoding for ambiguous input
I propose to implement simple multi-hot encoding which allows ambiguous input and outputs non-negative value. Let x_j be a realization of department of a student. Usually, we assume that x_j is def... More on github.com
🌐 github.com
2
January 2, 2019
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to perform one-hot encoding for multi categorical variables
How to Perform One-Hot Encoding For Multi Categorical Variables
February 3, 2025 - Learn multiple categorical variables using One-Hot Encoding in machine learning, including techniques for top-n frequent categories.
🌐
Google
developers.google.com › machine learning › categorical data: vocabulary and one-hot encoding
Categorical data: Vocabulary and one-hot encoding | Machine Learning | Google for Developers
The model learns a separate weight for each element of the feature vector. Note: In a true one-hot encoding, only one element has the value 1.0. In a variant known as multi-hot encoding, multiple values can be 1.0.
🌐
Reddit
reddit.com › r/learnmachinelearning › multi hot encoding question
r/learnmachinelearning on Reddit: Multi Hot Encoding question
October 11, 2023 -

If I understand Multi Hot encoding correctly basically we take arrays of different lengths, and turn them into arrays/lists of all the same length with each index of the new array/list being on(0) or off(1) for the various entries in the original array.

so

[1,3] would become [0, 1, 0, 1]

My question is, how do we account for if a number appears twice in the original array/list,

for example [1,3,3]

🌐
Google
docs.cloud.google.com › bigquery › the ml.multi_hot_encoder function
The ML.MULTI_HOT_ENCODER function | BigQuery | Google Cloud Documentation
SELECT f[OFFSET(0)] AS f0, ML.MULTI_HOT_ENCODER(f, 3, 1) OVER () AS output FROM ( SELECT ['a', 'b', 'b', 'c', NULL] AS f UNION ALL SELECT ['c', 'c', 'd', 'd', NULL] AS f ) ORDER BY f[OFFSET(0)];
Find elsewhere
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-one-hot-encoding
One Hot Encoding in Machine Learning - GeeksforGeeks
One-Hot Encoding creates a separate column for each category in the dataset. In the fruit example, when the fruit is Apple, the Fruit_Apple column gets the value 1 while the other fruit columns contain 0.
Published: May 29, 2026
🌐
Justrocketscience
justrocketscience.com › post › jax_multi_hot
Multi-hot encoding in JAX :: Blog
August 19, 2022 - JAX offers a range of practical functions, also for data preparation. One of these is jax.nn.one_hot, which performs a classic one-hot encoding. Unfortunately, I was not able find a suitable multi-hot equivalent for multi-label applications. However, it is also quite easy to implement the functionality directly in JAX: import jax.numpy as jnp from functools import partial @partial(jax.jit, static_argnames=("num_classes")) def multi_hot(labels, num_classes: int): return jnp.take(jnp.eye(num_classes), jnp.array(labels), axis=0).sum(axis=0)
Top answer
1 of 2
5

You can think of binary encoding as a compromise between label encoding and one-hot encoding. For distinct categories, label encoding introduces a false linear order that brings a lot of noise into the model (category 1 < category 2 < category 3....) . Binary encoding introduces false additive relationships between the categories (e.g. category 4 + category 1 = category 5 or 100 + 001 = 101) but fewer of them.

Therefore, binary will usually work better than label encoding, however only one-hot encoding will usually preserve the full information in the data.

Unless your algorithm (or computing power) is limited in the number of categories it can handle, one-hot encoding will be preferred over other encoding schemes. If you are limited, mean encoding is a powerful alternative because it transform a categorical feature into a numeric one (giving you the minimal number of inputs) while preserving the most important information in the data.

[Mean encoding replaces every category with its target mean. The mean encodings need to be constructed carefully on a separate dataset to avoid data leakage. If you want to reduce your input dimensions as much as possible, this can be helpful. Mean encoding can also be helpful as an additional feature and is very popular on Kaggle to squeeze out some extra performance. This is also known as target encoding or likelihood encoding.]

2 of 2
0

Do I understand correctly that a traditional binary approach to counting numbers is exactly what the Multi-Hot is? We can imagine a byte as a vector of 8 components, and each entry is either 0 or 1.0

I strongly believe so.

This means I can't use it for 255 distinct categories (or for relatively unrelated categories)

Is the effect as bad in Multi-Hot encoding, in particular in a binary approach?

I found a pretty interesting answer. It seems that binary can actually be used with classification tasks really well!
Have a look at the table on the last page of paper "A Comparative Study of Categorical Variable Encoding Techniques for Neural Network Classifiers" 2017

This would mean we can save the number of input neurons BY AN INCREDIBLE AMOUNT, instead of using traditional one-hot encoding.

Personally, I can get it intuitively: Label-encoding tells us to set a neuron with different values: 1,2,3,4... It's really easy for the network to linearly-interpolate from 1 to 2 and from 2 to 3, by using fractions. Thus, there is a really strong precidence between such input values, and the network will easily pick up on that. So we can't use Label-encoding for categories.

Contrary to that, Binary encoding exhibits a more "integer-like" behavior. In other words, it's not as blatantly evident how to linearly interpolate from 1 to 2 in binary form (from 0001 to 0010), which resembles one-hot approach too :)

🌐
ProjectPro
projectpro.io › recipes › one-hot-encoding-with-multiple-labels-in-python
One hot encoding for multi label classification - Projectpro
December 20, 2022 - So what we can do is we can make different columns acconding to the labels and assign bool values in it. This python source code does the following: 1. Converts categorical into numerical types. 2. Loads the important libraries and modules. 3. Implements multi label binarizer.
🌐
TheCVF
openaccess.thecvf.com › content_CVPRW_2020 › papers › w45 › Jaiswal_MUTE_Inter-Class_Ambiguity_Driven_Multi-Hot_Target_Encoding_for_Deep_Neural_CVPRW_2020_paper.pdf pdf
Inter-Class Ambiguity Driven Multi-Hot Target Encoding for ...
in models trained with one-hot encodings (c), get well-separated in ... Figure 2. MUTE generation. The inter-class ambiguity for a given data is used to compute weights between classes. These weights are used · in Algorithm 1 to optimally assign target codes to classes. The chosen set of MUTE are highlighted in green color. Details are in Section 3. obtain a multi...
🌐
Quora
quora.com › What-is-a-multi-hot-vector-in-machine-learning
What is a multi-hot vector in machine learning? - Quora
Answer (1 of 2): Multi hot vector is an artificial vector created on machine learning processes in order to represent categorical variables in a multidimensional space by encoding them into numerical values. For example, consider that we have a machine learning problem to predict whether a perso...
🌐
Ray
docs.ray.io › en › latest › data › api › doc › ray.data.preprocessors.MultiHotEncoder.html
ray.data.preprocessors.MultiHotEncoder — Ray 2.55.1
>>> encoder = MultiHotEncoder(columns=["genre"], max_categories={"genre": 3}) >>> encoder.fit_transform(ds).to_pandas() name genre 0 Shaolin Soccer [1, 1, 1] 1 Moana [1, 1, 0] 2 The Smartest Guys in the Room [0, 0, 0] >>> encoder.stats_ OrderedDict([('unique_values(genre)', {'comedy': 0, 'action': ...
🌐
GitHub
github.com › scikit-learn-contrib › category_encoders › issues › 161
Multi-hot encoding for ambiguous input · Issue #161 · scikit-learn-contrib/category_encoders
January 2, 2019 - I propose to implement simple multi-hot encoding which allows ambiguous input and outputs non-negative value. Let x_j be a realization of department of a student. Usually, we assume that x_j is def...
Author: scikit-learn-contrib
🌐
Keras
keras.io › api › layers › preprocessing_layers › categorical › category_encoding
Keras documentation: CategoryEncoding layer
If the last dimension is not size 1, will append a new dimension for the encoded output. - "multi_hot": Encodes each sample in the input into a single array of num_tokens size, containing a 1 for each vocabulary term present in the sample.
🌐
ResearchGate
researchgate.net › publication › 358098099_Effective_Multi-Hot_Encoding_and_Classifier_for_Lightweight_Scene_Text_Recognition_with_a_Large_Character_Set
(PDF) Effective Multi-Hot Encoding and Classifier for Lightweight Scene Text Recognition with a Large Character Set
August 1, 2022 - However, the conventional softmax-based one-hot classification module becomes a cumbersome obstacle when handling multi-languages or languages with large character set ( e.g. , Chinese) due to the rapid expansion of model parameters with the number of classes. To this end, we propose an Effective Multi-hot encoding and classification modUle (EMU) for scene text recognition in the scenario of multi-languages or languages with large character set.
🌐
Shiksha
shiksha.com › home › it & software › it & software articles › software tools articles › one hot encoding for multi categorical variables
One hot encoding for multi categorical variables - Shiksha Online
September 20, 2022 - One hot encoding can be used to handle multiple categorical categories also. In this blog we will learn this theoretical as well will implement python code with a practical example.
Top answer
1 of 2
1

Here’s a data.table solution in case anyone is interested:

library(data.table)
bf <- data.table(crops = c(1, 3, 345, 9562))

bf[, crops := strsplit(as.character(crops), "")]
cols <- sort(unique(unlist(bf$crops)))
bf[, (cols) := lapply(cols, \(col) sapply(crops, \(row) col %in% row))]
bf
##      crops     1     2     3     4     5     6     9
## 1:       1  TRUE FALSE FALSE FALSE FALSE FALSE FALSE
## 2:       3 FALSE FALSE  TRUE FALSE FALSE FALSE FALSE
## 3:   3,4,5 FALSE FALSE  TRUE  TRUE  TRUE FALSE FALSE
## 4: 9,5,6,2 FALSE  TRUE FALSE FALSE  TRUE  TRUE  TRUE

And a Base R solution:

bf <- data.frame(crops = c(1, 3, 345, 9562))
bf$crops <- strsplit(as.character(bf$crops), "")
cols <- sort(unique(unlist(bf$crops)))
for (col in cols) {
    bf[[col]] <- sapply(bf$crops, \(row) col %in% row)
}
bf
##        crops     1     2     3     4     5     6     9
## 1          1  TRUE FALSE FALSE FALSE FALSE FALSE FALSE
## 2          3 FALSE FALSE  TRUE FALSE FALSE FALSE FALSE
## 3    3, 4, 5 FALSE FALSE  TRUE  TRUE  TRUE FALSE FALSE
## 4 9, 5, 6, 2 FALSE  TRUE FALSE FALSE  TRUE  TRUE  TRUE
2 of 2
0

Here is a potential tidyverse solution:

library(dplyr)
library(tidyr)
bf <- data.frame(crops = c(1,3,345,9562))
bf %>%
  separate(crops, into = c('empty',
                           'firstcrop',
                           'secondcrop',
                           'thirdcrop',
                           'fourthcrop'),
           sep = "", extra = "merge") %>%
  select(-(empty))  %>%
  mutate(row = row_number()) %>%
  pivot_longer(-row) %>%
  na.omit() %>%
  pivot_wider(names_from = value,
              values_from = name) %>%
  select(-row) %>%
  mutate(across(everything(), ~+!is.na(.x))) %>%
  select(order(colnames(.)))
#> Warning: Expected 5 pieces. Missing pieces filled with `NA` in 3 rows [1, 2, 3].
#> # A tibble: 4 × 7
#>     `1`   `2`   `3`   `4`   `5`   `6`   `9`
#>   <int> <int> <int> <int> <int> <int> <int>
#> 1     1     0     0     0     0     0     0
#> 2     0     0     1     0     0     0     0
#> 3     0     0     1     1     1     0     0
#> 4     0     1     0     0     1     1     1

Created on 2022-03-14 by the reprex package (v2.0.1)