Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .
Examples:
'cat' = [0,0,0]
'dog' = [0,0,1]
'fish' = [0,1,0]
'bird' = [0,1,1]
'ant' = [1,0,0]
This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.
Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.
Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .
Examples:
'cat' = [0,0,0]
'dog' = [0,0,1]
'fish' = [0,1,0]
'bird' = [0,1,1]
'ant' = [1,0,0]
This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.
Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.
The accepted answer seems rather eccentric to me. I think that is rarely done, if ever, and will usually yield bad results.
There's a much more common, sensible use case for this. "Multi-hot encoding" doesn't seem to be a standard term, but I'm not sure there's any standard term. scikit-learn refers to a multi label binarizer.
This is simply used for multi label problems. That is, problems where more than one label can be associated with each example.
For example, say you are trying to detect whether certain types of animal are in a photo. Note that multiple types of animal can be in a single photo. Say the possible types of animal are ['cat', 'dog', 'fish', 'bird', 'ant']. A photo containing cats and dogs would be represented as [1, 1, 0, 0, 0].
I’m training a sentiment classification model of sorts that takes as input a sentence and then outputs however many categories it thinks that sentence belongs to. Often times it will only belong to one category, maybe “happy” or “sad,” but sometimes it could belong to several, like “happy, excited, nervous, etc.”
Right now I have the data in a pandas dataframe with the sentences in one column and then the categories that those sentences belong to in another column separated by commas:
SENTENCES SENTIMENTS
sentence 1 || sentiment 1, sentiment 2
sentence 2 || sentiment 1
sentence 3 || sentiment 2, sentiment 3, sentiment 6, sentiment 7
sentence 4 || sentiment 4, sentiment 5
I was able to find info on one-hot encoding, but I cannot for the life of me figure out how to multi-hot encode the sentiments column.
Does anyone know of a way to multi-hot encode a column in a Pandas dataframe?
Surely this is a common thing?
(Sorry for the bad representations of the data. I’m on my phone and haven’t yet figured out how to insert code into a post using the app.)
Edit: formatting
If I understand Multi Hot encoding correctly basically we take arrays of different lengths, and turn them into arrays/lists of all the same length with each index of the new array/list being on(0) or off(1) for the various entries in the original array.
so
[1,3] would become [0, 1, 0, 1]
My question is, how do we account for if a number appears twice in the original array/list,
for example [1,3,3]
I have a pandas data series of strings that each have a bunch of text symbols in them (I'll call them words for discussion's sake, but in my use case they aren't actually words). The series of strings is already parsed to give me a vocabulary of all the words found anywhere in the series.
I've taken that vocabulary (vocab1, a list of all the words in the vocabulary) and made a dict (vocab) to assign an index to each word:
vocab = {c:i for i,c in enumerate(vocab1)}I need to change these strings into a multi-hot encoded numpy array (for machine learning blah blah). The arrays are the size of the vocabulary. 1 at a given position in the array indicates the absence of the corresponding word in the string, while 0 indicates its absence. The strings typically only have a few words in them (which sometimes are redundantly repeated in the original data) but there are 30k different words in the vocabulary, and in the broader use case this goes up to 260k or so.
Anyway, vocablength is the number of words in the vocabulary. I'm using pandas.Series.transform to apply my function to the entire series, which is over 200k strings (and will be millions in the broader use case).
def multihot_codes(codes):
o = np.zeros((vocablength,),dtype=np.int32)
for k in codes.split():
o[vocab[k]] = 1
return o
df["codesmh"] = df["codess"].transform(multihot_codes)This takes close to forever and has significant disk usage. There's probably a faster way to do this. For example, Tensorflow generates binary one-hot vectors for each word and then does a row reduce on them, but if you dig into the code, they have a compiled C++ module that does the heavy lifting. Not exactly what I'm prepared to do here myself....
Any ideas on how to improve this, or if there's something already in numpy/scipy or another package that's suited for use here? Thanks in advance!
Considering the input data as:
import pandas as pd
data = {'Title': ['Movie 1', 'Movie 2', 'Movie 3', 'Movie 4', 'Movie 5'],
'column': ['Action, Fantasy', 'Fantasy, Drama', 'Action', 'Sci-Fi, Romance, Comedy', np.nan]}
df = pd.DataFrame(data)
df
Title column
0 Movie 1 Action, Fantasy
1 Movie 2 Fantasy, Drama
2 Movie 3 Action
3 Movie 4 Sci-Fi, Romance, Comedy
4 Movie 5 NaN
This code produces the desired output:
# treat null values
df['column'].fillna('NA', inplace = True)
# separate all genres into one list, considering comma + space as separators
genre = df['column'].str.split(', ').tolist()
# flatten the list
flat_genre = [item for sublist in genre for item in sublist]
# convert to a set to make unique
set_genre = set(flat_genre)
# back to list
unique_genre = list(set_genre)
# remove NA
unique_genre.remove('NA')
# create columns by each unique genre
df = df.reindex(df.columns.tolist() + unique_genre, axis=1, fill_value=0)
# for each value inside column, update the dummy
for index, row in df.iterrows():
for val in row.column.split(', '):
if val != 'NA':
df.loc[index, val] = 1
df.drop('column', axis = 1, inplace = True)
df
Title Action Fantasy Comedy Sci-Fi Drama Romance
0 Movie 1 1 1 0 0 0 0
1 Movie 2 0 1 0 0 1 0
2 Movie 3 1 0 0 0 0 0
3 Movie 4 0 0 1 1 0 1
4 Movie 5 0 0 0 0 0 0
UPDATE: I've added a null value into the test data, and treat it appropriately in the first line of the solution.
### Import libraries and load sample data
import numpy as np
import pandas as pd
data = {
'Movie 1': ['Action, Fantasy'],
'Movie 2': ['Fantasy, Drama'],
'Movie 3': ['Action'],
'Movie 4': ['Sci-Fi, Romance, Comedy'],
'Movie 5': ['NA'],
}
df = pd.DataFrame.from_dict(data, orient='index')
df.rename(columns={0:'column'}, inplace=True)
At this stage our DataFrame looks like this:
column
Movie 1 Action, Fantasy
Movie 2 Fantasy, Drama
Movie 3 Action
Movie 4 Sci-Fi, Romance, Comedy
Movie 5 NA
Now, the question we're asking is - does a given genre word ("sub-string") occur in 'column' for a given movie?
To do this we'll first need a list of genre words:
### Join every string in every row, split the result, pull out the unique values.
genres = np.unique(', '.join(df['column']).split(', '))
### Drop 'NA'
genres = np.delete(genres, np.where(genres == 'NA'))
Depending on how large your dataset is, this could be computationally costly. You mentioned that you know the unique values already. So you could just define the iterable 'genres' manually.
Getting the OneHotVectors:
for genre in genres:
df[genre] = df['column'].str.contains(genre).astype('int')
df.drop('column', axis=1, inplace=True)
We loop through each genre, we ask whether the genre exists in 'column', this returns a True or False, which is converted to 1 or 0 respectively - when we cast to type('int').
We end up with:
Action Comedy Drama Fantasy Romance Sci-Fi
Movie 1 1 0 0 1 0 0
Movie 2 0 0 1 1 0 0
Movie 3 1 0 0 0 0 0
Movie 4 0 1 0 0 1 1
Movie 5 0 0 0 0 0 0
You need to make your variables to be categorical and then you can use one hot encoding as shown:
In [18]: df1 = pd.DataFrame({"class":pd.Series(['2','1','3']).astype('category',categories=['1','2','3','4','5'])})
In [19]: df2 = pd.DataFrame({"class":pd.Series(['3','3','5']).astype('category',categories=['1','2','3','4','5'])})
In [20]: df_1 = pd.get_dummies(df1)
In [21]: df_2 = pd.get_dummies(df2)
In [22]: df_1.add(df_2).apply(lambda x: x * [i for i in range(1,len(df_1.columns)+1)], axis = 1).astype(int).rename_axis('id')
Out[22]:
class_1 class_2 class_3 class_4 class_5
id
0 0 2 3 0 0
1 1 0 3 0 0
2 0 0 3 0 5
Does this satisfy your problem as stated?
#!/usr/bin/python
input = [
(0, (2,3)),
(1, (1,3)),
(2, (3,5)),
]
maximum = max(reduce(lambda x, y: x+list(y[1]), input, []))
# Or ...
# maximum = 0
# for i, classes in input:
# maximum = max(maximum, *classes)
# print header.
print "\t".join(["id"] + ["class_%d" % i for i in range(1, 6)])
for i, classes in input:
print i,
for r in range(1, maximum+1):
print "\t",
if r in classes:
print float(r),
else:
print 0.0,
print
Output:
id class_1 class_2 class_3 class_4 class_5
0 0.0 2.0 3.0 0.0 0.0
1 1.0 0.0 3.0 0.0 0.0
2 0.0 0.0 3.0 0.0 5.0
Use get_dummies with DataFrame.set_index and aggregate max or sum:
#always 0,1 in output
df1 = pd.get_dummies(df.set_index('Name')['Class']).max(level=0).reset_index()
#if need count values
#df1 = pd.get_dummies(df.set_index('Name')['Class']).sum(level=0).reset_index()
print (df1)
Name FB GRS TWT
0 Aci 1 1 0
1 Dan 1 0 1
2 Ann 0 1 0
The other solutions were failing for me due to memory overflows when I have 315 samples and 1908 labels. Here are some more performant methods.
Using pandas pivot:
def multihotencode(data, samples_col, labels_col):
data = data.copy()
data['present'] = 1
multihot = data.pivot(index=samples_col, columns=labels_col, values='present')
return multihot.fillna(0).astype(int)
Using sklearn:
import pandas as pd
from sklearn.preprocessing import MultiLabelBinarizer
def multihotencode(data, samples_col, labels_col):
## Aggregate the labels into a tuple for each sample
tupled = data.groupby(samples_col).apply(lambda x: tuple(x[labels_col].to_numpy()))
## Get multi-hot encoded matrix.
mlb = MultiLabelBinarizer()
multihot = mlb.fit_transform(samplepgs)
## Reapply the index and column headers
return pd.DataFrame(multihot, index=tupled.index, columns=mlb.classes_)
Both methods give the same result:
df = pd.DataFrame([
['Aci', 'FB'],
['Dan', 'TWT'],
['Ann', 'GRS'],
['Aci', 'GRS'],
['Dan', 'FB']
], columns=['Name', 'Class'])
multihotencode(df, 'Name', 'Class')
>>>
FB GRS TWT
Name
Aci 1 1 0
Ann 0 1 0
Dan 1 0 1