Your original approach, without one-hot encoding, was doing what you wanted.
One-hot encoding is meant for inputs to many models, but outputs for only a few (e.g. training a neural network with cross-entropy loss). So these are only needed for some algorithm implementations, while others can do fine without it.
For output labels, a classifier like RandomForest is just fine with strings and multiple classes.
Answer from mcskinner on Stack OverflowYour original approach, without one-hot encoding, was doing what you wanted.
One-hot encoding is meant for inputs to many models, but outputs for only a few (e.g. training a neural network with cross-entropy loss). So these are only needed for some algorithm implementations, while others can do fine without it.
For output labels, a classifier like RandomForest is just fine with strings and multiple classes.
You don't have to make one hot encoding when using random forest in sklearn.
What you need is "label encoder", and your Y should looks like
from sklearn.preprocessing import LabelEncoder
y = ["A","B","D","A","C"]
le = LabelEncoder()
le.fit_transform(y)
# array([0, 1, 3, 0, 2], dtype=int64)
I tried to modified the sample code sklearn provided :
from sklearn.ensemble import RandomForestClassifier
import numpy as np
from sklearn.datasets import make_classification
>>> X, y = make_classification(n_samples=1000, n_features=4,
... n_informative=2, n_redundant=0,
... random_state=0, shuffle=False)
y = np.random.choice(["A","B","C","D"],1000)
print(y.shape)
>>> clf = RandomForestClassifier(max_depth=2, random_state=0)
>>> clf.fit(X, y)
>>> clf.classes_
# array(['A', 'B', 'C', 'D'], dtype='<U1')
Either process the y with label encoding or without, it both worked with RandomForestClassifier.
python - Multi-label one-hot encoding - Data Science Stack Exchange
machine learning - One hot encoding for multiple label(trainy) in .fit() method? - Data Science Stack Exchange
python - One-Hot Encoding of label not needed? - Stack Overflow
What exactly is multi-hot encoding and how is it different from one-hot? - Cross Validated
So now I am totally confused how people use one hot encoding for target variable in practice? please help me out.
This mostly comes down to the tool you are using. Sklearn, which I assume you are using, does not use one hot encoded target variables. So your y should be of dimension (1600, 1) where the classes are 0, 1, 2 and 3. Instead of applying one hot encoding you can use a LabelEncoder to get it on the correct format.
I suspect the reason for your confusion comes from having seen deep learning frameworks such as Tensorflow and Keras. With them you always one hot encode your target variable.
Short answer:
- Using sklearn: Label encode
- Using deep learning: one hot encode
Logistic regression is used for a binary response (two categories). Here are some things that you may consider doing:
1) If your target variables is normally distributed use ordinary least squares (OLS) regression and manually choose cutoffs for your categories (i.e. 0-25 = 1, 26-50 = 2, 51-75 = 3, 76-100 = 4, or similar)
2) Otherwise, you should consider multinomial logistic regression. Here is a great summary of multinomial logistic regression in R: https://stats.idre.ucla.edu/r/dae/multinomial-logistic-regression/
I think you might be confusing a multiclass (your case) with a multioutput classification.
In multiclass classification problems, your output should only be a single target column, and you'll be training the model to classify among the classes in that column. You'd have to split into separate target columns, in the case you had to predict n different classes per sample, which is not the case, you only want one of the targets per sample.
So for multiclass classification, there's no need to OneHotEncode the target, since you only want a single target column (which can also be categorical in SVC). What you do have to encode, either using OneHotEncoder or with some other encoders, is the categorical input features, which have to be numeric.
Also, SVC can deal with categorical targets, since it LabelEncode's them internally:
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y)
y_train_categorical = load_iris()['target_names'][y_train]
# array(['setosa', 'setosa', 'versicolor',...
sv = SVC()
sv.fit(X_train, y_train_categorical)
sv.classes_
# array(['setosa', 'versicolor', 'virginica'], dtype='<U10')
As far as I know, one hot encoding is never done on the output. You need to do one hot encoding on a feature so that the model never confuses that some color is greater than other colors. When you are computing the output the models use probability distributions based on classes. So there won't be any problem here.
In a nutshell, you should do one hot encoding only on the input features and not on the output classes.
Imagine your have five different classes e.g. ['cat', 'dog', 'fish', 'bird', 'ant']. If you would use one-hot-encoding you would represent the presence of 'dog' in a five-dimensional binary vector like [0,1,0,0,0]. If you would use multi-hot-encoding you would first label-encode your classes, thus having only a single number which represents the presence of a class (e.g. 1 for 'dog') and then convert the numerical labels to binary vectors of size .
Examples:
'cat' = [0,0,0]
'dog' = [0,0,1]
'fish' = [0,1,0]
'bird' = [0,1,1]
'ant' = [1,0,0]
This representation is basically the middle way between label-encoding, where you introduce false class relationships (0 < 1 < 2 < ... < 4, thus 'cat' < 'dog' < ... < 'ant') but only need a single value to represent class presence and one-hot-encoding, where you need a vector of size (which can be huge!) to represent all classes but have no false relationships.
Note: multi-hot-encoding introduces false additive relationships, e.g. [0,0,1] + [0,1,0] = [0,1,1] that is 'dog' + 'fish' = 'bird'. That is the price you pay for the reduced representation.
The accepted answer seems rather eccentric to me. I think that is rarely done, if ever, and will usually yield bad results.
There's a much more common, sensible use case for this. "Multi-hot encoding" doesn't seem to be a standard term, but I'm not sure there's any standard term. scikit-learn refers to a multi label binarizer.
This is simply used for multi label problems. That is, problems where more than one label can be associated with each example.
For example, say you are trying to detect whether certain types of animal are in a photo. Note that multiple types of animal can be in a single photo. Say the possible types of animal are ['cat', 'dog', 'fish', 'bird', 'ant']. A photo containing cats and dogs would be represented as [1, 1, 0, 0, 0].