🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › binary.html
Binary — Category Encoders 2.11.1 documentation
set_transform_request(*, override_return_df: bool | None | str = '$UNCHANGED$') → BinaryEncoder · Configure whether metadata should be requested to be passed to the transform method. Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()).
🌐
Medium
medium.com › @jaberi.mohamedhabib › encoding-categorical-variables-methods-and-techniques-in-pandas-scikit-learn-and-using-dummy-216ae2d5128d
Encoding Categorical Variables: Methods and Techniques in Pandas, Scikit-learn, and Using Dummy Function | by JABERI Mohamed Habib | Medium
September 27, 2024 - from sklearn.preprocessing import OrdinalEncoder # Sample data df = pd.DataFrame({'Size': ['Small', 'Medium', 'Large', 'Small', 'Large']}) # Ordinal Encoding oe = OrdinalEncoder(categories=[['Small', 'Medium', 'Large']]) df['Size_OrdinalEncoded'] ...
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › binary
Binary — Category Encoders 2.8.1 documentation
set_transform_request(*, override_return_df: bool | None | str = '$UNCHANGED$') → BinaryEncoder · Configure whether metadata should be requested to be passed to the transform method. Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()).
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › encoding-categorical-data-in-sklearn
Encoding Categorical Data in Sklearn - GeeksforGeeks
September 17, 2025 - Python · from sklearn.preprocessing import LabelEncoder le = LabelEncoder() df['class_encoded'] = le.fit_transform(df['class']) print("Class labels mapping:", dict(zip(le.classes_, le.transform(le.classes_)))) print(df[['class', 'class_encoded']].head()) Label Encoding · Now we will use One-Hot encoding which creates separate binary columns for each category, ideal for nominal data with no natural order.
🌐
GitHub
github.com › scikit-learn-contrib › category_encoders
GitHub - scikit-learn-contrib/category_encoders: A library of sklearn compatible categorical variable encoders · GitHub
An unsupervised example: from ... 51, 38], }) # use binary encoding to encode two categorical features enc = BinaryEncoder(cols=['gender', 'country']).fit(X) # transform the dataset numeric_dataset = enc.transform...
Author: scikit-learn-contrib
🌐
Feaz-book
feaz-book.com › categorical-binary
Feature Engineering A-Z | Binary Encoding – Feature Engineering A-Z
We are using the ames data set for examples. {category_encoders} provided the BinaryEncoder() method we can use. from feazdata import ames from sklearn.compose import ColumnTransformer from category_encoders.binary import BinaryEncoder ct = ColumnTransformer( [('binary', BinaryEncoder(), ['MS_Zoning'])], remainder="passthrough") ct.fit(ames) ColumnTransformer(remainder='passthrough', transformers=[('binary', BinaryEncoder(), ['MS_Zoning'])])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.LabelBinarizer.html
LabelBinarizer — scikit-learn 1.9.1 documentation
Possible type are ‘continuous’, ‘continuous-multioutput’, ‘binary’, ‘multiclass’, ‘multiclass-multioutput’, ‘multilabel-indicator’, and ‘unknown’. ... False otherwise. ... Function to perform the transform operation of LabelBinarizer with fixed classes. ... Encode categorical features using a one-hot aka one-of-K scheme. ... >>> from sklearn.preprocessing import LabelBinarizer >>> lb = LabelBinarizer() >>> lb.fit([1, 2, 6, 4, 2]) LabelBinarizer() >>> lb.classes_ array([1, 2, 4, 6]) >>> lb.transform([1, 6]) array([[1, 0, 0, 0], [0, 0, 0, 1]])
🌐
Datasciencehorizons
datasciencehorizons.com › handling-categorical-variables-scikit-learn-strategies-encoding-techniques
Handling Categorical Variables in scikit-learn: Strategies and Encoding Techniques – Data Science Horizons
For example, consider a “City” categorical feature with three possible values – “New York”, “Chicago”, and “Los Angeles”. Using one-hot encoding, this would get transformed into three separate binary columns: City_NewYork: 1 0 0 City_Chicago: 0 1 0 City_LosAngeles: 0 0 1 · ...
Top answer
1 of 5
46

A simple example which encodes an array using LabelEncoder, OneHotEncoder, LabelBinarizer is shown below.

I see that OneHotEncoder needs data in integer encoded form first to convert into its respective encoding which is not required in the case of LabelBinarizer.

from numpy import array
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import OneHotEncoder
from sklearn.preprocessing import LabelBinarizer

# define example
data = ['cold', 'cold', 'warm', 'cold', 'hot', 'hot', 'warm', 'cold', 
'warm', 'hot']
values = array(data)
print "Data: ", values
# integer encode
label_encoder = LabelEncoder()
integer_encoded = label_encoder.fit_transform(values)
print "Label Encoder:" ,integer_encoded

# onehot encode
onehot_encoder = OneHotEncoder(sparse=False)
integer_encoded = integer_encoded.reshape(len(integer_encoded), 1)
onehot_encoded = onehot_encoder.fit_transform(integer_encoded)
print "OneHot Encoder:", onehot_encoded

#Binary encode
lb = LabelBinarizer()
print "Label Binarizer:", lb.fit_transform(values)

Another good link which explains the OneHotEncoder is: Explain onehotencoder using python

There may be other valid differences between the two which experts can probably explain.

2 of 5
31

A difference is that you can use OneHotEncoder for multi column data, while not for LabelBinarizer and LabelEncoder.

from sklearn.preprocessing import LabelBinarizer, LabelEncoder, OneHotEncoder

X = [["US", "M"], ["UK", "M"], ["FR", "F"]]
OneHotEncoder().fit_transform(X).toarray()

# array([[0., 0., 1., 0., 1.],
#        [0., 1., 0., 0., 1.],
#        [1., 0., 0., 1., 0.]])
LabelBinarizer().fit_transform(X)
# ValueError: Multioutput target data is not supported with label binarization

LabelEncoder().fit_transform(X)
# ValueError: bad input shape (3, 2)
Find elsewhere
🌐
Packtpub
subscription.packtpub.com › book › data › 9781804611302 › 2 › ch02lvl1sec21 › performing-binary-encoding
Chapter 2: Encoding Categorical Variables | Python Feature Engineering Cookbook
In this recipe, we will learn how to perform binary encoding using Category Encoders. First, let’s import the necessary Python libraries and get the dataset ready: Import the required Python library, function, and class: import pandas as pd from sklearn.model_selection import train_test_split from category_encoders.binary import BinaryEncoder
🌐
GitHub
github.com › scikit-learn-contrib › category_encoders › blob › master › category_encoders › binary.py
category_encoders/category_encoders/binary.py at master · scikit-learn-contrib/category_encoders
Example · ------- >>> from category_encoders import * >>> import pandas as pd · >>> from sklearn.datasets import fetch_openml · >>> bunch = fetch_openml(name='house_prices', as_frame=True) >>> display_cols = [ ... 'Id',  ...
Author: scikit-learn-contrib
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders › _modules › category_encoders › binary.html
category_encoders.binary — Category Encoders 2.8.1 documentation
Example ------- >>> from category_encoders import * >>> import pandas as pd >>> from sklearn.datasets import fetch_openml >>> bunch = fetch_openml(name='house_prices', as_frame=True) >>> display_cols = [ ... 'Id', ... 'MSSubClass', ... 'MSZoning', ... 'LotFrontage', ... 'YearBuilt', ... 'Heating', ... 'CentralAir', ... ] >>> y = bunch.target >>> X = pd.DataFrame(bunch.data, columns=bunch.feature_names)[display_cols] >>> enc = BinaryEncoder(cols=['CentralAir', 'Heating']).fit(X, y) >>> numeric_dataset = enc.transform(X) >>> print(numeric_dataset.info()) <class 'pandas.core.frame.DataFrame'> Ran
🌐
Medium
garg-shelvi.medium.com › category-encoders-c2a9bb192f0a
How to Encode Categorical Data | by Shelvi Garg | Medium
July 12, 2022 - Sklearn Preprocessing · Python’s get_dummies · Binary Encoding · Frequency Encoding · Label Encoding · Ordinal Encoding · Categorical data is a type of data that is used to group information with similar characteristics while Numerical data is a type of data that expresses information in the form of numbers. Example: Gender ·
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.5.2 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.label_binarize.html
label_binarize — scikit-learn 1.9.0 documentation
... Class used to wrap the ... from sklearn.preprocessing import label_binarize >>> label_binarize([1, 6], classes=[1, 2, 4, 6]) array([[1, 0, 0, 0], [0, 0, 0, 1]])...
🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.preprocessing.LabelBinarizer.html
LabelBinarizer — scikit-learn 1.8.dev0 documentation
Possible type are ‘continuous’, ‘continuous-multioutput’, ‘binary’, ‘multiclass’, ‘multiclass-multioutput’, ‘multilabel-indicator’, and ‘unknown’. ... False otherwise. ... Function to perform the transform operation of LabelBinarizer with fixed classes. ... Encode categorical features using a one-hot aka one-of-K scheme. ... >>> from sklearn.preprocessing import LabelBinarizer >>> lb = LabelBinarizer() >>> lb.fit([1, 2, 6, 4, 2]) LabelBinarizer() >>> lb.classes_ array([1, 2, 4, 6]) >>> lb.transform([1, 6]) array([[1, 0, 0, 0], [0, 0, 0, 1]])
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.MultiLabelBinarizer.html
MultiLabelBinarizer — scikit-learn 1.9.1 documentation
Set to True if output binary array is desired in CSR sparse format. ... A copy of the classes parameter when provided. Otherwise it corresponds to the sorted set of classes found when fitting. ... Encode categorical features using a one-hot aka one-of-K scheme. ... >>> from sklearn.preprocessing import MultiLabelBinarizer >>> mlb = MultiLabelBinarizer() >>> mlb.fit_transform([(1, 2), (3,)]) array([[1, 1, 0], [0, 0, 1]]) >>> mlb.classes_ array([1, 2, 3])
Top answer
1 of 1
8

My question is WHY? I realize that this reduces the dimensionality of the output data, but what is the logic behind doing this?

Basically, the issue of categorical encoding is to make your algorithm it's dealing with categorical features. Therefore, several methods are available for doing it, including binary encoding. Actually, it's logic is close to the logic of One Hot Encoding (OHE), if you understood it.

For binary encoding, each unique label in your categorical vector is associated randomly to a number between (0) and (the number of unique labels-1). Now, you encode this number in base 2 and "transcript" the previous number in 0 and 1 through the newly created columns. As an example, let's say your dataset as three different labels: 'A', 'B' & 'C'.
The following correspondance is randomly built:

'A' -> 1 -> 01;

'B' -> 2 > 10;

'C' -> 0 -> 00.

Therefore, an example of encoding of a given dataset is:

index my_category enc_category_0 enc_category_1

0 A, 1, 0

1, B, 0, 1

2, C, 0, 0

3 A, 1, 0

Regarding the utility of it, as you said it's reduce the dimensionality. Besides, I guess it helps not having too much zeros in the encoded columns as with OHE. Here is an interesting post: https://medium.com/data-design/visiting-categorical-features-and-encoding-in-decision-trees-53400fa65931

How does taking the log2 determine how many digits are needed to represent the data? If you understood the working principle, you understand the use of the log2. Computing the log2 of a number retrives the necessary number of digits for a binary encoding of this number. Example: [log2(10)]=[3.32]=4, 4 digits are needed for binary encode 10.

For more info about the implementation and code example: http://contrib.scikit-learn.org/categorical-encoding/_modules/category_encoders/binary.html#BinaryEncoder

Hope I was clear,

Tchau