You just do the transform() on the test data (and do not fit the encoder again). The values that don't occur in "training" datasets would be encoded as 0 in all of the categories (as long as you won't change the handle_unknown parameter). For example:

import category_encoders as ce

train = pd.DataFrame({"var1": ["A", "B", "A", "B", "C"], "var2":["A", "A", "A", "A", "B"]})

encoder = ce.BinaryEncoder(cols = ['var1', 'var2'] , return_df = True)
x_train_data = encoder.fit_transform(train)

#   var1_0  var1_1  var2_0  var2_1
#0  0       1       0       1
#1  1       0       0       1
#2  0       1       0       1
#3  1       0       0       1
#4  1       1       1       0

test = pd.DataFrame({"var1": ["C", "D", "B"], "var2":["A", "C", "F"]})
x_test_data = encoder.transform(test)

#   var1_0  var1_1  var2_0  var2_1
#0  1       1       0       1
#1  0       0       0       0
#2  1       0       0       0

'D' doesn't occur in var1 in training data, so it was encoded as 0 0. 'C' and 'F'don't occur in var2 in training data, so they were both encoded as 0 0.

Answer from Daniel Wlazło on Stack Overflow
🌐
Packtpub
subscription.packtpub.com › book › data › 9781804611302 › 2 › ch02lvl1sec21 › performing-binary-encoding
Chapter 2: Encoding Categorical Variables | Python Feature Engineering Cookbook
Import the required Python library, function, and class: import pandas as pd from sklearn.model_selection import train_test_split from category_encoders.binary import BinaryEncoder
Starred by 79 users
Forked by 48 users
Languages: Python
🌐
YouTube
youtube.com › watch
How to perform Frequency Encoding & Binary Encoding | Python - YouTube
⭐️ Content Description ⭐️In this video, I have explained on how to perform frequency encoding and binary encoding in python. This encoding techniques are use...
Published: May 19, 2022
Top answer
1 of 2
2

Your categorical variable has two levels, so there is no actual difference between dummy-coding vs. simply entering the variable into the analysis. That is, to dummy code you would create one new variable with two values but your original variable is already one variable with two values. Dummy-coding is important for variables with more than two possible values. So, in this case the computer won't consider Pave > Grvl.

But if you have more than two variables then you should use dummy variables.

For your data, you can use pandas.get_dummies() or sklearn's one hot encoder to achieve your result.

2 of 2
0
  1. How to encode?

sklearn.preprocessing provides various classes for this purpose, LabelBinarizer is one of them.

  1. Wouldn't the computer interpret this as Pave > Grvl?

Consider an example, where people prefer Paved house in comparison to Graveled. Then their is a relationship between values and hence it should be treated as something like you have mentioned, other wise it should be independent values(refer next answer).

  1. How does the computer differentiate between binary and integer encoding?

As I mentioned above, if the categorical values have some relationship(as mentioned above), then in such a case it should be integer values(0,1,2 and so on), otherwise it should be binary. Binary representation will help us in presenting as an independent value to ML model (however it doesn't make much sense in this case as you just have 2 values). But consider an example where a feature have more than 2 categorical values. If they all are independent then it should be represented as binary value i.e in the form of OneHotEncoding(refer sklearn.preprocessing classes).

🌐
Python
docs.python.org › 3 › library › codecs.html
codecs — Codec registry and base classes
The following codecs provide str to bytes encoding and bytes-like object to str decoding, similar to the Unicode text encodings. Changed in version 3.8: “unicode_internal” codec is removed. The following codecs provide binary transforms: bytes-like object to bytes mappings.
🌐
YouTube
youtube.com › watch
Binary encoding - Encoding Techniques in Machine Learning - Python - YouTube
Watch Video to understand the overview of Binary Encoding technique in python for encoding categorical data.#binaryencodingmachinelearning #binaryencodingpyt...
Published: December 29, 2022
Find elsewhere
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders
Category Encoders — Category Encoders 2.11.1 documentation
import category_encoders as ce encoder = ce.BackwardDifferenceEncoder(cols=[...]) encoder = ce.BaseNEncoder(cols=[...]) encoder = ce.BinaryEncoder(cols=[...]) encoder = ce.CatBoostEncoder(cols=[...]) encoder = ce.CountEncoder(cols=[...]) encoder = ce.CountTargetEncoder(cols=[...]) encoder = ce.GLMMEncoder(cols=[...]) encoder = ce.GrayEncoder(cols=[...]) encoder = ce.HashingEncoder(cols=[...]) encoder = ce.HelmertEncoder(cols=[...]) encoder = ce.JamesSteinEncoder(cols=[...]) encoder = ce.LeaveOneOutEncoder(cols=[...]) encoder = ce.MEstimateEncoder(cols=[...]) encoder = ce.MultiHotEncoder(cols=[
🌐
Analytics Vidhya
analyticsvidhya.com › home › what are categorical data encoding methods | binary encoding
What are Categorical Data Encoding Methods | Binary Encoding
May 1, 2025 - Binary encoding efficiently converts categories into binary digits, making it suitable for high cardinality datasets. In contrast, one-hot encoding creates separate binary columns for each category.
🌐
Feaz-book
feaz-book.com › categorical-binary
Feature Engineering A-Z | Binary Encoding – Feature Engineering A-Z
We will be using the ames data set for these examples. The step_encoding_binary() function from the extrasteps package allows us to perform binary encoding.
🌐
Back2code
back2code.me › 2017 › 09 › numeric-and-binary-encoders-in-python
Numeric and Binary Encoders in Python - Back 2 Code
September 24, 2017 - Encode categorical variable into dummy/indicator (binary) variables: Pandas get_dummies and scikit-learn OneHotEncoder.
🌐
Medium
garg-shelvi.medium.com › category-encoders-c2a9bb192f0a
How to Encode Categorical Data | by Shelvi Garg | Medium
July 12, 2022 - category_encoders is an amazing python library that provides 15 different encoding schemes. One Hot Encoding · Label Encoding · Ordinal Encoding · Helmert Encoding · Binary Encoding · Frequency Encoding · Mean Encoding · Weight of Evidence Encoding · Probability Ratio Encoding ·
Top answer
1 of 1
8

My question is WHY? I realize that this reduces the dimensionality of the output data, but what is the logic behind doing this?

Basically, the issue of categorical encoding is to make your algorithm it's dealing with categorical features. Therefore, several methods are available for doing it, including binary encoding. Actually, it's logic is close to the logic of One Hot Encoding (OHE), if you understood it.

For binary encoding, each unique label in your categorical vector is associated randomly to a number between (0) and (the number of unique labels-1). Now, you encode this number in base 2 and "transcript" the previous number in 0 and 1 through the newly created columns. As an example, let's say your dataset as three different labels: 'A', 'B' & 'C'.
The following correspondance is randomly built:

'A' -> 1 -> 01;

'B' -> 2 > 10;

'C' -> 0 -> 00.

Therefore, an example of encoding of a given dataset is:

index my_category enc_category_0 enc_category_1

0 A, 1, 0

1, B, 0, 1

2, C, 0, 0

3 A, 1, 0

Regarding the utility of it, as you said it's reduce the dimensionality. Besides, I guess it helps not having too much zeros in the encoded columns as with OHE. Here is an interesting post: https://medium.com/data-design/visiting-categorical-features-and-encoding-in-decision-trees-53400fa65931

How does taking the log2 determine how many digits are needed to represent the data? If you understood the working principle, you understand the use of the log2. Computing the log2 of a number retrives the necessary number of digits for a binary encoding of this number. Example: [log2(10)]=[3.32]=4, 4 digits are needed for binary encode 10.

For more info about the implementation and code example: http://contrib.scikit-learn.org/categorical-encoding/_modules/category_encoders/binary.html#BinaryEncoder

Hope I was clear,

Tchau

🌐
GeeksforGeeks
geeksforgeeks.org › encoding-categorical-data-in-sklearn
Encoding Categorical Data in Sklearn - GeeksforGeeks
November 25, 2024 - Python · from sklearn.preprocessing ... Output: [2 0 1 0 2] One-Hot Encoding converts categorical data into a binary matrix, where each category is represented by a binary vector....
🌐
GitHub
github.com › chronoxor › FastBinaryEncoding
GitHub - chronoxor/FastBinaryEncoding: Fast Binary Encoding is ultra fast and universal serialization solution for C++, C#, Go, Java, JavaScript, Kotlin, Python, Ruby, Swift · GitHub
Fast Binary Encoding is ultra fast and universal serialization solution for C++, C#, Go, Java, JavaScript, Kotlin, Python, Ruby, Swift - chronoxor/FastBinaryEncoding
Author: chronoxor
🌐
KDnuggets
kdnuggets.com › crack-the-code-mastering-category-encoders-for-data-scientists
Crack the Code: Mastering Category Encoders for Data Scientists - KDnuggets
September 16, 2024 - The category_encoders package in Python provides a variety of encoding techniques that cater to different types of data and scenarios. You can install it using pip: ... One-hot encoding converts each category into a binary vector.