Most machine learning models accept only numerical variables. This is the reason behind why categorical variables are converted to number so the model can understand better.

Now lets address your second query lets look into what is one-hot encoding and dummy encoding and then see the difference

  • One hot Encoding: Take the example of column name Fruit which can have different types of fruits like Blackberry, Grape, Orange. Here each category is mapped to binary variable containing either 0 or 1. Widely utilized when features are nominal.

    Fruit Price (dollars per pound)
    Blackberry 3.82
    Grape 1.2
    Orange .64

    Post one hot encoding the table now looks as shown below

    One Hot Encoded table
    Blackberry Grape Orange Price (dollars per pound)
    1 0 0 3.82
    0 1 0 1.2
    0 0 1 .64
  • Dummy Encoding: similar to one hot encoding. While one hot encoding utilises N binary variables for N categories in a variable. Dummy encoding uses N-1 features to represent N labels/categories

    One Hot Coding Vs Dummy Coding
    Column One Hot Code Dummy Code
    Blackberry 100 10
    Grape 010 01
    Orange 001 00
Answer from Archana David on Stack Exchange
Top answer
1 of 3
13

Most machine learning models accept only numerical variables. This is the reason behind why categorical variables are converted to number so the model can understand better.

Now lets address your second query lets look into what is one-hot encoding and dummy encoding and then see the difference

  • One hot Encoding: Take the example of column name Fruit which can have different types of fruits like Blackberry, Grape, Orange. Here each category is mapped to binary variable containing either 0 or 1. Widely utilized when features are nominal.

    Fruit Price (dollars per pound)
    Blackberry 3.82
    Grape 1.2
    Orange .64

    Post one hot encoding the table now looks as shown below

    One Hot Encoded table
    Blackberry Grape Orange Price (dollars per pound)
    1 0 0 3.82
    0 1 0 1.2
    0 0 1 .64
  • Dummy Encoding: similar to one hot encoding. While one hot encoding utilises N binary variables for N categories in a variable. Dummy encoding uses N-1 features to represent N labels/categories

    One Hot Coding Vs Dummy Coding
    Column One Hot Code Dummy Code
    Blackberry 100 10
    Grape 010 01
    Orange 001 00
2 of 3
8

The purpose of one-hot encoding is to assign numbers to categorical variables which does not create a false, meaningless numerical pattern.

If you have categorical variables "Apple", "Orange", "Cherry", "Tomato" and you assign them numerical values 0, 1, 2, 3, then these numerical values have interpretations like "Cherry is between Tomato and Apple, but closer to Tomato" because 2 is between 0 and 3, but closer to 3. This is nonsense. It's bad nonsense, because algorithms to analyze this data (like regressions, or whatever) can pick up on it and read too much into it.

If you instead represent "Apple", "Orange", "Cherry", and "Tomato" as the 4-tuples (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), then you don't have this problem. Each coordinate is either 0 or 1, and measures the "Appleness" or the "Orangeness" or the "Cherriness" or the "Tomatoness" of your fruit. That's one-hot encoding.

As an example of this, suppose that the average apple weighs 200 grams, the average orange weighs 150 grams, the average cherry 30 grams, and the average tomato 100 grams. With one-hot encoding , this average weight is a linear function of the encoding: $200x_1 + 150x_2 + 30x_3 + 100x_4$. This is something a regression can figure out from data. With the 0, 1, 2, 3 encoding, there's no nice function that will give you the average weight of a fruit given its number.

Now, as for dummy encoding: one-hot encoding still has a problem, which is that the linear function is not unique. The function $100 + 100x_1 + 50x_2 - 70x_3$ gives the same values as the previous function at the points (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1). That's because each valid point satisfies .

(Again, this is not just a curiosity; this affects the way we analyze the data. For example, linear regressions behave badly when an -dimensional input doesn't actually range freely across all dimensions.)

Dummy encoding drops one of the coordinates, since it can be inferred from the other three, to avoid this issue. The four fruits might be encoded as (1,0,0), (0,1,0), (0,0,1), and (0,0,0).

🌐
Medium
medium.com › @javiersospedralegarda › cracking-the-categorical-dilemma-one-hot-encoding-vs-a9f3233b3e60
Cracking the Categorical Dilemma: One-Hot Encoding vs. Dummy Variables — Which is Your Machine Learning MVP? | by Javier Sospedra Legarda | Medium
July 26, 2024 - Imagine a “Color” feature with ... point with the color “Red” would be encoded as (1, 0, 0). ... Dummy encoding, like a savvy editor, trims the excess....
Discussions

python - What's the difference between dummy variable and one-hot encoding? - Stack Overflow
I'm making features for a machine learning model. I'm confused with dummy variable and one-hot encoding.For a instance,a category variable 'week' range 1-7.When using one-hot encoding, encode week ... More on stackoverflow.com
🌐 stackoverflow.com
regression - One-hot vs dummy encoding in Scikit-learn - Cross Validated
There are two different ways to encoding categorical variables. Say, one categorical variable has n values. One-hot encoding converts it into n variables, while dummy encoding converts it into n-1 More on stats.stackexchange.com
🌐 stats.stackexchange.com
July 16, 2016
What's the difference between one hot encoding and dummy encoding?
Dummy variable coding and one-hot encoding are exactly the same thing; the former term comes from statistics and the latter from computer science (borrowed from electronics). The dummy variable trap can happen regardless of what you call the technique, although it tends to be less of a concern in ML because (a) multicollinearity is only a big issue in linear regression, and can be minimized with regularization, and (b) in statistics, the aim is to determine the coefficients; in ML, it's to determine the predictions, so even if the coefficients are being warped by multicollinearity, it's of secondary concern. More on reddit.com
🌐 r/MLQuestions
2
2
November 6, 2016
regression - Problems with one-hot encoding vs. dummy encoding - Cross Validated
I am aware of the fact that categorical variables with k levels should be encoded with k-1 variables in dummy encoding (similarly for multi-valued categorical variables). I was wondering how much of a problem does a one-hot encoding (i.e. More on stats.stackexchange.com
🌐 stats.stackexchange.com
🌐
Towards Data Science
towardsdatascience.com › home › latest › encoding categorical variables: one-hot vs dummy encoding
Encoding Categorical Variables: One-hot vs Dummy Encoding | Towards Data Science
January 21, 2025 - To encode the same Color variable with three categories using the dummy encoding, we need to use only two dummy variables. ... Dummy encoding removes a duplicate category present in the one-hot encoding.
🌐
Medium
medium.com › data-scientists-diary › one-hot-encoding-vs-dummy-variables-d86f99e8ad5b
One Hot Encoding vs Dummy Variables | by Hey Amit | Data Scientist’s Diary | Medium
April 18, 2025 - They’re similar, but there’s a crucial difference that can save you a lot of headaches in certain situations. Definition: Dummy variables are a simplified version of one-hot encoding.
Top answer
1 of 2
11

In fact, there is no difference in the effect of the two approaches (rather wordings) on your regression.

In either case, you have to make sure that one of your dummies is left out (i.e. serves as base assumption) to avoid perfect multicollinearity among the set.

For instance, if you want to take the weekday of an observation into account, you only use 6 (not 7) dummies assuming the one left out to be the base variable. When using one-hot encoding, your weekday variable is present as a categorical value in one single column, effectively having the regression use the first of its values as the base.

2 of 2
4

Technically 6- a day week is enough to provide a unique mapping for a vocabulary of size 7:

 1. Sunday    [0,0,0,0,0,0]
 2. Monday    [1,0,0,0,0,0]
 3. Tuesday   [0,1,0,0,0,0]
 4. Wednesday [0,0,1,0,0,0]
 5. Thursday  [0,0,0,1,0,0]
 6. Friday    [0,0,0,0,1,0]
 7. Saturday  [0,0,0,0,0,1]

dummy coding is a more compact representation, it is preferred in statistical models that perform better when the inputs are linearly independent.

Modern machine learning algorithms, though, don’t require their inputs to be linearly independent and use methods such as L1 regularization to prune redundant inputs. The additional degree of freedom allows the framework to transparently handle a missing input in production as all zeros.

 1. Sunday    [0,0,0,0,0,0,1]
 2. Monday    [0,0,0,0,0,1,0]
 3. Tuesday   [0,0,0,0,1,0,0]
 4. Wednesday [0,0,0,1,0,0,0]
 5. Thursday  [0,0,1,0,0,0,0]
 6. Friday    [0,1,0,0,0,0,0]
 7. Saturday  [1,0,0,0,0,0,0]

 for missing values : [0,0,0,0,0,0,0]
Top answer
1 of 3
51

Scikit-learn's linear regression model allows users to disable intercept. So for one-hot encoding, should I always set fit_intercept=False? For dummy encoding, fit_intercept should always be set to True? I do not see any "warning" on the website.

For an unregularized linear model with one-hot encoding, yes, you need to set the intercept to be false or else incur perfect collinearity. sklearn also allows for a ridge shrinkage penalty, and in that case it is not necessary, and in fact you should include both the intercept and all the levels. For dummy encoding you should include an intercept, unless you have standardized all your variables, in which case the intercept is zero.

Since one-hot encoding generates more variables, does it have more degree of freedom than dummy encoding?

The intercept is an additional degree of freedom, so in a well specified model it all equals out.

For the second one, what if there are k categorical variables? k variables are removed in dummy encoding. Is the degree of freedom still the same?

You could not fit a model in which you used all the levels of both categorical variables, intercept or not. For, as soon as you have one-hot-encoded all the levels in one variable in the model, say with binary variables , then you have a linear combination of predictors equal to the constant vector

If you then try to enter all the levels of another categorical into the model, you end up with a distinct linear combination equal to a constant vector

and so you have created a linear dependency

So you must leave out a level in the second variable, and everything lines up properly.

Say, I have 3 categorical variables, each of which has 4 levels. In dummy encoding, 3*4-3=9 variables are built with one intercept. In one-hot encoding, 3*4=12 variables are built without an intercept. Am I correct?

The second thing does not actually work. The column design matrix you create will be singular. You need to remove three columns, one from each of three distinct categorical encodings, to recover non-singularity of your design.

2 of 3
7

To add a little to @MatthewDrury's answer regarding this question:

Say, I have 3 categorical variables, each of which has 4 levels. In dummy encoding, 3*4-3=9 variables are built with one intercept. In one-hot encoding, 3*4=12 variables are built without an intercept. Am I correct?

We can examine what the design matrix would look like with and without an intercept by using model.matrix from R.

With an intercept:

> df <- expand.grid(w = letters[1:4], x = letters[5:8], y = letters[9:12])
> model.matrix(~ w + x + y, df)
   (Intercept) wb wc wd xf xg xh yj yk yl
1            1  0  0  0  0  0  0  0  0  0
2            1  1  0  0  0  0  0  0  0  0
3            1  0  1  0  0  0  0  0  0  0
4            1  0  0  1  0  0  0  0  0  0
5            1  0  0  0  1  0  0  0  0  0
6            1  1  0  0  1  0  0  0  0  0
7            1  0  1  0  1  0  0  0  0  0
8            1  0  0  1  1  0  0  0  0  0
9            1  0  0  0  0  1  0  0  0  0
10           1  1  0  0  0  1  0  0  0  0
11           1  0  1  0  0  1  0  0  0  0
12           1  0  0  1  0  1  0  0  0  0
13           1  0  0  0  0  0  1  0  0  0
14           1  1  0  0  0  0  1  0  0  0
15           1  0  1  0  0  0  1  0  0  0
16           1  0  0  1  0  0  1  0  0  0
17           1  0  0  0  0  0  0  1  0  0
18           1  1  0  0  0  0  0  1  0  0
19           1  0  1  0  0  0  0  1  0  0
20           1  0  0  1  0  0  0  1  0  0
21           1  0  0  0  1  0  0  1  0  0
22           1  1  0  0  1  0  0  1  0  0
23           1  0  1  0  1  0  0  1  0  0
24           1  0  0  1  1  0  0  1  0  0
25           1  0  0  0  0  1  0  1  0  0
26           1  1  0  0  0  1  0  1  0  0
27           1  0  1  0  0  1  0  1  0  0
28           1  0  0  1  0  1  0  1  0  0
29           1  0  0  0  0  0  1  1  0  0
30           1  1  0  0  0  0  1  1  0  0
31           1  0  1  0  0  0  1  1  0  0
32           1  0  0  1  0  0  1  1  0  0
33           1  0  0  0  0  0  0  0  1  0
34           1  1  0  0  0  0  0  0  1  0
35           1  0  1  0  0  0  0  0  1  0
36           1  0  0  1  0  0  0  0  1  0
37           1  0  0  0  1  0  0  0  1  0
38           1  1  0  0  1  0  0  0  1  0
39           1  0  1  0  1  0  0  0  1  0
40           1  0  0  1  1  0  0  0  1  0
41           1  0  0  0  0  1  0  0  1  0
42           1  1  0  0  0  1  0  0  1  0
43           1  0  1  0  0  1  0  0  1  0
44           1  0  0  1  0  1  0  0  1  0
45           1  0  0  0  0  0  1  0  1  0
46           1  1  0  0  0  0  1  0  1  0
47           1  0  1  0  0  0  1  0  1  0
48           1  0  0  1  0  0  1  0  1  0
49           1  0  0  0  0  0  0  0  0  1
50           1  1  0  0  0  0  0  0  0  1
51           1  0  1  0  0  0  0  0  0  1
52           1  0  0  1  0  0  0  0  0  1
53           1  0  0  0  1  0  0  0  0  1
54           1  1  0  0  1  0  0  0  0  1
55           1  0  1  0  1  0  0  0  0  1
56           1  0  0  1  1  0  0  0  0  1
57           1  0  0  0  0  1  0  0  0  1
58           1  1  0  0  0  1  0  0  0  1
59           1  0  1  0  0  1  0  0  0  1
60           1  0  0  1  0  1  0  0  0  1
61           1  0  0  0  0  0  1  0  0  1
62           1  1  0  0  0  0  1  0  0  1
63           1  0  1  0  0  0  1  0  0  1
64           1  0  0  1  0  0  1  0  0  1

Without an intercept:

> model.matrix(~ w + x + y - 1, df)
   wa wb wc wd xf xg xh yj yk yl
1   1  0  0  0  0  0  0  0  0  0
2   0  1  0  0  0  0  0  0  0  0
3   0  0  1  0  0  0  0  0  0  0
4   0  0  0  1  0  0  0  0  0  0
5   1  0  0  0  1  0  0  0  0  0
6   0  1  0  0  1  0  0  0  0  0
7   0  0  1  0  1  0  0  0  0  0
8   0  0  0  1  1  0  0  0  0  0
9   1  0  0  0  0  1  0  0  0  0
10  0  1  0  0  0  1  0  0  0  0
11  0  0  1  0  0  1  0  0  0  0
12  0  0  0  1  0  1  0  0  0  0
13  1  0  0  0  0  0  1  0  0  0
14  0  1  0  0  0  0  1  0  0  0
15  0  0  1  0  0  0  1  0  0  0
16  0  0  0  1  0  0  1  0  0  0
17  1  0  0  0  0  0  0  1  0  0
18  0  1  0  0  0  0  0  1  0  0
19  0  0  1  0  0  0  0  1  0  0
20  0  0  0  1  0  0  0  1  0  0
21  1  0  0  0  1  0  0  1  0  0
22  0  1  0  0  1  0  0  1  0  0
23  0  0  1  0  1  0  0  1  0  0
24  0  0  0  1  1  0  0  1  0  0
25  1  0  0  0  0  1  0  1  0  0
26  0  1  0  0  0  1  0  1  0  0
27  0  0  1  0  0  1  0  1  0  0
28  0  0  0  1  0  1  0  1  0  0
29  1  0  0  0  0  0  1  1  0  0
30  0  1  0  0  0  0  1  1  0  0
31  0  0  1  0  0  0  1  1  0  0
32  0  0  0  1  0  0  1  1  0  0
33  1  0  0  0  0  0  0  0  1  0
34  0  1  0  0  0  0  0  0  1  0
35  0  0  1  0  0  0  0  0  1  0
36  0  0  0  1  0  0  0  0  1  0
37  1  0  0  0  1  0  0  0  1  0
38  0  1  0  0  1  0  0  0  1  0
39  0  0  1  0  1  0  0  0  1  0
40  0  0  0  1  1  0  0  0  1  0
41  1  0  0  0  0  1  0  0  1  0
42  0  1  0  0  0  1  0  0  1  0
43  0  0  1  0  0  1  0  0  1  0
44  0  0  0  1  0  1  0  0  1  0
45  1  0  0  0  0  0  1  0  1  0
46  0  1  0  0  0  0  1  0  1  0
47  0  0  1  0  0  0  1  0  1  0
48  0  0  0  1  0  0  1  0  1  0
49  1  0  0  0  0  0  0  0  0  1
50  0  1  0  0  0  0  0  0  0  1
51  0  0  1  0  0  0  0  0  0  1
52  0  0  0  1  0  0  0  0  0  1
53  1  0  0  0  1  0  0  0  0  1
54  0  1  0  0  1  0  0  0  0  1
55  0  0  1  0  1  0  0  0  0  1
56  0  0  0  1  1  0  0  0  0  1
57  1  0  0  0  0  1  0  0  0  1
58  0  1  0  0  0  1  0  0  0  1
59  0  0  1  0  0  1  0  0  0  1
60  0  0  0  1  0  1  0  0  0  1
61  1  0  0  0  0  0  1  0  0  1
62  0  1  0  0  0  0  1  0  0  1
63  0  0  1  0  0  0  1  0  0  1
64  0  0  0  1  0  0  1  0  0  1

We can see that when we use an intercept, model.matrix uses dummy encoding with each variable w, x, and y being turned into 3 dummy variables, plus an intercept column. So there is a total of 10 degrees of freedom.

When we don't use an intercept, model.matrix creates 4 dummy variables for w and 3 dummy variables for x and y (and no intercept column). So the number of degrees of freedom is still 10.

🌐
Kaggle
kaggle.com › questions-and-answers › 431615
Difference between One Hot Encoding and pandas dummies ? | Kaggle
One hot encoding would create three new columns: color_red, color_blue, and color_green. The color_red column would be 1 if the value of the color column is red, 0 otherwise. The color_blue column would be 1 if the value of the color column ...
Find elsewhere
🌐
LinkedIn
linkedin.com › pulse › dummy-variables-one-hot-encoding-subash-a
Dummy Variables & One Hot Encoding
August 1, 2023 - One-hot encoding, on the other hand, creates binary columns for each category, representing their presence or absence in the original data. ... To demonstrate the process, let's consider a dataset containing information about home prices in ...
🌐
Hultedtech
hultedtech.club › resources › shared-notes › statistics › one-hot-encoding-versus-dummy-encoding
One-hot encoding versus dummy encoding | Hult EdTech Club
November 21, 2023 - Dummy encoding is essentially the same as one-hot encoding, but with a slight difference in implementation that has implications for certain statistical analyses. In one-hot encoding, each category of a categorical variable is transformed into ...
🌐
Medium
medium.com › we-talk-data › one-hot-encoding-vs-dummy-encoding-08d5f2665ca1
One Hot Encoding vs. Dummy Encoding | by Hey Amit | We Talk Data | Medium
November 25, 2024 - Here’s the deal: One-Hot Encoding keeps all categories intact, creating a separate binary column for each. On the flip side, Dummy Encoding takes a more streamlined approach by dropping one category, often the first, to avoid what’s known ...
🌐
Codemia
codemia.io › home › knowledge hub › what's the difference between dummy variable and one-hot encoding?
What's the difference between dummy variable and one-hot encoding? | Codemia
September 23, 2025 - Dummy variable encoding and one-hot encoding both convert categorical variables into numeric format for machine learning models, but they differ in the number of columns created. One-hot encoding creates k binary columns for k categories. Dummy variable encoding creates k-1 columns, dropping ...
Top answer
1 of 3
10

The issue with representing a categorical variable that has $k$ levels with $k$ variables in regression is that, if the model also has a constant term, then the terms will be linearly dependent and hence the model will be unidentifiable. For example, if the model is $μ = a_0 + a_1X_1 + a_2X_2$ and $X_2 = 1 - X_1$, then any choice $(β_0, β_1, β_2)$ of the parameter vector is indistinguishable from $(β_0 + β_2,\; β_1 - β_2,\; 0)$. So although software may be willing to give you estimates for these parameters, they aren't uniquely determined and hence probably won't be very useful.

Penalization will make the model identifiable, but redundant coding will still affect the parameter values in weird ways, given the above.

The effect of a redundant coding on a decision tree (or ensemble of trees) will likely be to overweight the feature in question relative to others, since it's represented with an extra redundant variable and therefore will be chosen more often than it otherwise would be for splits.

2 of 3
6

I feel the best answer to this question is buried in the comments by @MatthewDrury, which states that there is a difference and that you should use the seemingly redundant column in any regularized approach. @MatthewDrury's reasoning is

[In regularized regression], the intercept is not penalized, so if you are inferring the effect of a level as not part of the intercept, its hard to say you are penalizing all levels equally. Instead, always include all the levels, so each is symmetric with respect to the penalty.

I think he's got a point.

🌐
Quora
quora.com › Is-one-hot-encoding-a-fancy-name-for-a-dummy-variable
Is one hot encoding a fancy name for a dummy variable? - Quora
Answer: No, they are separate concepts. Watch this. https://www.youtube.com/watch?v=qB55BXZX4LA&list=PLFMOofq-fah2NVygzYxBIOKiH9MjZramk&index=13 The one-hot encoding creates one binary variable for each category. The problem is that this representation includes redundancy. For example,...
🌐
DataCamp
campus.datacamp.com › courses › feature-engineering-for-machine-learning-in-python › creating-features
One-hot encoding and dummy variables | Python
To use categorical variables in a machine learning model, you first need to represent them in a quantitative way. The two most common approaches are to one-hot encode the variables using or to use dummy variables. In this exercise, you will create both types of encoding, and compare the created column sets.
🌐
Scribd
scribd.com › presentation › 516686040 › One-hot-encoding
One-Hot vs Dummy Encoding Explained | PDF | Categorical Variable | Computer Science
One-hot encoding creates a new binary feature for each unique category value and avoids any ordering implications. It is preferable when the categorical variable is nominal rather than ordinal.Read more ...
🌐
InfinityCodeX
infinitycodex.in › 2020 › 03 › categoricaldummy-varibles-one-hot.html
Categorical, Dummy Variables And One-Hot Encoding | (Data Science ss : 10.8) - InfinityCodeX
This is how we represented 4 Categories & 3 Dummy Variables. One-Hot Encoding : One-Hot Encoding transforms our Categorical Variables into Vectors of 0’s & 1’s the length of these vectors is equal to the number of classes or Categories that our model is expected to classify so if we are classifying whether the images were either of a Horse or a Donkey then our One-Hot Encoding vectors corresponds to these classes would be of length 2 since there are 2 categories total if we add another category such as Zebra so we could then classify whether images were Horse, Donkey or Zebra the our corresponding One-Hot Encoded vectors would each be of length 3 since we have now 3 categories.
🌐
Stack Exchange
ai.stackexchange.com › questions › 26747 › one-hot-encoding-vs-dummy-variables-best-practices-for-explainable-ai-xai
reference request - One hot encoding vs dummy variables best practices for explainable AI (XAI) - Artificial Intelligence Stack Exchange
March 10, 2021 - Dummy variables: each category ... for each record one-hot-encoding: similar to dummy variables, but one column is dropped, as its value can be derived from the other columns....
🌐
Medium
medium.com › @subashdhoni86 › dummy-variables-one-hot-encoding-ebf4f0391a2
Dummy Variables & One Hot Encoding | by Subash A | Medium
August 1, 2023 - One-hot encoding, on the other hand, creates binary columns for each category, representing their presence or absence in the original data. ... To demonstrate the process, let’s consider a dataset containing information about home prices in ...