I'm almost positive that you're encountering memory issues because str.get_dummies returns an array full of 1s and 0s, of datatype np.int64. This is quite different from the behavior of pd.get_dummies, which returns an array of values of datatype uint8.

This appears to be a known issue. However, there's been no update, nor fix, for the past year. Checking out the source code for str.get_dummies will indeed confirm that it is returning np.int64.

An 8 bit integer will take up 1 byte of memory, while a 64 bit integer will take up 8 bytes. I'm hopeful that memory problems can be avoided by finding an alternative way to one-hot encode Col2 which ensures the output are all 8 bit integers.

Here was my approach, beginning with your example:

df = pd.DataFrame({'Col1': ['X', 'Y', 'X'],
                   'Col2': ['a,b,c', 'a,b', 'b,d']})
df

    Col1    Col2
0   X       a,b,c
1   Y       a,b
2   X       b,d
  1. Since Col1 contains simple, non-delimited strings, we can easily one-hot encode it using pd.get_dummies:
df = pd.get_dummies(df, columns=['Col1'])
df

    Col2    Col1_X  Col1_Y
0   a,b,c        1       0
1   a,b          0       1
2   b,d          1       0

So far so good.

df['Col1_X'].values.dtype
dtype('uint8')
  1. Let's get a list of all unique substrings contained inside the comma-delimited strings in Col2:
vals = list(df['Col2'].str.split(',').values)
vals = [i for l in vals for i in l]
vals = list(set(vals))
vals.sort()
vals

['a', 'b', 'c', 'd']
  1. Now we can loop through the above list of values and use str.contains to create a new column for each value, such as 'a'. Each row in a new column will contain 1 if that row actually has the new column's value, such as 'a', inside its string in Col2. As we create each new column, we make sure to convert its datatype to uint8:
col='Col2'
for v in vals:
    n = col + '_' + v
    df[n] = df[col].str.contains(v)
    df[n] = df[n].astype('uint8')

df.drop(col, axis=1, inplace=True)
df

    Col1_X  Col1_Y  Col2_a  Col2_b  Col2_c  Col2_d
0        1       0       1       1       1       0
1        0       1       1       1       0       0
2        1       0       0       1       0       1

This results in a dataframe that meets your desired format. And thankfully, the integers in the four new columns that were one-hot encoded from Col2 only take up 1 byte each, as opposed to 8 bytes each.

df['Col2_a'].dtype
dtype('uint8')

If, on the outside chance, the above approach doesn't work. My advice would be to use str.get_dummies to one-hot encode Col2 in chunks of rows. Each time you do a chunk, you would convert its datatype from np.int64 to uint8, and then transform the chunk to a sparse matrix. You could eventually concatenate all chunks back together.

Answer from James Dellinger on Stack Overflow
Top answer
1 of 2
2

I'm almost positive that you're encountering memory issues because str.get_dummies returns an array full of 1s and 0s, of datatype np.int64. This is quite different from the behavior of pd.get_dummies, which returns an array of values of datatype uint8.

This appears to be a known issue. However, there's been no update, nor fix, for the past year. Checking out the source code for str.get_dummies will indeed confirm that it is returning np.int64.

An 8 bit integer will take up 1 byte of memory, while a 64 bit integer will take up 8 bytes. I'm hopeful that memory problems can be avoided by finding an alternative way to one-hot encode Col2 which ensures the output are all 8 bit integers.

Here was my approach, beginning with your example:

df = pd.DataFrame({'Col1': ['X', 'Y', 'X'],
                   'Col2': ['a,b,c', 'a,b', 'b,d']})
df

    Col1    Col2
0   X       a,b,c
1   Y       a,b
2   X       b,d
  1. Since Col1 contains simple, non-delimited strings, we can easily one-hot encode it using pd.get_dummies:
df = pd.get_dummies(df, columns=['Col1'])
df

    Col2    Col1_X  Col1_Y
0   a,b,c        1       0
1   a,b          0       1
2   b,d          1       0

So far so good.

df['Col1_X'].values.dtype
dtype('uint8')
  1. Let's get a list of all unique substrings contained inside the comma-delimited strings in Col2:
vals = list(df['Col2'].str.split(',').values)
vals = [i for l in vals for i in l]
vals = list(set(vals))
vals.sort()
vals

['a', 'b', 'c', 'd']
  1. Now we can loop through the above list of values and use str.contains to create a new column for each value, such as 'a'. Each row in a new column will contain 1 if that row actually has the new column's value, such as 'a', inside its string in Col2. As we create each new column, we make sure to convert its datatype to uint8:
col='Col2'
for v in vals:
    n = col + '_' + v
    df[n] = df[col].str.contains(v)
    df[n] = df[n].astype('uint8')

df.drop(col, axis=1, inplace=True)
df

    Col1_X  Col1_Y  Col2_a  Col2_b  Col2_c  Col2_d
0        1       0       1       1       1       0
1        0       1       1       1       0       0
2        1       0       0       1       0       1

This results in a dataframe that meets your desired format. And thankfully, the integers in the four new columns that were one-hot encoded from Col2 only take up 1 byte each, as opposed to 8 bytes each.

df['Col2_a'].dtype
dtype('uint8')

If, on the outside chance, the above approach doesn't work. My advice would be to use str.get_dummies to one-hot encode Col2 in chunks of rows. Each time you do a chunk, you would convert its datatype from np.int64 to uint8, and then transform the chunk to a sparse matrix. You could eventually concatenate all chunks back together.

2 of 2
1

I would like to give my solution as well. And I would like to thank @James-dellinger for the answer. So here is my approach

df = pd.DataFrame({'Col1': ['X', 'Y', 'X'],
               'Col2': ['a,b,c', 'a,b', 'b,d']})
df

  Col1  Col2
0   X   a,b,c
1   Y   a,b
2   X   b,d

I first split Col2 values and convert it into column values.

df= pd.DataFrame(df['Col2'].str.split(',',3).tolist(),columns = ['Col1','Col2','Col3'])

df

   Col1 Col2 Col3
0   a   b    c
1   a   b    None
2   b   d    None

Then I applied dummy creation on this dataframe without giving any prefix.

df=pd.get_dummies(df, prefix="")

df

    _a  _b  _b  _d  _c
0   1   0   1   0   1
1   1   0   1   0   0
2   0   1   0   1   0

Now to get the desired result we can sum up all the duplicate columns.

df.groupby(level=0, axis=1).sum()

df

    _a  _b  _c  _d
0   1   1   1   0
1   1   1   0   0
2   0   1   0   1

For Col1 we can directly create dummy variables using pd.get_dummies() and store it into different dataframe suppose col1_df. We can concat both columns using pd.concat([df,col1_df], axis=1, sort=False)

🌐
Medium
medium.com › @techwithpraisejames › how-to-turn-categorical-variables-into-numbers-using-python-pandas-get-dummies-method-5a7d0ae0b3a3
How to Convert Categorical Variables To Numbers Using Python Pandas get_dummies() Method | by Praise James | Medium
September 11, 2025 - These dummy variables allow machine learning algorithms to interpret and utilize categorical variables effectively. In this guide, you will learn how to use Python Pandas get_dummies() method to create dummy variables.
🌐
Pandas
pandas.pydata.org › docs › reference › api › pandas.get_dummies.html
pandas.get_dummies — pandas 3.0.6 documentation - PyData |
pandas.get_dummies(data, prefix=None, prefix_sep='_', dummy_na=False, columns=None, sparse=False, drop_first=False, dtype=None)[source]# Convert categorical variable into dummy/indicator variables.
Top answer
1 of 7
54

It's been a few years, so this may well not have been in the pandas toolkit back when this question was originally asked, but this approach seems a little easier to me. idxmax will return the index corresponding to the largest element (i.e. the one with a 1). We do axis=1 because we want the column name where the 1 occurs.

EDIT: I didn't bother making it categorical instead of just a string, but you can do that the same way as @Jeff did by wrapping it with pd.Categorical (and pd.Series, if desired).

In [1]: import pandas as pd

In [2]: s = pd.Series(['a', 'b', 'a', 'c'])

In [3]: s
Out[3]: 
0    a
1    b
2    a
3    c
dtype: object

In [4]: dummies = pd.get_dummies(s)

In [5]: dummies
Out[5]: 
   a  b  c
0  1  0  0
1  0  1  0
2  1  0  0
3  0  0  1

In [6]: s2 = dummies.idxmax(axis=1)

In [7]: s2
Out[7]: 
0    a
1    b
2    a
3    c
dtype: object

In [8]: (s2 == s).all()
Out[8]: True

EDIT in response to @piRSquared's comment: This solution does indeed assume there's one 1 per row. I think this is usually the format one has. pd.get_dummies can return rows that are all 0 if you have drop_first=True or if there are NaN values and dummy_na=False (default) (any cases I'm missing?). A row of all zeros will be treated as if it was an instance of the variable named in the first column (e.g. a in the example above).

If drop_first=True, you have no way to know from the dummies dataframe alone what the name of the "first" variable was, so that operation isn't invertible unless you keep extra information around; I'd recommend leaving drop_first=False (default).

Since dummy_na=False is the default, this could certainly cause problems. Please set dummy_na=True when you call pd.get_dummies if you want to use this solution to invert the "dummification" and your data contains any NaNs. Setting dummy_na=True will always add a "nan" column, even if that column is all 0s, so you probably don't want to set this unless you actually have NaNs. A nice approach might be to set dummies = pd.get_dummies(series, dummy_na=series.isnull().any()). What's also nice is that idxmax solution will correctly regenerate your NaNs (not just a string that says "nan").

It's also worth mentioning that setting drop_first=True and dummy_na=False means that NaNs become indistinguishable from an instance of the first variable, so this should be strongly discouraged if your dataset may contain any NaN values.

2 of 7
27
In [46]: s = Series(list('aaabbbccddefgh')).astype('category')

In [47]: s
Out[47]: 
0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
dtype: category
Categories (8, object): [a < b < c < d < e < f < g < h]

In [48]: df = pd.get_dummies(s)

In [49]: df
Out[49]: 
    a  b  c  d  e  f  g  h
0   1  0  0  0  0  0  0  0
1   1  0  0  0  0  0  0  0
2   1  0  0  0  0  0  0  0
3   0  1  0  0  0  0  0  0
4   0  1  0  0  0  0  0  0
5   0  1  0  0  0  0  0  0
6   0  0  1  0  0  0  0  0
7   0  0  1  0  0  0  0  0
8   0  0  0  1  0  0  0  0
9   0  0  0  1  0  0  0  0
10  0  0  0  0  1  0  0  0
11  0  0  0  0  0  1  0  0
12  0  0  0  0  0  0  1  0
13  0  0  0  0  0  0  0  1

In [50]: x = df.stack()

# I don't think you actually need to specify ALL of the categories here, as by definition
# they are in the dummy matrix to start (and hence the column index)
In [51]: Series(pd.Categorical(x[x!=0].index.get_level_values(1)))
Out[51]: 
0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
Name: level_1, dtype: category
Categories (8, object): [a < b < c < d < e < f < g < h]

So I think we need a function to 'do' this as it seems to be a natural operations. Maybe get_categories(), see here

🌐
GeeksforGeeks
geeksforgeeks.org › convert-a-categorical-variable-into-dummy-variables
Convert A Categorical Variable Into Dummy Variables | GeeksforGeeks
December 11, 2020 - Firstly, we have to understand what are Categorical variables in pandas. Categorical are the datatype available in pandas library of python. A categorical variable takes only a fixed category (usually fixed number) of values.
🌐
Seaborn Line Plots
marsja.se › home › programming › python › how to use pandas get_dummies to create dummy variables in python
How to use Pandas get_dummies to Create Dummy Variables in Python
August 22, 2023 - To convert your categorical variables to dummy variables in Python, you can use Pandas get_dummies() method. For example, if you have the categorical variable “Gender” in your dataframe called “df” you can use the following code to make ...
🌐
Pandas
pandas.pydata.org › pandas-docs › stable › reference › api › pandas.get_dummies.html
pandas.get_dummies — pandas 2.2.2 documentation - PyData |
pandas.get_dummies(data, prefix=None, prefix_sep='_', dummy_na=False, columns=None, sparse=False, drop_first=False, dtype=None)[source]# Convert categorical variable into dummy/indicator variables.
Find elsewhere
Top answer
1 of 1
1

Nope. The column from priority or cluster won't be misinterpreted as the third column of severity.

Here's answer to how reference is kept:

in pandas.get_dummies there is a parameter i.e. drop_first allows you whether to keep or remove the reference (whether to keep k or k-1 dummies out of k categorical levels).

Please note drop_first = False meaning that the reference is not dropped and k dummies created out of k categorical levels! You set drop_first = True, then it will drop the reference column after encoding.

Here's link to one hot encoding.

As in your case severity has 3 categories S1, S2 and S3. After creating dummies one of these categories will always be 1 and others 0.

for s1 it will be [1,0,0], s2 will be [0,1,0] and s3 will be [0,0,1]

Now if you drop the column for category s1.

The values will be [0,0] if severity is S1

[1,0] if severity is S2

[0,1] if severity is S3.

So there is no information loss here and your model has one less column to deal with. That's why it is always recommended to keep drop_first parameter as True.

Edit :

After applying the dummies you will get columns like:

severity_S1   severity_S2   severity_S3  

  1              0              0                  # when value is S1
  0              1              0                  # when value is S2  
  0              0              1                  # when value is S3

pandas.get_dummies() drops the 1st column after creating the above references. So in your data will be like below:

 severity_S2   severity_S3

   0              0                  # when value is S1
   1              0                  # when value is S2  
   0              1                  # when value is S3

For all there variables your final data will look like below: I'm using short column names due to space issue:

s2  s3  p2  p3  B  C  D
0   0   1   0   1  0  0     # For row with S1, P2 and B
0   1   0   1   0  1  0     # For row with S3, P3 and C
1   0   0   0   0  0  1     # For row with S2, P1 and D
1   0   0   0   0  0  0     # For row with S2, P1 and A
🌐
LOST
lost-stats.github.io › Data_Manipulation › Creating_Dummy_Variables › creating_dummy_variables.html
Creating Dummy Variables | LOST
Several python libraries have functions to turn categorical variables into dummies, including pandas, scikit-learn (where it is called OneHotEncoder), and statsmodels (where it is called categorical). This example uses pandas get_dummies function. import pandas as pd # Create a dataframe df ...
🌐
GeeksforGeeks
geeksforgeeks.org › python › how-to-create-dummy-variables-in-python-with-pandas
How to Create Dummy Variables in Python with Pandas? - GeeksforGeeks
July 23, 2025 - Each new column corresponds to one category and contains a 1 if that category is present in a row and 0 otherwise. pandas.get_dummies(data, prefix=None, prefix_sep='_', ...) ... Returns: A new DataFrame with dummy/indicator variables.
🌐
Edureka Community
edureka.co › home › community › categories › machine learning › how to convert categorical variable into dummy...
How to convert categorical variable into dummy variable | Edureka Community
May 7, 2020 - Hi Guys, I am trying to create one Machine Learning model . But my dataset contains one string ... values into dummy variable. How can I do that?
🌐
Note.nkmk.me
note.nkmk.me › home › python › pandas
pandas: Get dummy variables with pd.get_dummies() | note.nkmk.me
January 17, 2024 - In pandas, the pd.get_dummies() function converts categorical variables to dummy variables. ... This function can convert data categorized by strings, such as gender, to a format like 0 for male and 1 for female.
🌐
Arab Psychology
scales.arabpsychology.com › home › how to easily convert categorical data to dummy variables with pandas get_dummies
How To Easily Convert Categorical Data To Dummy Variables With Pandas Get_dummies
December 5, 2025 - The Pandas get_dummies function is a fundamental tool used in data preprocessing, designed specifically to handle categorical variables. Its primary purpose is to convert nominal categorical data into a numerical format suitable for statistical ...
🌐
ProjectPro
projectpro.io › recipes › convert-categorical-variables-into-numerical-variables-in-python
How to convert categorical variables into numerical in Python? -
September 6, 2023 - We can assign numbers for each category, but it may not be effective when the difference between the categories can not be measured. This can be done by making new features according to the categories with bool values. For this, we will be using ...
🌐
Medium
medium.com › @urvashilluniya › convert-multiple-categorical-columns-into-numeric-columns-in-single-line-of-code-577bab825635
Convert Multiple Categorical Data Columns to Numerical Data Columns using Dummy Variables | by Urvashi Jaitley | Medium
February 6, 2019 - One of the methods to create dummy variables involves following steps: 1) creating dummy variables for each of the columns, 2) concatenate the new columns to the main data frame, 3) drop corresponding categorical columns.