Try this:

import pandas as pd

file = open("your.csv", "r")

data = pd.read_csv(file, sep = ",")

gender = {'male': 1,'female': 0}

data.Gender = [gender[item] for item in data.Gender]
print(data)

Or

data.Gender[data.Gender == 'male'] = 1
data.Gender[data.Gender == 'female'] = 0
print(data)
Answer from Nidhin Sajeev on Stack Overflow
🌐
Stack Overflow
stackoverflow.com › questions › 45018601 › decimal-to-binary-in-dataframe-pandas
python - Decimal TO Binary in Dataframe Pandas - Stack Overflow
You want them to be string representations of binary numbers? ... Save this answer. ... Show activity on this post. ... the thing is format doesn't work on strings so you need to convert your inputs to integers before getting their binary string representation ('04b' is just to have the representation on 4 bits).
🌐
TutorialsPoint
tutorialspoint.com › python-convert-pandas-dataframe-to-binary-data
Python - Convert Pandas DataFrame to binary data
Use the get_dummies() and set the column which you want to convert to binary form. Here, we want the Result in “Pass” and “Fail” form to be visible. Therefore, we will set the “Result” column − ... import pandas as pd # Create DataFrame dataFrame = pd.DataFrame( { "Student": ['Jack', ...
🌐
GeeksforGeeks
geeksforgeeks.org › how-to-convert-categorical-data-to-binary-data-in-python
How to convert categorical data to binary data in Python? - GeeksforGeeks
January 17, 2022 - So for that, we have to the inbuilt function of Pandas i.e. get_dummies() as shown: Here we use get_dummies() for only Gender column because here we want to convert Categorical Data to Binary data only for Gender Column.
🌐
YouTube
youtube.com › watch
How to Convert Categorical Values to Binary (0 and 1) in Python with Pandas) - YouTube
This video explains How to Convert Categorical Values to Binary values (Python and Pandas) with Jupyter NotebookHow to build a simple Neural Network - https...
Published: September 27, 2019
Top answer
1 of 2
4

If performance is important, use numpy with this solution:

d = df['Col_B'].values
m = 2
df[['Col_C','Col_D']]  = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
print (df)
  Col_A  Col_B  Col_C  Col_D
0     a      1      1      0
1     b      2      0      1
2     c      0      0      0

Performance (about 1000 times faster):

df = pd.DataFrame([['a', 1], ['b', 2], ['c', 0]], columns=["Col_A", "Col_B"])


df = pd.concat([df] * 1000, ignore_index=True)

In [162]: %%timeit
     ...: df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
     ...: 
609 ms ± 14.5 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

In [163]: %%timeit
     ...: d = df['Col_B'].values
     ...: m = 2
     ...: df[['Col_C','Col_D']]  = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
     ...: 
618 µs ± 26.2 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
2 of 2
2

apply is the method you are looking for.

df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))

does the trick.

I benchmarked it on 3000 rows and it is faster than the for cycle method you mention (0.5 seconds vs 3 seconds). But generally the speed won't be much faster since it still needs to apply the function for each row separately.

from time import time
start = time()
for i in range(0,len(df)):
    df.loc[i,'Col_C'],df.loc[i,'Col_D'] = list( (bin(df.loc[i,'Col_B'])[2:].zfill(2) ) )
print(time() - start)
# 3.4339962005615234

start = time()
df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
print(time() - start)
# 0.5619983673095703

Note: I am using python 3, so e.g. bin(1) returns '0b1' and thus I use bin(x)[2:] to get rid of the '0b' part.

Find elsewhere
Top answer
1 of 3
2

You can also work with numpy, much faster than pandas.

Edit: faster numpy using view Couple of tricks here:

  • Work only with the column of interest
  • Convert the underlaying array to uint16, to ensure compatibility with any integer input
  • Swapbytes to have a proper H,L order (at least on my architecture)
  • Split H,L without actually moving any data with view
  • Run unpackbits and reshape accordingly

My machine requires a byteswap to have the bytes of the uint16 in the proper place. Note that this aproach requires to have the data as int16/uint16, while the other one would work for int64 as well.

import pandas as pd
import numpy as np

df = pd.DataFrame({'EVENT_ID': [ 4162, 4161, 4160, 4159,4158, 4157, 4156, 4155, 4154]}, dtype='uint16')

zz=np.unpackbits(df.EVENT_ID.values.astype('uint16').byteswap().view('uint8')).reshape(-1,16)
df3 = pd.concat([df,pd.DataFrame(zz)],axis=1)

print(f"{df3 =}")

df3 =   EVENT_ID  0  1  2  3  4  5  6  7  8  9  10  11  12  13  14  15
0      4162  0  0  0  1  0  0  0  0  0  1   0   0   0   0   1   0
1      4161  0  0  0  1  0  0  0  0  0  1   0   0   0   0   0   1
2      4160  0  0  0  1  0  0  0  0  0  1   0   0   0   0   0   0
3      4159  0  0  0  1  0  0  0  0  0  0   1   1   1   1   1   1
4      4158  0  0  0  1  0  0  0  0  0  0   1   1   1   1   1   0
5      4157  0  0  0  1  0  0  0  0  0  0   1   1   1   1   0   1
6      4156  0  0  0  1  0  0  0  0  0  0   1   1   1   1   0   0
7      4155  0  0  0  1  0  0  0  0  0  0   1   1   1   0   1   1
8      4154  0  0  0  1  0  0  0  0  0  0   1   1   1   0   1   0

older proposed method:

lh = np.unpackbits((df.values & 0xFF).astype('uint8')).reshape(-1,8)
uh = np.unpackbits((df.values >> 8).astype('uint8')).reshape(-1,8)

df2 = pd.concat([df, pd.DataFrame(np.concatenate([uh,lh],axis=1),index=df.index)],axis=1)

Benchmark: numpy is orders of magnitude faster than pandas" For 1M points:

  • numpy view: 35ms for 1million uint64 points
  • numpy low/high: 50ms
  • pandas list bin: 1.78s
  • pandas apply format + list: 1.97s
  • pandas apply lambda: 6.08s
df = pd.DataFrame({'EVENT_ID': (np.random.random(int(1e6))*65000).astype('uint16')})

pandas apply format list

In [13]: %timeit df2 = df.join(pd.DataFrame(df['EVENT_ID'].apply('{0:b}'.format).apply(list).tolist()))
1.97 s ± 42.5 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

pandas list bin

In [10]: %%timeit
    ...: binary_values = pd.DataFrame([list(bin(x)[2:]) for x in df['EVENT_ID']])
    ...: df2 = df.join(binary_values)
    ...:
    ...:
1.78 s ± 53.9 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

pandas3 apply lambda

In [5]: %%timeit
   ...: for i in range(16):
   ...:     df[f"bit{i}"] = df["EVENT_ID"].apply(lambda x: x & 1 << i).astype(bool).astype(int)
   ...:
6.08 s ± 65.8 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

numpy

In [14]: %%timeit
    ...: lh = np.unpackbits((df.values & 0xFF).astype('uint8')).reshape(-1,8)
    ...: uh = np.unpackbits((df.values >> 8).astype('uint8')).reshape(-1,8)
    ...: df3=pd.concat([df, pd.DataFrame(np.concatenate([uh,lh],axis=1),index=df.index)],axis=1)
    ...:
    ...:
49.9 ms ± 232 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)
2 of 3
1

Edit

Although my idea is there, this implementation is in fact much slower than all the other answers. Please see @ZaeroDivide's answer and comments below.

Original Answer

I don't think using the bin function and working with the str type is particularly efficient. Please consider using bitmasks.

for i in range(16):
    df[f"bit{i}"] = df["EVENT_ID"].apply(lambda x: x & 1 << i).astype(bool).astype(int)

Testing with your data, I have the following results

   EVENT_ID              B  bit0  bit1  bit2  bit3  bit4  bit5  bit6  bit7  \
0      4162  1000001000010     0     1     0     0     0     0     1     0   
1      4161  1000001000001     1     0     0     0     0     0     1     0   
2      4160  1000001000000     0     0     0     0     0     0     1     0   
3      4159  1000000111111     1     1     1     1     1     1     0     0   
4      4158  1000000111110     0     1     1     1     1     1     0     0   
5      4157  1000000111101     1     0     1     1     1     1     0     0   
6      4156  1000000111100     0     0     1     1     1     1     0     0   
7      4155  1000000111011     1     1     0     1     1     1     0     0   
8      4154  1000000111010     0     1     0     1     1     1     0     0   

   bit8  bit9  bit10  bit11  bit12  bit13  bit14  bit15  
0     0     0      0      0      1      0      0      0  
1     0     0      0      0      1      0      0      0  
2     0     0      0      0      1      0      0      0  
3     0     0      0      0      1      0      0      0  
4     0     0      0      0      1      0      0      0  
5     0     0      0      0      1      0      0      0  
6     0     0      0      0      1      0      0      0  
7     0     0      0      0      1      0      0      0  
8     0     0      0      0      1      0      0      0 
🌐
Stack Overflow
stackoverflow.com › questions › 45763493 › vectorized-method-of-converting-binary-to-int-for-pandas-dataframe-series
python - Vectorized method of converting binary to int for pandas dataframe/series - Stack Overflow
thanks, I've just been reading a lot that using .str methods, and .np methods will always be faster than using apply and lambda and wondered if there was a similar trick I could use here for binary -> int 2017-08-18T19:21:10.043Z+00:00 ... @MattTakao I'm not aware of any that convert base-2 strings into integers though...
🌐
TutorialsPoint
tutorialspoint.com › article › how-to-convert-categorical-data-to-binary-data-in-python
How to convert categorical data to binary data in Python?
April 18, 2023 - This transformation is useful because ... require numerical inputs, rather than categorical inputs. Binary encoding is a common approach that converts each unique category in a categorical variable into a separate binary column, where a value of 1 indicates the presence of the category and 0 indicates its absence. The simplest method to convert categorical data to binary format is using the pd.get_dummies() function. Let's see how it works ? import pandas as pd # Create ...
Top answer
1 of 2
6

We can do this by a chain of actions:

  1. first we convert the hexadecimal number to an int with .apply(int, base=16);
  2. next we convert this to binary data, with .apply(bin);
  3. next we chunk off the first two characters with .str[2:];
  4. then we obtain the last three characters with .str[-3:]; and
  5. finally we again interpret these as ints, with .apply(int, base=2).

So:

>>> df.Data.apply(int, base=16).apply(bin).str[2:].str[-3:].apply(int, base=2)
0    2
1    3
2    3
3    7
4    7
5    0
6    3
Name: Data, dtype: int64

We can however use another strategy here:

  1. we first convert the hexadecimal number to an int; and
  2. then we apply a bitwise and with 0b111.

for example:

>>> df.Data.apply(int, base=16) & 0b111
0    2
1    3
2    3
3    7
4    7
5    0
6    3
Name: Data, dtype: int64

The second attempt is not only simpler, but faster as well, approximately by 66%:

>>> timeit(first_strategy, number=10000)
6.962630775000434
>>> timeit(second_strategy, number=10000)
2.330652763019316

for a dataframe that repeats the sample data 100 times, we get:

>>> timeit(first_strategy, number=10000)
17.603060900000855
>>> timeit(second_strategy, number=10000)
5.901462858979357

this is again 66% faster.

2 of 2
2

You can use:

df.Data.apply(lambda v: int(format(int(v, 16), '08b')[-3:], 2))

Which gives you:

0    2
1    3
2    3
3    7
4    7
5    0
6    3
Name: Data, dtype: int64

Those steps are:

  • Take your original data and convert it to decimal using int(number, 16) (base 16 is hex) (int('1A', 16) == 26)
  • Take that number and format it as a binary string format(number, '08b') gives you an character string of 0/1's zero filled on the left (format(26, '08b') == '00011010')
  • Take the last 3 characters of that string [-3:] ('010') and convert it to decimal with a base 2, int(binary_string[-3:], 2) gives you: 2
🌐
Finxter
blog.finxter.com › 5-best-ways-to-convert-pandas-dataframe-to-binary-data-in-python
5 Best Ways to Convert Pandas DataFrame to Binary Data in Python – Be on the Right Side of Change
The to_parquet function is used to convert a DataFrame into the binary Parquet format, which is a highly efficient, columnar storage file format. This method is advantageous when working with big data and analytics tools, as it can significantly reduce file size and improve read and write ...
🌐
GitHub
github.com › pandas-dev › pandas › issues › 59207
ENH: Extend to_numeric to Convert Hexadecimal, Octal, and Binary Strings with Prefixes · Issue #59207 · pandas-dev/pandas
July 8, 2024 - pandas.to_numeric is a versatile function for converting various data types to numeric values. However, it currently does not support the direct conversion of strings representing numbers in different bases (hexadecimal, octal, and binary) that use standard prefixes.
Author: pandas-dev
🌐
Pandas
pandas.pydata.org › pandas-docs › version › 1.3 › reference › api › pandas.DataFrame.html
pandas.DataFrame — pandas 1.3.5 documentation
The primary pandas data structure. ... Dict can contain Series, arrays, constants, dataclass or list-like objects. If data is a dict, column order follows insertion-order. Changed in version 0.25.0: If data is a list of dicts, column order follows insertion-order. ... Index to use for resulting ...
🌐
Practical Business Python
pbpython.com › categorical-encoding.html
Guide to Encoding Categorical Values in Python - Practical Business Python
Scikit-learn also supports binary encoding by using the OneHotEncoder. We use a similar process as above to transform the data but the process of creating a pandas DataFrame adds a couple of extra steps.