Try this:
import pandas as pd
file = open("your.csv", "r")
data = pd.read_csv(file, sep = ",")
gender = {'male': 1,'female': 0}
data.Gender = [gender[item] for item in data.Gender]
print(data)
Or
data.Gender[data.Gender == 'male'] = 1
data.Gender[data.Gender == 'female'] = 0
print(data)
Answer from Nidhin Sajeev on Stack OverflowTry this:
import pandas as pd
file = open("your.csv", "r")
data = pd.read_csv(file, sep = ",")
gender = {'male': 1,'female': 0}
data.Gender = [gender[item] for item in data.Gender]
print(data)
Or
data.Gender[data.Gender == 'male'] = 1
data.Gender[data.Gender == 'female'] = 0
print(data)
You can do the conversion as you load the file:
d = pandas.read_csv('yourfile.csv', converters={'Gender': lambda x: int(x == 'Male')})
The converters argument takes a dictionary whose keys are the column names (or indices), and the value is a function to call for each item. The function must return the converted value.
The other way to do it is to convert it once you have the dataframe, as @DJK pointed in their comment:
data['Gender'] = (data['Gender'] == 'Male').astype(int)
In [6]: (df['number'] < 15).astype(int)
Out[6]:
0 1
1 0
2 1
3 0
4 0
5 1
6 0
7 1
8 0
Name: number, dtype: int32
In [7]: df['binary'] = (df['number'] < 15).astype(int)
In [8]: df
Out[8]:
number binary
0 12 1
1 89 0
2 12 1
3 56 0
4 62 0
5 2 1
6 657 0
7 5 1
8 73 0
You can convert it to boolean then multiply by 1.
import pandas as pd
df = pd.DataFrame({'number': [12, 89, 12, 56, 62, 2, 657, 5, 73]})
df['binary'] = (df.number < 15)*1
If performance is important, use numpy with this solution:
d = df['Col_B'].values
m = 2
df[['Col_C','Col_D']] = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
print (df)
Col_A Col_B Col_C Col_D
0 a 1 1 0
1 b 2 0 1
2 c 0 0 0
Performance (about 1000 times faster):
df = pd.DataFrame([['a', 1], ['b', 2], ['c', 0]], columns=["Col_A", "Col_B"])
df = pd.concat([df] * 1000, ignore_index=True)
In [162]: %%timeit
...: df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
...:
609 ms ± 14.5 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
In [163]: %%timeit
...: d = df['Col_B'].values
...: m = 2
...: df[['Col_C','Col_D']] = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
...:
618 µs ± 26.2 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
apply is the method you are looking for.
df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
does the trick.
I benchmarked it on 3000 rows and it is faster than the for cycle method you mention (0.5 seconds vs 3 seconds). But generally the speed won't be much faster since it still needs to apply the function for each row separately.
from time import time
start = time()
for i in range(0,len(df)):
df.loc[i,'Col_C'],df.loc[i,'Col_D'] = list( (bin(df.loc[i,'Col_B'])[2:].zfill(2) ) )
print(time() - start)
# 3.4339962005615234
start = time()
df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
print(time() - start)
# 0.5619983673095703
Note: I am using python 3, so e.g. bin(1) returns '0b1' and thus I use bin(x)[2:] to get rid of the '0b' part.
Well, one way i like to handle this problem (which is a common problem, at least in daily job life) is to convert each possibility in a column with binary value. Let me elaborate a bit. Let's say you have your column animals with 3 possibilities : dog, cat, and horse. You explode your column in 3 differents columns : colDog, colCat and colHorse. And you fill your new columns based on the value of the column animals. For example : if you have dog in the first row, you put 1 in the column colDog, etc.
The problem with handling categorical data with numerical value instead of binary is that you create a hierarchical order between your values. If dog is 1, cat is 2 and horse is 3, then horse will have more impact than cat and dog. Or i think you just want to represent your categories.
Sure its called label encoding, example from that page:
le = preprocessing.LabelEncoder()
le.fit(["paris", "paris", "tokyo", "amsterdam"])
LabelEncoder()
list(le.classes_)
['amsterdam', 'paris', 'tokyo']
le.transform(["tokyo", "tokyo", "paris"])
array([2, 2, 1]...)
list(le.inverse_transform([2, 2, 1]))
['tokyo', 'tokyo', 'paris']
df.round
>>> df.round()
np.round
>>> np.round(df)
astype
>>> df.ge(0.5).astype(int)
All which yield
0 1 2
0 0.0 0.0 1.0
1 0.0 1.0 1.0
2 1.0 0.0 1.0
3 0.0 1.0 0.0
Note: round works here because it automatically sets the threshold for .5 between two integers. For custom thresholds, use the 3rd solution
Or you can use np.where() and assign the values to the underlying array:
df[:]=np.where(df<0.5,0,1)
0 1 2
0 0 0 1
1 0 1 1
2 1 0 1
3 0 1 0
You can also work with numpy, much faster than pandas.
Edit: faster numpy using view Couple of tricks here:
- Work only with the column of interest
- Convert the underlaying array to
uint16, to ensure compatibility with any integer input - Swapbytes to have a proper H,L order (at least on my architecture)
- Split H,L without actually moving any data with
view - Run
unpackbitsand reshape accordingly
My machine requires a byteswap to have the bytes of the uint16 in the proper place. Note that this aproach requires to have the data as int16/uint16, while the other one would work for int64 as well.
import pandas as pd
import numpy as np
df = pd.DataFrame({'EVENT_ID': [ 4162, 4161, 4160, 4159,4158, 4157, 4156, 4155, 4154]}, dtype='uint16')
zz=np.unpackbits(df.EVENT_ID.values.astype('uint16').byteswap().view('uint8')).reshape(-1,16)
df3 = pd.concat([df,pd.DataFrame(zz)],axis=1)
print(f"{df3 =}")
df3 = EVENT_ID 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
0 4162 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 0
1 4161 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 1
2 4160 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 0
3 4159 0 0 0 1 0 0 0 0 0 0 1 1 1 1 1 1
4 4158 0 0 0 1 0 0 0 0 0 0 1 1 1 1 1 0
5 4157 0 0 0 1 0 0 0 0 0 0 1 1 1 1 0 1
6 4156 0 0 0 1 0 0 0 0 0 0 1 1 1 1 0 0
7 4155 0 0 0 1 0 0 0 0 0 0 1 1 1 0 1 1
8 4154 0 0 0 1 0 0 0 0 0 0 1 1 1 0 1 0
older proposed method:
lh = np.unpackbits((df.values & 0xFF).astype('uint8')).reshape(-1,8)
uh = np.unpackbits((df.values >> 8).astype('uint8')).reshape(-1,8)
df2 = pd.concat([df, pd.DataFrame(np.concatenate([uh,lh],axis=1),index=df.index)],axis=1)
Benchmark: numpy is orders of magnitude faster than pandas" For 1M points:
- numpy view: 35ms for 1million
uint64points - numpy low/high: 50ms
- pandas list bin: 1.78s
- pandas apply format + list: 1.97s
- pandas apply lambda: 6.08s
df = pd.DataFrame({'EVENT_ID': (np.random.random(int(1e6))*65000).astype('uint16')})
pandas apply format list
In [13]: %timeit df2 = df.join(pd.DataFrame(df['EVENT_ID'].apply('{0:b}'.format).apply(list).tolist()))
1.97 s ± 42.5 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
pandas list bin
In [10]: %%timeit
...: binary_values = pd.DataFrame([list(bin(x)[2:]) for x in df['EVENT_ID']])
...: df2 = df.join(binary_values)
...:
...:
1.78 s ± 53.9 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
pandas3 apply lambda
In [5]: %%timeit
...: for i in range(16):
...: df[f"bit{i}"] = df["EVENT_ID"].apply(lambda x: x & 1 << i).astype(bool).astype(int)
...:
6.08 s ± 65.8 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
numpy
In [14]: %%timeit
...: lh = np.unpackbits((df.values & 0xFF).astype('uint8')).reshape(-1,8)
...: uh = np.unpackbits((df.values >> 8).astype('uint8')).reshape(-1,8)
...: df3=pd.concat([df, pd.DataFrame(np.concatenate([uh,lh],axis=1),index=df.index)],axis=1)
...:
...:
49.9 ms ± 232 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)
Edit
Although my idea is there, this implementation is in fact much slower than all the other answers. Please see @ZaeroDivide's answer and comments below.
Original Answer
I don't think using the bin function and working with the str type is particularly efficient. Please consider using bitmasks.
for i in range(16):
df[f"bit{i}"] = df["EVENT_ID"].apply(lambda x: x & 1 << i).astype(bool).astype(int)
Testing with your data, I have the following results
EVENT_ID B bit0 bit1 bit2 bit3 bit4 bit5 bit6 bit7 \
0 4162 1000001000010 0 1 0 0 0 0 1 0
1 4161 1000001000001 1 0 0 0 0 0 1 0
2 4160 1000001000000 0 0 0 0 0 0 1 0
3 4159 1000000111111 1 1 1 1 1 1 0 0
4 4158 1000000111110 0 1 1 1 1 1 0 0
5 4157 1000000111101 1 0 1 1 1 1 0 0
6 4156 1000000111100 0 0 1 1 1 1 0 0
7 4155 1000000111011 1 1 0 1 1 1 0 0
8 4154 1000000111010 0 1 0 1 1 1 0 0
bit8 bit9 bit10 bit11 bit12 bit13 bit14 bit15
0 0 0 0 0 1 0 0 0
1 0 0 0 0 1 0 0 0
2 0 0 0 0 1 0 0 0
3 0 0 0 0 1 0 0 0
4 0 0 0 0 1 0 0 0
5 0 0 0 0 1 0 0 0
6 0 0 0 0 1 0 0 0
7 0 0 0 0 1 0 0 0
8 0 0 0 0 1 0 0 0
you can use map -
a = {'education' : 1,'cinema' : 0}
train_docs['Class'] = train_docs['Class'].map(a)
When you use .apply method of pandas.Series given function should accept what that pandas.Series is holding, in this case strs,
def numconv(a):
return a.map({'education' : 1,'cinema' : 0})
will not work if a is str, you might repair it by changing to
def numconv(a):
return {'education' : 1,'cinema' : 0}[a]
or in this case just use .replace method of pandas.Series like so
import pandas as pd
df = pd.DataFrame({"Class":["education","education","education","cinema","cinema"]})
df["Class"].replace({'education' : 1,'cinema' : 0},inplace=True)
print(df)
output
Class
0 1
1 1
2 1
3 0
4 0
Beware that function I proposed will fail if any other value will appears, whilst .replace ignores unknown values
As the contents of the columns are integers rather than string, the int(i, 2) function cannot be used directly on the integer i, or else it will throw an error. E.g.
int(11, 2)
TypeError: int() can't convert non-string with explicit base
You have to firstly convert the columns to string and then convert string to decimal values. For example, assuming column 'val', use the code:
df['val'] = df['val'].astype(str).map(lambda x: int(x, 2))
astype(str)convert the column type to string, thenuse the
int()function withinmap()to apply to each element in the column
Test codes and output:
data = { 'val': [11, 101, 1001, 1101]}
df = pd.DataFrame(data)
print(df)
Output: (before conversion)
val
0 11
1 101
2 1001
3 1101
df['val'] = df['val'].astype(str).map(lambda x: int(x, 2))
print(df)
Output: (after conversion)
val
0 3
1 5
2 9
3 13
You can use int with base set as 2.
Example:
>>> int('0011', 2)
3
from functools import partial
bin_to_int = partial(int, base=2)
s = pd.Series(['0011', '101', '1010'])
s.apply(bin_to_int)
0 3
1 5
2 10
dtype: int64
We can do this by a chain of actions:
- first we convert the hexadecimal number to an
intwith.apply(int, base=16); - next we convert this to binary data, with
.apply(bin); - next we chunk off the first two characters with
.str[2:]; - then we obtain the last three characters with
.str[-3:]; and - finally we again interpret these as
ints, with.apply(int, base=2).
So:
>>> df.Data.apply(int, base=16).apply(bin).str[2:].str[-3:].apply(int, base=2)
0 2
1 3
2 3
3 7
4 7
5 0
6 3
Name: Data, dtype: int64
We can however use another strategy here:
- we first convert the hexadecimal number to an
int; and - then we apply a bitwise and with
0b111.
for example:
>>> df.Data.apply(int, base=16) & 0b111
0 2
1 3
2 3
3 7
4 7
5 0
6 3
Name: Data, dtype: int64
The second attempt is not only simpler, but faster as well, approximately by 66%:
>>> timeit(first_strategy, number=10000)
6.962630775000434
>>> timeit(second_strategy, number=10000)
2.330652763019316
for a dataframe that repeats the sample data 100 times, we get:
>>> timeit(first_strategy, number=10000)
17.603060900000855
>>> timeit(second_strategy, number=10000)
5.901462858979357
this is again 66% faster.
You can use:
df.Data.apply(lambda v: int(format(int(v, 16), '08b')[-3:], 2))
Which gives you:
0 2
1 3
2 3
3 7
4 7
5 0
6 3
Name: Data, dtype: int64
Those steps are:
- Take your original data and convert it to decimal using
int(number, 16)(base 16 is hex) (int('1A', 16)==26) - Take that number and format it as a binary string
format(number, '08b')gives you an character string of 0/1's zero filled on the left (format(26, '08b')=='00011010') - Take the last 3 characters of that string
[-3:]('010') and convert it to decimal with a base 2,int(binary_string[-3:], 2)gives you:2
