df.round
>>> df.round()
np.round
>>> np.round(df)
astype
>>> df.ge(0.5).astype(int)
All which yield
0 1 2
0 0.0 0.0 1.0
1 0.0 1.0 1.0
2 1.0 0.0 1.0
3 0.0 1.0 0.0
Note: round works here because it automatically sets the threshold for .5 between two integers. For custom thresholds, use the 3rd solution
df.round
>>> df.round()
np.round
>>> np.round(df)
astype
>>> df.ge(0.5).astype(int)
All which yield
0 1 2
0 0.0 0.0 1.0
1 0.0 1.0 1.0
2 1.0 0.0 1.0
3 0.0 1.0 0.0
Note: round works here because it automatically sets the threshold for .5 between two integers. For custom thresholds, use the 3rd solution
Or you can use np.where() and assign the values to the underlying array:
df[:]=np.where(df<0.5,0,1)
0 1 2
0 0 0 1
1 0 1 1
2 1 0 1
3 0 1 0
How to convert Pandas dataframe into a binary format?
how to convert pandas dataframe to binary file in python - Stack Overflow
python - How to change values in a column into binary? - Stack Overflow
how shall I save data in DataFrame formation into a "binary filed"?
I know how to save it as CSV but I want to save it as binary. How can I possibly do it?
EDIT: I used parquet with pyarrow as the engine. Thank you all for the feedback!
Try this:
import pandas as pd
file = open("your.csv", "r")
data = pd.read_csv(file, sep = ",")
gender = {'male': 1,'female': 0}
data.Gender = [gender[item] for item in data.Gender]
print(data)
Or
data.Gender[data.Gender == 'male'] = 1
data.Gender[data.Gender == 'female'] = 0
print(data)
You can do the conversion as you load the file:
d = pandas.read_csv('yourfile.csv', converters={'Gender': lambda x: int(x == 'Male')})
The converters argument takes a dictionary whose keys are the column names (or indices), and the value is a function to call for each item. The function must return the converted value.
The other way to do it is to convert it once you have the dataframe, as @DJK pointed in their comment:
data['Gender'] = (data['Gender'] == 'Male').astype(int)
If performance is important, use numpy with this solution:
d = df['Col_B'].values
m = 2
df[['Col_C','Col_D']] = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
print (df)
Col_A Col_B Col_C Col_D
0 a 1 1 0
1 b 2 0 1
2 c 0 0 0
Performance (about 1000 times faster):
df = pd.DataFrame([['a', 1], ['b', 2], ['c', 0]], columns=["Col_A", "Col_B"])
df = pd.concat([df] * 1000, ignore_index=True)
In [162]: %%timeit
...: df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
...:
609 ms ± 14.5 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
In [163]: %%timeit
...: d = df['Col_B'].values
...: m = 2
...: df[['Col_C','Col_D']] = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
...:
618 µs ± 26.2 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
apply is the method you are looking for.
df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
does the trick.
I benchmarked it on 3000 rows and it is faster than the for cycle method you mention (0.5 seconds vs 3 seconds). But generally the speed won't be much faster since it still needs to apply the function for each row separately.
from time import time
start = time()
for i in range(0,len(df)):
df.loc[i,'Col_C'],df.loc[i,'Col_D'] = list( (bin(df.loc[i,'Col_B'])[2:].zfill(2) ) )
print(time() - start)
# 3.4339962005615234
start = time()
df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
print(time() - start)
# 0.5619983673095703
Note: I am using python 3, so e.g. bin(1) returns '0b1' and thus I use bin(x)[2:] to get rid of the '0b' part.
You mean "one-hot" encoding?
Say you have the following dataset:
import pandas as pd
df = pd.DataFrame([
['green', 1, 10.1, 0],
['red', 2, 13.5, 1],
['blue', 3, 15.3, 0]])
df.columns = ['color', 'size', 'prize', 'class label']
df

Now, you have multiple options ...
A) The Tedious Approach
color_mapping = {
'green': (0,0,1),
'red': (0,1,0),
'blue': (1,0,0)}
df['color'] = df['color'].map(color_mapping)
df

import numpy as np
y = df['class label'].values
X = df.iloc[:, :-1].values
X = np.apply_along_axis(func1d= lambda x: np.array(list(x[0]) + list(x[1:])), axis=1, arr=X)
print('Class labels:', y)
print('\nFeatures:\n', X)
Yielding:
Class labels: [0 1 0]
Features:
[[ 0. 0. 1. 1. 10.1]
[ 0. 1. 0. 2. 13.5]
[ 1. 0. 0. 3. 15.3]]
B) Scikit-learn's DictVectorizer
from sklearn.feature_extraction import DictVectorizer
dvec = DictVectorizer(sparse=False)
X = dvec.fit_transform(df.transpose().to_dict().values())
X
Yielding:
array([[ 0. , 0. , 1. , 0. , 10.1, 1. ],
[ 1. , 0. , 0. , 1. , 13.5, 2. ],
[ 0. , 1. , 0. , 0. , 15.3, 3. ]])
C) Pandas' get_dummies
pd.get_dummies(df)

It seems that you are using scikit-learn's DictVectorizer to convert the categorical values to binary. In that case, to store the result along with the new column names, you can construct a new DataFrame with values from vec_x and columns from DV.get_feature_names(). Then, store the DataFrame to disk (e.g. with to_csv()) instead of the numpy array.
Alternatively, it is also possible to use pandas to do the encoding directly with the get_dummies function:
import pandas as pd
data = pd.DataFrame({'T': ['A', 'B', 'C', 'D', 'E']})
res = pd.get_dummies(data)
res.to_csv('output.csv')
print res
Output:
T_A T_B T_C T_D T_E
0 1 0 0 0 0
1 0 1 0 0 0
2 0 0 1 0 0
3 0 0 0 1 0
4 0 0 0 0 1
Pandas now offers a wide variety of formats:
Format Type Data Description Reader Writer
text CSV read_csv to_csv
text JSON read_json to_json
text HTML read_html to_html
text Local clipboard read_clipboard to_clipboard
binary MS Excel read_excel to_excel
binary HDF5 Format read_hdf to_hdf
binary Feather Format read_feather to_feather
binary Parquet Format read_parquet to_parquet
binary Msgpack read_msgpack to_msgpack
binary Stata read_stata to_stata
binary SAS read_sas
binary Python Pickle Format read_pickle to_pickle
SQL SQL read_sql to_sql
SQL Google Big Query read_gbq to_gbq
For small to medium sized files, I prefer CSV, as properly-formatted CSV can store arbitrary string data, is human readable, and is as dirt-simple as any format can be while achieving the previous two goals.
Unfortunately, for more complex problems, the choice is harder.
If I were on Amazon AWS, I would consider using parquet. However, I do not have any experience with this format.
I can no longer favor the pickle format. Although the pickle format claims long term stability, it allows arbitrary code execution. And even if no code is actually stored in the pickle, ALL pickles execute code just to unpickle them, which is what lead the folks at huggingface to recommend first pickle-tools for scanning pickles and then safetensors as an alternative to pickles. But for saving and loading pandas tables, safetensors is not an option, as they are designed for large multidimensional arrays of floating-point values ("tensors") not for tabular data.
I do not recommend using tofile(). tofile() is best for quick file storage where you do not expect the file to be used on a different machine where the data may have a different endianness (big-/little-endian).
I no longer favor the HDF5 format. It has serious risks for long-term archival since it is fairly complex. It has a 150 page specification, and only one 300,000 line C implementation.
If you have advice on stable, secure, binary formats for saving your own pandas data, please share them! In the meantime, I think CSV for small data and zipped CSVs to save disk space or network bandwidth when needed may be the best way to go. I personally try to avoid any pandas format that requires more than plain-text to store, if possible.
It isn't clear to me if the DataFrame is a view or a copy, but assuming it is a copy, you can use the to_records method of the DataFrame.
This gives you back a record array that you can then put to disk using tofile.
e.g.
df_records = pd.DataFrame(records)
# do some stuff
new_recarray = df_records.to_records()
new_recarray.tofile("myfile.npy")
The data will reside in memory as packed bytes with the format described by the recarray dtype.