df.round

>>> df.round()

np.round

>>> np.round(df)

astype

>>> df.ge(0.5).astype(int)

All which yield

     0    1    2
0  0.0  0.0  1.0
1  0.0  1.0  1.0
2  1.0  0.0  1.0
3  0.0  1.0  0.0

Note: round works here because it automatically sets the threshold for .5 between two integers. For custom thresholds, use the 3rd solution

Answer from rafaelc on Stack Overflow
Discussions

How to convert Pandas dataframe into a binary format?
the next generation of pandas will use Arrow. you can use it today as well https://arrow.apache.org/docs/python/pandas.html More on reddit.com
🌐 r/Python
8
6
January 22, 2018
how to convert pandas dataframe to binary file in python - Stack Overflow
import numpy as np import pandas as pd def get_values_for_frequency(freq): # sampling information Fs = 100# sample rate no of samppes per second T = 1/Fs # sampling period %sample per More on stackoverflow.com
🌐 stackoverflow.com
python - How to change values in a column into binary? - Stack Overflow
New to python and I am stuck at this. My CSV file contain this: Sr,Gender 1,Male 2,Male 3,Female Now I want to convert the Gender values into binary so the the file will look something like: Sr,G... More on stackoverflow.com
🌐 stackoverflow.com
how shall I save data in DataFrame formation into a "binary filed"?
You have to first create the excel ... in your Binary field `file`. However, if you want to edit an already existing excel spreadsheet on your computer you can simply do this: import pandas as pd # Create a Pandas dataframe from the data. df = pd.DataFrame({'Data': [10, 20, 30, 20, 15, 30, 45]}) # Create a Pandas Excel writer using XlsxWriter as the engine. writer = pd.ExcelWriter('pandas_simple.xlsx', engine='xlsxwriter') # Convert the dataframe ... More on odoo.com
🌐 odoo.com
2
0
November 21, 2020
🌐
Finxter
blog.finxter.com › 5-best-ways-to-convert-pandas-dataframe-to-binary-data-in-python
5 Best Ways to Convert Pandas DataFrame to Binary Data in Python – Be on the Right Side of Change
The ‘df_from_binary’ will be identical to the original ‘df’ once read from the binary file, demonstrating the effectiveness of the pickle method for binary conversion. The to_parquet function is used to convert a DataFrame into the binary Parquet format, which is a highly efficient, ...
🌐
YouTube
youtube.com › watch
Converting DataFrame Values to Binary Using Pandas - YouTube
Learn how to reassign values to specific columns and merge them with the rest of your DataFrame using pandas in Python. This guide provides a clear explanati...
Published: September 29, 2025
Views: 0
Find elsewhere
🌐
Stack Overflow
stackoverflow.com › questions › 28854821 › converting-dataframe-from-hex-to-binary-in-using-python
pandas - converting dataframe from Hex to binary in using python - Stack Overflow
If I am right, you first need to concat three columns in a string value = str(col1)+str(col2)+str(col3) and then use the method to convert it in binary.
🌐
GeeksforGeeks
geeksforgeeks.org › how-to-convert-categorical-data-to-binary-data-in-python
How to convert categorical data to binary data in Python? - GeeksforGeeks
January 17, 2022 - NumPy and Pandas are two powerful libraries in the Python ecosystem for data manipulation and analysis. Converting a DataFrame column to a NumPy array is a common operation when you need to perform array-based operations on the data.
🌐
Stack Overflow
stackoverflow.com › questions › 45018601 › decimal-to-binary-in-dataframe-pandas
python - Decimal TO Binary in Dataframe Pandas - Stack Overflow
You want them to be string representations of binary numbers? ... Save this answer. ... Show activity on this post. ... the thing is format doesn't work on strings so you need to convert your inputs to integers before getting their binary string representation ('04b' is just to have the representation on 4 bits).
🌐
TutorialsPoint
tutorialspoint.com › article › how-to-convert-categorical-data-to-binary-data-in-python
How to convert categorical data to binary data in Python?
April 18, 2023 - The simplest method to convert categorical data to binary format is using the pd.get_dummies() function. Let's see how it works ? import pandas as pd # Create a sample DataFrame with categorical data data = {'Gender': ['Male', 'Female', 'Male', ...
Top answer
1 of 2
4

If performance is important, use numpy with this solution:

d = df['Col_B'].values
m = 2
df[['Col_C','Col_D']]  = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
print (df)
  Col_A  Col_B  Col_C  Col_D
0     a      1      1      0
1     b      2      0      1
2     c      0      0      0

Performance (about 1000 times faster):

df = pd.DataFrame([['a', 1], ['b', 2], ['c', 0]], columns=["Col_A", "Col_B"])


df = pd.concat([df] * 1000, ignore_index=True)

In [162]: %%timeit
     ...: df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
     ...: 
609 ms ± 14.5 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

In [163]: %%timeit
     ...: d = df['Col_B'].values
     ...: m = 2
     ...: df[['Col_C','Col_D']]  = pd.DataFrame((((d[:,None] & (1 << np.arange(m)))) > 0).astype(int))
     ...: 
618 µs ± 26.2 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
2 of 2
2

apply is the method you are looking for.

df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))

does the trick.

I benchmarked it on 3000 rows and it is faster than the for cycle method you mention (0.5 seconds vs 3 seconds). But generally the speed won't be much faster since it still needs to apply the function for each row separately.

from time import time
start = time()
for i in range(0,len(df)):
    df.loc[i,'Col_C'],df.loc[i,'Col_D'] = list( (bin(df.loc[i,'Col_B'])[2:].zfill(2) ) )
print(time() - start)
# 3.4339962005615234

start = time()
df[['Col_C','Col_D']] = df['Col_B'].apply(lambda x: pd.Series(list(bin(x)[2:].zfill(2))))
print(time() - start)
# 0.5619983673095703

Note: I am using python 3, so e.g. bin(1) returns '0b1' and thus I use bin(x)[2:] to get rid of the '0b' part.

Top answer
1 of 2
30

You mean "one-hot" encoding?

Say you have the following dataset:

import pandas as pd
df = pd.DataFrame([
            ['green', 1, 10.1, 0], 
            ['red', 2, 13.5, 1], 
            ['blue', 3, 15.3, 0]])

df.columns = ['color', 'size', 'prize', 'class label']
df

Now, you have multiple options ...

A) The Tedious Approach

color_mapping = {
           'green': (0,0,1),
           'red': (0,1,0),
           'blue': (1,0,0)}

df['color'] = df['color'].map(color_mapping)
df

import numpy as np
y = df['class label'].values
X = df.iloc[:, :-1].values
X = np.apply_along_axis(func1d= lambda x: np.array(list(x[0]) + list(x[1:])), axis=1, arr=X)

print('Class labels:', y)
print('\nFeatures:\n', X)

Yielding:

Class labels: [0 1 0]

Features:
 [[  0.    0.    1.    1.   10.1]
 [  0.    1.    0.    2.   13.5]
 [  1.    0.    0.    3.   15.3]]

B) Scikit-learn's DictVectorizer

from sklearn.feature_extraction import DictVectorizer
dvec = DictVectorizer(sparse=False)

X = dvec.fit_transform(df.transpose().to_dict().values())
X

Yielding:

array([[  0. ,   0. ,   1. ,   0. ,  10.1,   1. ],
       [  1. ,   0. ,   0. ,   1. ,  13.5,   2. ],
       [  0. ,   1. ,   0. ,   0. ,  15.3,   3. ]])

C) Pandas' get_dummies

pd.get_dummies(df)

2 of 2
14

It seems that you are using scikit-learn's DictVectorizer to convert the categorical values to binary. In that case, to store the result along with the new column names, you can construct a new DataFrame with values from vec_x and columns from DV.get_feature_names(). Then, store the DataFrame to disk (e.g. with to_csv()) instead of the numpy array.

Alternatively, it is also possible to use pandas to do the encoding directly with the get_dummies function:

import pandas as pd
data = pd.DataFrame({'T': ['A', 'B', 'C', 'D', 'E']})
res = pd.get_dummies(data)
res.to_csv('output.csv')
print res

Output:

   T_A  T_B  T_C  T_D  T_E
0    1    0    0    0    0
1    0    1    0    0    0
2    0    0    1    0    0
3    0    0    0    1    0
4    0    0    0    0    1
🌐
GitHub
gist.github.com › fonnesbeck › 629f0e46420633b365c57783d17dda79
Writing a formated binary file from a Pandas Dataframe · GitHub
Save fonnesbeck/629f0e46420633b365c57783d17dda79 to your computer and use it in GitHub Desktop. Download ZIP · Writing a formated binary file from a Pandas Dataframe · Raw · binary_df.py · This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below.
🌐
YouTube
youtube.com › watch
How to Convert Categorical Values to Binary (0 and 1) in Python with Pandas) - YouTube
This video explains How to Convert Categorical Values to Binary values (Python and Pandas) with Jupyter NotebookHow to build a simple Neural Network - https...
Published: September 27, 2019
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › pyspark.sql › api › pyspark.sql.functions.to_binary.html
pyspark.sql.functions.to_binary — PySpark 4.2.0 documentation
>>> import pyspark.sql.functions as sf >>> df = spark.createDataFrame([("abc",)], ["e"]) >>> df.select(sf.try_to_binary(df.e, sf.lit("utf-8")).alias('r')).collect() [Row(r=b'abc')] Example 2: Convert string to a timestamp without encoding specified
Top answer
1 of 2
11

Pandas now offers a wide variety of formats:

Format Type Data Description     Reader         Writer
text        CSV                  read_csv       to_csv
text        JSON                 read_json      to_json
text        HTML                 read_html      to_html
text        Local clipboard      read_clipboard to_clipboard
binary      MS Excel             read_excel     to_excel
binary      HDF5 Format          read_hdf       to_hdf
binary      Feather Format       read_feather   to_feather
binary      Parquet Format       read_parquet   to_parquet
binary      Msgpack              read_msgpack   to_msgpack
binary      Stata                read_stata     to_stata
binary      SAS                  read_sas    
binary      Python Pickle Format read_pickle    to_pickle
SQL         SQL                  read_sql       to_sql
SQL         Google Big Query     read_gbq       to_gbq

For small to medium sized files, I prefer CSV, as properly-formatted CSV can store arbitrary string data, is human readable, and is as dirt-simple as any format can be while achieving the previous two goals.

Unfortunately, for more complex problems, the choice is harder.

If I were on Amazon AWS, I would consider using parquet. However, I do not have any experience with this format.

I can no longer favor the pickle format. Although the pickle format claims long term stability, it allows arbitrary code execution. And even if no code is actually stored in the pickle, ALL pickles execute code just to unpickle them, which is what lead the folks at huggingface to recommend first pickle-tools for scanning pickles and then safetensors as an alternative to pickles. But for saving and loading pandas tables, safetensors is not an option, as they are designed for large multidimensional arrays of floating-point values ("tensors") not for tabular data.

I do not recommend using tofile(). tofile() is best for quick file storage where you do not expect the file to be used on a different machine where the data may have a different endianness (big-/little-endian).

I no longer favor the HDF5 format. It has serious risks for long-term archival since it is fairly complex. It has a 150 page specification, and only one 300,000 line C implementation.

If you have advice on stable, secure, binary formats for saving your own pandas data, please share them! In the meantime, I think CSV for small data and zipped CSVs to save disk space or network bandwidth when needed may be the best way to go. I personally try to avoid any pandas format that requires more than plain-text to store, if possible.

2 of 2
6

It isn't clear to me if the DataFrame is a view or a copy, but assuming it is a copy, you can use the to_records method of the DataFrame.

This gives you back a record array that you can then put to disk using tofile.

e.g.

df_records = pd.DataFrame(records)
# do some stuff
new_recarray = df_records.to_records()
new_recarray.tofile("myfile.npy")

The data will reside in memory as packed bytes with the format described by the recarray dtype.