You can use this solution implemented to pandas by Series.apply:

from Crypto.Cipher import XOR
import base64

def encrypt(key, plaintext):
  cipher = XOR.new(key)
  return base64.b64encode(cipher.encrypt(plaintext))

def decrypt(key, ciphertext):
  cipher = XOR.new(key)
  return cipher.decrypt(base64.b64decode(ciphertext))

load['Encoded_Column'] = load['F'].apply(lambda x: encrypt('password',x))
load['Decoded_Column'] = (load['Encoded_Column'].apply(lambda x: decrypt('password', x))
                                                .str.decode("utf-8"))
print (load)
   A  B  C  D  E  F Encoded_Column Decoded_Column
0  a  4  7  1  5  a        b'EQ=='              a
1  b  5  8  3  3  a        b'EQ=='              a
2  c  4  9  5  6  a        b'EQ=='              a
3  d  5  4  4  9  b        b'Eg=='              b
4  e  5  2  2  2  b        b'Eg=='              b
5  f  4  0  0  4  b        b'Eg=='              b

Another solution:

import base64
def encode(key, clear):
    enc = []
    for i in range(len(clear)):
        key_c = key[i % len(key)]
        enc_c = chr((ord(clear[i]) + ord(key_c)) % 256)
        enc.append(enc_c)
    return base64.urlsafe_b64encode("".join(enc).encode()).decode()

def decode(key, enc):
    dec = []
    enc = base64.urlsafe_b64decode(enc).decode()
    for i in range(len(enc)):
        key_c = key[i % len(key)]
        dec_c = chr((256 + ord(enc[i]) - ord(key_c)) % 256)
        dec.append(dec_c)
    return "".join(dec)

load['Encoded_Column'] = load['F'].apply(lambda x: encode('password',x))
load['Decoded_Column'] = load['Encoded_Column'].apply(lambda x: decode('password', x))

Or use list comprehension:

load['Encoded_Column'] = [encode('password',x) for x in load['F']]
load['Decoded_Column'] = [decode('password', x) for x in load['Encoded_Column']]

print (load)
   A  B  C  D  E  F Encoded_Column Decoded_Column
0  a  4  7  1  5  a           w5E=              a
1  b  5  8  3  3  a           w5E=              a
2  c  4  9  5  6  a           w5E=              a
3  d  5  4  4  9  b           w5I=              b
4  e  5  2  2  2  b           w5I=              b
5  f  4  0  0  4  b           w5I=              b
Answer from jezrael on Stack Overflow
Top answer
1 of 2
3

You can use this solution implemented to pandas by Series.apply:

from Crypto.Cipher import XOR
import base64

def encrypt(key, plaintext):
  cipher = XOR.new(key)
  return base64.b64encode(cipher.encrypt(plaintext))

def decrypt(key, ciphertext):
  cipher = XOR.new(key)
  return cipher.decrypt(base64.b64decode(ciphertext))

load['Encoded_Column'] = load['F'].apply(lambda x: encrypt('password',x))
load['Decoded_Column'] = (load['Encoded_Column'].apply(lambda x: decrypt('password', x))
                                                .str.decode("utf-8"))
print (load)
   A  B  C  D  E  F Encoded_Column Decoded_Column
0  a  4  7  1  5  a        b'EQ=='              a
1  b  5  8  3  3  a        b'EQ=='              a
2  c  4  9  5  6  a        b'EQ=='              a
3  d  5  4  4  9  b        b'Eg=='              b
4  e  5  2  2  2  b        b'Eg=='              b
5  f  4  0  0  4  b        b'Eg=='              b

Another solution:

import base64
def encode(key, clear):
    enc = []
    for i in range(len(clear)):
        key_c = key[i % len(key)]
        enc_c = chr((ord(clear[i]) + ord(key_c)) % 256)
        enc.append(enc_c)
    return base64.urlsafe_b64encode("".join(enc).encode()).decode()

def decode(key, enc):
    dec = []
    enc = base64.urlsafe_b64decode(enc).decode()
    for i in range(len(enc)):
        key_c = key[i % len(key)]
        dec_c = chr((256 + ord(enc[i]) - ord(key_c)) % 256)
        dec.append(dec_c)
    return "".join(dec)

load['Encoded_Column'] = load['F'].apply(lambda x: encode('password',x))
load['Decoded_Column'] = load['Encoded_Column'].apply(lambda x: decode('password', x))

Or use list comprehension:

load['Encoded_Column'] = [encode('password',x) for x in load['F']]
load['Decoded_Column'] = [decode('password', x) for x in load['Encoded_Column']]

print (load)
   A  B  C  D  E  F Encoded_Column Decoded_Column
0  a  4  7  1  5  a           w5E=              a
1  b  5  8  3  3  a           w5E=              a
2  c  4  9  5  6  a           w5E=              a
3  d  5  4  4  9  b           w5I=              b
4  e  5  2  2  2  b           w5I=              b
5  f  4  0  0  4  b           w5I=              b
2 of 2
0
import pandas as pd
import binascii

load = pd.DataFrame({'A':list('abcdef'),
                   'B':[4,5,4,5,5,4],
                   'C':[7,8,9,4,2,0],
                   'D':[1,3,5,4,2,0],
                   'E':[5,3,6,9,2,4],
                   'F':[binascii.hexlify(x.encode()) for x in 'aaabbb']
                    })

   A  B  C  D  E      F
0  a  4  7  1  5  b'61'
1  b  5  8  3  3  b'61'
2  c  4  9  5  6  b'61'
3  d  5  4  4  9  b'62'
4  e  5  2  2  2  b'62'
5  f  4  0  0  4  b'62'


# decode
binascii.unhexlify(load.loc[1]['F']).decode('utf-8') -->> 'a'

example

print(binascii.hexlify('HelloWorld'.encode())) --> b'48656c6c6f576f726c64'

print(binascii.unhexlify('48656c6c6f576f726c64'.encode())) --> b'HelloWorld'
๐ŸŒ
DataScience Made Simple
datasciencemadesimple.com โ€บ home โ€บ encode and decode a column of a dataframe in python โ€“ pandas
Encode and decode a column of a dataframe in python - pandas - DataScience Made Simple
February 4, 2023 - encode() function with codec โ€˜base64โ€™ and error handling scheme โ€˜strictโ€™ is used along with the map() function to encode a column of a dataframe and it is stored in the column named quarter_encoded as shown above so the resultant dataframe will be
๐ŸŒ
Practical Business Python
pbpython.com โ€บ categorical-encoding.html
Guide to Encoding Categorical Values in Python - Practical Business Python
A common alternative approach is called one hot encoding (but also goes by several different names shown below). Despite the different names, the basic strategy is to convert each category value into a new column and assigns a 1 or 0 (True/False) value to the column. This has the benefit of not weighting a value improperly but does have the downside of adding more columns to the data set. Pandas supports this feature using get_dummies.
๐ŸŒ
KDnuggets
kdnuggets.com โ€บ 2023 โ€บ 07 โ€บ pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - We use the get_dummies method and pass the original data frame as data input. In columns, we pass a list containing only the categorical_column header. df_encoded = pd.get_dummies(df, columns=['categorical_column', ])
๐ŸŒ
Medium
medium.com โ€บ analytics-vidhya โ€บ categorical-encoding-with-pandas-get-dummies-d6f1ae6a3e06
Categorical Encoding with Pandas: get_dummies | by Samuel Kehinde Ayo | Analytics Vidhya | Medium
September 17, 2021 - Pandas uses the object data type to indicate categorical variables/columns because there are categorical (non-numerical) columns and we need to transform them. For this, we will implement get_dummies.
๐ŸŒ
Skytowner
skytowner.com โ€บ explore โ€บ encoding_categorical_variables_in_pandas
Encoding categorical variables in Pandas
To encode categorical variables, either using one-hot encoding or dummy coding, use Pandas get_dummies(~) method.
Find elsewhere
๐ŸŒ
Pandas
pandas.pydata.org โ€บ pandas-docs โ€บ version โ€บ 0.17.0 โ€บ generated โ€บ pandas.Series.str.encode.html
pandas.Series.str.encode โ€” pandas 0.17.0 documentation
Enter search terms or a module, class or function name ยท Encode character string in the Series/Index to some other encoding using indicated encoding. Equivalent to str.encode()
Top answer
1 of 2
1

One way would be to use df.replace(). You would avoid changing numeric columns this way.

df.replace('[A-Za-z]','N', regex=True).replace('\d','D', regex=True)

Full example with a numeric column called D, A non-numeric called N and TransDetails.

import pandas as pd

data = '''\
D,N,TransDetails
1,ABC,NEFT-PUNB0315500-JITENDER SING
1,123,NEFT-UTIB0CCH274-VIRENDER KUMA
1,123,NEFT-UTIB0CCH274-SUNITA DEVI
1,123,NEFT-PUNB0315500-AMLASH KUMAR
1,123,NEFT-PUNB0109800-FARIDUDDEN
1,123,NEFT-PUNB0109800-IDREESH
1,123,NEFT-PUNB0315500-BUDDHU
1,123,NEFT-UTIB0CCH274-SAKIL AHAMAD
1,123,NEFT-UTIB0CCH274-NAIM AHAMAD
1,123,NEFT-UTIB0CCH274-SALIM AHAMAD
1,123,NEFT-UTIB0CCH274-NADIM AHAMAD'''

fileobj = pd.compat.StringIO(data) # or 'path/to/csv'
df = pd.read_csv(fileobj)
df = df.replace('[A-Za-z]','N', regex=True).replace('\d','D', regex=True)
print(df)

Returns:

    D    N                    TransDetails
0   1  NNN  NNNN-NNNNDDDDDDD-NNNNNNNN NNNN
1   1  DDD  NNNN-NNNNDNNNDDD-NNNNNNNN NNNN
2   1  DDD    NNNN-NNNNDNNNDDD-NNNNNN NNNN
3   1  DDD   NNNN-NNNNDDDDDDD-NNNNNN NNNNN
4   1  DDD     NNNN-NNNNDDDDDDD-NNNNNNNNNN
5   1  DDD        NNNN-NNNNDDDDDDD-NNNNNNN
6   1  DDD         NNNN-NNNNDDDDDDD-NNNNNN
7   1  DDD   NNNN-NNNNDNNNDDD-NNNNN NNNNNN
8   1  DDD    NNNN-NNNNDNNNDDD-NNNN NNNNNN
9   1  DDD   NNNN-NNNNDNNNDDD-NNNNN NNNNNN
10  1  DDD   NNNN-NNNNDNNNDDD-NNNNN NNNNNN
2 of 2
1

You can use a regular expression to handle the replacement.

df['TransDetails'] = df['TransDetails'].str.replace('[A-Za-z]', 'N')
df['TransDetails'] = df['TransDetails'].str.replace('\d', 'D')

df
# returns:
                      TransDetails
0   NNNN-NNNNDDDDDDD-NNNNNNNN NNNN
1   NNNN-NNNNDNNNDDD-NNNNNNNN NNNN
2     NNNN-NNNNDNNNDDD-NNNNNN NNNN
3    NNNN-NNNNDDDDDDD-NNNNNN NNNNN
4      NNNN-NNNNDDDDDDD-NNNNNNNNNN
5         NNNN-NNNNDDDDDDD-NNNNNNN
6          NNNN-NNNNDDDDDDD-NNNNNN
7    NNNN-NNNNDNNNDDD-NNNNN NNNNNN
8     NNNN-NNNNDNNNDDD-NNNN NNNNNN
9    NNNN-NNNNDNNNDDD-NNNNN NNNNNN
10   NNNN-NNNNDNNNDDD-NNNNN NNNNNN
๐ŸŒ
Altcademy
altcademy.com โ€บ blog โ€บ how-to-one-hot-encode-a-column-in-pandas
How to one hot encode a column in Pandas - Altcademy.com
January 16, 2024 - Pandas simplifies the one hot encoding process with a function called get_dummies. This function automatically converts all categorical variables in a DataFrame to one hot encoded vectors. Here's how you can use get_dummies to one hot encode the 'Pet' column:
๐ŸŒ
Linux find Examples
queirozf.com โ€บ entries โ€บ one-hot-encoding-a-feature-on-a-pandas-dataframe-an-example
One-Hot Encoding a Feature on a Pandas Dataframe: Examples
September 14, 2020 - import pandas pd df = pd.DataFrame({'country': ['russia', 'germany', 'australia','korea','germany']}) pd.get_dummies(df['country'], prefix='country') ... By default, the get_dummies() does not do dummy encoding, but one-hot encoding. To produce an actual dummy encoding from your data, use drop_first=True (not that 'australia' is missing from the columns)
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ one-hot-encoding-from-a-pandas-column-containing-a-list
One-Hot-Encoding from a Pandas Column Containing a List - GeeksforGeeks
August 22, 2024 - It converts categorical variables into a binary matrix representation, where each category is represented by a separate column. This article will guide you through the process of one-hot encoding a Pandas column containing a list of elements, a common scenario in data analysis and machine learning.
Top answer
1 of 16
609

You can easily do this though,

df.apply(LabelEncoder().fit_transform)

EDIT2:

In scikit-learn 0.20, the recommended way is

OneHotEncoder().fit_transform(df)

as the OneHotEncoder now supports string input. Applying OneHotEncoder only to certain columns is possible with the ColumnTransformer.

EDIT:

Since this original answer is over a year ago, and generated many upvotes (including a bounty), I should probably extend this further.

For inverse_transform and transform, you have to do a little bit of hack.

from collections import defaultdict
d = defaultdict(LabelEncoder)

With this, you now retain all columns LabelEncoder as dictionary.

# Encoding the variable
fit = df.apply(lambda x: d[x.name].fit_transform(x))

# Inverse the encoded
fit.apply(lambda x: d[x.name].inverse_transform(x))

# Using the dictionary to label future data
df.apply(lambda x: d[x.name].transform(x))

MOAR EDIT:

Using Neuraxle's FlattenForEach step, it's possible to do this as well to use the same LabelEncoder on all the flattened data at once:

FlattenForEach(LabelEncoder(), then_unflatten=True).fit_transform(df)

For using separate LabelEncoders depending for your columns of data, or if only some of your columns of data needs to be label-encoded and not others, then using a ColumnTransformer is a solution that allows for more control on your column selection and your LabelEncoder instances.

2 of 16
132

As mentioned by larsmans, LabelEncoder() only takes a 1-d array as an argument. That said, it is quite easy to roll your own label encoder that operates on multiple columns of your choosing, and returns a transformed dataframe. My code here is based in part on Zac Stewart's excellent blog post found here.

Creating a custom encoder involves simply creating a class that responds to the fit(), transform(), and fit_transform() methods. In your case, a good start might be something like this:

import pandas as pd
from sklearn.preprocessing import LabelEncoder
from sklearn.pipeline import Pipeline

# Create some toy data in a Pandas dataframe
fruit_data = pd.DataFrame({
    'fruit':  ['apple','orange','pear','orange'],
    'color':  ['red','orange','green','green'],
    'weight': [5,6,3,4]
})

class MultiColumnLabelEncoder:
    def __init__(self,columns = None):
        self.columns = columns # array of column names to encode

    def fit(self,X,y=None):
        return self # not relevant here

    def transform(self,X):
        '''
        Transforms columns of X specified in self.columns using
        LabelEncoder(). If no columns specified, transforms all
        columns in X.
        '''
        output = X.copy()
        if self.columns is not None:
            for col in self.columns:
                output[col] = LabelEncoder().fit_transform(output[col])
        else:
            for colname,col in output.iteritems():
                output[colname] = LabelEncoder().fit_transform(col)
        return output

    def fit_transform(self,X,y=None):
        return self.fit(X,y).transform(X)

Suppose we want to encode our two categorical attributes (fruit and color), while leaving the numeric attribute weight alone. We could do this as follows:

MultiColumnLabelEncoder(columns = ['fruit','color']).fit_transform(fruit_data)

Which transforms our fruit_data dataset from

to

Passing it a dataframe consisting entirely of categorical variables and omitting the columns parameter will result in every column being encoded (which I believe is what you were originally looking for):

MultiColumnLabelEncoder().fit_transform(fruit_data.drop('weight',axis=1))

This transforms

to

.

Note that it'll probably choke when it tries to encode attributes that are already numeric (add some code to handle this if you like).

Another nice feature about this is that we can use this custom transformer in a pipeline:

encoding_pipeline = Pipeline([
    ('encoding',MultiColumnLabelEncoder(columns=['fruit','color']))
    # add more pipeline steps as needed
])
encoding_pipeline.fit_transform(fruit_data)
๐ŸŒ
CodeSignal
codesignal.com โ€บ learn โ€บ courses โ€บ cleaning-and-transforming-data-with-pandas โ€บ lessons โ€บ encoding-categorical-variables-using-python
Encoding Categorical Variables Using Python
Encoding using map: We use the map method to replace each label in the Gender column based on our specified dictionary: {'Male': 1, 'Female': 0}. This dictionary tells Python to encode Male as 1 and Female as 0.
๐ŸŒ
TutorialsPoint
tutorialspoint.com โ€บ python_pandas โ€บ python_pandas_series_str_encode_method.htm
Pandas Series.str.encode() Method
This example demonstrates how to use the Series.str.encode() method to encode a column of strings in a DataFrame using the 'utf-8' encoding. import pandas as pd # Create a DataFrame with a column of strings df = pd.DataFrame({ 'COLUMN1': ['', ...