🌐
Practical Business Python
pbpython.com › categorical-encoding.html
Guide to Encoding Categorical Values in Python - Practical Business Python
A common alternative approach is called one hot encoding (but also goes by several different names shown below). Despite the different names, the basic strategy is to convert each category value into a new column and assigns a 1 or 0 (True/False) value to the column. This has the benefit of not weighting a value improperly but does have the downside of adding more columns to the data set. Pandas supports this feature using get_dummies.
🌐
Medium
medium.com › @jaberi.mohamedhabib › encoding-categorical-variables-methods-and-techniques-in-pandas-scikit-learn-and-using-dummy-216ae2d5128d
Encoding Categorical Variables: Methods and Techniques in Pandas, Scikit-learn, and Using Dummy Function | by JABERI Mohamed Habib | Medium
September 27, 2024 - Label Encoding is the simplest form of encoding categorical variables. It assigns a unique integer value to each category. This technique works well when the categories have an inherent ordinal relationship (e.g., Low, Medium, High), but it ...
🌐
KDnuggets
kdnuggets.com › 2023 › 07 › pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - For the categorical column, we can break it down into multiple columns. For this, we use pandas.get_dummies() method. It takes the following arguments: To better understand the function, let us work on one-hot encoding the dummy dataset.
🌐
Saturn Cloud
saturncloud.io › blog › a-list-of-pandas-readcsv-encoding-options
A List of Pandas readcsv Encoding Options | Saturn Cloud Blog
May 1, 2026 - UTF-8 is the most widely used encoding format for text data. It supports all characters in the Unicode standard and is compatible with ASCII. To specify UTF-8 encoding in Pandas, use the encoding='utf-8' parameter.
Find elsewhere
🌐
Turing
turing.com › kb › convert-categorical-data-in-pandas-and-scikit-learn
How to Convert Categorical Data in Pandas and Scikit-learn
We generally use one-hot encoding to solve the disadvantage of label encoding. The strategy is to convert each category into a column and assign it a 1 or 0 value. It is a process of creating dummy variables. ... Import pandas as pd #Creating a dataframe Df = pd.Dataframe({‘City’ : [‘Delhi’,’Mumbai’,’Hydrabad’,’Chennai’,’Bangalore’,’Delhi’,’Hydrabad’,’Banglore’,’Delhi’]})
🌐
Seaborn
deeplearningnerds.com › pandas-encode-ordinal-categorical-features
Pandas - Ordinal Encoding
November 16, 2023 - In order to do this, we use the replace() method of Pandas. ... data = { "language": ["Python", "Python", "Java", "JavaScript"], "framework": ["Django", "FastAPI", "Spring", "ReactJS"], "users": [20000, 9000, 7000, 5000], "popularity": ["High", ...
🌐
TutorialsPoint
tutorialspoint.com › python_pandas › python_pandas_series_str_encode_method.htm
Pandas Series.str.encode() Method
This example demonstrates how to use the Series.str.encode() method to encode a column of strings in a DataFrame using the 'utf-8' encoding. import pandas as pd # Create a DataFrame with a column of strings df = pd.DataFrame({ 'COLUMN1': ['', '', ''] }) # Encode strings using 'utf-8' encoding result = df['COLUMN1'].str.encode('utf-8') print("Input DataFrame:") print(df) print("\nDataFrame column after calling str.encode('utf-8'):") print(result)
Top answer
1 of 2
3

You can use this solution implemented to pandas by Series.apply:

from Crypto.Cipher import XOR
import base64

def encrypt(key, plaintext):
  cipher = XOR.new(key)
  return base64.b64encode(cipher.encrypt(plaintext))

def decrypt(key, ciphertext):
  cipher = XOR.new(key)
  return cipher.decrypt(base64.b64decode(ciphertext))

load['Encoded_Column'] = load['F'].apply(lambda x: encrypt('password',x))
load['Decoded_Column'] = (load['Encoded_Column'].apply(lambda x: decrypt('password', x))
                                                .str.decode("utf-8"))
print (load)
   A  B  C  D  E  F Encoded_Column Decoded_Column
0  a  4  7  1  5  a        b'EQ=='              a
1  b  5  8  3  3  a        b'EQ=='              a
2  c  4  9  5  6  a        b'EQ=='              a
3  d  5  4  4  9  b        b'Eg=='              b
4  e  5  2  2  2  b        b'Eg=='              b
5  f  4  0  0  4  b        b'Eg=='              b

Another solution:

import base64
def encode(key, clear):
    enc = []
    for i in range(len(clear)):
        key_c = key[i % len(key)]
        enc_c = chr((ord(clear[i]) + ord(key_c)) % 256)
        enc.append(enc_c)
    return base64.urlsafe_b64encode("".join(enc).encode()).decode()

def decode(key, enc):
    dec = []
    enc = base64.urlsafe_b64decode(enc).decode()
    for i in range(len(enc)):
        key_c = key[i % len(key)]
        dec_c = chr((256 + ord(enc[i]) - ord(key_c)) % 256)
        dec.append(dec_c)
    return "".join(dec)

load['Encoded_Column'] = load['F'].apply(lambda x: encode('password',x))
load['Decoded_Column'] = load['Encoded_Column'].apply(lambda x: decode('password', x))

Or use list comprehension:

load['Encoded_Column'] = [encode('password',x) for x in load['F']]
load['Decoded_Column'] = [decode('password', x) for x in load['Encoded_Column']]

print (load)
   A  B  C  D  E  F Encoded_Column Decoded_Column
0  a  4  7  1  5  a           w5E=              a
1  b  5  8  3  3  a           w5E=              a
2  c  4  9  5  6  a           w5E=              a
3  d  5  4  4  9  b           w5I=              b
4  e  5  2  2  2  b           w5I=              b
5  f  4  0  0  4  b           w5I=              b
2 of 2
0
import pandas as pd
import binascii

load = pd.DataFrame({'A':list('abcdef'),
                   'B':[4,5,4,5,5,4],
                   'C':[7,8,9,4,2,0],
                   'D':[1,3,5,4,2,0],
                   'E':[5,3,6,9,2,4],
                   'F':[binascii.hexlify(x.encode()) for x in 'aaabbb']
                    })

   A  B  C  D  E      F
0  a  4  7  1  5  b'61'
1  b  5  8  3  3  b'61'
2  c  4  9  5  6  b'61'
3  d  5  4  4  9  b'62'
4  e  5  2  2  2  b'62'
5  f  4  0  0  4  b'62'


# decode
binascii.unhexlify(load.loc[1]['F']).decode('utf-8') -->> 'a'

example

print(binascii.hexlify('HelloWorld'.encode())) --> b'48656c6c6f576f726c64'

print(binascii.unhexlify('48656c6c6f576f726c64'.encode())) --> b'HelloWorld'