How to encode and decode a column in python pandas? - Stack Overflow
python - Pandas - String values encoding - Stack Overflow
Pandas DataFrame.to_csv() struggling with encoding
python - encoding text columns in pandas data frame - Stack Overflow
You can use this solution implemented to pandas by Series.apply:
from Crypto.Cipher import XOR
import base64
def encrypt(key, plaintext):
cipher = XOR.new(key)
return base64.b64encode(cipher.encrypt(plaintext))
def decrypt(key, ciphertext):
cipher = XOR.new(key)
return cipher.decrypt(base64.b64decode(ciphertext))
load['Encoded_Column'] = load['F'].apply(lambda x: encrypt('password',x))
load['Decoded_Column'] = (load['Encoded_Column'].apply(lambda x: decrypt('password', x))
.str.decode("utf-8"))
print (load)
A B C D E F Encoded_Column Decoded_Column
0 a 4 7 1 5 a b'EQ==' a
1 b 5 8 3 3 a b'EQ==' a
2 c 4 9 5 6 a b'EQ==' a
3 d 5 4 4 9 b b'Eg==' b
4 e 5 2 2 2 b b'Eg==' b
5 f 4 0 0 4 b b'Eg==' b
Another solution:
import base64
def encode(key, clear):
enc = []
for i in range(len(clear)):
key_c = key[i % len(key)]
enc_c = chr((ord(clear[i]) + ord(key_c)) % 256)
enc.append(enc_c)
return base64.urlsafe_b64encode("".join(enc).encode()).decode()
def decode(key, enc):
dec = []
enc = base64.urlsafe_b64decode(enc).decode()
for i in range(len(enc)):
key_c = key[i % len(key)]
dec_c = chr((256 + ord(enc[i]) - ord(key_c)) % 256)
dec.append(dec_c)
return "".join(dec)
load['Encoded_Column'] = load['F'].apply(lambda x: encode('password',x))
load['Decoded_Column'] = load['Encoded_Column'].apply(lambda x: decode('password', x))
Or use list comprehension:
load['Encoded_Column'] = [encode('password',x) for x in load['F']]
load['Decoded_Column'] = [decode('password', x) for x in load['Encoded_Column']]
print (load)
A B C D E F Encoded_Column Decoded_Column
0 a 4 7 1 5 a w5E= a
1 b 5 8 3 3 a w5E= a
2 c 4 9 5 6 a w5E= a
3 d 5 4 4 9 b w5I= b
4 e 5 2 2 2 b w5I= b
5 f 4 0 0 4 b w5I= b
import pandas as pd
import binascii
load = pd.DataFrame({'A':list('abcdef'),
'B':[4,5,4,5,5,4],
'C':[7,8,9,4,2,0],
'D':[1,3,5,4,2,0],
'E':[5,3,6,9,2,4],
'F':[binascii.hexlify(x.encode()) for x in 'aaabbb']
})
A B C D E F
0 a 4 7 1 5 b'61'
1 b 5 8 3 3 b'61'
2 c 4 9 5 6 b'61'
3 d 5 4 4 9 b'62'
4 e 5 2 2 2 b'62'
5 f 4 0 0 4 b'62'
# decode
binascii.unhexlify(load.loc[1]['F']).decode('utf-8') -->> 'a'
example
print(binascii.hexlify('HelloWorld'.encode())) --> b'48656c6c6f576f726c64'
print(binascii.unhexlify('48656c6c6f576f726c64'.encode())) --> b'HelloWorld'
Heya,
I got a DataFrame filled with strings in the first column and the rest consisting of integers (except for the headers).
Now when I export this dataframe to a csv file, and the strings contain German Umlauts (ä,ö,ü or something like ß), the exported csv file has weird looking strings at these indices.
Like "für" became "für".
As far as I know the default encoding for to_csv() is utf-8, which means it should work fine? But I also tried the parameter encoding='utf-8', same results.
What am I doing wrong here?
Encoding
As the first error says, codecs is not callable. In fact is the name of the module.
You probably want:
data['text'] = data.apply(lambda row:
codecs.encode(row['text'], 'utf-8'), axis=1)
Tokenization
The error raised by word_tokenize is due to the fact that the function is used on the previously encoded string: codecs.encode renders the text into a bytes literal string.
From the codecs doc:
Most standard codecs are text encodings, which encode text to bytes, but there are also codecs provided that encode text to text, and bytes to bytes.
word_tokenize doesn't work with bytes literar, like the error says (last line of your error traceback).
If you remove the encoding passage it will work.
About your worries on the video: the prefix u means unicode.1
The prefix b means bytes literal.2 This is the prefix of the strings if you print your dataframe after the use of codecs.encode.
In python 3 (I see from the traceback that your version is 3.6) the default string type is Unicode, so the u is redundant and often not shown, but the strings are already unicode.
So I'm quite sure you are safe: you can safely not use codecs.encode.
You could even do something simpler:
df['text'] = df['text'].str.encode('utf-8')
Ref: https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.encode.html