How does hex to base64 encoding actually work?
So basically, you take the bit representation of the hex string (each hex char is 4 bits), then you take those bits 6 at time (values from 0 to 63) and convert them into an ascii char: values 0-25 to uppercase characters, 26-51 to lowercase characters, 52-61 to digits, 62 to + and 63 to / (or - and _ in urls)
More on reddit.compython base64 to hex - Stack Overflow
Hex to Base64 conversion in Python - Stack Overflow
Base64 to Hex encode/decode
How do I convert Base64 to Hex?
Can I format the Hex output?
Is the conversion secure?
Not sure if this is the right sub but I'm trying to understand how a hex encoded string formated as a base64 string. I understand the concept just not how it's actually implemented. Thanks.
So basically, you take the bit representation of the hex string (each hex char is 4 bits), then you take those bits 6 at time (values from 0 to 63) and convert them into an ascii char: values 0-25 to uppercase characters, 26-51 to lowercase characters, 52-61 to digits, 62 to + and 63 to / (or - and _ in urls)
A hexadecimal string is a representation of some arbitrary binary data. Each hex digit can have one of 16 values, so it's precisely equivalent to four bits of data. This means that each pair of hex digits can represent a single byte. This property also makes it useful for humans working with the binary data, because the conversion is something you can learn to work out in your head.
Base64 is a different, somewhat more efficient representation of some arbitrary binary data. Because there are more legal digits in a base64 encoding, each digit encodes six bits of data. This results in a shorter encoded string than in hex.
The conversion process is quite conceptually simple. First, allocate a sufficiently-sized buffer to hold your data. Then, decode the hex string. Then, re-encode it as base64.
The actual implementation of a base64 encoder is a fun exercise in bit-twiddling, but too detailed to get into here. At a high level, you maintain an index into a bit position in the input (usually a byte index and an offset). You use various masks and shifts to grab six-bit chunks; for each such chunk, you emit an output byte. Pad as required at the end, and you're done.
The best way for converting base64 to hex string is:
# Python 2
>>> base64.b64decode('woidjw==').encode('hex')
# Python 3
>>> base64.b64decode('woidjw==').hex()
'c2889d8f'
You can also try it just like this:
>>> base64.b64decode('woidjw==')
but I am not a fan of the output:
'\xc2\x88\x9d\x8f'
As far as your original request goes, there must be something wrong with your initial data, as it does not result in data that you expected:
>>> base64.b64decode('AAMkADk0ZjU4ODc1LTY1MzAtNDdhZS04NGU5LTAwYjE2Mzg5NDA1ZABGAAAAAAAZS9Y2rt6uTJgnyUZSiNf0BwC6iam6EuExS4FgbbOF87exAAAAdGVuAAC6iam6EuExS4FgbbOF87exAAAxj5dhAAA=').encode('hex')
'0003240039346635383837352d363533302d343761652d383465392d30306231363338393430356400460000000000194bd636aedeae4c9827c9465288d7f40700ba89a9ba12e1314b81606db385f3b7b100000074656e0000ba89a9ba12e1314b81606db385f3b7b10000318f97610000'
For others who wants prefix such as \x you can try the following code:
base64_response = "aGVsbG8h"
base64_decoded_response = base64.b64decode(base64_response)
print(''.join([rf'\x{byte:02x}' for byte in base64_decoded_response]))
'\x68\x65\x6c\x6c\x6f\x21'
Use xxd with the -r argument (and possibly the -p argument) to convert from hex to plain binary/octets and base64 to convert the binary/octet form to base64.
For a file:
cat file.dat | xxd -r -p | base64
For a string of hex numbers:
echo "6F0AD0BFEE7D4B478AFED096E03CD80A" | xxd -r -p | base64
Well, it depends on the exact formatting of your data. But you can do it with a simple shell scripts:
echo "obase=10; ibase=16; `cat in.dat`" | bc | base64 > out.dat
Modify as needed depending on your data.
Edit 26 Aug 2020: As suggested by Ali in the comments, using codecs.encode(b, "base64") would result in extra line breaks for MIME syntax. Only use this method if you do want those line breaks formatting.
For a plain Base64 encoding/decoding, use base64.b64encode and base64.b64decode. See the answer from Ali for details.
In Python 3, arbitrary encodings including Hex and Base64 has been moved to codecs module. To get a Base64 str from a hex str:
import codecs
hex = "10000000000002ae"
b64 = codecs.encode(codecs.decode(hex, 'hex'), 'base64').decode()
from base64 import b64encode, b64decode
# hex -> base64
s = 'cafebabe'
b64 = b64encode(bytes.fromhex(s)).decode()
print('cafebabe in base64:', b64)
# base64 -> hex
s2 = b64decode(b64.encode()).hex()
print('yv66vg== in hex is:', s2)
assert s == s2
This prints:
cafebabe in base64: yv66vg== yv66vg== in hex is: cafebabe
The relevant functions in the documentation, hex to base64:
- b64encode
- bytes.fromhex
- bytes.decode
Base64 to hex:
- b64decode
- str.encode
- bytes.hex
I don't understand why many of the other answers are making it so complicated. For example the most upvoted answer as of Aug 26, 2020:
- There is no need for the
codecsmodule here. - The
codecsmodule usesbase64.encodebytes(s)under the hood (see reference here), so it converts to multiline MIME base64, so you get a new line after every 76 bytes of output. Unless you are sending it in e-mail, it is most likely not what you want.
As for specifying 'utf-8' when encoding a string, or decoding bytes: It adds unnecessary noise. Python 3 uses utf-8 encoding for strings by default. It is not a coincidence that the writers of the standard library made the default encoding of the encode/decode methods also utf-8, so that you don't have to needlessly specify the utf-8 encoding over and over again.
If I understand this correctly, I think the requirement is to translate a base64 encoded string to a hex string in blocks of 8 bytes (16 hex digits). If so, od -t x8 -An, after the base64 decoding will get you there:
$ echo -n "7WkoOEfwfTTioxG6CatHBw==" | base64 -d | od -t x8 -An
347df047382869ed 0747ab09ba11a3e2
$
Output the hex code without newline:
echo "<BASE64>" | base64 -d | hexdump -v -e '/1 "%02x" '