hexdump -ve '1/1 "%02x"'
xxd -p | tr -d '\n'
If you get tired of writing this every time, create an alias.
Answer from grawity on Stack Exchangehexdump - Parsing Hex dump - Stack Overflow
c - Transform hexadecimal information to binary using a Linux command - Stack Overflow
bash - Can xxd be used to output the binary representation of hex number , not a string? - Unix & Linux Stack Exchange
bash - Decode a base64 string and encode it as hex using xxd - Stack Overflow
How do I decode hex to UTF-8 or ASCII correctly?
Why does my hex string fail to decode?
Can I decode odd-length hex?
hexdump -ve '1/1 "%02x"'
xxd -p | tr -d '\n'
If you get tired of writing this every time, create an alias.
How to easily convert to/from plain machine-readable hexadecimal data
In brief.
$ xxd -plain test.txt > test.hex $ xxd -plain -revert test.hex test2.txt $ diff test.txt test2.txt $
Explanation:
$ xxd -plain test.txt > test.hex
This writes a hex encoding of the data in test.txt into new file test.hex.
The -p or -plain option makes xxd use "plain" hex format with no spaces between pairs of hex digits (i.e. no spaces between byte values). This converts "abc ABC" to "61626320414243". Without the -p it would convert the text to a 16-bit word oriented traditional hexdump format, which is arguably easier to read but less compact and therefore less suitable as a transmission format and slightly harder to reverse.
$ xxd -plain -revert text.hex test2.txt
This uses the -r or -revert option for reverse operation.
The -plain option is used again to indicate that the input hex file is in plain format.
I make the output filename different from the original filename so we can later compare the results with the original file.
$ diff test.txt test2.txt
$
The diff command outputs nothing - this means there is no difference between the original and reconstituted file contents.
I'm tired of digging of some special format strings
Use alias or declare functions in your .profile to create mnemonics so you don't have to remember or dig about in man pages.
or just remember -plain and -revert.
Wrapped output
Yes, there are new-line characters in the output. You want to avoid that.
You can use the -c or -cols option to specify how long you want the output lines to be to attempt to avoid line-wrapping of the output. -c 0 gives the default length and the man page suggests 256 is the limit but it seems to work beyond that.
$ xxd -plain -cols 9999 test.txt > test.hex
$ wc test.txt test.hex
121 880 4603 test.txt
1 1 9207 test.hex
The wc wordcount command tells us how many lines, words and characters are in each file.
So 121 lines (880 words, 4603 bytes) of ASCII text were encoded as 1 line of hex digits.
As @user786653 suggested, use the xxd(1) program:
xxd -r -p input.txt output.bin
Python stdlib solution
If for some unfathomably enterprisey reason you can't sudo apt install xxd, it is easy to reimplement it in Python as per: How to create python bytes object from long hex string? with:
xxd2() ( python -c "import sys;import fileinput;sys.stdout.buffer.write(bytes.fromhex(''.join(fileinput.input(sys.argv[1:]))))" "$@" )
which works both with files and stdin:
printf 01ab | xxd2
printf '01 ab' | xxd2
or:
printf 01ab > myfile.hex
xxd2 myfile.hex
Here's the script with better indentation:
import sys
import fileinput
sys.stdout.buffer.write(
bytes.fromhex(
''.join(
fileinput.input(sys.argv[1:])
)
)
)
The bytes.fromhex function ignores whitespaces and newlines since Python 3.7, so it works regardless of the indentation details of the format, as per docs: https://docs.python.org/3.12/library/stdtypes.html#bytes.fromhex
Changed in version 3.7: bytes.fromhex() now skips all ASCII whitespace in the string, not just spaces.
Tested on Python 3.12.3, Ubuntu 24.04.
echo '0A' produces three characters: 0 A NL; xxd -b will then print those three characters in binary. If you wanted just the single byte whose value is 10 (i.e. hexadecimal A), you could write (in bash):
echo -n $'\x0A'
^ ^ ^
| | |
| | +-- `\x` indicates a hexadecimal escape
| +----- Inside a $' string, escapes are interpreted
+------- -n suppresses the trailing newline
A better alternative would be printf '\x0A'; printf interprets escape sequences in the format string, and does not output implicit newlines. (For a completely Posix-compatible solution, you would need an octal escape: printf '\012'. printf should work on any Posix-compatible shell but hexadecimal escapes are an extension.) Yet another bash possibility is echo -n -e '\x0A'; the (non-standard) -e flag asks echo to interpret escape sequences.
echo '0A' | xxd -b won't output the equivalent of hex 0A, because xxd doesn't know that you intend 0A to be a hex number rather than two characters. It just takes its input as a series of characters, regardless of what those characters are.
Endianness does not affect bytes. The order of bits inside a byte is entirely conceptual until the byte is transmitted over a serial line and even then it is only visible with an oscilloscope or something similar.
If you want the binary output for a string of hex digits, xxd -r -p. E.g.:
$ echo '0A0B0C0D' | xxd -r -p | xxd -b
0000000: 00001010 00001011 00001100 00001101 ....
converts 0A0B0C0D into a four bytes of binary (first call to xxd), and then converts it back to be printable (second call). You say you want a binary output, but the examples you're trying for are a printable representation.
I don't know of anything where endianness is ambiguous at the nibble level, as you imply in your second example. The conversion in xxd is a pure byte at a time, not assuming they represent any particular multi-byte number.
Convert small base64 encoded string to hexadecimal
1. Reducing forks
Accepted answer do offer a solution with a lot of forks! I hate useless forks!
There is my short base64 to hexadecimal converter:
b64toHex() {
local _arLines
mapfile -t _arLines < <(base64 -d <<< "$2" | xxd -p)
printf -v "${1:-hexString}" %s "${_arLines[@]}"
}
In order to avoid fork like myVar=$(myFunc args), this function will only populate a variable and won't print out anything.
b64toHex myVar "OQbb8rXnj/DwvglW018uP/1tqldwiJMbjxBhX7ZqwTw="
echo $myVar
3906dbf2b5e78ff0f0be0956d35f2e3ffd6daa577088931b8f10615fb66ac13c
Then you could use:
b64toHex iv_hex "${iv}"
b64toHex key_hex "${key}"
See at bottom of this for execution time comparison.
2. Pure bash way:
This don't depend on xxd or base64 to be installed. (And without forks, will
be significantly quicker than running 3 forks! Keep in mind, if you plan to run this repetitively!)
From this: Bash script - decode encoded string to byte array, on my website: base64decoder.sh.txt, base64decoder.sh, with a very small modification**:
**At line 41:
- ((_ar==0)) && printf -v _res %b "${_res[@]/#/\\x}"
+ ((_ar==0)) && printf -v _res %s "${_res[@]}"
First prepare a read-only array as:
declare -a B64=( {A..Z} {a..z} {0..9} + / '=' )
declare -Ai 'B64R=()'
for i in "${!B64[@]}"; do B64R["${B64[i]}"]=i%64; done
declare -r B64R
unset B64
Then b64ToHex function:
b64ToHex() {
local _4B _Tail _hVal _v _opt OPTIND
local -i iFd _24b _ar
while getopts "av:" _opt; do case $_opt in
a) _ar=1;; v) _v=${OPTARG};; *) return 1;; esac; done
shift $((OPTIND-1))
if [[ $_v ]];then local -n _res=${_v}; else local _res; fi
if [[ $1 ]]; then exec {iFd}<<<"$1" # Open Input FD from string
else exec {iFd}<&0 ; fi # Open Input FD from STDIN
_res=()
while read -rn4 -u $iFd _4B; do
if [[ "$_4B" ]]; then
_Tail=$_4B
_24b=" B64R['${_4B::1}'] << 18 | B64R['${_4B:1:1}'] << 12 |
B64R['${_4B:2:1}'] << 6 | B64R['${_4B:3:1}'] "
printf -v _hval %02x\ $((_24b>>16)) $((_24b>>8&255)) $((_24b&255))
read -ra _hval <<<"$_hval"
_res+=("${_hval[@]}")
fi
done
exec {iFd}<&-
_Tail=${_Tail##*([^=])}
while [[ $_Tail ]]; do
unset "_res[-1]"
_Tail=${_Tail:1}
done
((_ar==0)) && printf -v _res %s "${_res[@]}" && _res=("${_res[0]}")
[[ -z $_v ]] && echo "${_res[@]}"
}
If bash loop are known to be slow, doing a loop over only 32 byte will be significantly quicker and less system expansive than running four forks!
Usage from STDIN:
b64ToHex <<<"OQbb8rXnj/DwvglW018uP/1tqldwiJMbjxBhX7ZqwTw="
3906dbf2b5e78ff0f0be0956d35f2e3ffd6daa577088931b8f10615fb66ac13c
From an argument:
b64ToHex "OQbb8rXnj/DwvglW018uP/1tqldwiJMbjxBhX7ZqwTw="
3906dbf2b5e78ff0f0be0956d35f2e3ffd6daa577088931b8f10615fb66ac13c
Assign a variable:
b64ToHex -v someVar "OQbb8rXnj/DwvglW018uP/1tqldwiJMbjxBhX7ZqwTw="
echo "$someVar"
3906dbf2b5e78ff0f0be0956d35f2e3ffd6daa577088931b8f10615fb66ac13c
Then from, to variables:
b64ToHex -v iv_hex "${iv}"
b64ToHex -v key_hex "${key}"
Note: this is done without any fork.
For fun: retrieving original base64 string, with bash V5.1+, you could:
shopt -s extglob
printf %b ${someVar//??/\\x& } | base64
OQbb8rXnj/DwvglW018uP/1tqldwiJMbjxBhX7ZqwTw=
3. Pure bash way, but usign bash V5.2+
Same script, but by using patsub_replacement and mapfile, I could do this without any bash loop!
declare -a B64=( {A..Z} {a..z} {0..9} + / '=' )
printf -v _b64_tstr '["\44{B64[%d]}"]=%%d%%%%64 ' {0..64}
# shellcheck disable=SC2059 # format is variable.
printf -v _b64_tstr "$_b64_tstr" {0..64}
declare -Ai "B64R=($_b64_tstr)"
unset B64 _b64_tstr
declare -r B64R
b64ToHex52() {
local _line _Tail _v _opt OPTIND _resArry
local -i iFd _ar
while getopts "av:" _opt; do case $_opt in
a) _ar=1;; v) _v=${OPTARG};; *) return 1;; esac; done
shift $((OPTIND-1))
if [[ $_v ]];then local -n _res=${_v}; else local _res; fi
if [[ $1 ]]; then exec {iFd}<<<"$1" # Open Input FD from string
else exec {iFd}<&0 ; fi # Open Input FD from STDIN
mapfile -tu $iFd _lines
read -ra _resArry <<<"${_lines[*]//?/& }"
printf -v _tmpStr '"B64R[%s]<<18|B64R[%s]<<12|B64R[%s]<<6|B64R[%s]" ' \
"${_resArry[@]}"
local -ia "_tmpArry=($_tmpStr)"
printf -v _tmpStr '%06x' "${_tmpArry[@]}"
read -ra _res <<<"${_tmpStr//??/& }"
exec {iFd}<&-
_Tail=${_lines[-1]##*([^=])}
_res=("${_res[@]::${#_res[@]}-${#_Tail}}")
((_ar==0)) && printf -v _res %s "${_res[@]}" && _res=("${_res[0]}")
[[ -z $_v ]] && echo "${_res[@]}"
}
4. Execution time comparison
Well, now a little comparison test by doing repetitively same conversion to compute execution time.
I now have 4 functions:
b64toHexMy version usingbase64andxxdb64ToHexMy pure bash version (notice upperT)b64ToHex52My pure bash using bash version 5.2+base64_to_hexfrom accepted answer.
Here's my little test function:
testB64decoders(){
local TIMEFORMAT='r %3lR, u %3lU, s %3lS, p %P' bunch \
inString="${2:-SGVsbG8gd29ybGQhIFRoaXMgaXMgYSB0ZXN0IHN0cmluZy4=}"
printf -v bunch '%*s' ${1:-100} ''
mapfile -t bunch <<<"${bunch// /$'\n'}"
printf ' - %-29s: ' "3 fork (base64 | xxd)";
time for i in "${bunch[@]}"; do b64toHex hx "$inString"; done
printf ' - %-29s: ' "Pure bash";
time for i in "${bunch[@]}"; do b64ToHex -v Hx "$inString"; done;
printf ' - %-29s: ' "Pure bash V5.2+";
time for i in "${bunch[@]}"; do b64ToHex52 -v Hx5 "$inString"; done;
printf ' - %-29s: ' "5 fork =\$(echo| base64 | xxd)";
time for i in "${bunch[@]}"; do hex=$(base64_to_hex "$inString"); done;
[[ $hx == "$hex" ]] && [[ $Hx == "$hx" ]] && [[ $Hx5 == "$hx" ]] &&
printf -v inString '%b' ${hx//??/\\x&} &&
printf 'Hopefully result strings are same (%s).\n' "${inString@Q}"
}
Let's show with a small thousand of operation:
testB64decoders 1000
Could produce something like:
- 3 fork (base64 | xxd) : r 0m2.032s, u 0m2.315s, s 0m0.699s, p 148.32
- Pure bash : r 0m0.965s, u 0m0.735s, s 0m0.213s, p 98.14
- Pure bash V5.2+ : r 0m0.571s, u 0m0.477s, s 0m0.088s, p 98.91
- 5 fork =$(echo| base64 | xxd): r 0m2.433s, u 0m3.242s, s 0m0.989s, p 173.92
Where r for real, u: user, s: system time and p: cpu percentage is: 100 * ( u + s ) / r
- pure bash method is quicker,
- pure bash method using bash 5.2+ is significantly quicker,
- version using
xxdandbase64with only two fork will even be a little quicker than - version using four forks (a subshell to run three more forks to
xxd,base64andtr). They are the slowest, and yes: two more forks do have system footprint.
Note: on a multicore system, user time is bigger than real time, thanks to parallelization
On my Raspberry-Pi II model B, I had to reduce my test down to 20 loops.
testB64decoders 20
Did produce on my RPi II:
- 3 fork (base64 | xxd) … xxd -p may have \n chars in the output so you need to remove them:
base64_to_hex()
{
echo "$1" | base64 --decode | xxd -p | tr -d '\n'
}
Use your base64 string as example:
$ hex=$( base64_to_hex OQbb8rXnj/DwvglW018uP/1tqldwiJMbjxBhX7ZqwTw= )
$ echo $hex
3906dbf2b5e78ff0f0be0956d35f2e3ffd6daa577088931b8f10615fb66ac13c
First and foremost, I just want to say that I'm relatively new to Bash and I'm taking a pen testing course. This is the command that I need help to decipher because I don't understand how it works:
echo '4f537b32393736656137393964633562333261343431353535313735303937346230317d0a0a' | xxd -r -p
I think that this command pipes in hex code to xxd, which is a utility used to make a hexdump or do the reverse (as explainshell says). The -r option tells xxd to convert the hexdump into binary and the -p option tells it to output the binary string in plain hexdump style. However, the output is not in binary. It seems to be in (I don't know the technical term for it) regular language (OS) and contains a hex string ({2976ea799dc5b32a4415551750974b01}):
OS{2976ea799dc5b32a4415551750974b01}
Why is the output not in binary?