Decode it.
>>> b'a string'.decode('ascii')
'a string'
To get bytes from string, encode it.
>>> 'a string'.encode('ascii')
b'a string'
Answer from falsetru on Stack OverflowYour bytes simply cannot be decoded as utf-8, just as the error message tells you.
utf-8 is the default encoding parameter of decode - and the best way to put in the correct encoding value is to know the encoding - otherwise you'll have to guess.
And guessing is probably what the website does, too, by trying the most common encodings, until one does not throw an exception:
def decodeAscii(bin_string):
binary_int = int(bin_string, 2);
byte_number = binary_int.bit_length() + 7 // 8
binary_array = binary_int.to_bytes(byte_number, "big")
ascii_text = "Bin string cannot be decoded"
for enc in ['utf-8', 'ascii', 'ansi']:
try:
ascii_text = binary_array.decode(encoding=enc)
break
except:
pass
print(ascii_text)
s = "01010011101100000110010101101100011011000110111101110100011010000110010101110010011001010110100001101111011101110111100101101111011101010110010001101111011010010110111001100111011010010110110101100110011010010110111001100101011000010111001001100101011110010110111101110101011001100110100101101110011001010101000000000000"
decodeAscii(s)
Output:
SยฐellotherehowyoudoingimfineareyoufineP
But there's no guarantee that you find the "correct" encoding by guessing.
Your binary string is just not a valid ascii or utf-8 string. You can tell decode to ignore invalid sequences by saying
ascii_text = binary_array.decode(errors='ignore')
You are looping over the individual characters of the input message, but you need to instead look for groups of 9 characters (2 times 4 binary digits and the space). Your mapping has keys like '0100 1001', not '0' and '1' and ' '
The simplest approach (albeit a bit brittle) would be to loop over indices in steps of 10 characters (1 extra for the space between the characters), then grab 9 characters:
for i in xrange(0, len(messageDecode), 10):
group = messageDecode[i:i + 9]
print inverseBINARY[group],
The xrange() object produces integers 10 apart; so 0, 10, 20, etc. The messageDecode string is then sliced to grab 9 characters starting at that index, so messageDecode[0:9] and messageDecode[10:19], messageDecode[20:29], etc.
A more robust approach would be to remove all spaces and grab blocks every 8 characters; that'd leave room for extra spaces in between, but you do have to re-insert that space to match your keys:
messageDecode = messageDecode.replace(' ', '')
for i in xrange(0, len(messageDecode), 8):
group = messageDecode[i:i + 4] + ' ' + messageDecode[i + 4:i + 8]
print inverseBINARY[group],
or you could perhaps not include the spaces in your inverseBINARY mapping here:
inverseBINARY = {v.replace(' ', ''): k for k, v in BINARY.items()}
and then simply slice every 8 characters:
messageDecode = messageDecode.replace(' ', '')
for i in xrange(0, len(messageDecode), 8):
group = messageDecode[i:i + 8]
print inverseBINARY[group],
if you want to decode binary why not use native functions as the binary number and chr ?
>>> print chr(0b01000010)
B
EDIT
Ok then, this is how i would solve that:
from string import letters, punctuation
encode_data = {letter:bin(ord(letter)) for letter in letters+punctuation+' '}
decode_data = {bin(ord(letter)):letter for letter in letters+punctuation+' '}
def encode(message):
return [encode_data[letter] for letter in message]
def decode(table):
return [decode_data[item] for item in table]
encoded = encode('hello there')
print decode(encoded) # ['h', 'e', 'l', 'l', 'o', ' ', 't', 'h', 'e', 'r', 'e']