Here's an example of doing it the first way that Patrick mentioned: convert the bitstring to an int and take 8 bits at a time. The natural way to do that generates the bytes in reverse order. To get the bytes back into the proper order I use extended slice notation on the bytearray with a step of -1: b[::-1].
def bitstring_to_bytes(s):
v = int(s, 2)
b = bytearray()
while v:
b.append(v & 0xff)
v >>= 8
return bytes(b[::-1])
s = "0110100001101001"
print(bitstring_to_bytes(s))
Clearly, Patrick's second way is more compact. :)
However, there's a better way to do this in Python 3: use the int.to_bytes method:
def bitstring_to_bytes(s):
return int(s, 2).to_bytes((len(s) + 7) // 8, byteorder='big')
If len(s) is guaranteed to be a multiple of 8, then the first arg of .to_bytes can be simplified:
return int(s, 2).to_bytes(len(s) // 8, byteorder='big')
This will raise OverflowError if len(s) is not a multiple of 8, which may be desirable in some circumstances.
Another option is to use double negation to perform ceiling division. For integers a & b, floor division using //
n = a // b
gives the integer n such that
n <= a/b < n + 1
Eg,
47 // 10 gives 4, and
-47 // 10 gives -5. So
-(-47 // 10) gives 5, effectively performing ceiling division.
Thus in bitstring_to_bytes we could do:
return int(s, 2).to_bytes(-(-len(s) // 8), byteorder='big')
However, not many people are familiar with this efficient & compact idiom, so it's generally considered to be less readable than
return int(s, 2).to_bytes((len(s) + 7) // 8, byteorder='big')
Answer from PM 2Ring on Stack OverflowHere's an example of doing it the first way that Patrick mentioned: convert the bitstring to an int and take 8 bits at a time. The natural way to do that generates the bytes in reverse order. To get the bytes back into the proper order I use extended slice notation on the bytearray with a step of -1: b[::-1].
def bitstring_to_bytes(s):
v = int(s, 2)
b = bytearray()
while v:
b.append(v & 0xff)
v >>= 8
return bytes(b[::-1])
s = "0110100001101001"
print(bitstring_to_bytes(s))
Clearly, Patrick's second way is more compact. :)
However, there's a better way to do this in Python 3: use the int.to_bytes method:
def bitstring_to_bytes(s):
return int(s, 2).to_bytes((len(s) + 7) // 8, byteorder='big')
If len(s) is guaranteed to be a multiple of 8, then the first arg of .to_bytes can be simplified:
return int(s, 2).to_bytes(len(s) // 8, byteorder='big')
This will raise OverflowError if len(s) is not a multiple of 8, which may be desirable in some circumstances.
Another option is to use double negation to perform ceiling division. For integers a & b, floor division using //
n = a // b
gives the integer n such that
n <= a/b < n + 1
Eg,
47 // 10 gives 4, and
-47 // 10 gives -5. So
-(-47 // 10) gives 5, effectively performing ceiling division.
Thus in bitstring_to_bytes we could do:
return int(s, 2).to_bytes(-(-len(s) // 8), byteorder='big')
However, not many people are familiar with this efficient & compact idiom, so it's generally considered to be less readable than
return int(s, 2).to_bytes((len(s) + 7) // 8, byteorder='big')
>>> zero_one_string = "0110100001101001"
>>> int(zero_one_string, 2).to_bytes((len(zero_one_string) + 7) // 8, 'big')
b'hi'
It returns bytes object that is an immutable sequence of bytes. If you want to get a bytearray -- a mutable sequence of bytes -- then just call bytearray(b'hi').
How to convert string to byte array in Python - Stack Overflow
python - How to convert string to byte arrays? - Stack Overflow
python - How to convert string to binary? - Stack Overflow
How do I convert a string to bytes?
You can convert to an int and use the to_bytes method:
s="00000000000000001011000001000010"
def bitstring_to_bytes(s):
return int(s, 2).to_bytes(len(s) // 8, byteorder='big')
print(bitstring_to_bytes(s))
>>>b'\x00\x00\xb0B'
And to get a float:
import struct
struct.unpack('f', bitstring_to_bytes(s))
>>>(88.0,)
From the docs:
Using unsigned char type:
import struct
def bitstring_to_bytes(s):
v = int(s, 2)
b = bytearray()
while v:
b.append(v & 0xff)
v >>= 8
return bytes(b[::-1])
s = "00000000000000001011000001000010"
r = bitstring_to_bytes(s)
print(struct.unpack('2B', r))
OUTPUT:
(176, 66)
Just use a bytearray() which is a list of bytes.
Python2:
s = "ABCD"
b = bytearray()
b.extend(s)
Python3:
s = "ABCD"
b = bytearray()
b.extend(map(ord, s))
By the way, don't use str as a variable name since that is builtin.
encode function can help you here, encode returns an encoded version of the string
In [44]: str = "ABCD"
In [45]: [elem.encode("hex") for elem in str]
Out[45]: ['41', '42', '43', '44']
or you can use array module
In [49]: import array
In [50]: print array.array('B', "ABCD")
array('B', [65, 66, 67, 68])
Python 2.6 and later have a bytearray type which may be what you're looking for. Unlike strings, it is mutable, i.e., you can change individual bytes "in place" rather than having to create a whole new string. It has a nice mix of the features of lists and strings. And it also makes your intent clear, that you are working with arbitrary bytes rather than text.
Perhaps you want this (Python 2):
>>> map(ord,'hello')
[104, 101, 108, 108, 111]
For a Unicode string this would return Unicode code points:
>>> map(ord,u'Hello, ้ฉฌๅ
')
[72, 101, 108, 108, 111, 44, 32, 39532, 20811]
But encode it to get byte values for the encoding:
>>> map(ord,u'Hello, ้ฉฌๅ
'.encode('chinese'))
[72, 101, 108, 108, 111, 44, 32, 194, 237, 191, 203]
>>> map(ord,u'Hello, ้ฉฌๅ
'.encode('utf8'))
[72, 101, 108, 108, 111, 44, 32, 233, 169, 172, 229, 133, 139]
Something like this?
>>> st = "hello world"
>>> ' '.join(format(ord(x), 'b') for x in st)
'1101000 1100101 1101100 1101100 1101111 100000 1110111 1101111 1110010 1101100 1100100'
#using `bytearray`
>>> ' '.join(format(x, 'b') for x in bytearray(st, 'utf-8'))
'1101000 1100101 1101100 1101100 1101111 100000 1110111 1101111 1110010 1101100 1100100'
If by binary you mean bytes type, you can just use encode method of the string object that encodes your string as a bytes object using the passed encoding type. You just need to make sure you pass a proper encoding to encode function.
In [9]: "hello world".encode('ascii')
Out[9]: b'hello world'
In [10]: byte_obj = "hello world".encode('ascii')
In [11]: byte_obj
Out[11]: b'hello world'
In [12]: byte_obj[0]
Out[12]: 104
Otherwise, if you want them in form of zeros and ones --binary representation-- as a more pythonic way you can first convert your string to byte array then use bin function within map :
>>> st = "hello world"
>>> map(bin,bytearray(st))
['0b1101000', '0b1100101', '0b1101100', '0b1101100', '0b1101111', '0b100000', '0b1110111', '0b1101111', '0b1110010', '0b1101100', '0b1100100']
Or you can join it:
>>> ' '.join(map(bin,bytearray(st)))
'0b1101000 0b1100101 0b1101100 0b1101100 0b1101111 0b100000 0b1110111 0b1101111 0b1110010 0b1101100 0b1100100'
Note that in python3 you need to specify an encoding for bytearray function :
>>> ' '.join(map(bin,bytearray(st,'utf8')))
'0b1101000 0b1100101 0b1101100 0b1101100 0b1101111 0b100000 0b1110111 0b1101111 0b1110010 0b1101100 0b1100100'
You can also use binascii module in python 2:
>>> import binascii
>>> bin(int(binascii.hexlify(st),16))
'0b110100001100101011011000110110001101111001000000111011101101111011100100110110001100100'
hexlify return the hexadecimal representation of the binary data then you can convert to int by specifying 16 as its base then convert it to binary with bin.
Suppose I something like
s = "GW\x25\001"
How do I convert that string to bytes, interpreting the backslashes as escapes? In other words, the resulting byte array should be of length 4.
UPDATE: Hmmm, I was taking the string from sys.argv[1], which seems to complicate things and not make it turn out as expected. So I'm still not sure what the answer is.
Another way to do this is by using the bitstring module:
>>> from bitstring import BitArray
>>> input_str = '0xff'
>>> c = BitArray(hex=input_str)
>>> c.bin
'0b11111111'
And if you need to strip the leading 0b:
>>> c.bin[2:]
'11111111'
The bitstring module isn't a requirement, as jcollado's answer shows, but it has lots of performant methods for turning input into bits and manipulating them. You might find this handy (or not), for example:
>>> c.uint
255
>>> c.invert()
>>> c.bin[2:]
'00000000'
etc.
What about something like this?
>>> bin(int('ff', base=16))
'0b11111111'
This will convert the hexadecimal string you have to an integer and that integer to a string in which each byte is set to 0/1 depending on the bit-value of the integer.
As pointed out by a comment, if you need to get rid of the 0b prefix, you can do it this way:
>>> bin(int('ff', base=16))[2:]
'11111111'
... or, if you are using Python 3.9 or newer:
>>> bin(int('ff', base=16)).removeprefix('0b')
'11111111'
Note: using lstrip("0b") here will lead to 0 integer being converted to an empty string. This is almost always not what you want to do.