The string is a literal.
3>> bin(int('0b1010', 2))
'0b1010'
3>> bin(int('1010', 2))
'0b1010'
3>> 0b1010
10
3>> int('0b1010', 2)
10
Answer from Ignacio Vazquez-Abrams on Stack OverflowThe string is a literal.
3>> bin(int('0b1010', 2))
'0b1010'
3>> bin(int('1010', 2))
'0b1010'
3>> 0b1010
10
3>> int('0b1010', 2)
10
Try the following code:
#!python3
def fn(s, base=10):
prefix = s[0:2]
if prefix == '0x':
base = 16
elif prefix == '0b':
base = 2
return bin(int(s, base))
print(fn('15'))
print(fn('0xF'))
print(fn('0b1111'))
If you are sure you have s = "'0b010111'" and you only want to get 010111, then you can just slice the middle like:
s = s[2:-1]
i.e. from index 2 to the one before the last.
But as Ignacio and Antti wrote, numbers are abstract. The 0b11 is one of the string representations of the number 3 the same ways as 3 is another string representation of the number 3.
The repr() always returns a string. The only thing that can be done with the repr result is to strip the apostrophes -- because the repr adds the apostrophes to the string representation to emphasize it is the string representation. If you want a binary representation of a number (as a string without apostrophes) then bin() is the ultimate answer.
python - How to convert string to binary? - Stack Overflow
how do i convert string to binary string and the reverse too?
String to binary
Converting a string which represents binary to binary python - Stack Overflow
Something like this?
>>> st = "hello world"
>>> ' '.join(format(ord(x), 'b') for x in st)
'1101000 1100101 1101100 1101100 1101111 100000 1110111 1101111 1110010 1101100 1100100'
#using `bytearray`
>>> ' '.join(format(x, 'b') for x in bytearray(st, 'utf-8'))
'1101000 1100101 1101100 1101100 1101111 100000 1110111 1101111 1110010 1101100 1100100'
If by binary you mean bytes type, you can just use encode method of the string object that encodes your string as a bytes object using the passed encoding type. You just need to make sure you pass a proper encoding to encode function.
In [9]: "hello world".encode('ascii')
Out[9]: b'hello world'
In [10]: byte_obj = "hello world".encode('ascii')
In [11]: byte_obj
Out[11]: b'hello world'
In [12]: byte_obj[0]
Out[12]: 104
Otherwise, if you want them in form of zeros and ones --binary representation-- as a more pythonic way you can first convert your string to byte array then use bin function within map :
>>> st = "hello world"
>>> map(bin,bytearray(st))
['0b1101000', '0b1100101', '0b1101100', '0b1101100', '0b1101111', '0b100000', '0b1110111', '0b1101111', '0b1110010', '0b1101100', '0b1100100']
Or you can join it:
>>> ' '.join(map(bin,bytearray(st)))
'0b1101000 0b1100101 0b1101100 0b1101100 0b1101111 0b100000 0b1110111 0b1101111 0b1110010 0b1101100 0b1100100'
Note that in python3 you need to specify an encoding for bytearray function :
>>> ' '.join(map(bin,bytearray(st,'utf8')))
'0b1101000 0b1100101 0b1101100 0b1101100 0b1101111 0b100000 0b1110111 0b1101111 0b1110010 0b1101100 0b1100100'
You can also use binascii module in python 2:
>>> import binascii
>>> bin(int(binascii.hexlify(st),16))
'0b110100001100101011011000110110001101111001000000111011101101111011100100110110001100100'
hexlify return the hexadecimal representation of the binary data then you can convert to int by specifying 16 as its base then convert it to binary with bin.
is there any proper method to convert a string in utf-8 encoded formant to binary string format and the reverse too. like example
data = "this string".encode("utf-8")
data = func(data)
print(data) #want to have a string as output like "0101010101111010101010......."
I want a function that takes for example 'Hello brother! ๐'(a combination of any unicode symbols) as input and outputs the binary equivalent using utf-8 so here:
'01001000 01100101 01101100 01101100 01101111 00100000 01100010 01110010 01101111 01110100 01101000 01100101 01110010 00100001 00100000 11110000 10011111 10011000 10001010'
The output should also be a string.
import struct
with open("foo.bin", 'wb') as f:
f.write(struct.pack('h', 0b0010010110101))
Will take 2 bytes (16 bits) as a short integer (h). You can define your own format string using the struct module, but I am not sure you will be able to get under the byte size.
EDIT
As per your comment, here's a bit of context:
When writing something in a file, it is always converted to binary. Character are encoded using some rule, called encoding (such as ASCII) where each character is mapped to a number, itself represented in binary. This way, the number 00100100 (36) and the character '$' are the same thing. '$' is represented by 36 on the file, and the software layers between you (such as an editor) will render every '00100100' it encounters as character '$'.
Now when you write the string '00100100' into a file, it will print the characters '0', '1' etc.... So the string '00100100' is represented by the binary number 110000110000110001110000110000110001110000110000. This is necessary because the input being a string, you need an unambiguous way of representing all possible 8-characters long strings, not only the ones representing 0s and 1s.
The Python API for writing files is always writing strings, i.e. it will perform this conversion string -> binary number automatically, and I don't know any way to override that. What you can do however is generate the string such that its binary representation is the actual binary string you wanted to write: if you want to write the number 00100100 in a file, you can just write f.write('$'), which is effectively the same thing.
That is exactly what the 'struct' module performs: it generates a string of bytes, or characters, which exactly match the number you are providing them.
In my example, I give it the number 0b0010010110101, and tell it to encode it as a short integer, i.e. on two bytes. If you execute struct.pack('h', 1205) in the Python interpreter, it will print out the two characters (bytes) \xb5\x04 which correspond to this number in 'byte-base', i.e. base 256 (with big-endian convention). Indeed:
>>> 0x04 * 256 + 0xb5
1205
Just like you can represent any decimal number in base 10 (e.g. 36), base 16 (e.g. 0x24), base 2 (e.g. 0b100100), you can also represent it in base 256 via the ASCII encoding (e.g. '$'). Struct does exactly that, also providing a convenient 'fmt' string convention for the type of data you are writing. You can also do it directly by converting each of your bytes into the corresponding character:
def encode(binary):
# Aligning on bytes
binary = '0' * (8 - len(binary) % 8) + binary
# Generating the corresponding character for each
# byte encountered
return ''.join(chr(int('0b' + binary[i:i+8], base = 2))
for i in xrange(0, len(binary), 8))
This is a very crude and not super efficient way of proceeding, but it does convert every byte into its corresponding character, and returns the corresponding string, which you can directly write into a file:
>>> encode('001001001010100100100100100111110010101110100')
'\x04\x95$\x93\xe5t'
And indeed, writing this to a file produces 6 bytes, corresponding to the 6 characters:
with open("foo.bin", 'wb') as f:
f.write('\x04\x95$\x93\xe5t')
>>> os.path.getsize("foo.bin")
6L
struct modules performs exactly the same thing, except with a fixed format, and in a more efficient fashion. Instead of getting the chr corresponding to the integer,
def encode2(binary):
rawbytes = []
while binary > 0:
binary, byte = divmod(binary, 256)
rawbytes.append(byte)
fmt_string = '%sB' % len(rawbytes)
print "Encoding %s into %s bytes (%s)" % (rawbytes, len(rawbytes), fmt_string)
return struct.pack(fmt_string, *rawbytes)
>>> encode2(0b001001001010100100100100100111110010101110100)
Encoding [116L, 229L, 147L, 36L, 149L, 4L] into 6 bytes (6B)
't\xe5\x93$\x95\x04'
(Notice that these are the same character outputted as in encode. The only difference is the order, depending on endianness of the conversion).
You can then decode these character using struct as well, and the same format string:
>>> bytes = struct.unpack('6B', 't\xe5\x93$\x95\x04')
>>> bytes
(116, 229, 147, 36, 149, 4)
>>> bin(sum(x * 256 ** i for i, x in enumerate(bytes)))
'0b1001001010100100100100100111110010101110100'
Which is our original number.
Bottom line is: Python file API can only process characters, which are effectively bytes. There might be some magic way of writing individual bits to a file, but I wouldn't count too much on that, as this introduces its own world of problems, and bytes are more than sufficient in 99% of cases. To write binary data, represent it in base 256, and convert each of its b256 digits to the corresponding character. The binary representation of this string is, by definition, your original number.
binascii can be used.
import binascii
a = "1010"
b = "10"
c = "00"
data = a + b + c
hex_string = hex(int(data, 2))[2:] #remove '0x'
with open('foo', 'wb') as f:
f.write(binascii.unhexlify(hex_string))
The hex_string should be even, so you need to add one bit to "0010010110101" to make unhexlify work properly.