Python 3 strings contain Unicode code points, not "UTF-8 characters". You can use ord() to get the Unicode code point value, and .encode() to convert it to UTF-8 bytes. Then format each byte as 8-digit binary text, and .join() them together. Example:
# starting and ending code points for 1-, 2-, 3- and 4-byte UTF-8.
s1 = '\x00\x7f\x80\u07ff\u0800\uffff\U00010000\U0010FFFF'
# some printable characters in each range
s2 = 'Aü马🎂'
def utf8_bin(u):
# format as 8-digit binary, join each byte with space
return ' '.join([f'{i:08b}' for i in u.encode()])
for u in s1:
col1 = f'U+{ord(u):04X}' # format Unicode codepoint, leading zeros if <4 digits.
print(f'{col1:8} {utf8_bin(u)}')
print()
for u in s2:
col1 = f'U+{ord(u):04X}'
print(f'{col1:8} {u} {utf8_bin(u)}')
Output:
U+0000 00000000
U+007F 01111111
U+0080 11000010 10000000
U+07FF 11011111 10111111
U+0800 11100000 10100000 10000000
U+FFFF 11101111 10111111 10111111
U+10000 11110000 10010000 10000000 10000000
U+10FFFF 11110100 10001111 10111111 10111111
U+0041 A 01000001
U+00FC ü 11000011 10111100
U+9A6C 马 11101001 10101001 10101100
U+1F382 🎂 11110000 10011111 10001110 10000010
Answer from Mark Tolonen on Stack OverflowCleaner version:
>>> test_string = '1101100110000110110110011000001011011000101001111101100010101000'
>>> print ('%x' % int(test_string, 2)).decode('hex').decode('utf-8')
نقاب
Inverse (from @Robᵩ's comment):
>>> '{:b}'.format(int(u'نقاب'.encode('utf-8').encode('hex'), 16))
1: '1101100110000110110110011000001011011000101001111101100010101000'
Well, the idea I have is:
1. Split the string into octets
2. Convert the octet to hexadecimal using int and later chr
3. Join them and decode the utf-8 string into Unicode
This code works for me, but I'm not sure what does it print because I don't have utf-8 in my console (Windows :P ).
s = '1101100110000110110110011000001011011000101001111101100010101000'
u = "".join([chr(int(x,2)) for x in [s[i:i+8]
for i in range(0,len(s), 8)
]
])
d = u.decode('utf-8')
Hope this helps!
utf 8 - How to convert utf8 into binary - Stack Overflow
writing unicode to binary file in python - Stack Overflow
python - Convert UTF-8 as string of binary 0 and 1s to code point - Stack Overflow
utf 8 - python - convert binary data to utf-8 - Stack Overflow
is there any proper method to convert a string in utf-8 encoded formant to binary string format and the reverse too. like example
data = "this string".encode("utf-8")
data = func(data)
print(data) #want to have a string as output like "0101010101111010101010......."
Unicode != UTF-8. UTF-8 is a binary encoding of Unicode, so just write the UTF-8 string just as you would an ASCII string. No need to pack an encoded string either. It's already "just a bunch of bytes".
# coding: utf8
import struct
text = u'我是美国人。'
encoded_text = text.encode('utf8')
# proof packing is redundant...
format = '{0}s'.format(len(encoded_text))
packed_text = struct.pack(format,encoded_text)
print encoded_text == packed_text # result: True
So just encode your Unicode strings and append them to the file after writing your packed ints.
unicode.encode('utf-8') will return a byte string encoded in UTF-8; just check for the length before packing.