How about this:
>>> ''.join('{:08b}'.format(b) for b in 'سلام'.encode('utf8'))
'1101100010110011110110011000010011011000101001111101100110000101'
This iterates over the encoded bytes object, where you get an integer in the range 0..255 for each iteration.
Then the integer is formatted in binary notation with zero padding up to 8 digits.
Then glue everything together with str.join().
For the inverse, the approach given in an answer from the question you linked to can be adapted to Python 3 as follows (s is the output of the above example, ie. a str of 0s and 1s):
>>> import re
>>> bytes(int(b, 2) for b in re.split('(........)', s) if b).decode('utf8')
'سلام'
Answer from lenz on Stack Overflowwriting unicode to binary file in python - Stack Overflow
string - Convert Python str/unicode object to binary/hex blob - Stack Overflow
python - Write a unicode character to a file in a binary way - Stack Overflow
python - Binary Data To Unicode - Stack Overflow
Unicode != UTF-8. UTF-8 is a binary encoding of Unicode, so just write the UTF-8 string just as you would an ASCII string. No need to pack an encoded string either. It's already "just a bunch of bytes".
# coding: utf8
import struct
text = u'我是美国人。'
encoded_text = text.encode('utf8')
# proof packing is redundant...
format = '{0}s'.format(len(encoded_text))
packed_text = struct.pack(format,encoded_text)
print encoded_text == packed_text # result: True
So just encode your Unicode strings and append them to the file after writing your packed ints.
unicode.encode('utf-8') will return a byte string encoded in UTF-8; just check for the length before packing.
You could try bitarray:
>>> import bitarray
>>> b = bitarray.bitarray()
>>> b.fromstring('a')
>>> b
bitarray('01100001')
>>> b.to01()
'01100001'
>>> b.fromstring('pples')
>>> b.tostring()
'apples'
>>> b.to01()
'011000010111000001110000011011000110010101110011'
Quite simple and don't require modules from pypi:
def strbin(s):
return ''.join(format(ord(i),'0>8b') for i in s)
You'll need Python 2.6+ to use that.
You are confusing Unicode with encodings. An encoding is a standard that represents text as within the confines of individual values in the range of 0-255 (bytes), while Unicode is a standard that describes codepoints representing textual glyphs. The two are related but not the same thing.
The Unicode standard includes several encodings. UTF-16 is one such encoding that uses 2 bytes per codepoint, but it is not the only encoding included in the standard. UTF-8 is another such encoding, and it uses a variable number of bytes per codepoint.
Your file, however, is written using ASCII, the default codec used by Python 2 when you do not specify an explicit encoding. If you expected to see 2 bytes per codepoint, encode to UTF-16 explicitly:
fin.write(u'\x40'.encode('utf16-le')
This writes UTF-16 in little endian byte order; there is also a utf16-be codec. Normally, for multi-byte encodings like UTF-16 or UTF32, you'd also include a BOM, or Byte Order Mark; it is included automatically when you write UTF-16 without picking any endianes.
fin.write(u'\x40'.encode('utf16')
I strongly urge you to study up on Unicode, codecs and Python before you continue:
The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!) by Joel Spolsky
The Python Unicode HOWTO
Pragmatic Unicode by Ned Batchelder
- Character numbers from U+0000 to U+007F (US-ASCII repertoire) correspond to octets 00 to 7F (7 bit US-ASCII values). A direct consequence is that a plain ASCII string is also a valid UTF-8 string.
- UTF-8, a transformation format of ISO 10646
If you want to post binary data use the base64 encoding.
http://docs.python.org/library/base64.html
"Binary data" is not text, therefore converting it to a unicode is meaningless. If there is text embedded in the binary data then extract it first and decode using the encoding given in the specification for the data format.
Cleaner version:
>>> test_string = '1101100110000110110110011000001011011000101001111101100010101000'
>>> print ('%x' % int(test_string, 2)).decode('hex').decode('utf-8')
نقاب
Inverse (from @Robᵩ's comment):
>>> '{:b}'.format(int(u'نقاب'.encode('utf-8').encode('hex'), 16))
1: '1101100110000110110110011000001011011000101001111101100010101000'
Well, the idea I have is:
1. Split the string into octets
2. Convert the octet to hexadecimal using int and later chr
3. Join them and decode the utf-8 string into Unicode
This code works for me, but I'm not sure what does it print because I don't have utf-8 in my console (Windows :P ).
s = '1101100110000110110110011000001011011000101001111101100010101000'
u = "".join([chr(int(x,2)) for x in [s[i:i+8]
for i in range(0,len(s), 8)
]
])
d = u.decode('utf-8')
Hope this helps!