How about this:

>>> ''.join('{:08b}'.format(b) for b in 'سلام'.encode('utf8'))
'1101100010110011110110011000010011011000101001111101100110000101'

This iterates over the encoded bytes object, where you get an integer in the range 0..255 for each iteration. Then the integer is formatted in binary notation with zero padding up to 8 digits. Then glue everything together with str.join().

For the inverse, the approach given in an answer from the question you linked to can be adapted to Python 3 as follows (s is the output of the above example, ie. a str of 0s and 1s):

>>> import re
>>> bytes(int(b, 2) for b in re.split('(........)', s) if b).decode('utf8')
'سلام'
Answer from lenz on Stack Overflow
Discussions

writing unicode to binary file in python - Stack Overflow
I'm wondering how to write unicode (utf-8) to a binary file. Here's the background: I've got a 40 byte header (10 ints), and a table with a variable number of triple-int structs. Writing these wa... More on stackoverflow.com
🌐 stackoverflow.com
March 23, 2021
string - Convert Python str/unicode object to binary/hex blob - Stack Overflow
Is there an easy way to get some str/unicode object represented as a big binary number (or an hex one)? I've been reading some answers to related questions but none of them works for my scenario. I More on stackoverflow.com
🌐 stackoverflow.com
python - Write a unicode character to a file in a binary way - Stack Overflow
The Unicode standard includes several encodings. UTF-16 is one such encoding that uses 2 bytes per codepoint, but it is not the only encoding included in the standard. UTF-8 is another such encoding, and it uses a variable number of bytes per codepoint. Your file, however, is written using ASCII, the default codec used by Python 2 when you do not specify an explicit encoding. If you expected to ... More on stackoverflow.com
🌐 stackoverflow.com
August 17, 2013
python - Binary Data To Unicode - Stack Overflow
It could be bytes from which Unicode can be decoded. 2011-02-23T18:09:06.83Z+00:00 ... @S.Lott: If the extraction process is just using the whole thing as-is then so be it. But I stand by my answer. 2011-02-23T18:14:08.38Z+00:00 ... You should stand by your answer. However, you could also consider extending it to cover the most common case of getting binary data from a file. 2011-02-23T18:45:42.163Z+00:00 ... Your answer makes complete sense to me. I'm in a Python ... More on stackoverflow.com
🌐 stackoverflow.com
🌐
Python
docs.python.org › 3 › library › binascii.html
binascii — Convert between binary and ASCII
Changed in version 3.3: ASCII-only unicode strings are now accepted by the a2b_* functions. The binascii module defines the following functions: ... Convert a single line of uuencoded data back to binary and return the binary data.
🌐
Learning Python
learning-python.com › strings30.html
Strings in 3.X: Unicode and Binary Data
At a more concrete level, the Python language provides multiple string data types to represent content in your script: both textual data—integer code-point values of decoded Unicode characters in memory, as well as binary data—raw byte values, including text that is in encoded form.
🌐
Online Tools
onlinetools.com › unicode › convert-unicode-to-binary
Convert Unicode to Binary – Online Unicode Tools
It supports the most popular Unicode encodings (such as UTF-8, UTF-16, UTF-32, UCS-2, and UCS-4) and it works with emoji characters. You can also customize the binary output format by enabling binary padding and spacing. Created by encoding gurus from team Browserling. ... Can't convert. ... Chain with... ... This tool cannot be chained.
🌐
Delft Stack
delftstack.com › home › howto › python › string to binary python
How to Convert a String to Binary in Python | Delft Stack
February 2, 2024 - This article will discuss some methods to convert a string to its binary representation in Python. We use the ord() function that translates the Unicode point of the string to a corresponding integer.
🌐
Real Python
realpython.com › python-encodings-guide
Unicode & Character Encodings in Python: A Painless Guide – Real Python
May 20, 2019 - Python 3 source code is assumed to be UTF-8 by default. This means that you don’t need # -*- coding: UTF-8 -*- at the top of .py files in Python 3. All text (str) is Unicode by default. Encoded Unicode text is represented as binary data (bytes).
Find elsewhere
Top answer
1 of 3
7

You are confusing Unicode with encodings. An encoding is a standard that represents text as within the confines of individual values in the range of 0-255 (bytes), while Unicode is a standard that describes codepoints representing textual glyphs. The two are related but not the same thing.

The Unicode standard includes several encodings. UTF-16 is one such encoding that uses 2 bytes per codepoint, but it is not the only encoding included in the standard. UTF-8 is another such encoding, and it uses a variable number of bytes per codepoint.

Your file, however, is written using ASCII, the default codec used by Python 2 when you do not specify an explicit encoding. If you expected to see 2 bytes per codepoint, encode to UTF-16 explicitly:

fin.write(u'\x40'.encode('utf16-le')

This writes UTF-16 in little endian byte order; there is also a utf16-be codec. Normally, for multi-byte encodings like UTF-16 or UTF32, you'd also include a BOM, or Byte Order Mark; it is included automatically when you write UTF-16 without picking any endianes.

fin.write(u'\x40'.encode('utf16')

I strongly urge you to study up on Unicode, codecs and Python before you continue:

  • The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!) by Joel Spolsky

  • The Python Unicode HOWTO

  • Pragmatic Unicode by Ned Batchelder

2 of 3
1
  • Character numbers from U+0000 to U+007F (US-ASCII repertoire) correspond to octets 00 to 7F (7 bit US-ASCII values). A direct consequence is that a plain ASCII string is also a valid UTF-8 string.
  • UTF-8, a transformation format of ISO 10646
🌐
Note.nkmk.me
note.nkmk.me › home › python
Convert Between Unicode Code Point and Character: chr, ord | note.nkmk.me
January 31, 2024 - When converting a hexadecimal string to an integer using int(), the second argument can be 0 if the string is prefixed with 0x. For more details on handling hexadecimal numbers and strings, refer to the following article. Convert binary, octal, decimal, and hexadecimal in Python · Unicode code points are often represented as U+XXXX.
🌐
Gitbooks
hacktec.gitbooks.io › effective-python › content › en › Chapter1 › Item3.html
Item 3: Know the Differences Between bytes, str, and unicode · Effective Python
The most common encoding is UTF-8. Importantly, str instances in Python 3 and unicode instances in Python 2 do not have an associated binary encoding. To convert Unicode characters to binary data, you must use the encode method.
🌐
Educative
educative.io › answers › how-to-convert-string-to-binary-in-python
How to convert string to binary in Python
Strings are an array of Unicode code characters. Binary is a base-2 number system consisting of 0’s and 1’s that computers understand. The computer sees strings in binary format, i.e., ‘H’=1001000. The string, as seen by the computer, is a binary number that is an ASCCI value( Decimal number) of the string converted to binary.
🌐
Quora
quora.com › How-do-you-convert-a-string-to-binary-in-Python
How to convert a string to binary in Python - Quora
Answer (1 of 2): You can convert a string to binary by converting each character to its Unicode representation, then converting it to binary. Let us take an example: [code]>>> s = "hello" [/code]Understanding bin(), ord(): * bin() is an in-built function in Python that takes in integer x and ...
🌐
Sticky Bits
blog.feabhas.com › home › python 3 unicode and byte strings
Python 3 Unicode and Byte Strings - Sticky Bits - Powered by FeabhasSticky Bits – Powered by Feabhas
February 21, 2019 - A notable difference between Python 2 and Python 3 is that character data is stored using Unicode instead of bytes. It is quite likely that when migrating existing code and writing new code you may be unaware of this change as most string algorithms will work with either type of representation; but you cannot intermix the two. If you are working with web service libraries such as urllib (formerly urllib2) and requests, network sockets, binary files, or serial I/O with pySerial you will find that data is now stored as byte strings.
🌐
Emacsos
blog.emacsos.com › unicode-in-python.html
Handling Unicode Strings in Python - Yuanle's blog
August 25, 2016 - Once you need to do IO, you need a binary representation of the string. Typical IO includes reading from and writing to console, files, and network sockets. Unicode string literal, byte literal and their types are different in python 2 and python 3, as shown in the following table.
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-unicode-to-bytes-in-python
Convert Unicode to Bytes in Python - GeeksforGeeks
July 23, 2025 - In this example, the Unicode string is encoded into a byte string using the UTF-16 encoding with the `encode()` method. The resulting `byte_string_utf16` contains the UTF-16 encoded representation of the original string, which is then printed to the console.
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-unicode-string-to-a-string-in-python
Convert Unicode String to a Byte String in Python - GeeksforGeeks
July 23, 2025 - In this example, a Unicode string containing English and Chinese characters is encoded to a byte string using UTF-8 encoding. The resulting `bytes_representation` is printed, demonstrating the transformation of the mixed-language Unicode string into its byte representation suitable for storage or transmission in a UTF-8 encoded format. ... unicode_string = "Hello, 你好" bytes_representation = unicode_string.encode('utf-8') print(bytes_representation) ... In this example, the Unicode string is encoded into a byte string using UTF-16 encoding, resulting in a sequence of bytes that represents the mixed-language string.