How about this:

>>> ''.join('{:08b}'.format(b) for b in 'سلام'.encode('utf8'))
'1101100010110011110110011000010011011000101001111101100110000101'

This iterates over the encoded bytes object, where you get an integer in the range 0..255 for each iteration. Then the integer is formatted in binary notation with zero padding up to 8 digits. Then glue everything together with str.join().

For the inverse, the approach given in an answer from the question you linked to can be adapted to Python 3 as follows (s is the output of the above example, ie. a str of 0s and 1s):

>>> import re
>>> bytes(int(b, 2) for b in re.split('(........)', s) if b).decode('utf8')
'سلام'
Answer from lenz on Stack Overflow
Discussions

string - Convert Python str/unicode object to binary/hex blob - Stack Overflow
Is there an easy way to get some str/unicode object represented as a big binary number (or an hex one)? I've been reading some answers to related questions but none of them works for my scenario. I More on stackoverflow.com
🌐 stackoverflow.com
python - Write a unicode character to a file in a binary way - Stack Overflow
The Unicode standard includes several encodings. UTF-16 is one such encoding that uses 2 bytes per codepoint, but it is not the only encoding included in the standard. UTF-8 is another such encoding, and it uses a variable number of bytes per codepoint. Your file, however, is written using ASCII, the default codec used by Python 2 when you do not specify an explicit encoding. If you expected to ... More on stackoverflow.com
🌐 stackoverflow.com
August 17, 2013
string - converting binary to utf-8 in python - Stack Overflow
Hmmm, I'm somewhat suspicious of unichr. Because OP says his binary is already utf-8. utf-8 has variable character length, so I just used chr to join the raw bytes in a string and decode them later into Unicode. 2013-10-08T19:09:16.187Z+00:00 ... @JoranBeasley - I disagree, assuming Python2. More on stackoverflow.com
🌐 stackoverflow.com
python - Binary Data To Unicode - Stack Overflow
It could be bytes from which Unicode can be decoded. 2011-02-23T18:09:06.83Z+00:00 ... @S.Lott: If the extraction process is just using the whole thing as-is then so be it. But I stand by my answer. 2011-02-23T18:14:08.38Z+00:00 ... You should stand by your answer. However, you could also consider extending it to cover the most common case of getting binary data from a file. 2011-02-23T18:45:42.163Z+00:00 ... Your answer makes complete sense to me. I'm in a Python ... More on stackoverflow.com
🌐 stackoverflow.com
🌐
Learning Python
learning-python.com › strings30.html
Strings in 3.X: Unicode and Binary Data
At a more concrete level, the Python language provides multiple string data types to represent content in your script: both textual data—integer code-point values of decoded Unicode characters in memory, as well as binary data—raw byte values, including text that is in encoded form.
🌐
Python
docs.python.org › 3 › library › binascii.html
binascii — Convert between binary and ASCII
Changed in version 3.3: ASCII-only unicode strings are now accepted by the a2b_* functions. The binascii module defines the following functions: ... Convert a single line of uuencoded data back to binary and return the binary data.
🌐
Online Tools
onlinetools.com › unicode › convert-unicode-to-binary
Convert Unicode to Binary – Online Unicode Tools
It supports the most popular Unicode encodings (such as UTF-8, UTF-16, UTF-32, UCS-2, and UCS-4) and it works with emoji characters. You can also customize the binary output format by enabling binary padding and spacing. Created by encoding gurus from team Browserling. ... Can't convert. ... Chain with... ... This tool cannot be chained.
🌐
Delft Stack
delftstack.com › home › howto › python › string to binary python
How to Convert a String to Binary in Python | Delft Stack
February 2, 2024 - This article will discuss some methods to convert a string to its binary representation in Python. We use the ord() function that translates the Unicode point of the string to a corresponding integer.
🌐
Real Python
realpython.com › python-encodings-guide
Unicode & Character Encodings in Python: A Painless Guide – Real Python
May 20, 2019 - Python 3 source code is assumed to be UTF-8 by default. This means that you don’t need # -*- coding: UTF-8 -*- at the top of .py files in Python 3. All text (str) is Unicode by default. Encoded Unicode text is represented as binary data (bytes).
Find elsewhere
🌐
Note.nkmk.me
note.nkmk.me › home › python
Convert Between Unicode Code Point and Character: chr, ord | note.nkmk.me
January 31, 2024 - When converting a hexadecimal string to an integer using int(), the second argument can be 0 if the string is prefixed with 0x. For more details on handling hexadecimal numbers and strings, refer to the following article. Convert binary, octal, decimal, and hexadecimal in Python · Unicode code points are often represented as U+XXXX.
Top answer
1 of 3
7

You are confusing Unicode with encodings. An encoding is a standard that represents text as within the confines of individual values in the range of 0-255 (bytes), while Unicode is a standard that describes codepoints representing textual glyphs. The two are related but not the same thing.

The Unicode standard includes several encodings. UTF-16 is one such encoding that uses 2 bytes per codepoint, but it is not the only encoding included in the standard. UTF-8 is another such encoding, and it uses a variable number of bytes per codepoint.

Your file, however, is written using ASCII, the default codec used by Python 2 when you do not specify an explicit encoding. If you expected to see 2 bytes per codepoint, encode to UTF-16 explicitly:

fin.write(u'\x40'.encode('utf16-le')

This writes UTF-16 in little endian byte order; there is also a utf16-be codec. Normally, for multi-byte encodings like UTF-16 or UTF32, you'd also include a BOM, or Byte Order Mark; it is included automatically when you write UTF-16 without picking any endianes.

fin.write(u'\x40'.encode('utf16')

I strongly urge you to study up on Unicode, codecs and Python before you continue:

  • The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!) by Joel Spolsky

  • The Python Unicode HOWTO

  • Pragmatic Unicode by Ned Batchelder

2 of 3
1
  • Character numbers from U+0000 to U+007F (US-ASCII repertoire) correspond to octets 00 to 7F (7 bit US-ASCII values). A direct consequence is that a plain ASCII string is also a valid UTF-8 string.
  • UTF-8, a transformation format of ISO 10646
🌐
Gitbooks
hacktec.gitbooks.io › effective-python › content › en › Chapter1 › Item3.html
Item 3: Know the Differences Between bytes, str, and unicode · Effective Python
The most common encoding is UTF-8. Importantly, str instances in Python 3 and unicode instances in Python 2 do not have an associated binary encoding. To convert Unicode characters to binary data, you must use the encode method.
🌐
Educative
educative.io › answers › how-to-convert-string-to-binary-in-python
How to convert string to binary in Python
Strings are an array of Unicode code characters. Binary is a base-2 number system consisting of 0’s and 1’s that computers understand. The computer sees strings in binary format, i.e., ‘H’=1001000. The string, as seen by the computer, is a binary number that is an ASCCI value( Decimal number) of the string converted to binary.
🌐
Quora
quora.com › How-do-you-convert-a-string-to-binary-in-Python
How to convert a string to binary in Python - Quora
Answer (1 of 2): You can convert a string to binary by converting each character to its Unicode representation, then converting it to binary. Let us take an example: [code]>>> s = "hello" [/code]Understanding bin(), ord(): * bin() is an in-built function in Python that takes in integer x and ...
🌐
Sticky Bits
blog.feabhas.com › home › python 3 unicode and byte strings
Python 3 Unicode and Byte Strings - Sticky Bits - Powered by FeabhasSticky Bits – Powered by Feabhas
February 21, 2019 - A notable difference between Python 2 and Python 3 is that character data is stored using Unicode instead of bytes. It is quite likely that when migrating existing code and writing new code you may be unaware of this change as most string algorithms will work with either type of representation; but you cannot intermix the two. If you are working with web service libraries such as urllib (formerly urllib2) and requests, network sockets, binary files, or serial I/O with pySerial you will find that data is now stored as byte strings.
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-unicode-to-bytes-in-python
Convert Unicode to Bytes in Python - GeeksforGeeks
July 23, 2025 - In this example, the Unicode string is encoded into a byte string using the UTF-16 encoding with the `encode()` method. The resulting `byte_string_utf16` contains the UTF-16 encoded representation of the original string, which is then printed to the console.
🌐
Emacsos
blog.emacsos.com › unicode-in-python.html
Handling Unicode Strings in Python - Yuanle's blog
August 25, 2016 - Once you need to do IO, you need a binary representation of the string. Typical IO includes reading from and writing to console, files, and network sockets. Unicode string literal, byte literal and their types are different in python 2 and python 3, as shown in the following table.
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
These code points will then turn back into the same bytes when the surrogateescape error handler is used to encode the data and write it back out. One section of Mastering Python 3 Input/Output, a PyCon 2010 talk by David Beazley, discusses text processing and binary data handling. The PDF slides for Marc-André Lemburg’s presentation “Writing Unicode-aware Applications in Python” discuss questions of character encodings as well as how to internationalize and localize an application.
Top answer
1 of 10
174

Something like this?

>>> st = "hello world"
>>> ' '.join(format(ord(x), 'b') for x in st)
'1101000 1100101 1101100 1101100 1101111 100000 1110111 1101111 1110010 1101100 1100100'

#using `bytearray`
>>> ' '.join(format(x, 'b') for x in bytearray(st, 'utf-8'))
'1101000 1100101 1101100 1101100 1101111 100000 1110111 1101111 1110010 1101100 1100100'
2 of 10
133

If by binary you mean bytes type, you can just use encode method of the string object that encodes your string as a bytes object using the passed encoding type. You just need to make sure you pass a proper encoding to encode function.

In [9]: "hello world".encode('ascii')                                                                                                                                                                       
Out[9]: b'hello world'

In [10]: byte_obj = "hello world".encode('ascii')                                                                                                                                                           

In [11]: byte_obj                                                                                                                                                                                           
Out[11]: b'hello world'

In [12]: byte_obj[0]                                                                                                                                                                                        
Out[12]: 104

Otherwise, if you want them in form of zeros and ones --binary representation-- as a more pythonic way you can first convert your string to byte array then use bin function within map :

>>> st = "hello world"
>>> map(bin,bytearray(st))
['0b1101000', '0b1100101', '0b1101100', '0b1101100', '0b1101111', '0b100000', '0b1110111', '0b1101111', '0b1110010', '0b1101100', '0b1100100']
 

Or you can join it:

>>> ' '.join(map(bin,bytearray(st)))
'0b1101000 0b1100101 0b1101100 0b1101100 0b1101111 0b100000 0b1110111 0b1101111 0b1110010 0b1101100 0b1100100'

Note that in python3 you need to specify an encoding for bytearray function :

>>> ' '.join(map(bin,bytearray(st,'utf8')))
'0b1101000 0b1100101 0b1101100 0b1101100 0b1101111 0b100000 0b1110111 0b1101111 0b1110010 0b1101100 0b1100100'

You can also use binascii module in python 2:

>>> import binascii
>>> bin(int(binascii.hexlify(st),16))
'0b110100001100101011011000110110001101111001000000111011101101111011100100110110001100100'

hexlify return the hexadecimal representation of the binary data then you can convert to int by specifying 16 as its base then convert it to binary with bin.