Cleaner version:

>>> test_string = '1101100110000110110110011000001011011000101001111101100010101000'
>>> print ('%x' % int(test_string, 2)).decode('hex').decode('utf-8')
نقاب

Inverse (from @Robᵩ's comment):

>>> '{:b}'.format(int(u'نقاب'.encode('utf-8').encode('hex'), 16))
1: '1101100110000110110110011000001011011000101001111101100010101000'
Answer from Igonato on Stack Overflow
🌐
Online Tools
onlinetools.com › utf8 › convert-binary-to-utf8
Convert Binary Bits to UTF8 – Online UTF8 Tools
World's simplest browser-based binary to UTF8 converter. Just import your binary bits in the editor on the left and you will instantly get UTF8 strings on the right. Free, quick, and very powerful. Import bits – get UTF8.
Discussions

command line - Convert binary encoding that head and Notepad can read to UTF-8 - Unix & Linux Stack Exchange
I have a CSV file which is in binary character set but I have to convert to UTF-8 to process in HDFS (Hadoop). I have used the below command to check characterset. file -bi filename.csv Output : More on unix.stackexchange.com
🌐 unix.stackexchange.com
August 18, 2016
utf 8 - How to convert binary to utf-8 and vice-versa in R? - Stack Overflow
I want to code a shinyApp which can convert a string from binary (0 and 1, not hex) to utf-8 and vice-versa. More on stackoverflow.com
🌐 stackoverflow.com
utf 8 - python - convert binary data to utf-8 - Stack Overflow
import base64 encoded = base64.b64encode(image_binary_data) ... Encoding means converting strings to storable bytes. And Decoding means converting bytes to readable strings. The data in your code is already encoded. ... Image cannot be converted into something like charters in utf8. ... Find the answer to your question by asking. Ask question ... See similar questions with these tags. ... 1 How to convert a string representation of 'bytes' to its corresponding 'utf-8... More on stackoverflow.com
🌐 stackoverflow.com
Convert binary data ASCII to UTF-8 | OutSystems
I have a problem with the BinaryData API · In the ConvertEncoding method More on outsystems.com
🌐 outsystems.com
People also ask

What does this binary to UTF-8 converter do?
It converts binary bit strings into readable UTF-8 text by grouping 8 bits into bytes and decoding them as UTF-8 characters.
🌐
easyprotools.com
easyprotools.com › home › utf-8 tools › convert binary bits to utf-8
Convert Binary Bits to UTF-8 - Free Binary Decoder
Can I convert UTF-8 text back to binary?
Yes. Click the UTF-8 to Bin button for instant reverse conversion.
🌐
easyprotools.com
easyprotools.com › home › utf-8 tools › convert binary bits to utf-8
Convert Binary Bits to UTF-8 - Free Binary Decoder
Does it handle multi-byte UTF-8 characters?
Yes. Supports 1-byte ASCII through 4-byte emoji and supplementary Unicode characters.
🌐
easyprotools.com
easyprotools.com › home › utf-8 tools › convert binary bits to utf-8
Convert Binary Bits to UTF-8 - Free Binary Decoder
🌐
Stack Overflow
stackoverflow.com › questions › 41909888 › how-to-convert-a-binary-string-to-utf-8-string-in-java
file - How to convert a Binary String to UTF-8 String in java? - Stack Overflow
I think this is what you needed docs.oracle.com/javase/7/docs/api/java/lang/… and in reverse process you can simply create a new String instance from bytes. ... @SabirKhan i'm sorry but it's not a duplicate, i can't afford to write extra bits to make every character 16 bits binary value. ... Exactly where did you see that UTF-8 requires 16 bits per character?
🌐
AskPython
askpython.com › python › examples › binary-to-utf8-conversion
How to Convert Binary Data to UTF-8 in Python - AskPython
April 10, 2025 - Base64 produces UTF-8 compatible text output from any binary input. import base64 data = b"\x12\xab" # Base64 encode b64_data = base64.b64encode(data) # Now decode base64 text from UTF-8 to string text = b64_data.decode("utf-8") print(text) ...
🌐
DNS Checker
dnschecker.org › binary-to-text-translator.php
Binary to Text Converter | Translate Binary Code Online
Enter or paste your binary string into our tool and get the translation in seconds. ... Your rating helps us improve DNSChecker. ... Thanks for your feedback! ... RAID Calculator HTTP Headers Check Check Website Operating System MD5 & Base64 Generator Multi URL Opener · Our binary code translator is used to convert binary code into text. There are different encoding systems that you can select to convert the binary code to, such as Unicode, UTF-8, UTF-16, Windows 1252, and so on.
Top answer
1 of 4
18

"binary" isn't an encoding (character-set name). iconv needs an encoding name to do its job.

The file utility doesn't give useful information when it doesn't recognize the file format. It could be UTF-16 for example, without a byte-encoding-mark (BOM). notepad reads that. The same applies to UTF-8 (and head would display that since your terminal may be set to UTF-8 encoding, and it would not care about a BOM).

If the file is UTF-16, your terminal would display that using head because most of the characters would be ASCII (or even Latin-1), making the "other" byte of the UTF-16 characters a null.

In either case, the lack of BOM will (depending on the version of file) confuse it. But other programs may work, because these file formats can be used with Microsoft Windows as well as portable applications that may run on Windows.

To convert the file to UTF-8, you have to know which encoding it uses, and what the name for that encoding is with iconv. If it is already UTF-8, then whether you add a BOM (at the beginning) is optional. UTF-16 has two flavors, according to which byte is first. Or you could even have UTF-32. iconv -l lists these:

ISO-10646/UTF-8/
ISO-10646/UTF8/
UTF-7//
UTF-8//
UTF-16//
UTF-16BE//
UTF-16LE//
UTF-32//
UTF-32BE//
UTF-32LE//
UTF7//
UTF8//
UTF16//
UTF16BE//
UTF16LE//
UTF32//
UTF32BE//
UTF32LE//

"LE" and "BE" refer to little-end and big-end for the byte-order. Windows uses the "LE" flavors, and iconv likely assumes that for the flavors lacking "LE" or "BE".

You can see this using an octal (sic) dump:

$ od -bc big-end
0000000 000 124 000 150 000 165 000 040 000 101 000 165 000 147 000 040
         \0   T  \0   h  \0   u  \0      \0   A  \0   u  \0   g  \0    
0000020 000 061 000 070 000 040 000 060 000 065 000 072 000 060 000 061
         \0   1  \0   8  \0      \0   0  \0   5  \0   :  \0   0  \0   1
0000040 000 072 000 065 000 067 000 040 000 105 000 104 000 124 000 040
         \0   :  \0   5  \0   7  \0      \0   E  \0   D  \0   T  \0    
0000060 000 062 000 060 000 061 000 066 000 012
         \0   2  \0   0  \0   1  \0   6  \0  \n
0000072

$ od -bc little-end
0000000 124 000 150 000 165 000 040 000 101 000 165 000 147 000 040 000
          T  \0   h  \0   u  \0      \0   A  \0   u  \0   g  \0      \0
0000020 061 000 070 000 040 000 060 000 065 000 072 000 060 000 061 000
          1  \0   8  \0      \0   0  \0   5  \0   :  \0   0  \0   1  \0
0000040 072 000 065 000 067 000 040 000 105 000 104 000 124 000 040 000
          :  \0   5  \0   7  \0      \0   E  \0   D  \0   T  \0      \0
0000060 062 000 060 000 061 000 066 000 012 000
          2  \0   0  \0   1  \0   6  \0  \n  \0
0000072

Assuming UTF-16LE, you could convert using

iconv -f UTF-16LE// -t UTF-8// <input >output
2 of 4
3

strings (from binutils) succeeds to "print the strings of printable characters in files" when both iconv and recode failed as well, with file still reporting the content as binary data:

$ file -i /tmp/textFile
/tmp/textFile: application/octet-stream; charset=binary

$ chardetect /tmp/textFile
/tmp/textFile: utf-8 with confidence 0.99

$ iconv -f utf-8 -t utf-8 /tmp/textFile -o /tmp/textFile.iconv
$ file -i /tmp/textFile.iconv
/tmp/textFile.iconv: application/octet-stream; charset=binary

$ cp /tmp/textFile /tmp/textFile.recode ; recode utf-8 /tmp/textFile.recode
$ file -i /tmp/textFile.recode 
/tmp/textFile.recode: application/octet-stream; charset=binary

$ strings /tmp/textFile > /tmp/textFile.strings
$ file -i /tmp/textFile.strings
/tmp/textFile.strings: text/plain; charset=us-ascii
Find elsewhere
🌐
Online Tools
onlinetools.com › binary › convert-binary-to-utf8
Convert Binary to UTF8 – Online Binary Tools
Free online binary to UTF8 converter. Just load your binary numbers and they will automatically get converted to UTF8 characters. There are no ads, popups or nonsense, just an awesome binary digits to UTF8 symbols converter. Load binary, get UTF8.
🌐
TheToolApp
thetoolapp.com › utilities › convert-binary-to-utf8
Change Binary To Utf8 Online
March 15, 2026 - Paste binary bytes and get back the UTF-8 text they represent — plain ASCII, accented letters, CJK characters, even 4-byte emoji. Feed it a continuous string or space-separated 8-bit groups; the decoder handles both and flags malformed sequences with the exact byte position.
🌐
Easyprotools
easyprotools.com › home › utf-8 tools › convert binary bits to utf-8
Convert Binary Bits to UTF-8 - Free Binary Decoder
Our UTF-8 binary decoder online handles all of these cases correctly. When you enter the binary string 01001000 01100101 01101100, the tool groups each space-separated 8-bit sequence, converts each to its byte value, and passes the resulting byte array through a UTF-8 decoder to produce "Hel".
Top answer
1 of 1
5

You don't need a loop in your encode function

All the functions you're using are already vectorised so you can just as easily do:

encode2 <- function(message) {
    charToRaw(message) |>
        rawToBits() |>
        as.integer() |>
        rev()
}

x <- "some random string"
all(encode(x) == encode2(x))
# [1] TRUE

Build in handling of other locales

It's probably safer to use enc2utf8(), in case you're running in a different locale. As you explain in the comments you need to rev(), I would suggest something like:

encode3 <- function(message) {
    message |>
        enc2utf8() |>
        charToRaw() |>
        rawToBits() |>
        as.integer() |>
        rev()
}

message <- "café naïve — ∑ ∆ ∫ √ π λ 東京 😀 नमस्ते ✔ שלום ♥ ♦ ♣ ♠"
encoded <- encode3(message)
encoded
#  [1] 1 0 1 0 0 0 0 0 1 0 0 1 1 0 0 1 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 1 0 0 0 1 1 1 0 0 1 1 0 0 1 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 1 0 0 1 1 0 1 0 0 1 1 0 0 1 1 1 [83] 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 1 0 0 1 0 1 1 0 0 1 1 0 0 1 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 0 1 1 1 0 1 1 1 0 1 0 1 1 1 1 0 0 1 0 1 0 1 1 1 0 1 0 1 1 1 1 0 0 1 [165] 1 1 0 0 1 1 0 1 0 1 1 1 1 0 1 0 1 0 0 1 1 1 0 1 0 1 1 1 0 0 1 0 0 0 0 0 1 0 0 1 0 1 0 0 1 0 0 1 1 1 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 1 1 1 0 1 0 0 1 [247] 0 1 1 1 1 0 0 0 0 0 1 0 1 0 0 1 0 0 1 0 1 0 0 1 0 0 1 1 1 0 0 0 0 0 1 0 0 0 1 1 0 1 1 0 1 0 0 1 0 1 1 1 1 0 0 0 0 0 1 0 1 1 1 0 0 0 1 0 1 0 0 1 0 0 1 1 1 0 0 0 0 0 [329] 1 0 1 0 1 1 1 0 1 0 1 0 0 1 0 0 1 1 1 0 0 0 0 0 1 0 1 0 1 0 0 0 1 0 1 0 0 1 0 0 1 1 1 0 0 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 0 0 1 0 0 1 1 0 0 0 1 0 0 1 1 1 1 1 1 1 [411] 1 1 0 0 0 0 0 0 1 0 0 0 0 0 1 0 1 0 1 1 0 0 1 0 1 1 1 0 1 0 1 1 1 0 0 1 0 0 1 0 1 1 0 0 0 1 1 0 0 1 1 1 0 1 1 1 1 0 0 1 1 0 0 0 1 0 0 0 0 0 1 0 1 1 1 0 1 1 1 1 0 0 [493] 1 1 1 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 0 0 1 1 0 0 1 1 1 1 0 0 1 0 0 0 0 0 1 0 0 1 1 0 1 0 1 0 0 0 1 0 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 1 0 1 0 1 1 1 0 0 0 1 0 [575] 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 1 0 1 0 0 0 1 0 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 1 0 0 1 0 0 0 1 1 0 0 0 1 0 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 [657] 1 0 0 1 0 1 0 0 1 0 0 0 0 0 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 0 0 1 1 0 0 1 0 1 0 1 1 1 0 1 1 0 1 0 1 0 1 1 1 1 1 1 0 0 0 0 1 1 0 1 1 0 0 0 0 1 0 1 1 0 1 1 1 0 0 0 [739] 1 0 0 0 0 0 1 0 1 0 1 0 0 1 1 1 0 0 0 0 1 1 0 1 1 0 0 1 1 0 0 1 1 0 0 0 0 1 0 1 1 0 0 0 1 1

Decoding requires adding the structure that you stripped

The key decoding step is that you lose byte structure after your encode() function converts from raw to integer. Let's look at an example with "A" (65 in ASCII or \u0041). You can see that as.integer() gives us a vector of bits:

"A" |>
    charToRaw() |>  # 41
    rawToBits() |> # 01 00 00 00 00 00 01 00
    as.integer() # 1 0 0 0 0 0 1 0

Each of the eight bits that made up one byte are flattened into separate integers, so you can't reconstruct the original raw value without re-grouping them into bytes. You can use packBits() after as.raw() to get these back:

decode <- function(message) {
    message |>
        rev() |>
        as.raw() |>
        packBits(type = "raw") |>
        rawToChar()
}

This allows us to decode() the message by reversing the order of encoding operations:

decode(encoded)
# [1] "café naïve — ∑ ∆ ∫ √ π λ 東京 😀 नमस्ते ✔ שלום ♥ ♦ ♣ ♠"

decode(encoded) == message
# [1] TRUE
🌐
Wikipedia
en.wikipedia.org › wiki › UTF-8
UTF-8 - Wikipedia
2 days ago - UTF-8 supports all 1,112,064 valid Unicode code points using a variable-width encoding of one to four one-byte (8-bit) code units. Code points with lower numerical values, which tend to occur more frequently, are encoded using fewer bytes. It was designed for backward compatibility with ASCII: the first 128 characters of Unicode, which correspond one-to-one with ASCII, are encoded using a single byte with the same binary ...
🌐
HubSpot
blog.hubspot.com › home › website › what is utf-8 encoding? a walkthrough for non-programmers
What is UTF-8 encoding? A walkthrough for non-programmers
November 20, 2025 - UTF-8 is an encoding system for Unicode. UTF-8 stands for “Unicode Transformation Format - 8 bits.” It can translate any Unicode character to a matching unique binary string, and can also translate the binary string back to a Unicode character.
🌐
PlanetCalc
planetcalc.com › 9033
Online calculator: UTF-8 encoded string
The calculator converts an input string to a UTF-8 encoded binary/decimal/hexadecimal dump and vice versa.
🌐
Hixie
software.hixie.ch › utilities › cgi › unicode-decoder › utf8-decoder
utf8-decoder
Space-separated list of binary numbers, e.g. 01010101 10101010 11110000 ... Raw UTF-8 encoded text, but interpreted as Windows-1252.
🌐
RapidTables
rapidtables.com › convert › number › binary-to-ascii.html
Binary to Text Translator
Enter binary numbers with any prefix/postfix/delimiter and press the Convert button. (E.g: 01000101 01111000 01100001 01101101 01110000 01101100 01100101): * ASCII text encoding uses fixed 1 byte for each character. ** UTF-8 text encoding uses a variable number of bytes for each character.
🌐
AutoIt
autoitscript.com › autoit3 › docs › functions › StringToBinary.htm
Function StringToBinary
$dBinary = StringToBinary($sString, $SB_UTF16BE) ; Convert the UTF16-BE binary string back into a string. $sConverted = BinaryToString($dBinary, $SB_UTF16BE) ; Display the resulsts. DisplayResults($sString, $dBinary, $sConverted, "UTF16-BE") ; Convert the original UTF-8 string to an UTF-8 binary ...
🌐
GitHub
github.com › nodejs › node › issues › 41379
Converting binary data to UTF-8 changes the data · Issue #3674 · nodejs/help
January 2, 2022 - let buffer = fs.readFileSync("file.pdf") let byteLength = Buffer.byteLength(buffer) // **Valid pdf, Length: 1440177** let utf8String = buffer.toString("utf8") let bufferAgain = Buffer.from(utf8String,"utf8") // **Bad pdf, Length: 2551916* buffer.equals(bufferAgain) // **Gives false** ... Buffer of PDF file must be converted into string utf8, then converted again in the same inital buffer
Author: nodejs