The logic of encoding Unicode in UTF-8 is basically:

  • Up to 4 bytes per character can be used. The fewest number of bytes possible is used.
  • Characters up to U+007F are encoded with a single byte.
  • For multibyte sequences, the number of leading 1 bits in the first byte gives the number of bytes for the character. The rest of the bits of the first byte can be used to encode bits of the character.
  • The continuation bytes begin with 10, and the other 6 bits encode bits of the character.

Here's a function I wrote a while back for encoding a JavaScript UTF-16 string in UTF-8:

function toUTF8Array(str) {
    var utf8 = [];
    for (var i=0; i < str.length; i++) {
        var charcode = str.charCodeAt(i);
        if (charcode < 0x80) utf8.push(charcode);
        else if (charcode < 0x800) {
            utf8.push(0xc0 | (charcode >> 6), 
                      0x80 | (charcode & 0x3f));
        }
        else if (charcode < 0xd800 || charcode >= 0xe000) {
            utf8.push(0xe0 | (charcode >> 12), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
        // surrogate pair
        else {
            i++;
            // UTF-16 encodes 0x10000-0x10FFFF by
            // subtracting 0x10000 and splitting the
            // 20 bits of 0x0-0xFFFFF into two halves
            charcode = 0x10000 + (((charcode & 0x3ff)<<10)
                      | (str.charCodeAt(i) & 0x3ff));
            utf8.push(0xf0 | (charcode >>18), 
                      0x80 | ((charcode>>12) & 0x3f), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
    }
    return utf8;
}
Answer from Joni on Stack Overflow
Top answer
1 of 10
84

The logic of encoding Unicode in UTF-8 is basically:

  • Up to 4 bytes per character can be used. The fewest number of bytes possible is used.
  • Characters up to U+007F are encoded with a single byte.
  • For multibyte sequences, the number of leading 1 bits in the first byte gives the number of bytes for the character. The rest of the bits of the first byte can be used to encode bits of the character.
  • The continuation bytes begin with 10, and the other 6 bits encode bits of the character.

Here's a function I wrote a while back for encoding a JavaScript UTF-16 string in UTF-8:

function toUTF8Array(str) {
    var utf8 = [];
    for (var i=0; i < str.length; i++) {
        var charcode = str.charCodeAt(i);
        if (charcode < 0x80) utf8.push(charcode);
        else if (charcode < 0x800) {
            utf8.push(0xc0 | (charcode >> 6), 
                      0x80 | (charcode & 0x3f));
        }
        else if (charcode < 0xd800 || charcode >= 0xe000) {
            utf8.push(0xe0 | (charcode >> 12), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
        // surrogate pair
        else {
            i++;
            // UTF-16 encodes 0x10000-0x10FFFF by
            // subtracting 0x10000 and splitting the
            // 20 bits of 0x0-0xFFFFF into two halves
            charcode = 0x10000 + (((charcode & 0x3ff)<<10)
                      | (str.charCodeAt(i) & 0x3ff));
            utf8.push(0xf0 | (charcode >>18), 
                      0x80 | ((charcode>>12) & 0x3f), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
    }
    return utf8;
}
2 of 10
48

JavaScript Strings are stored in UTF-16. To get UTF-8, you'll have to convert the String yourself.

One way is to mix encodeURIComponent(), which will output UTF-8 bytes URL-encoded, with unescape, as mentioned on ecmanaut.

var utf8 = unescape(encodeURIComponent(str));

var arr = [];
for (var i = 0; i < utf8.length; i++) {
    arr.push(utf8.charCodeAt(i));
}
🌐
MDN Web Docs
developer.mozilla.org › en-US › docs › Web › API › TextEncoder › encodeInto
TextEncoder: encodeInto() method - Web APIs | MDN
June 28, 2025 - The ratio for these characters is 2, because they take 4 bytes in UTF-8 and 2 in UTF-16. If the output allocation (typically within Wasm heap) is expected to be short-lived, it makes sense to allocate s.length * 3 bytes for the output, in which case the first conversion attempt is guaranteed ...
🌐
Dirask
dirask.com › posts › JavaScript-convert-string-to-bytes-array-UTF-8-1XkbEj
JavaScript - convert string to bytes array (UTF-8)
// ONLINE-RUNNER:browser; const ... result.push(character.charCodeAt(0)); } } return result; }; // Usage example: const bytes = toBytes('Some text here...'); // converts string to UTF-8 bytes console.log(bytes); // [83, 111, 109, 101, ...
🌐
Burke
kevin.burke.dev › kevin › node-js-string-encoding
Let’s talk about Javascript string encoding | Kevin Burke
September 1, 2017 - Where the cent character is the UTF-8 encoded byte sequence "\xc2\xa2". When Node starts and you try to reference x in your program, it will be re-encoded as a UTF-16 string. If you type the literal characters: ... This will be turned into the UTF-16 string "\xc2\x00\xa2\x00". So be careful to mind your inputs and outputs. Encoding in Node is extremely confusing, and difficult to get right. It helps, though, when you realize that Javascript string types will always be encoded as UTF-16, and most of the other places strings in RAM interact with sockets, files, or byte arrays, the string gets re-encoded as UTF-8.
🌐
GitHub
gist.github.com › vbabak › c47a67ff5e89cab8954097eeb2a33fb8
Convert JavaScript utf-8 string to bytes array. · GitHub
Convert JavaScript utf-8 string to bytes array. GitHub Gist: instantly share code, notes, and snippets.
🌐
Designcise
designcise.com › web › tutorial › how-to-convert-a-javascript-string-to-a-byte-array
How to Convert a JavaScript String to a Byte Array? - Designcise
April 9, 2023 - In JavaScript, you can convert ... In this example, the TextEncoder.encode() method takes a string as input and returns a Uint8Array object containing the UTF-8 encoded bytes of the string....
🌐
GeeksforGeeks
geeksforgeeks.org › javascript-program-to-convert-string-to-bytes
JavaScript Program to Convert String to Bytes | GeeksforGeeks
June 14, 2024 - This approach is particularly useful for server-side JavaScript running in a Node.js environment. Example: In this example, we use the Buffer.from() method to convert a string into bytes using UTF-8 encoding.
🌐
GitHub
gist.github.com › lihnux › 2aa4a6f5a9170974f6aa
Javascript Convert String to Byte Array · GitHub
function unpack(str) { var bytes = []; for(var i = 0; i < str.length; i++) { var char = str.charCodeAt(i); bytes.push(char >>> 8); bytes.push(char & 0xFF); } return bytes; } ... function toUTF8Array(str) { let utf8 = []; for (let i = 0; i < str.length; i++) { let charcode = str.charCodeAt(i); if (charcode < 0x80) utf8.push(charcode); else if (charcode < 0x800) { utf8.push(0xc0 | (charcode >> 6), 0x80 | (charcode & 0x3f)); } else if (charcode < 0xd800 || charcode >= 0xe000) { utf8.push(0xe0 | (charcode >> 12), 0x80 | ((charcode>>6) & 0x3f), 0x80 | (charcode & 0x3f)); } // surrogate pair else {
Find elsewhere
🌐
Evan Hahn
evanhahn.com › working-with-utf8-bytes-of-javascript-strings
Working with the UTF-8 bytes of JavaScript strings - Evan Hahn
June 10, 2023 - I took advantage of JavaScript’s built-in TextEncoder, which turns a string into a Uint8Array of the string’s bytes. new TextEncoder().encode("hi 🌍"); // => Uint8Array(7) [104, 105, 32, 240, 159, 140, 141] You can use TextDecoder to reverse the process. const bytes = new Uint8Array([240, 159, 145, 139, 32, 104, 105]); new TextDecoder().decode(bytes); // => "👋 hi" That’s it! If you’re curious, I also wrote up how to do this for UTF-16 and UTF-32, which are more complicated.
🌐
EyeHunts
tutorial.eyehunts.com › home › javascript string to byte array | convert to example code
JavaScript string to byte array | Convert to Example code
December 7, 2021 - JavaScript Strings are stored in UTF-16. To get UTF-8, you’ll have to convert the String yourself. HTML example code. <!DOCTYPE HTML> <html> <body> <script> var str = "Hello"; var bytes = []; var bytesv2 = []; for (var i = 0; i < str.length; ...
🌐
GitHub
gist.github.com › joni › 3760795
toUTF8Array: Javascript function for encoding a string in UTF8. · GitHub
var htmlString = '<div>Your html á é í ó ú</div>'; var arrayUTF8 = toUTF8Array(htmlString); //Your function var byteNumbers = new Uint8Array(arrayUTF8.length); for (var i = 0; i < arrayUTF8.length; i++) { byteNumbers[i] = arrayUTF8[i]; } var blob = new Blob([byteNumbers], {type: 'text/html;charset=UTF-8;' }); FileSaver.saveAs(blob, 'yourfile.doc'); ... @bkdotcom looks like there should not be that extra i+1 in str.charCodeAt, because there is i++; right after else. WDYT? ... } else if (charcode < 0x800) { // ... } else if (charcode < 0xd800 || charcode >= 0xe000) { // ^ never true given previous if · **UPD:** Here is a similar function inside google closure library: [stringToUtf8ByteArray()](https://github.com/google/closure-library/blob/8598d87242af59aac233270742c8984e2b2bdbe0/closure/goog/crypt/crypt.js#L117-L143).
🌐
Rogueamoeba
weblog.rogueamoeba.com › 2017 › 02 › 27 › javascript-correctly-converting-a-byte-array-to-a-utf-8-string
JavaScript: Correctly Converting a Byte Array to a UTF-8 String
February 27, 2017 - This sounds simple enough, but there’s a catch: Many metadata strings require special handling. An easy example of this is accented characters, as seen in band names ranging from Queensrÿche to Sigur Rós. When converting this text to a stream of bytes, the special characters need to be encoded with something like UTF-8.
🌐
xjavascript
xjavascript.com › blog › how-to-convert-utf8-string-to-byte-array
How to Convert a UTF8 String to a Byte Array in JavaScript: Handling Multi-Byte Characters with charCodeAt — xjavascript.com
This blog will demystify the process of converting a UTF-8 string to a byte array in JavaScript using charCodeAt(), explaining how to handle multi-byte characters, surrogate pairs, and edge cases.
🌐
Bobby Hadz
bobbyhadz.com › blog › convert-string-to-byte-array-in-javascript
How to convert a String to a Byte Array in JavaScript | bobbyhadz
... Copied!const utf8EncodeText ... // ] console.log(byteArray); ... The TextEncoder() constructor creates a TextEncoder object that is used to generate a byte stream with UTF-8 encoding....
🌐
npm
npmjs.com › package › utf8-string-bytes
utf8-string-bytes - npm
December 13, 2017 - Latest version: 1.0.3, last published: 8 years ago. Start using utf8-string-bytes in your project by running `npm i utf8-string-bytes`. There are 5 other projects in the npm registry using utf8-string-bytes.
      » npm install utf8-string-bytes
    
Published: Dec 13, 2017
Version: 1.0.3
Author: leonardosnt
🌐
SSOJet
ssojet.com › character-encoding-decoding › utf-8-in-javascript-in-browser
UTF-8 in JavaScript in Browser | Encoding Standards for Programming Languages
JavaScript strings are internally represented using UTF-16. When you need to work with UTF-8 encoded data, such as sending data over a network or saving it to a file, you'll use the TextEncoder and TextDecoder APIs.
🌐
npm
npmjs.com › package › utf8-bytes
utf8-bytes - npm
November 30, 2013 - Return an array of integers from 0 through 255, inclusive, representing the bytes in the unicode string str.
      » npm install utf8-bytes
    
Published: Nov 30, 2013
Version: 0.0.1
Author: James Halliday
Top answer
1 of 4
8

You don't need to write a full-on UTF-8 encoder; there is a much easier JS idiom to convert a Unicode string into a string of bytes representing UTF-8 code units:

unescape(encodeURIComponent(str))

(This works because the odd encoding used by escape/unescape uses %xx hex sequences to represent ISO-8859-1 characters with that code, instead of UTF-8 as used by URI-component escaping. Similarly decodeURIComponent(escape(bytes)) goes in the other direction.)

So if you want an Array out it would be:

function toUTF8Array(str) {
    var utf8= unescape(encodeURIComponent(str));
    var arr= new Array(utf8.length);
    for (var i= 0; i<utf8.length; i++)
        arr[i]= utf8.charCodeAt(i);
    return arr;
}
2 of 4
3

You can use this function (gist):

function toUTF8Array(str) {
    var utf8 = [];
    for (var i=0; i < str.length; i++) {
        var charcode = str.charCodeAt(i);
        if (charcode < 0x80) utf8.push(charcode);
        else if (charcode < 0x800) {
            utf8.push(0xc0 | (charcode >> 6), 
                      0x80 | (charcode & 0x3f));
        }
        else if (charcode < 0xd800 || charcode >= 0xe000) {
            utf8.push(0xe0 | (charcode >> 12), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
        else {
            // let's keep things simple and only handle chars up to U+FFFF...
            utf8.push(0xef, 0xbf, 0xbd); // U+FFFE "replacement character"
        }
    }
    return utf8;
}

Example of use:

>>> toUTF8Array("中€")
[228, 184, 173, 226, 130, 172]

If you want negative numbers for values over 127, like Java's byte-to-int conversion does, you have to tweak the constants and use

            utf8.push(0xffffffc0 | (charcode >> 6), 
                      0xffffff80 | (charcode & 0x3f));

and

            utf8.push(0xffffffe0 | (charcode >> 12), 
                      0xffffff80 | ((charcode>>6) & 0x3f), 
                      0xffffff80 | (charcode & 0x3f));