The logic of encoding Unicode in UTF-8 is basically:

  • Up to 4 bytes per character can be used. The fewest number of bytes possible is used.
  • Characters up to U+007F are encoded with a single byte.
  • For multibyte sequences, the number of leading 1 bits in the first byte gives the number of bytes for the character. The rest of the bits of the first byte can be used to encode bits of the character.
  • The continuation bytes begin with 10, and the other 6 bits encode bits of the character.

Here's a function I wrote a while back for encoding a JavaScript UTF-16 string in UTF-8:

function toUTF8Array(str) {
    var utf8 = [];
    for (var i=0; i < str.length; i++) {
        var charcode = str.charCodeAt(i);
        if (charcode < 0x80) utf8.push(charcode);
        else if (charcode < 0x800) {
            utf8.push(0xc0 | (charcode >> 6), 
                      0x80 | (charcode & 0x3f));
        }
        else if (charcode < 0xd800 || charcode >= 0xe000) {
            utf8.push(0xe0 | (charcode >> 12), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
        // surrogate pair
        else {
            i++;
            // UTF-16 encodes 0x10000-0x10FFFF by
            // subtracting 0x10000 and splitting the
            // 20 bits of 0x0-0xFFFFF into two halves
            charcode = 0x10000 + (((charcode & 0x3ff)<<10)
                      | (str.charCodeAt(i) & 0x3ff));
            utf8.push(0xf0 | (charcode >>18), 
                      0x80 | ((charcode>>12) & 0x3f), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
    }
    return utf8;
}
Answer from Joni on Stack Overflow
🌐
Dirask
dirask.com › posts › JavaScript-convert-string-to-bytes-array-UTF-8-1XkbEj
JavaScript - convert string to bytes array (UTF-8)
// ONLINE-RUNNER:browser; const toBytes = (text) => { const result = []; for (let i = 0; i < text.length; i += 1) { const hi = text.charCodeAt(i); if (hi < 0x0080) { // code point range: U+0000 - U+007F // bytes: 0xxxxxxx result.push(hi); continue; } if (hi < 0x0800) { // code point range: U+0080 - U+07FF // bytes: 110xxxxx 10xxxxxx result.push(0xC0 | hi >> 6, 0x80 | hi & 0x3F); continue; } if (hi < 0xD800 || hi >= 0xE000 ) { // code point range: U+0800 - U+FFFF // bytes: 1110xxxx 10xxxxxx 10xxxxxx result.push(0xE0 | hi >> 12, 0x80 | hi >> 6 & 0x3F, 0x80 | hi & 0x3F); continue; } i += 1; if (i
Top answer
1 of 10
84

The logic of encoding Unicode in UTF-8 is basically:

  • Up to 4 bytes per character can be used. The fewest number of bytes possible is used.
  • Characters up to U+007F are encoded with a single byte.
  • For multibyte sequences, the number of leading 1 bits in the first byte gives the number of bytes for the character. The rest of the bits of the first byte can be used to encode bits of the character.
  • The continuation bytes begin with 10, and the other 6 bits encode bits of the character.

Here's a function I wrote a while back for encoding a JavaScript UTF-16 string in UTF-8:

function toUTF8Array(str) {
    var utf8 = [];
    for (var i=0; i < str.length; i++) {
        var charcode = str.charCodeAt(i);
        if (charcode < 0x80) utf8.push(charcode);
        else if (charcode < 0x800) {
            utf8.push(0xc0 | (charcode >> 6), 
                      0x80 | (charcode & 0x3f));
        }
        else if (charcode < 0xd800 || charcode >= 0xe000) {
            utf8.push(0xe0 | (charcode >> 12), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
        // surrogate pair
        else {
            i++;
            // UTF-16 encodes 0x10000-0x10FFFF by
            // subtracting 0x10000 and splitting the
            // 20 bits of 0x0-0xFFFFF into two halves
            charcode = 0x10000 + (((charcode & 0x3ff)<<10)
                      | (str.charCodeAt(i) & 0x3ff));
            utf8.push(0xf0 | (charcode >>18), 
                      0x80 | ((charcode>>12) & 0x3f), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
    }
    return utf8;
}
2 of 10
48

JavaScript Strings are stored in UTF-16. To get UTF-8, you'll have to convert the String yourself.

One way is to mix encodeURIComponent(), which will output UTF-8 bytes URL-encoded, with unescape, as mentioned on ecmanaut.

var utf8 = unescape(encodeURIComponent(str));

var arr = [];
for (var i = 0; i < utf8.length; i++) {
    arr.push(utf8.charCodeAt(i));
}
🌐
MDN Web Docs
developer.mozilla.org › en-US › docs › Web › API › TextEncoder › encodeInto
TextEncoder: encodeInto() method - Web APIs | MDN
June 28, 2025 - However, it is sometimes useful to make the output start at a particular index. The solution is TypedArray.prototype.subarray(): ... const encoder = new TextEncoder(); function encodeIntoAtPosition(string, u8array, position) { return encoder.encodeInto( string, position ?
🌐
GeeksforGeeks
geeksforgeeks.org › javascript-program-to-convert-string-to-bytes
JavaScript Program to Convert String to Bytes | GeeksforGeeks
June 14, 2024 - This approach is particularly useful for server-side JavaScript running in a Node.js environment. Example: In this example, we use the Buffer.from() method to convert a string into bytes using UTF-8 encoding.
🌐
Evan Hahn
evanhahn.com › working-with-utf8-bytes-of-javascript-strings
Working with the UTF-8 bytes of JavaScript strings - Evan Hahn
June 10, 2023 - I took advantage of JavaScript’s built-in TextEncoder, which turns a string into a Uint8Array of the string’s bytes. new TextEncoder().encode("hi 🌍"); // => Uint8Array(7) [104, 105, 32, 240, 159, 140, 141] You can use TextDecoder to reverse the process. const bytes = new Uint8Array([240, ...
🌐
Java2s
java2s.com › example › nodejs › string › string-to-utf8-byte-array.html
String to UTF8 Byte Array - Node.js String
if (typeof String.prototype.toUTF8ByteArray != 'function') String.prototype.toUTF8ByteArray = function() { var bytes = []; var s = unescape(encodeURIComponent(this)); for (var i = 0; i < s.length; i++) { var c = s.charCodeAt(i); bytes.push(c);/* w w w. j ava 2 s . c o m*/ } return bytes; }; ...
🌐
Designcise
designcise.com › web › tutorial › how-to-convert-a-javascript-string-to-a-byte-array
How to Convert a JavaScript String to a Byte Array? - Designcise
April 9, 2023 - In JavaScript, you can convert a string to an array of bytes by using the TextEncoder API, for example, in the following way: function toBytesArray(str) { const encoder = new TextEncoder(); return encoder.encode(str); } console.log(toBytesA...
🌐
Bobby Hadz
bobbyhadz.com › blog › convert-string-to-byte-array-in-javascript
How to convert a String to a Byte Array in JavaScript | bobbyhadz
Call the encode() method on the object to convert the string to a byte array. ... Copied!const utf8EncodeText = new TextEncoder(); const str = 'bobbyhadz.com'; const byteArray = utf8EncodeText.encode(str); // Uint8Array(13) [ // 98, 111, 98, ...
Find elsewhere
🌐
GitHub
gist.github.com › vbabak › c47a67ff5e89cab8954097eeb2a33fb8
Convert JavaScript utf-8 string to bytes array. · GitHub
Convert JavaScript utf-8 string to bytes array. GitHub Gist: instantly share code, notes, and snippets.
🌐
Burke
kevin.burke.dev › kevin › node-js-string-encoding
Let’s talk about Javascript string encoding | Kevin Burke
September 1, 2017 - Where the cent character is the UTF-8 encoded byte sequence "\xc2\xa2". When Node starts and you try to reference x in your program, it will be re-encoded as a UTF-16 string. If you type the literal characters: ... This will be turned into the UTF-16 string "\xc2\x00\xa2\x00".
🌐
Techie Delight
techiedelight.com › home › java › convert a string to bytes in javascript
Convert a string to bytes in JavaScript | Techie Delight
July 7, 2026 - Here are some of the most common functions: One way is to use the TextEncoder() constructor, which creates a TextEncoder object that can encode a string into a byte stream with UTF-8 encoding.
🌐
EyeHunts
tutorial.eyehunts.com › home › javascript string to byte array | convert to example code
JavaScript string to byte array | Convert to Example code
December 7, 2021 - <!DOCTYPE HTML> <html> <body> <script> var str = "Hello"; var bytes = []; var bytesv2 = []; for (var i = 0; i < str.length; ++i) { var code = str.charCodeAt(i); bytes = bytes.concat([code]); bytesv2 = bytesv2.concat([code & 0xff, code / 256 ...
🌐
GitHub
gist.github.com › joni › 3760795
toUTF8Array: Javascript function for encoding a string in UTF8. · GitHub
var htmlString = '<div>Your html á é í ó ú</div>'; var arrayUTF8 = toUTF8Array(htmlString); //Your function var byteNumbers = new Uint8Array(arrayUTF8.length); for (var i = 0; i < arrayUTF8.length; i++) { byteNumbers[i] = arrayUTF8[i]; ...
🌐
xjavascript
xjavascript.com › blog › how-to-convert-utf8-string-to-byte-array
How to Convert a UTF8 String to a Byte Array in JavaScript: Handling Multi-Byte Characters with charCodeAt — xjavascript.com
In JavaScript, strings are internally represented using UTF-16, a character encoding that uses 16-bit "code units" to represent most Unicode characters. However, when working with binary data—such as sending data over a network, writing to a file, or interacting with low-level APIs—we often need to convert strings to UTF-8, a variable-length encoding that uses 1–4 bytes per character.
🌐
Ixti
ixti.net › development › node.js › 2011 › 10 › 26 › get-utf-8-string-from-array-of-bytes-in-node-js.html
Get UTF-8 string from array of bytes in Node.JS * ixti's personal scratchpad
October 26, 2011 - We can get array of UTF8 bytes representation of a string with following snippet: function getBytes(str) { var bytes = [], char; str = encodeURI(str); while (str.length) { char = str.slice(0, 1); str = str.slice(1); if ('%' !== char) { bytes.push(char.charCodeAt(0)); } else { char = str.slice(0, ...
Top answer
1 of 4
8

You don't need to write a full-on UTF-8 encoder; there is a much easier JS idiom to convert a Unicode string into a string of bytes representing UTF-8 code units:

unescape(encodeURIComponent(str))

(This works because the odd encoding used by escape/unescape uses %xx hex sequences to represent ISO-8859-1 characters with that code, instead of UTF-8 as used by URI-component escaping. Similarly decodeURIComponent(escape(bytes)) goes in the other direction.)

So if you want an Array out it would be:

function toUTF8Array(str) {
    var utf8= unescape(encodeURIComponent(str));
    var arr= new Array(utf8.length);
    for (var i= 0; i<utf8.length; i++)
        arr[i]= utf8.charCodeAt(i);
    return arr;
}
2 of 4
3

You can use this function (gist):

function toUTF8Array(str) {
    var utf8 = [];
    for (var i=0; i < str.length; i++) {
        var charcode = str.charCodeAt(i);
        if (charcode < 0x80) utf8.push(charcode);
        else if (charcode < 0x800) {
            utf8.push(0xc0 | (charcode >> 6), 
                      0x80 | (charcode & 0x3f));
        }
        else if (charcode < 0xd800 || charcode >= 0xe000) {
            utf8.push(0xe0 | (charcode >> 12), 
                      0x80 | ((charcode>>6) & 0x3f), 
                      0x80 | (charcode & 0x3f));
        }
        else {
            // let's keep things simple and only handle chars up to U+FFFF...
            utf8.push(0xef, 0xbf, 0xbd); // U+FFFE "replacement character"
        }
    }
    return utf8;
}

Example of use:

>>> toUTF8Array("中€")
[228, 184, 173, 226, 130, 172]

If you want negative numbers for values over 127, like Java's byte-to-int conversion does, you have to tweak the constants and use

            utf8.push(0xffffffc0 | (charcode >> 6), 
                      0xffffff80 | (charcode & 0x3f));

and

            utf8.push(0xffffffe0 | (charcode >> 12), 
                      0xffffff80 | ((charcode>>6) & 0x3f), 
                      0xffffff80 | (charcode & 0x3f));
Top answer
1 of 2
19

You can use TextEncoder which is part of the Encoding Living Standard. According to the Encoding API entry from the Chromium Dashboard, it shipped in Firefox and will ship in Chrome 38. There is also a text-encoding polyfill available.

The JavaScript code sample below returns a Uint8Array filled with the values you expect.

var s = "test.message";
var encoder = new TextEncoder();
encoder.encode(s);
// [116, 101, 115, 116, 46, 109, 101, 115, 115, 97, 103, 101]
2 of 2
10

JavaScript has no concept of character encoding for String, everything is in UTF-16. Most of time time the value of a char in UTF-16 matches UTF-8, so you can forget it's any different.

There are more optimal ways to do this but

function s(x) {return x.charCodeAt(0);}
"test.message".split('').map(s);
// [116, 101, 115, 116, 46, 109, 101, 115, 115, 97, 103, 101]

So what is unescape(encodeURIComponent(str)) doing? Let's look at each individually,

  1. encodeURIComponent is converting every character in str which is illegal or has a meaning in URI Syntax into a URI escaped version so that there is no problem using it as a key or value in the search component of a URI, for example encodeURIComponent('&='); // "%26%3D" Notice how this is now a 6 character long String.
  2. unescape is actually depreciated, but it does a similar job to decodeURI or decodeURIComponent (the reverse of encodeURIComponent). If we look in the ES5 spec we can see 11. Let c be the character whose code unit value is the integer represented by the four hexadecimal digits at positions k+2, k+3, k+4, and k+5 within Result(1).
    So, 4 digits is 2 bytes is "UTF-8", however as I mentioned, all Strings are UTF-16, so it's really a UTF-16 string limiting itself to UTF-8.
🌐
GitHub
gist.github.com › lihnux › 2aa4a6f5a9170974f6aa
Javascript Convert String to Byte Array · GitHub
function unpack(str) { var bytes = []; for(var i = 0; i < str.length; i++) { var char = str.charCodeAt(i); bytes.push(char >>> 8); bytes.push(char & 0xFF); } return bytes; } ... function toUTF8Array(str) { let utf8 = []; for (let i = 0; i < ...
🌐
Rogueamoeba
weblog.rogueamoeba.com › 2017 › 02 › 27 › javascript-correctly-converting-a-byte-array-to-a-utf-8-string
JavaScript: Correctly Converting a Byte Array to a UTF-8 String
February 27, 2017 - This sounds simple enough, but there’s a catch: Many metadata strings require special handling. An easy example of this is accented characters, as seen in band names ranging from Queensrÿche to Sigur Rós. When converting this text to a stream of bytes, the special characters need to be encoded with something like UTF-8.