Single Buffer

If you have a single Buffer you can use its toString method that will convert all or part of the binary contents to a string using a specific encoding. It defaults to utf8 if you don't provide a parameter, but I've explicitly set the encoding in this example.

var req = http.request(reqOptions, function(res) {
    ...

    res.on('data', function(chunk) {
        var textChunk = chunk.toString('utf8');
        // process utf8 text chunk
    });
});

Streamed Buffers

If you have streamed buffers like in the question above where the first byte of a multi-byte UTF8-character may be contained in the first Buffer (chunk) and the second byte in the second Buffer then you should use a StringDecoder. :

var StringDecoder = require('string_decoder').StringDecoder;

var req = http.request(reqOptions, function(res) {
    ...
    var decoder = new StringDecoder('utf8');

    res.on('data', function(chunk) {
        var textChunk = decoder.write(chunk);
        // process utf8 text chunk
    });
});

This way bytes of incomplete characters are buffered by the StringDecoder until all required bytes were written to the decoder.

Answer from Biggie on Stack Overflow
Top answer
1 of 2
319

Single Buffer

If you have a single Buffer you can use its toString method that will convert all or part of the binary contents to a string using a specific encoding. It defaults to utf8 if you don't provide a parameter, but I've explicitly set the encoding in this example.

var req = http.request(reqOptions, function(res) {
    ...

    res.on('data', function(chunk) {
        var textChunk = chunk.toString('utf8');
        // process utf8 text chunk
    });
});

Streamed Buffers

If you have streamed buffers like in the question above where the first byte of a multi-byte UTF8-character may be contained in the first Buffer (chunk) and the second byte in the second Buffer then you should use a StringDecoder. :

var StringDecoder = require('string_decoder').StringDecoder;

var req = http.request(reqOptions, function(res) {
    ...
    var decoder = new StringDecoder('utf8');

    res.on('data', function(chunk) {
        var textChunk = decoder.write(chunk);
        // process utf8 text chunk
    });
});

This way bytes of incomplete characters are buffered by the StringDecoder until all required bytes were written to the decoder.

2 of 2
-5
var fs = require("fs");

function readFileLineByLine(filename, processline) {
    var stream = fs.createReadStream(filename);
    var s = "";
    stream.on("data", function(data) {
        s += data.toString('utf8');
        var lines = s.split("\n");
        for (var i = 0; i < lines.length - 1; i++)
            processline(lines[i]);
        s = lines[lines.length - 1];
    });

    stream.on("end",function() {
        var lines = s.split("\n");
        for (var i = 0; i < lines.length; i++)
            processline(lines[i]);
    });
}

var linenumber = 0;
readFileLineByLine(filename, function(line) {
    console.log(++linenumber + " -- " + line);
});
Discussions

Buffer.toString('utf8') appears to use wtf-8
The byte sequence 237, 166, 164 is not valid utf8, since it encodes a surrogate code point, which is not a valid unicode scalar value. So Buffer.from([237, 166, 164]).toString('utf8') should error. But instead, it returns a string, effectively implementing wtf-8 rather than utf-8. More on github.com
🌐 github.com
10
October 5, 2018
Buffer.toString('utf8') seems to change the output
I would've expected both console.log to echo true. Is it something I am missing? The documentation of buffer doesn't seem to specify that utf8 stringification could lead to a loss of data. More on github.com
🌐 github.com
2
January 29, 2016
Buffer to UTF8 String conversion DoS in node.js and io.js
Links to mentioned blog would be nice. More on reddit.com
🌐 r/netsec
7
79
July 6, 2015
NodeJS inserting unwanted characters into my string, stdout showing garbage!
It looks to me like you are getting the 'European UTF-8 encoding that can stick a c2 in the first byte and the hex value in the second byte". since JavaScript/Node uses UTF-16 it is sending the 16 bit characters to your C program. I'm getting the same results as you on Linux, but the same code on Windows just prints the 3 shading characters and the b0b1b2 as expected. I'm not sure why its different. However, if what you want is exact binary output without encoding, try using Buffers instead of strings. look here http://resources.mpi-inf.mpg.de/d5/teaching/ss05/is05/oracle/server.920/a96529/appb.htm and look at table B-2. this link has an interesting warning : https://xtermjs.org/docs/guides/encoding/ Caveat: Watch out for automatic UTF-8 conversions done by some nodejs interfaces with string type, always go with the buffer variant if possible. edit: on my windows system the default code page is 437. since linux doesn't use codepages, the locale command shows 'en-us.UTF-8'. that could account for the node to c string conversion encoding as utf-8 using the European conversion for characters higher than 0x7F More on reddit.com
🌐 r/node
7
0
November 16, 2020
🌐
Node.js
nodejs.org › api › buffer.html
Buffer | Node.js v26.10.0 Documentation
Empty value (string, Uint8Array, Buffer) is coerced to 0. offset <integer> Number of bytes to skip before starting to fill buf. Default: 0. end <integer> Where to stop filling buf (not inclusive). Default: buf.length. encoding <string> The encoding for value if value is a string. Default: 'utf8'.
🌐
amanhimself.dev
amanhimself.dev › home › blog › converting a buffer to json and utf8 strings in nodejs
Converting a Buffer to JSON and Utf8 Strings in Nodejs | amanhimself.dev
August 10, 2017 - .toString() is not the only way to convert a buffer to a string. Also, it by defaults converts to a utf-8 format string. The other way to convert a buffer to a string is using StringDecoder core module from Nodejs API.
🌐
GitHub
github.com › nodejs › node › issues › 23280
Buffer.toString('utf8') appears to use wtf-8 · Issue #23280 · nodejs/node
October 5, 2018 - The byte sequence 237, 166, 164 is not valid utf8, since it encodes a surrogate code point, which is not a valid unicode scalar value. So Buffer.from([237, 166, 164]).toString('utf8') should error. But instead, it returns a string, effectively implementing wtf-8 rather than utf-8.
Author: nodejs
🌐
MojoAuth
mojoauth.com › character-encoding-decoding › utf-8-encoding--nodejs
What Is UTF-8 in Node.js? Encoding Guide | MojoAuth Dev Guides
Encode a string to UTF-8 in Node.js with Buffer.from(str, 'utf8') and decode with buf.toString('utf8').
🌐
GitHub
github.com › nodejs › node › issues › 4942
Buffer.toString('utf8') seems to change the output · Issue #4942 · nodejs/node
January 29, 2016 - I have the following snippet var base = '3e39c403407257f245f490ed08149ada'; var buffer = new Buffer(base, 'hex'); console.log(buffer.equals(new Buffer(buffer.toString('utf8'), 'utf8'))); // yields false console.log(buffer.equals(new Buff...
Author: nodejs
🌐
Toptal
toptal.com › external-blogs › adeva › convert-node-js-buffer-to-string
How to Convert Node.js Buffer to String | Toptal®
May 26, 2026 - const b = Buffer.alloc(10); b.write('test', 4, 'utf8'); console.log(b); // <Buffer 00 00 00 00 74 65 73 74 00 00> In a nutshell, it’s easy to convert a Buffer object to a string using the toString() method.
Find elsewhere
🌐
Mastering JS
masteringjs.io › tutorials › node › buffer-to-string
Using the Buffer `toString()` Function in Node.js - Mastering JS
August 21, 2020 - By default, toString() converts the buffer to a string using UTF8 encoding.
🌐
npm
npmjs.com › package › utf8-buffer
utf8-buffer - npm
Only UTF-8 strings with a max of 4 bytes per character are supported. BOM is kept untouched. Invalid characters are replaced with Unicode Character 'REPLACEMENT CHARACTER' (U+FFFD).
      » npm install utf8-buffer
    
Published: Dec 31, 2019
Version: 1.0.0
🌐
GeeksforGeeks
geeksforgeeks.org › node.js › node-js-buffer-tostring-method
Node.js Buffer.toString() Method - GeeksforGeeks
October 13, 2021 - Its default value is 'utf8'. start: The beginning index of the buffer data from which encoding has to be start. Its default value is 0. end: The last index of the buffer data up to which encoding has to be done. Its default value is Buffer.length. Return Value: It returns decoded string from buffer to string according to specified character encoding.
🌐
Node.js
nodejs.org › docs › v0.5.9 › api › buffers.html
buffers - Node.js v0.5.9 Manual & Documentation
Here are the different string ... a null character ('\0' or '\u0000') into 0x20 (character code of a space). If you want to convert a null character into 0x00, you should use 'utf8'....
🌐
Node.js
nodejs.org › docs › v0.5.0 › api › buffers.html
buffers - Node.js Manual & Documentation
Converting between Buffers and JavaScript string objects requires an explicit encoding method. Here are the different string encodings; 'ascii' - for 7 bit ASCII data only. This encoding method is very fast, and will strip the high bit if set. 'utf8' - Multi byte encoded Unicode characters.
🌐
Melvin George
melvingeorge.me › blog › convert-string-buffer-nodejs
How to convert a String to Buffer and vice versa in Node.js | MELVIN GEORGE
September 19, 2020 - // buffer const str = "Hey. this is a string!"; const buff = Buffer.from(str, "utf-8"); // convert buffer to string const resultStr = buff.toString(); console.log(resultStr); //Hey.
🌐
HackerNoon
hackernoon.com › https-medium-com-amanhimself-converting-a-buffer-to-json-and-utf8-strings-in-nodejs-2150b1e3de57
Converting a Buffer to JSON and Utf8 Strings in Nodejs | HackerNoon
August 10, 2017 - Nodejs and browser based JavaScript differ because Node has a way to handle binary data even before the ES6 draft came up with ArrayBuffer. In Node, Buffer class is the primary data structure used with most I/O operations. It is a raw binary data that is allocated outside the V8 heap and once allocated, cannot be resized.
🌐
Reddit
reddit.com › r/netsec › buffer to utf8 string conversion dos in node.js and io.js
r/netsec on Reddit: Buffer to UTF8 String conversion DoS in node.js and io.js
July 6, 2015 - Interesting to see some nodejs vulns finally. Should be a lot of stuff in there ... Node.js is totally safe! It doesn't have vulnerabilities, just like OS X doesn't get viruses... ... Program received signal SIGSEGV, Segmentation fault. 0x0000000000b56dab in unibrow::Utf8DecoderBase::WriteUtf16Slow(unsigned char const, unsigned short, unsigned int) ()
🌐
Readthedocs
node.readthedocs.io › en › latest › api › buffer
Buffer - node - Read the Docs
Decodes and returns a string from buffer data encoded using the specified character set encoding. If encoding is undefined or null, then encoding defaults to 'utf8'. The start and end parameters default to 0 and buffer.length when undefined.
🌐
Educative
educative.io › answers › what-is-the-nodejs-buffertostring-method
What is the Node.js Buffer.toString() method?
In the code above, we decode our buffer to a string. Note that we do not specify any encoding. Now, let’s specify some encoding. // Create a Buffer · const ourBuffer = Buffer.from("edpresso"); // log out Buffer · console.log(ourBuffer); // decode Buffer · console.log(ourBuffer.toString("utf8")) console.log(ourBuffer.toString("utf16le")) console.log(ourBuffer.toString("latin1")) Run ·
🌐
Google Groups
groups.google.com › g › nodejs › c › 89JMthbHCvw
buffer toString with partial utf8 character?
September 3, 2014 - Or will it have garbage at the end of the string or throw an exception? ... Either email addresses are anonymous for this group or you need the view member email addresses permission to view the original message ... I seem to remember it converts any unknown codes to the Unicode character 65533 (�). Giving it a try is easy! node -e 'console.log(new Buffer([255]).toString().charCodeAt(0))'