If you really want "\u00c2\u00a9" as the output, give json a Unicode string as input.

>>> print json.dumps(u'\xc2\xa9')
"\u00c2\u00a9"

You can generate this Unicode string from the raw bytes:

s = unicode('©', 'utf-8').encode('utf-8')
s2 = u''.join(unichr(ord(c)) for c in s)

I think what you really want is "\xc2\xa9" as the output, but I'm not sure how to generate that yet.

Answer from Mark Ransom on Stack Overflow
Top answer
1 of 1
1

As you said in comments, you solved your issue using:

jsonContent = json.dumps(request.json)
guid_object = hashlib.sha1(jsonContent.encode('utf-8'))

But it's important to understand why this works. Flask sends you unicode() for non-ASCII, and str() for ASCII. Dumping the result using JSON will give you consistent results since it abstracts away the internal Python representation, just as if you only had unicode().

Python 2

In Python 2 (the Python version you're using), you don't need .encode('utf-8') because the default value of ensure_ascii of json.dumps() is True. When you send non-ASCII data to json.dumps(), it will use JSON escape sequences to actually dump ASCII: no need to encode to UTF-8. Also, since the Zen of Python says that "Explicit is better than implicit", even if ensure_ascii is already True, you could specify it:

jsonContent = json.dumps(request.json, ensure_ascii=True)
guid_object = hashlib.sha1(jsonContent)

Python 3

In Python 3 however, this would no longer work. Inded, json.dumps() returns unicode in Python 3, even if everything in the unicode string is ASCII. But hashlib.sha1 only works on bytes. You need to make the conversion explicit, even if the ASCII encoding is all you need:

jsonContent = json.dumps(request.json, ensure_ascii=True)
guid_object = hashlib.sha1(jsonContent.encode('ascii'))

This is why Python 3 is a better language: it forces you to be more explicit about the text you use, whether it is str (Unicode) or bytes. This avoids many, many problems down the road.

🌐
GeeksforGeeks
geeksforgeeks.org › python › python-encode-unicode-and-non-ascii-characters-into-json
Python Encode Unicode and non-ASCII characters into JSON - GeeksforGeeks
July 23, 2025 - By default, the JSON module encodes Unicode objects (such as str and Unicode) into the \u escape sequence when generating JSON data. However, you can serialize Unicode objects into UTF-8 JSON strings by using the json.dumps() function with the encoding parameter set to 'UTF-8'.
Top answer
1 of 2
9

You have UTF-8 JSON data:

>>> import json
>>> data = {'content': u'\u4f60\u597d'}
>>> json.dumps(data, indent=1, ensure_ascii=False)
u'{\n "content": "\u4f60\u597d"\n}'
>>> json.dumps(data, indent=1, ensure_ascii=False).encode('utf8')
'{\n "content": "\xe4\xbd\xa0\xe5\xa5\xbd"\n}'
>>> print json.dumps(data, indent=1, ensure_ascii=False).encode('utf8')
{
 "content": "你好"
}

My terminal just happens to be configured to handle UTF-8, so printing the UTF-8 bytes to my terminal produced the desired output.

However, if your terminal is not set up for such output, it is your terminal that then shows 'wrong' characters:

>>> print json.dumps(data, indent=1,  ensure_ascii=False).encode('utf8').decode('latin1')
{
 "content": "你好"
}

Note how I decoded the data to Latin-1 to deliberately mis-read the UTF-8 bytes.

This isn't a Python problem; this is a problem with how you are handling the UTF-8 bytes in whatever tool you used to read these bytes.

2 of 2
4

in python2, it works; however in python3 print will output like:

>>> b'{\n "content": "\xe4\xbd\xa0\xe5\xa5\xbd"\n}'

do not use encode('utf8'):

>>> print(json.dumps(data, indent=1, ensure_ascii=False))
{
 "content": "你好"
}

or use sys.stdout.buffer.write instead of print:

>>> import sys
>>> import json
>>> data = {'content': u'\u4f60\u597d'}
>>> sys.stdout.buffer.write(json.dumps(data, indent=1, 
ensure_ascii=False).encode('utf8') + b'\n')
{
 "content": "你好"
}

see Write UTF-8 to stdout, regardless of the console's encoding

🌐
GitHub
gist.github.com › 2558970
Python 2.7: Save unicode string to file not using '\uXXXX' but using UTF-8. · GitHub
Python 2.7: Save unicode string to file not using '\uXXXX' but using UTF-8. - json-utf8.py
Top answer
1 of 13
1373

Use the ensure_ascii=False switch to json.dumps(), then encode the value to UTF-8 manually:

>>> json_string = json.dumps("ברי צקלה", ensure_ascii=False).encode('utf8')
>>> json_string
b'"\xd7\x91\xd7\xa8\xd7\x99 \xd7\xa6\xd7\xa7\xd7\x9c\xd7\x94"'
>>> print(json_string.decode())
"ברי צקלה"

If you are writing to a file, just use json.dump() and leave it to the file object to encode:

with open('filename', 'w', encoding='utf8') as json_file:
    json.dump("ברי צקלה", json_file, ensure_ascii=False)

Caveats for Python 2

For Python 2, there are some more caveats to take into account. If you are writing this to a file, you can use io.open() instead of open() to produce a file object that encodes Unicode values for you as you write, then use json.dump() instead to write to that file:

with io.open('filename', 'w', encoding='utf8') as json_file:
    json.dump(u"ברי צקלה", json_file, ensure_ascii=False)

Do note that there is a bug in the json module where the ensure_ascii=False flag can produce a mix of unicode and str objects. The workaround for Python 2 then is:

with io.open('filename', 'w', encoding='utf8') as json_file:
    data = json.dumps(u"ברי צקלה", ensure_ascii=False)
    # unicode(data) auto-decodes data to unicode if str
    json_file.write(unicode(data))

In Python 2, when using byte strings (type str), encoded to UTF-8, make sure to also set the encoding keyword:

>>> d={ 1: "ברי צקלה", 2: u"ברי צקלה" }
>>> d
{1: '\xd7\x91\xd7\xa8\xd7\x99 \xd7\xa6\xd7\xa7\xd7\x9c\xd7\x94', 2: u'\u05d1\u05e8\u05d9 \u05e6\u05e7\u05dc\u05d4'}

>>> s=json.dumps(d, ensure_ascii=False, encoding='utf8')
>>> s
u'{"1": "\u05d1\u05e8\u05d9 \u05e6\u05e7\u05dc\u05d4", "2": "\u05d1\u05e8\u05d9 \u05e6\u05e7\u05dc\u05d4"}'
>>> json.loads(s)['1']
u'\u05d1\u05e8\u05d9 \u05e6\u05e7\u05dc\u05d4'
>>> json.loads(s)['2']
u'\u05d1\u05e8\u05d9 \u05e6\u05e7\u05dc\u05d4'
>>> print json.loads(s)['1']
ברי צקלה
>>> print json.loads(s)['2']
ברי צקלה
2 of 13
163

To write to a file

import codecs
import json

with codecs.open('your_file.txt', 'w', encoding='utf-8') as f:
    json.dump({"message":"xin chào việt nam"}, f, ensure_ascii=False)

To print to stdout

import json
print(json.dumps({"message":"xin chào việt nam"}, ensure_ascii=False))
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-a-string-to-utf-8-in-python
Convert a String to Utf-8 in Python - GeeksforGeeks
July 23, 2025 - Below, are the methods for How To Convert A String To Utf-8 In Python. ... The most straightforward way to convert a string to UTF-8 in Python is by using the encode method.
🌐
PYnative
pynative.com › home › python › json › python encode unicode and non-ascii characters as-is into json
Python Encode Unicode and non-ASCII characters as-is into JSON
May 14, 2021 - import json sampleDict= { "string1": "明彦", "string2": u"\u00f8" } with open("unicodeFile.json", "w", encoding='utf-8') as write_file: json.dump(sampleDict, write_file, ensure_ascii=False) print("Done writing JSON serialized Unicode Data as-is into file") with open("unicodeFile.json", "r", encoding='utf-8') as read_file: print("Reading JSON serialized Unicode data from file") sampleData = json.load(read_file) print("Decoded JSON serialized Unicode data") print(sampleData["string1"], sampleData["string1"])Code language: Python (python)
Find elsewhere
🌐
Python
docs.python.org › 3.0 › library › json.html
json — JSON encoder and decoder — Python v3.0.1 documentation
March 20, 2010 - If encoding is not None, then all input strings will be transformed into unicode using that encoding prior to JSON-encoding. The default is UTF-8.
Top answer
1 of 3
9

Requirements

  • Make sure your python files are encoded in UTF-8. Or else your non-ascii characters will become question marks, ?. Notepad++ has excellent encoding options for this.

  • Make sure that you have the appropriate fonts included. If you want to display Japanese characters then you need to install Japanese fonts.

  • Make sure that your IDE supports displaying unicode characters. Otherwise you might get an UnicodeEncodeError error thrown.

Example:

UnicodeEncodeError: 'charmap' codec can't encode characters in position 22-23: character maps to <undefined>

PyScripter works for me. It's included with "Portable Python" at http://portablepython.com/wiki/PortablePython3.2.1.1

  • Make sure you're using Python 3+, since this version offers better unicode support.

Problem

json.dumps() escapes unicode characters.

Solution

Read the update at the bottom. Or...

Replace each escaped characters with the parsed unicode character.

I created a simple lambda function called getStringWithDecodedUnicode that does just that.

import re   
getStringWithDecodedUnicode = lambda str : re.sub( '\\\\u([\da-f]{4})', (lambda x : chr( int( x.group(1), 16 ) )), str )

Here's getStringWithDecodedUnicode as a regular function.

def getStringWithDecodedUnicode( value ):
    findUnicodeRE = re.compile( '\\\\u([\da-f]{4})' )
    def getParsedUnicode(x):
        return chr( int( x.group(1), 16 ) )

    return  findUnicodeRE.sub(getParsedUnicode, str( value ) )

Example

testJSONWithUnicode.py (Using PyScripter as the IDE)

import re
import json
getStringWithDecodedUnicode = lambda str : re.sub( '\\\\u([\da-f]{4})', (lambda x : chr( int( x.group(1), 16 ) )), str )

data = {"Japan":"日本"}
jsonString = json.dumps( data )
print( "json.dumps({0}) = {1}".format( data, jsonString ) )
jsonString = getStringWithDecodedUnicode( jsonString )
print( "Decoded Unicode: %s" % jsonString )

Output

json.dumps({'Japan': '日本'}) = {"Japan": "\u65e5\u672c"}
Decoded Unicode: {"Japan": "日本"}

Update

Or... just pass ensure_ascii=False as an option for json.dumps.

Note: You need to meet the requirements that I outlined at the beginning or else this isn't going to work.

import json
data = {'navn': 'Åge', 'stilling': 'Lærling'}
result = json.dumps(d, ensure_ascii=False)
print( result ) # prints '{"stilling": "Lærling", "navn": "Åge"}'
2 of 3
6

encode_ascii=False is the best solution IMHO.

If you are using Python2.7, here is example python file :

#!/usr/bin/env python
# -*- coding: utf-8 -*-
# example.py
from __future__ import unicode_literals
from json import dumps as json_dumps
d = {'navn': 'Åge', 'stilling': 'Lærling'}
print json_dumps(d, ensure_ascii=False).encode('utf-8')
🌐
Readthedocs
simplejson.readthedocs.io › en › v3.17.0
simplejson — JSON encoder and decoder — simplejson 3.17.0 documentation
If encoding is not None, then all input bytes objects in Python 3 and 8-bit strings in Python 2 will be transformed into unicode using that encoding prior to JSON-encoding. The default is 'utf-8'. If encoding is None, then all bytes objects will be passed to the default function in Python 3
🌐
Medium
paul-d-chuang.medium.com › python-json-dumps-ensure-ascii-false-to-keep-the-original-non-ascii-characters-d496f250245b
Python json.dumps() results in Unicode escape sequence by default | by Paul Chuang | Medium
April 25, 2025 - import json # Example body with non-ASCII characters body = { "key": "你" } # Escaped JSON (with ensure_ascii=True, which is the default) escaped_json = json.dumps(body) print("Escaped JSON:") print(escaped_json) # Non-Escaped JSON (with ensure_ascii=False) non_escaped_json = json.dumps(body, ensure_ascii=False) print("\nNon-Escaped JSON:") print(non_escaped_json) # Encode both JSON strings in UTF-8 encoded_escaped_json = escaped_json.encode("utf-8") encoded_non_escaped_json = non_escaped_json.encode("utf-8") # Print the encoded versions print("\nEncoded Escaped JSON (UTF-8):", encoded_escaped_json) print("Encoded Non-Escaped JSON (UTF-8):", encoded_non_escaped_json) # Compare if both encoded byte sequences are the same print("\nAre the encoded byte sequences the same?", encoded_escaped_json == encoded_non_escaped_json)
🌐
Jsonic
jsonic.io › home › guides › json utf-8 encoding
JSON UTF-8 Encoding: \uXXXX Escapes, BOM, Charset — Jsonic
May 23, 2026 - Some libraries default to escaping non-ASCII characters anyway (Python json.dumps does this, jq has --ascii-output) to produce JSON that survives ASCII-only pipelines, but the resulting file means the same thing as the raw-UTF-8 version. Parsers must accept both forms and produce the same in-memory string. \uXXXX is JSON&apos;s escape syntax for any Unicode codepoint in the Basic Multilingual Plane (U+0000 through U+FFFF).
Top answer
1 of 4
7

Your text is already encoded and you need to tell this to Python by using a b prefix in your string but since you're using json and the input needs to be string you have to decode your encoded text manually. Since your input is not byte you can use 'raw_unicode_escape' encoding to convert the string to byte without encoding and prevent the open method to use its own default encoding. Then you can simply use aforementioned approach to get the desired result.

Note that since you need to do the encoding and decoding your have to read file content and perform the encoding on loaded string, then you should use json.loads() instead of json.load().

In [168]: with open('test.json', encoding='raw_unicode_escape') as f:
     ...:     d = json.loads(f.read().encode('raw_unicode_escape').decode())
     ...:     

In [169]: d
Out[169]: {'sender_name': 'Horníková'}
2 of 4
7

The JSON you are reading was written incorrectly and the Unicode strings decoded from it will have to be re-encoded with the wrong encoding used, then decoded with the correct encoding.

Here's an example:

#!python3
import json

# The bad JSON you have
bad_json = r'{"sender_name": "Horn\u00c3\u00adkov\u00c3\u00a1"}'
print('bad_json =',bad_json)

# The wanted result from json.loads()
wanted = {'sender_name':'Horníková'}

# What correctly written JSON should look like
good_json = json.dumps(wanted)
print('good_json =',good_json)

# What you get when loading the bad JSON.
got = json.loads(bad_json)
print('wanted =',wanted)
print('got =',got)

# How to correct the mojibake string
corrected_sender = got['sender_name'].encode('latin1').decode('utf8')
print('corrected_sender =',corrected_sender)

Output:

bad_json = {"sender_name": "Horn\u00c3\u00adkov\u00c3\u00a1"}
good_json = {"sender_name": "Horn\u00edkov\u00e1"}
wanted = {'sender_name': 'Horníková'}
got = {'sender_name': 'HornÃ\xadková'}
corrected_sender = Horníková