First, obtain the integer value from the string of a, noting that a is expressed in hexadecimal:

a_int = int(a, 16)

Next, convert this int to a character. In python 2 you need to use the unichr method to do this, because the chr method can only deal with ASCII characters:

a_chr = unichr(a_int)

Whereas in python 3 you can just use the chr method for any character:

a_chr = chr(a_int)

So, in python 3, the full command is:

a_chr = chr(int(a, 16))
Answer from Rob Bricheno on Stack Overflow
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-hex-to-string-in-python
Convert Hex to String in Python - GeeksforGeeks
July 23, 2025 - This built-in method converts a hex string into bytes. Use .decode() to convert it to a UTF-8 string.
Discussions

utf 8 - Read hex characters and convert them to utf-8 using python 3 - Stack Overflow
I have a file data.txt containing this string: M\xc3\xbchle\x0astra\xc3\x9fe Now the file needs to be read and the hex code interpreted as utf-8. So far this is my try: #!/usr/bin/python3 impor... More on stackoverflow.com
🌐 stackoverflow.com
April 16, 2017
utf 8 - Python how to decode unicode with hex characters - Stack Overflow
I have extracted a string from web crawl script as following: u'\xe3\x80\x90\xe4\xb8\xad\xe5\xad\x97\xe3\x80\x91' I want to decode u'\xe3\x80\x90\xe4\xb8\xad\xe5\xad\x97\xe3\x80\x91' with utf-8. W... More on stackoverflow.com
🌐 stackoverflow.com
Convert unicode codepoint to UTF8 hex in python - Stack Overflow
I want to convert a number of unicode codepoints read from a file to their UTF8 encoding. e.g I want to convert the string 'FD9B' to the string 'EFB69B'. I can do this manually using string liter... More on stackoverflow.com
🌐 stackoverflow.com
utf 8 - python encode/decode hex string to utf-8 string - Stack Overflow
Please check your hex string, it simply looks wrong ... @giacomo-catenazzi can you please explain it a litle more ? your commet is a litle bit to short for me to full understand the problem. ... Your data is encoded as UTF-8, which means that you sometimes have to look at more than one byte to get one character. The easiest way to do this is probably to decode your string into a sequence of bytes, and then decode those bytes into a string. Python ... More on stackoverflow.com
🌐 stackoverflow.com
People also ask

Is hex to ASCII different from hex to UTF-8?
For bytes 00–7F they are identical, since UTF-8 was designed to be ASCII-compatible. They differ once non-ASCII data appears: UTF-8 uses two-to-four-byte sequences for non-ASCII scalar values. Accented letters, symbols, and emoji require an encoding beyond ASCII; UTF-8 is the usual choice for modern interchange, but the source format must determine the decoder.
🌐
hextoascii.co
hextoascii.co › home › articles › decode hex to utf-8 without garbled text
Decode Hex to UTF-8 Without Garbled Text — hextoascii.co
Why does my hex decode to `é` instead of `é`?
Because the UTF-8 bytes c3 a9 were decoded one at a time as Latin-1 or Windows-1252, splitting a single two-byte character into à and ©. This is called mojibake. Decode the bytes as UTF-8 in a single operation and the two bytes combine correctly into é.
🌐
hextoascii.co
hextoascii.co › home › articles › decode hex to utf-8 without garbled text
Decode Hex to UTF-8 Without Garbled Text — hextoascii.co
Can I tell which encoding a hex string uses just by looking?
Not with certainty. All bytes ≤ 7F are compatible with ASCII and UTF-8 but could belong to many other formats. Valid UTF-8 lead/continuation patterns make UTF-8 plausible, and a leading ef bb bf is a strong UTF-8 signature, but short byte strings can be valid under several encodings. Prefer declared metadata, protocol specifications, or knowledge of the source over guessing from bytes alone.
🌐
hextoascii.co
hextoascii.co › home › articles › decode hex to utf-8 without garbled text
Decode Hex to UTF-8 Without Garbled Text — hextoascii.co
🌐
GitHub
gist.github.com › ngelik › 46b9fc25800f0333dd6e
Python: hex to uft-8 convertation · GitHub
In python (3.11) things have changed a bit. decode is now on bytes not str and you need parens, so something like · > print(bytes.fromhex('D0A6D0B5D0BDD182D180').decode('utf-8')) Центр
🌐
GitHub
github.com › maximilianh › maxtools › blob › master › lib › unicodeConvert.py
maxtools/lib/unicodeConvert.py at master · maximilianh/maxtools
"""Convert EUC hex (e.g. "d2bb") to UTF8 hex (e.g. "e4 b8 80")""" · utf8 = euc_to_python(euchex).encode("utf-8") · utf8 = repr(utf8)[1:-1].replace("\\x", " ").strip() · return utf8 · · #TODO: make these conversions go ...
Author: maximilianh
Top answer
1 of 3
6

The problem with

msg = u'\xe3\x80\x90\xe4\xb8\xad\xe5\xad\x97\xe3\x80\x91'
result = msg.decode('utf8')

is that you are trying to decode Unicode. That doesn't really make sense. You can encode from Unicode to some type of encoding, or you can decode a byte string to Unicode.

When you do

msg.decode('utf8')

Python 2 sees that msg is Unicode. It knows that it can't decode Unicode so it "helpfully" assumes that you want to encode msg with the default ASCII codec so the result of that transformation can be decoded to Unicode using the UTF-8 codec. Python 3 behaves much more sensibly: that code would simply fail with

AttributeError: 'str' object has no attribute 'decode'

The technique given in kennytm's answer:

msg.encode('latin1').decode('utf-8')

works because the Unicode codepoints less than 256 correspond directly to the characters in the Latin1 encoding (aka ISO 8859-1).

Here's some Python 2 code that illustrates this:

for i in xrange(256):
    lat = chr(i)
    uni = unichr(i)
    assert lat == uni.encode('latin1')
    assert lat.decode('latin1') == uni

And here is the equivalent Python 3 code:

for i in range(256):
    lat = bytes([i])
    uni = chr(i)
    assert lat == uni.encode('latin1')
    assert lat.decode('latin1') == uni

You may find this article helpful: Pragmatic Unicode, which was written by SO veteran Ned Batchelder.

Unless you are forced to use Python 2 I strongly advise you to switch to Python 3. It will make handling Unicode far less painful.

2 of 3
4
  1. Perhaps you should fix the crawl script instead, a Unicode string should contain u'【中字】' (u'\u3010\u4e2d\u5b57\u3011') already, instead of the raw UTF-8 bytes.

  2. To convert msg to the correct encoding, first you need to turn the wrong Unicode string back to byte string (encode it as Latin-1), then decode it as UTF-8:

    >>> print msg.encode('latin1').decode('utf-8')
    【中字】
    
Find elsewhere
🌐
Hextoascii
hextoascii.co › home › articles › decode hex to utf-8 without garbled text
Decode Hex to UTF-8 Without Garbled Text — hextoascii.co
May 22, 2026 - Fix: Decode with utf-8-sig in Python to strip a leading BOM automatically, or slice off the first three bytes if present. When a hex string won't decode cleanly, run through this in order:
🌐
Medium
medium.com › the-dark-grimoire › decoding-utf-8-hex-strings-in-python-5a22b06d7018
Decoding UTF-8 Hex Strings in Python | by Yen Wang | The Dark Grimoire | Medium
May 17, 2026 - Each pair of hex digits becomes one byte. "e2a081" becomes b'\xe2\xa0\x81'. .decode("utf-8"). UTF-8 is a variable-width encoding. Single ASCII characters occupy one byte; characters in the range U+0080 to U+07FF use two bytes; U+0800 to U+FFFF ...
🌐
Reddit
reddit.com › r/learnpython › how do i deal with hex utf-8 bytes?
r/learnpython on Reddit: How do I deal with Hex UTF-8 bytes?
October 2, 2021 -

So here's the thing, I'm currently doing some Wikipedia web scrapping and one of the main libs I'm using doesn't handle Hex UTF-8 very well, and some of the links I'm trying to access from Wikipedia contain Hex UTF-8.

For example, the Wikipedia article for the History the US from 1991 to 2008 link is: https://en.wikipedia.org/wiki/History_of_the_United_States_(1991%E2%80%932008)

But once you actually access it, your browser will decode the %E2%80%93 to an actual Unicode character, the en dash (–), so you'll see the link on your web browser as: https://en.wikipedia.org/wiki/History_of_the_United_States_(1991–2008)

So, what I'm looking for is a way to substitute the part hex part of my string (%E2%80%93) for the actual Unicode character (–)

🌐
FavTutor
favtutor.com › blogs › hex-to-string-in-python
Convert Hex to String in Python | 6 Methods with Code
September 16, 2023 - Python offers a built-in method called fromhex() for converting a hex string into a bytes object. We can use this method to convert hex to string in Python. ... hex_string = "57656c636f6d6520746f204661765475746f72" bytes_obj = bytes.fromhex...
🌐
Python
docs.python.org › 3.3 › howto › unicode.html
Unicode HOWTO — Python 3.3.7 documentation
September 19, 2017 - >>> "\N{GREEK CAPITAL LETTER DELTA}" # Using the character name '\u0394' >>> "\u0394" # Using a 16-bit hex value '\u0394' >>> "\U00000394" # Using a 32-bit hex value '\u0394' In addition, one can create a string using the decode() method of bytes. This method takes an encoding argument, such as UTF-8, and optionally an errors argument. The errors argument specifies the response when the input string can’t be converted according to the encoding’s rules.
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
>>> "\N{GREEK CAPITAL LETTER DELTA}" # Using the character name '\u0394' >>> "\u0394" # Using a 16-bit hex value '\u0394' >>> "\U00000394" # Using a 32-bit hex value '\u0394' In addition, one can create a string using the decode() method of bytes. This method takes an encoding argument, such as UTF-8, and optionally an errors argument. The errors argument specifies the response when the input string can’t be converted according to the encoding’s rules.
Top answer
1 of 2
5

Python 3 does the right thing of inserting a character into a str which is string of characters, not a byte sequence.

UTF8 is the default encoding. If you need to insert a byte, a different encoding where that character is represented as a byte is needed.

$ PYTHONIOENCODING=iso-8859-1 python3 -c 'print("\x80")' | xxd
00000000: 800a

PYTHONIOENCODING

If this is set before running the interpreter, it overrides the encoding used for stdin/stdout/stderr, in the syntax encodingname:errorhandler. Both the encodingname and the :errorhandler parts are optional and have the same meaning as in str.encode().

2 of 2
4

If you want to output raw bytes in Python 3 you shouldn't be using the print function, since it's for outputting text in your default encoding. Instead, you can use sys.stdout.buffer.write.

ASCII is a 7 bit encoding, so if your so-called ASCII contains characters like b'\x80' it's not legal ASCII. Perhaps your data is actually encoded with iso-8859-1, aka latin-1, or it could be the closely-related Windows variant cp1252. To do this kind of thing correctly you need to determine the actual encoding that was used to create the data.

If you want to output "Test\x80Test2\x81" and have the hex dump look like this:

00000000  54 65 73 74 80 54 65 73  74 32 81                 |Test.Test2.|

You can do

import sys
s = "Test\x80Test2\x81"
sys.stdout.buffer.write(s.encode('latin1'))

This works because Latin-1 is a subset of Unicode. Here's a quick demo:

import binascii

a = ''.join([chr(i) for i in range(256)])
b = a.encode('latin1')
print(binascii.hexlify(b))

output

b'000102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f202122232425262728292a2b2c2d2e2f303132333435363738393a3b3c3d3e3f404142434445464748494a4b4c4d4e4f505152535455565758595a5b5c5d5e5f606162636465666768696a6b6c6d6e6f707172737475767778797a7b7c7d7e7f808182838485868788898a8b8c8d8e8f909192939495969798999a9b9c9d9e9fa0a1a2a3a4a5a6a7a8a9aaabacadaeafb0b1b2b3b4b5b6b7b8b9babbbcbdbebfc0c1c2c3c4c5c6c7c8c9cacbcccdcecfd0d1d2d3d4d5d6d7d8d9dadbdcdddedfe0e1e2e3e4e5e6e7e8e9eaebecedeeeff0f1f2f3f4f5f6f7f8f9fafbfcfdfeff'

However, if you're actually working with binary data then you shouldn't be storing it in text strings in the first place, you should be using bytes, or possibly bytearray. The sane way to produce the b bytes string from my previous example is to do

b = bytes(range(256))

And if you have a bytes object like b"Test\x80Test2\x81" you can dump those bytes to stdout with

sys.stdout.buffer.write(b"Test\x80Test2\x81")
🌐
Teleport
goteleport.com › home › resources › tools › utf-8 encoder | instantly transform text to utf-8 hex
UTF-8 Encoder | Instantly Transform Text to UTF-8 Hex | Teleport
For text heavy in non-ASCII characters, UTF-8 can be less space-efficient than encodings like UTF-16. And in certain scenarios, the variable-width nature of UTF-8 can complicate string processing and indexing operations. Python has excellent built-in support for UTF-8.