First, obtain the integer value from the string of a, noting that a is expressed in hexadecimal:
a_int = int(a, 16)
Next, convert this int to a character. In python 2 you need to use the unichr method to do this, because the chr method can only deal with ASCII characters:
a_chr = unichr(a_int)
Whereas in python 3 you can just use the chr method for any character:
a_chr = chr(a_int)
So, in python 3, the full command is:
a_chr = chr(int(a, 16))
Answer from Rob Bricheno on Stack Overflowutf 8 - Read hex characters and convert them to utf-8 using python 3 - Stack Overflow
utf 8 - Python how to decode unicode with hex characters - Stack Overflow
Convert unicode codepoint to UTF8 hex in python - Stack Overflow
utf 8 - python encode/decode hex string to utf-8 string - Stack Overflow
Is hex to ASCII different from hex to UTF-8?
Why does my hex decode to `é` instead of `é`?
Can I tell which encoding a hex string uses just by looking?
Try(Python 3.x):
import codecs
codecs.decode("707974686f6e2d666f72756d2e696f", "hex").decode('utf-8')
From here.
Assuming Python 2.6,
>>> print('kitap ara\xfet\xfdrmas\xfd'.decode('iso-8859-9'))
kitap araştırması
>>> 'kitap ara\xfet\xfdrmas\xfd'.decode('iso-8859-9').encode('utf-8')
'kitap ara\xc5\x9ft\xc4\xb1rmas\xc4\xb1'
The problem with
msg = u'\xe3\x80\x90\xe4\xb8\xad\xe5\xad\x97\xe3\x80\x91'
result = msg.decode('utf8')
is that you are trying to decode Unicode. That doesn't really make sense. You can encode from Unicode to some type of encoding, or you can decode a byte string to Unicode.
When you do
msg.decode('utf8')
Python 2 sees that msg is Unicode. It knows that it can't decode Unicode so it "helpfully" assumes that you want to encode msg with the default ASCII codec so the result of that transformation can be decoded to Unicode using the UTF-8 codec. Python 3 behaves much more sensibly: that code would simply fail with
AttributeError: 'str' object has no attribute 'decode'
The technique given in kennytm's answer:
msg.encode('latin1').decode('utf-8')
works because the Unicode codepoints less than 256 correspond directly to the characters in the Latin1 encoding (aka ISO 8859-1).
Here's some Python 2 code that illustrates this:
for i in xrange(256):
lat = chr(i)
uni = unichr(i)
assert lat == uni.encode('latin1')
assert lat.decode('latin1') == uni
And here is the equivalent Python 3 code:
for i in range(256):
lat = bytes([i])
uni = chr(i)
assert lat == uni.encode('latin1')
assert lat.decode('latin1') == uni
You may find this article helpful: Pragmatic Unicode, which was written by SO veteran Ned Batchelder.
Unless you are forced to use Python 2 I strongly advise you to switch to Python 3. It will make handling Unicode far less painful.
Perhaps you should fix the crawl script instead, a Unicode string should contain
u'【中字】'(u'\u3010\u4e2d\u5b57\u3011') already, instead of the raw UTF-8 bytes.To convert
msgto the correct encoding, first you need to turn the wrong Unicode string back to byte string (encode it as Latin-1), then decode it as UTF-8:>>> print msg.encode('latin1').decode('utf-8') 【中字】
Use the built-in function chr() to convert the number to character, then encode that:
>>> chr(int('fd9b', 16)).encode('utf-8')
'\xef\xb6\x9b'
This is the string itself. If you want the string as ASCII hex, you'd need to walk through and convert each character c to hex, using hex(ord(c)) or similar.
Note: If you are still stuck with Python 2, you can use unichr() instead.
here's a complete solution:
>>> ''.join(['{0:x}'.format(ord(x)) for x in unichr(int('FD9B', 16)).encode('utf-8')]).upper()
'EFB69B'
Your data is encoded as UTF-8, which means that you sometimes have to look at more than one byte to get one character. The easiest way to do this is probably to decode your string into a sequence of bytes, and then decode those bytes into a string. Python has built-in features for both:
value = bytes.fromhex("54 C3 BC").decode("utf-8")
The problem is the result of the string
"54 C3 BC 72 20 6F 66 66 65 6E 20 4B 6C 69 6D 61"
is indeed
Tür offen Klima
The proper hex string that result in "Tür offen Klima" is actually:
"54 FC 72 20 6F 66 66 65 6E 20 4B 6C 69 6D 61"
Therefore, the code below would generate the result you expected:
value = ""
for i in "54 FC 72 20 6F 66 66 65 6E 20 4B 6C 69 6D 61".split(" "):
value += chr(int(i, 16))
print(value)
So here's the thing, I'm currently doing some Wikipedia web scrapping and one of the main libs I'm using doesn't handle Hex UTF-8 very well, and some of the links I'm trying to access from Wikipedia contain Hex UTF-8.
For example, the Wikipedia article for the History the US from 1991 to 2008 link is: https://en.wikipedia.org/wiki/History_of_the_United_States_(1991%E2%80%932008)
But once you actually access it, your browser will decode the %E2%80%93 to an actual Unicode character, the en dash (–), so you'll see the link on your web browser as: https://en.wikipedia.org/wiki/History_of_the_United_States_(1991–2008)
So, what I'm looking for is a way to substitute the part hex part of my string (%E2%80%93) for the actual Unicode character (–)
Where ever you decoded the original string, it was likely decoded with latin-1 or a close relative. Since latin-1 is the first 256 codepoints of Unicode, this works:
>>> s = u'Gaga\xe2\x80\x99s'
>>> s.encode('latin-1').decode('utf8')
u'Gaga\u2019s'
s = u'Gaga\xe2\x80\x99s'
t = u'Gaga\u2019s'
x = s.encode('raw-unicode-escape').decode('utf-8')
assert x==t
print(x)
yields
Gaga’s
When you do string.encode('utf-8'), it changes to hex notation.
But if you print it, you will get original unicode string.
If you want the hex notation you can get it like this with repr() function:
>>> print u'\u041a\u0418\u0421\u0410'.encode('utf-8')
КИСА
>>> print repr(u'\u041a\u0418\u0421\u0410'.encode('utf-8'))
'\xd0\x9a\xd0\x98\xd0\xa1\xd0\x90'
you can also try:
print "hex_signature : ",'\\X'.join(x.encode("hex") for x in signature)
The join function is used with the separator '\X' so that for each byte to hex conversion the \X is inserted. The join function is done for each byte of the variable signature in a loop. Everything is joined/concatenated and printed.
Python 3 does the right thing of inserting a character into a str which is string of characters, not a byte sequence.
UTF8 is the default encoding. If you need to insert a byte, a different encoding where that character is represented as a byte is needed.
$ PYTHONIOENCODING=iso-8859-1 python3 -c 'print("\x80")' | xxd
00000000: 800a
PYTHONIOENCODING
If this is set before running the interpreter, it overrides the encoding used for stdin/stdout/stderr, in the syntax encodingname:errorhandler. Both the encodingname and the :errorhandler parts are optional and have the same meaning as in str.encode().
If you want to output raw bytes in Python 3 you shouldn't be using the print function, since it's for outputting text in your default encoding. Instead, you can use sys.stdout.buffer.write.
ASCII is a 7 bit encoding, so if your so-called ASCII contains characters like b'\x80' it's not legal ASCII. Perhaps your data is actually encoded with iso-8859-1, aka latin-1, or it could be the closely-related Windows variant cp1252. To do this kind of thing correctly you need to determine the actual encoding that was used to create the data.
If you want to output "Test\x80Test2\x81" and have the hex dump look like this:
00000000 54 65 73 74 80 54 65 73 74 32 81 |Test.Test2.|
You can do
import sys
s = "Test\x80Test2\x81"
sys.stdout.buffer.write(s.encode('latin1'))
This works because Latin-1 is a subset of Unicode. Here's a quick demo:
import binascii
a = ''.join([chr(i) for i in range(256)])
b = a.encode('latin1')
print(binascii.hexlify(b))
output
b'000102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f202122232425262728292a2b2c2d2e2f303132333435363738393a3b3c3d3e3f404142434445464748494a4b4c4d4e4f505152535455565758595a5b5c5d5e5f606162636465666768696a6b6c6d6e6f707172737475767778797a7b7c7d7e7f808182838485868788898a8b8c8d8e8f909192939495969798999a9b9c9d9e9fa0a1a2a3a4a5a6a7a8a9aaabacadaeafb0b1b2b3b4b5b6b7b8b9babbbcbdbebfc0c1c2c3c4c5c6c7c8c9cacbcccdcecfd0d1d2d3d4d5d6d7d8d9dadbdcdddedfe0e1e2e3e4e5e6e7e8e9eaebecedeeeff0f1f2f3f4f5f6f7f8f9fafbfcfdfeff'
However, if you're actually working with binary data then you shouldn't be storing it in text strings in the first place, you should be using bytes, or possibly bytearray. The sane way to produce the b bytes string from my previous example is to do
b = bytes(range(256))
And if you have a bytes object like b"Test\x80Test2\x81" you can dump those bytes to stdout with
sys.stdout.buffer.write(b"Test\x80Test2\x81")