Another solution is to use the name aliases and print them using the string literal \N
print('\N{grinning face with smiling eyes}')
Current list of name aliases can be found at: https://unicode.org/Public/15.0.0/ucd/UnicodeData-15.0.0d6.txt
Or the Folder for everything here: https://unicode.org/Public/
Answer from nth-attempt on Stack OverflowAnother solution is to use the name aliases and print them using the string literal \N
print('\N{grinning face with smiling eyes}')
Current list of name aliases can be found at: https://unicode.org/Public/15.0.0/ucd/UnicodeData-15.0.0d6.txt
Or the Folder for everything here: https://unicode.org/Public/
>>> print u'\U0001f604'.encode('unicode-escape')
\U0001f604
What is Unicode Emoji?
Converting emojis to Unicode and vice versa in python 3 - Stack Overflow
python 3.x - Python3 emoji characters as unicode - Stack Overflow
Emoji's in Python???
What is the purpose of Unicode Emoji?
What is the difference between Unicode Emoji and ASCII?
Can I use Unicode Emoji in programming?
'😀' is already a Unicode object. UTF-8 is not Unicode, it's a byte encoding for Unicode. To get the codepoint number of a Unicode character, you can use the ord function. And to print it in the form you want you can format it as hex. Like this:
s = '😀'
print('U+{:X}'.format(ord(s)))
output
U+1F600
If you have Python 3.6+, you can make it even shorter (and more efficient) by using an f-string:
s = '😀'
print(f'U+{ord(s):X}')
BTW, if you want to create a Unicode escape sequence like '\U0001F600' there's the 'unicode-escape' codec. However, it returns a bytes string, and you may wish to convert that back to text. You could use the 'UTF-8' codec for that, but you might as well just use the 'ASCII' codec, since it's guaranteed to only contain valid ASCII.
s = '😀'
print(s.encode('unicode-escape'))
print(s.encode('unicode-escape').decode('ASCII'))
output
b'\\U0001f600'
\U0001f600
I suggest you take a look at this short article by Stack Overflow co-founder Joel Spolsky The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!).
sentence = "Head-Up Displays (HUD)💻 for #automotive🚗 sector\n \nThe #UK-based #startup🚀 Envisics got €42 million #funding💰 from l… "
print("normal sentence - ", sentence)
uc_sentence = sentence.encode('unicode-escape')
print("\n\nunicode represented sentence - ", uc_sentence)
decoded_sentence = uc_sentence.decode('unicode-escape')
print("\n\ndecoded sentence - ", decoded_sentence)
output
normal sentence - Head-Up Displays (HUD)💻 for #automotive🚗 sector
The #UK-based #startup🚀 Envisics got €42 million #funding💰 from l…
unicode represented sentence - b'Head-Up Displays (HUD)\\U0001f4bb for #automotive\\U0001f697 sector\\n \\nThe #UK-based #startup\\U0001f680 Envisics got \\u20ac42 million #funding\\U0001f4b0 from l\\u2026 '
decoded sentence - Head-Up Displays (HUD)💻 for #automotive🚗 sector
The #UK-based #startup🚀 Envisics got €42 million #funding💰 from l…
Here's some code that will take any character that maps into two UTF-16 words and convert it to a hex sequence.
s = '\U0001f62c \U0001f60e hello'
def pairup(b):
return [(b[i] << 8 | b[i+1]) for i in range(0, len(b), 2)]
def utf16(c):
e = c.encode('utf_16_be')
return ''.join(chr(x) for x in pairup(e))
u = ''.join(utf16(c) for c in s)
print(repr(u))
print(u[0] == '\ud83d' and u[1] == '\ude2c')
print(len(u))
'\ud83d\ude2c \ud83d\ude0e hello'
True
11
I thought this was going to be a no-brainer, but it turned out to be trickier than I expected. Especially since I didn't understand the problem properly the first time through.
It is not clear why do you need it but here's how you could represent non-BMP Unicode characters as surrogate pairs:
#!/usr/bin/env python3
import re
def as_surrogates(astral):
b = astral.group().encode('utf-16be')
return ''.join([b[i:i+2].decode('utf-16be', 'surrogatepass')
for i in range(0, len(b), 2)])
s = '\U0001f62c \U0001f60e hello'
u = re.sub(r'[^\u0000-\uFFFF]+', as_surrogates, s)
print(ascii(u))
assert u.encode('utf-16', 'surrogatepass').decode('utf-16') == s
Output
'\ud83d\ude2c \ud83d\ude0e hello'
I recently discovered that I can store an emoji as a variable and respectively use that.
E.G fruits = ['🍌', '🍒', '🍐', '🍈', '🍇', '🍊', '🍉']
When I looked it up it said I have to download an emoji package but I never did this. So why can these be taken in as a character?