To include Unicode characters in your Python source code, you can use Unicode escape characters in the form \u0123 in your string. In Python 2.x, you also need to prefix the string literal with 'u'.

Here's an example running in the Python 2.x interactive console:

>>> print u'\u0420\u043e\u0441\u0441\u0438\u044f'
Россия

In Python 2, prefixing a string with 'u' declares them as Unicode-type variables, as described in the Python Unicode documentation.

In Python 3, the 'u' prefix is now optional:

>>> print('\u0420\u043e\u0441\u0441\u0438\u044f')
Россия

If running the above commands doesn't display the text correctly for you, perhaps your terminal isn't capable of displaying Unicode characters.

These examples use Unicode escapes (\u...), which allows you to print Unicode characters while keeping your source code as plain ASCII. This can help when working with the same source code on different systems. You can also use Unicode characters directly in your Python source code (e.g. print u'Россия' in Python 2), if you are confident all your systems handle Unicode files properly.

For information about reading Unicode data from a file, see this answer:

Character reading from file in Python

Answer from Matt Ryall on Stack Overflow
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
To summarize the previous section: a Unicode string is a sequence of code points, which are numbers from 0 through 0x10FFFF (1,114,111 decimal). This sequence of code points needs to be represented in memory as a set of code units, and code ...
🌐
GeeksforGeeks
geeksforgeeks.org › python › working-with-unicode-in-python
Working with Unicode in Python - GeeksforGeeks
July 23, 2025 - Below, code prints False because Python strings do not consider the two characters identical. Normalization becomes crucial when dealing with combined characters, as shown in the example with strings s1 and s2. ... Python's unicodedata module provides the normalize() function for normalizing Unicode strings.
Discussions

Explain it like I'm five: Python and Unicode?
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character. More on reddit.com
🌐 r/Python
60
106
June 12, 2013
Malicious Actors Use Unicode Support in Python to Evade Detection
i thought it would be some sort of complex obfuscation but they're simply using weird characters to avoid automated code inspection, it's not a python issue really More on reddit.com
🌐 r/Python
67
350
March 23, 2023
Julia in VS Code: how to have Unicode support
You can install the Julia extension, which includes the Julia language server. It provides code suggestions, type information and also lets you use REPL-like LaTeX for symbols. More on reddit.com
🌐 r/Julia
11
8
October 23, 2021
Replacing literal '\u****' in string with corresponding Unicode character
Maybe you have to do this: s = 'blah\\x2Ddude' s.encode().decode('unicode-escape') print(s) 'blah-dude' There's going to be some encoding or decoding that'll make it prettier. The unicode-escape codec can transform embedded Unicode escapes. The string needs to be a byte string, however, hence the .encode() first. (from google) More on reddit.com
🌐 r/learnpython
4
8
March 13, 2020
Top answer
1 of 10
175

To include Unicode characters in your Python source code, you can use Unicode escape characters in the form \u0123 in your string. In Python 2.x, you also need to prefix the string literal with 'u'.

Here's an example running in the Python 2.x interactive console:

>>> print u'\u0420\u043e\u0441\u0441\u0438\u044f'
Россия

In Python 2, prefixing a string with 'u' declares them as Unicode-type variables, as described in the Python Unicode documentation.

In Python 3, the 'u' prefix is now optional:

>>> print('\u0420\u043e\u0441\u0441\u0438\u044f')
Россия

If running the above commands doesn't display the text correctly for you, perhaps your terminal isn't capable of displaying Unicode characters.

These examples use Unicode escapes (\u...), which allows you to print Unicode characters while keeping your source code as plain ASCII. This can help when working with the same source code on different systems. You can also use Unicode characters directly in your Python source code (e.g. print u'Россия' in Python 2), if you are confident all your systems handle Unicode files properly.

For information about reading Unicode data from a file, see this answer:

Character reading from file in Python

2 of 10
54

Print a unicode character in Python:

Print a unicode character directly from python interpreter:

el@apollo:~$ python
Python 2.7.3
>>> print u'\u2713'
✓

Unicode character u'\u2713' is a checkmark. The interpreter prints the checkmark on the screen.

Print a unicode character from a python script:

Put this in test.py:

#!/usr/bin/python
print("here is your checkmark: " + u'\u2713');

Run it like this:

el@apollo:~$ python test.py
here is your checkmark: ✓

If it doesn't show a checkmark for you, then the problem could be elsewhere, like the terminal settings or something you are doing with stream redirection.

Store unicode characters in a file:

Save this to file: foo.py:

#!/usr/bin/python -tt
# -*- coding: utf-8 -*-
import codecs
import sys 
UTF8Writer = codecs.getwriter('utf8')
sys.stdout = UTF8Writer(sys.stdout)
print(u'e with obfuscation: é')

Run it and pipe output to file:

python foo.py > tmp.txt

Open tmp.txt and look inside, you see this:

el@apollo:~$ cat tmp.txt 
e with obfuscation: é

Thus you have saved unicode e with a obfuscation mark on it to a file.

🌐
Real Python
realpython.com › python-encodings-guide
Unicode & Character Encodings in Python: A Painless Guide – Real Python
May 20, 2019 - UTF-8 as well as its lesser-used cousins, UTF-16 and UTF-32, are encoding formats for representing Unicode characters as binary data of one or more bytes per character. We’ll discuss UTF-16 and UTF-32 in a moment, but UTF-8 has taken the largest ...
🌐
DigitalOcean
digitalocean.com › community › tutorials › how-to-work-with-unicode-in-python
How To Work with Unicode in Python | DigitalOcean
The tutorial will cover the basics of Unicode in Python and how Python interprets Unicode characters. It covers the concepts of unicodedata and how to use th…
🌐
GitHub
gist.github.com › seanh › 0a56cd528714496625662dd9136d0cd3
Unicode in Python · GitHub
A from __future__ import unicode_literals turns literal strings into unicode instead of byte strings. In either Python 2 or Python 3 you can force a literal byte string with b"..." or force a literal unicode string with u"...".
🌐
Linode
linode.com › docs › guides › how-to-use-unicode-in-python3
Using Unicode in Python 3 | Linode Docs
March 20, 2023 - To properly understand how Python manages Unicode, you need to understand character processing. Computer files are written using a specific character set. A character set is a collection of characters used within a language or domain. For instance, the written English language maps to a character set containing 26 upper and lower case letters, along with punctuation marks.
Find elsewhere
🌐
Python Cheat Sheet
pythonsheets.com › notes › basic › python-unicode.html
Unicode — Python Cheat Sheet
The main goal of this cheat sheet is to collect some common snippets which are related to Unicode. In Python 3, strings are represented by Unicode instead of bytes.
🌐
Real Python
realpython.com › ref › glossary › unicode
Unicode | Python Glossary – Real Python
These encoding schemes determine ... Fixed-length encoding (4 bytes), simple but space-inefficient · In Python, all strings are Unicode by default....
🌐
Python Reference
python-reference.readthedocs.io › en › latest › docs › functions › unicode.html
unicode — Python Reference (The Right Way) 0.1 documentation
If encoding and/or errors are given, unicode() will decode the object which can either be an 8-bit string or a character buffer using the codec for encoding. The encoding parameter is a string giving the name of an encoding; if the encoding is not known, LookupError is raised.
🌐
Reddit
reddit.com › r/python › explain it like i'm five: python and unicode?
r/Python on Reddit: Explain it like I'm five: Python and Unicode?
June 12, 2013 -

I am seriously confused. And whenever I think I got it, I see some - in my opinion - inconsistent behavior. Can it be consistently explained or is it more art than science?

When do I have to encode/decode("UTF-8")? What does it do exactly? Whats so special about unicode("abc"), or is it identical to u"abc"?

Why, if I'm using a HTML-encoding of UTF8, a python-script with encoding-UTF-8 and a UTF-8 capable shell and have them all interact, do I have to randomly start adding the above functions until stuff accidentally doesn't break anymore? :)

My problem is that while I can code quite well, I have no formal computer science education and don't tend to think in bytes.

Top answer
1 of 5
83
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character.
2 of 5
22
To answer your specific questions: when you encode("UTF-8") you are converting a unicode string to a byte string. It should be called on unicode strings. When you decode("UTF-8") you are converting a byte string to a unicode string. It should be called on byte strings. unicode("abc") is the same as u"abc": they both create a unicode string with three characters. Most of the confusion comes from the fact that python 2 plays fast and loose with unicode strings. It will try and convert between them for you when you mix them together, which yields unexpected results. Python 3 has much more sane behavior: it forces you to encode or decode explicitly to convert between the two. Basically what you need to do to avoid most problems and confusion is to do your encoding/decoding at the input/output boundaries of your program. Decode as soon as you get a byte string from external sources, use unicode strings throughout the program, and encode it just before it leaves.
🌐
Tutorialspoint
tutorialspoint.com › python › python_unicode_system.htm
Python - Unicode System
The default encoding for Python source code is UTF-8. Hence, string may contain literal representation of a Unicode character (3/4) or its Unicode value (\u00BE). var = "3/4" print (var) var = "\u00BE" print (var) This above code will produce ...
🌐
GitHub
gist.github.com › arrowtype › 713dad14fe9a574d58d1aab61ba9b2f0
The basics of working with unicode values in Python · GitHub
I'm not sure you'll need it for what you do, but all Unicode code points have names, and Python can tell you what they are: >>> import unicodedata >>> unicodedata.name("\U0001EE01") 'ARABIC MATHEMATICAL BEH' >>> unicodedata.name("\U0001F4A9") 'PILE OF POO' Copy link · Copy Markdown · I recommend unicodedata2 (https://github.com/mikekap/unicodedata2) instead of the standard library module unicodedata, as the latter one is often not the latest.
🌐
freeCodeCamp
freecodecamp.org › news › a-beginner-friendly-guide-to-unicode-d6d45a903515
A Beginner-Friendly Guide to Unicode in Python
July 18, 2018 - Together, these characters comprise the Unicode character set. Code points are typically written in hexadecimal and prefixed with U+ to denote the connection to Unicode, representing characters from:
🌐
Python
docs.python.org › 3 › c-api › unicode.html
Unicode Objects and Codecs — Python 3.14.7 documentation
UTF-8 representation is created on demand and cached in the Unicode object. ... The Py_UNICODE representation has been removed since Python 3.12 with deprecated APIs.
🌐
B-List
b-list.org › weblog › 2017 › sep › 05 › how-python-does-unicode
How Python does Unicode - James Bennett
September 5, 2017 - To create a str in Python 2, you can use the str() built-in, or string-literal syntax, like so: my_string = 'This is my string.'. To create an instance of unicode, you can use the unicode() built-in, or prefix a string literal with a u, like so: my_unicode = u'This is my Unicode string.'.
🌐
Python for Network Engineers
pyneng.readthedocs.io › en › latest › book › 16_unicode › python_3_unicode.html
Unicode in Python 3 - Python for network engineers
Function ord() returns value of Unicode code for character: ... Bytes are an immutable sequence of bytes. Bytes are denoted in the same way as strings but with addition of letter b before string: In [30]: b1 = b'\xd0\xb4\xd0\xb0' In [31]: b2 = b"\xd0\xb4\xd0\xb0" In [32]: b3 = b'''\xd0\xb4...
🌐
My Docs
codebay.ai › welcome to codebay › glossaries › miscellaneous › unicode
Unicode | Codebay
April 7, 2024 - Unicode is a standard in computing that allows computers to consistently represent and manipulate text expressed in any of the world’s writing systems. In Python, the term “Unicode” refers to the built-in data type that holds such text.