Update: Python 3

In Python 3, Unicode strings are the default. The type str is a collection of Unicode code points, and the type bytes is used for representing collections of 8-bit integers (often interpreted as ASCII characters).

Here is the code from the question, updated for Python 3:

>>> my_str = 'A unicode \u018e string \xf1' # no need for "u" prefix
# the escape sequence "\u" denotes a Unicode code point (in hex)
>>> my_str
'A unicode Ǝ string ñ'
# the Unicode code points U+018E and U+00F1 were displayed
# as their corresponding glyphs
>>> my_bytes = my_str.encode('utf-8') # convert to a bytes object
>>> my_bytes
b'A unicode \xc6\x8e string \xc3\xb1'
# the "b" prefix means a bytes literal
# the escape sequence "\x" denotes a byte using its hex value
# the code points U+018E and U+00F1 were encoded as 2-byte sequences
>>> my_str2 = my_bytes.decode('utf-8') # convert back to str
>>> my_str2 == my_str
True

Working with files:

>>> f = open('foo.txt', 'r') # text mode (Unicode)
>>> # the platform's default encoding (e.g. UTF-8) is used to decode the file
>>> # to set a specific encoding, use open('foo.txt', 'r', encoding="...")
>>> for line in f:
>>>     # here line is a str object

>>> f = open('foo.txt', 'rb') # "b" means binary mode (bytes)
>>> for line in f:
>>>     # here line is a bytes object

Historical answer: Python 2

In Python 2, the str type was a collection of 8-bit characters (like Python 3's bytes type). The English alphabet can be represented using these 8-bit characters, but symbols such as Ω, и, ±, and ♠ cannot.

Unicode is a standard for working with a wide range of characters. Each symbol has a code point (a number), and these code points can be encoded (converted to a sequence of bytes) using a variety of encodings.

UTF-8 is one such encoding. The low code points are encoded using a single byte, and higher code points are encoded as sequences of bytes.

To allow working with Unicode characters, Python 2 has a unicode type which is a collection of Unicode code points (like Python 3's str type). The line ustring = u'A unicode \u018e string \xf1' creates a Unicode string with 20 characters.

When the Python interpreter displays the value of ustring, it escapes two of the characters (Ǝ and ñ) because they are not in the standard printable range.

The line s = unistring.encode('utf-8') encodes the Unicode string using UTF-8. This converts each code point to the appropriate byte or sequence of bytes. The result is a collection of bytes, which is returned as a str. The size of s is 22 bytes, because two of the characters have high code points and are encoded as a sequence of two bytes rather than a single byte.

When the Python interpreter displays the value of s, it escapes four bytes that are not in the printable range (\xc6, \x8e, \xc3, and \xb1). The two pairs of bytes are not treated as single characters like before because s is of type str, not unicode.

The line t = unicode(s, 'utf-8') does the opposite of encode(). It reconstructs the original code points by looking at the bytes of s and parsing byte sequences. The result is a Unicode string.

The call to codecs.open() specifies utf-8 as the encoding, which tells Python to interpret the contents of the file (a collection of bytes) as a Unicode string that has been encoded using UTF-8.

Answer from tom on Stack Overflow
🌐
Microsoft Learn
learn.microsoft.com › en-us › windows › win32 › api › ntdef › ns-ntdef-_unicode_string
_UNICODE_STRING (ntdef.h) - Win32 apps | Microsoft Learn
February 22, 2024 - Pointer to a buffer used to contain a string of wide characters. The UNICODE_STRING structure is used to pass Unicode strings.
Top answer
1 of 2
61

Update: Python 3

In Python 3, Unicode strings are the default. The type str is a collection of Unicode code points, and the type bytes is used for representing collections of 8-bit integers (often interpreted as ASCII characters).

Here is the code from the question, updated for Python 3:

>>> my_str = 'A unicode \u018e string \xf1' # no need for "u" prefix
# the escape sequence "\u" denotes a Unicode code point (in hex)
>>> my_str
'A unicode Ǝ string ñ'
# the Unicode code points U+018E and U+00F1 were displayed
# as their corresponding glyphs
>>> my_bytes = my_str.encode('utf-8') # convert to a bytes object
>>> my_bytes
b'A unicode \xc6\x8e string \xc3\xb1'
# the "b" prefix means a bytes literal
# the escape sequence "\x" denotes a byte using its hex value
# the code points U+018E and U+00F1 were encoded as 2-byte sequences
>>> my_str2 = my_bytes.decode('utf-8') # convert back to str
>>> my_str2 == my_str
True

Working with files:

>>> f = open('foo.txt', 'r') # text mode (Unicode)
>>> # the platform's default encoding (e.g. UTF-8) is used to decode the file
>>> # to set a specific encoding, use open('foo.txt', 'r', encoding="...")
>>> for line in f:
>>>     # here line is a str object

>>> f = open('foo.txt', 'rb') # "b" means binary mode (bytes)
>>> for line in f:
>>>     # here line is a bytes object

Historical answer: Python 2

In Python 2, the str type was a collection of 8-bit characters (like Python 3's bytes type). The English alphabet can be represented using these 8-bit characters, but symbols such as Ω, и, ±, and ♠ cannot.

Unicode is a standard for working with a wide range of characters. Each symbol has a code point (a number), and these code points can be encoded (converted to a sequence of bytes) using a variety of encodings.

UTF-8 is one such encoding. The low code points are encoded using a single byte, and higher code points are encoded as sequences of bytes.

To allow working with Unicode characters, Python 2 has a unicode type which is a collection of Unicode code points (like Python 3's str type). The line ustring = u'A unicode \u018e string \xf1' creates a Unicode string with 20 characters.

When the Python interpreter displays the value of ustring, it escapes two of the characters (Ǝ and ñ) because they are not in the standard printable range.

The line s = unistring.encode('utf-8') encodes the Unicode string using UTF-8. This converts each code point to the appropriate byte or sequence of bytes. The result is a collection of bytes, which is returned as a str. The size of s is 22 bytes, because two of the characters have high code points and are encoded as a sequence of two bytes rather than a single byte.

When the Python interpreter displays the value of s, it escapes four bytes that are not in the printable range (\xc6, \x8e, \xc3, and \xb1). The two pairs of bytes are not treated as single characters like before because s is of type str, not unicode.

The line t = unicode(s, 'utf-8') does the opposite of encode(). It reconstructs the original code points by looking at the bytes of s and parsing byte sequences. The result is a Unicode string.

The call to codecs.open() specifies utf-8 as the encoding, which tells Python to interpret the contents of the file (a collection of bytes) as a Unicode string that has been encoded using UTF-8.

2 of 2
-5

Python supports the string type and the unicode type. A string is a sequence of chars while a unicode is a sequence of "pointers". The unicode is an in-memory representation of the sequence and every symbol on it is not a char but a number (in hex format) intended to select a char in a map. So a unicode var does not have encoding because it does not contain chars.

🌐
TensorFlow
tensorflow.org › text › unicode strings
Unicode strings | Text | TensorFlow
July 19, 2024 - Unicode is a standard encoding ... a unique integer code point between 0 and 0x10FFFF. A Unicode string is a sequence of zero or more code points....
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.8 documentation
To summarize the previous section: a Unicode string is a sequence of code points, which are numbers from 0 through 0x10FFFF (1,114,111 decimal). This sequence of code points needs to be represented in memory as a set of code units, and code ...
🌐
Reddit
reddit.com › r/python › explain it like i'm five: python and unicode?
r/Python on Reddit: Explain it like I'm five: Python and Unicode?
June 12, 2013 -

I am seriously confused. And whenever I think I got it, I see some - in my opinion - inconsistent behavior. Can it be consistently explained or is it more art than science?

When do I have to encode/decode("UTF-8")? What does it do exactly? Whats so special about unicode("abc"), or is it identical to u"abc"?

Why, if I'm using a HTML-encoding of UTF8, a python-script with encoding-UTF-8 and a UTF-8 capable shell and have them all interact, do I have to randomly start adding the above functions until stuff accidentally doesn't break anymore? :)

My problem is that while I can code quite well, I have no formal computer science education and don't tend to think in bytes.

Top answer
1 of 5
83
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character.
2 of 5
22
To answer your specific questions: when you encode("UTF-8") you are converting a unicode string to a byte string. It should be called on unicode strings. When you decode("UTF-8") you are converting a byte string to a unicode string. It should be called on byte strings. unicode("abc") is the same as u"abc": they both create a unicode string with three characters. Most of the confusion comes from the fact that python 2 plays fast and loose with unicode strings. It will try and convert between them for you when you mix them together, which yields unexpected results. Python 3 has much more sane behavior: it forces you to encode or decode explicitly to convert between the two. Basically what you need to do to avoid most problems and confusion is to do your encoding/decoding at the input/output boundaries of your program. Decode as soon as you get a byte string from external sources, use unicode strings throughout the program, and encode it just before it leaves.
Find elsewhere
🌐
Unicode
unicode-org.github.io › icu › userguide › strings
Chars and Strings | ICU Documentation
The Unicode standard defines a default encoding based on 16-bit code units. Since ICU has moved to C++11, it uses the standard type char16_t. Previously, ICU defined its own UChar type to be an unsigned 16-bit integer type. char16_t or UChar is the base type for character arrays for strings in ICU.
🌐
MetaCPAN
metacpan.org › pod › Unicode::String
Unicode::String - String of Unicode characters (UTF-16BE) - metacpan.org
The string passed should be in the ISO-8859-1 encoding. The sample "µm" is "\xB5m" in this encoding. Characters outside the "\x00" .. "\xFF" range are simply removed from the return value of the latin1() method. If you want more control over the mapping from Unicode to ISO-8859-1, use the Unicode::Map8 class.
🌐
Medium
medium.com › free-code-camp › a-beginner-friendly-guide-to-unicode-d6d45a903515
A Beginner-Friendly Guide to Unicode 😎
July 18, 2018 - Internally, Sublime Text still represents each of our “Windows-1258 decoded” characters as a Unicode code point, as we see below when we fire up the console: >>> view.encoding() 'Vietnamese (Windows 1258)'# Python 3 strings are "immutable sequences of Unicode code points" >>> type(view.substr(0)) <class 'str'>>>> view.substr(0) 'đ' >>> view.substr(1) 'Ÿ' >>> view.substr(2) '˜' >>> view.substr(3) '®'>>> ['U+x' % ord(view.substr(x)) for x in range(0, 4)] ['U+0111', 'U+0178', 'U+02dc', 'U+00ae']
🌐
Microsoft Learn
learn.microsoft.com › en-us › windows › win32 › api › subauth › ns-subauth-unicode_string
UNICODE_STRING (subauth.h) - Win32 apps | Microsoft Learn
February 22, 2024 - The UNICODE_STRING structure is used by various Local Security Authority (LSA) functions to specify a Unicode string.
🌐
Hex
hex.pm › packages › unicode_string
unicode_string | Hex
Unicode locale-aware case folding, case mapping (upcase, downcase and titlecase) case-insensitive equality as well as word, line, grapheme and sentence breaking
🌐
JUCE
juce.com › home › technical deep dive: unicode literals
Technical Deep Dive: Unicode Literals - JUCE
April 29, 2024 - Clearly MSVC has not encoded the string the way we might have expected. It turns out that Unicode string literals only refer to what the standard calls the “execution character set”. That is the coded character set (or encoding) that will be used to store the string in the binary.
🌐
OSR
community.osr.com › ntdev
Displaying a UNICODE_STRING - NTDEV - OSR Developer Community
March 18, 2008 - Hi there. I am fairly new to driver development and have been fighting with this problem for a few days now…Might be really obvious and probably is in my experience but I cannot get a complete UNICODE string to be displayed. Every time I try to output the data I only get the first character ?
🌐
Giodicanio
giodicanio.com › tag › unicode_string
UNICODE_STRING – Giovanni Dicanio's Blog
May 23, 2023 - A brief introduction to the UNICODE_STRING structure used in Windows kernel mode programming.
🌐
Geoffchappell
geoffchappell.com › studies › windows › km › ntoskrnl › inc › shared › ntdef › unicode_string.htm
UNICODE_STRING
The UNICODE_STRING structure keeps the address and size of a Unicode string, presumably to save on passing them as separate arguments for subsequent work with the string and to save on repeated re-reading of the whole string to rediscover its size.
🌐
GitHub
github.com › hillu › go-ntdll › blob › master › unicode_string.go
go-ntdll/unicode_string.go at master · hillu/go-ntdll
typedef struct _UNICODE_STRING { USHORT Length; USHORT MaximumLength; PWSTR Buffer; } UNICODE_STRING, *PUNICODE_STRING; */ · // String converts the UTF-16-encoded string stored in a UnicodeString · // to UTF-8 and returns that as a Go string. func (u UnicodeString) String() string { var s []uint16 ·
Author: hillu
🌐
Unicode
home.unicode.org
Unicode – The World Standard for Text and Emoji
The Unicode Consortium 611 Gateway Blvd. Suite 120 South San Francisco, CA 94080 USA +1-408-401-8915 · Issues & Submissions Contact Form
🌐
Pinvoke
pinvoke.dev › foundation › unicode_string
UNICODE_STRING | P/Invoke
public struct UNICODE_STRING { public ushort Length; public ushort MaximumLength; public PWSTR Buffer; }