The standard module unicodedata defines a lot of properties, but not everything. A quick peek at its source confirms this.

Fortunately unicodedata.txt, the data file where this comes from, is not hard to parse. Each line consists of exactly 15 elements, ; separated, which makes it ideal for parsing. Using the description of the elements on ftp://ftp.unicode.org/Public/3.0-Update/UnicodeData-3.0.0.html, you can create a few classes to encapsulate the data. I've taken the names of the class elements from that list; the meaning of each of the elements is explained on that same page.

Make sure to download ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt and ftp://ftp.unicode.org/Public/UNIDATA/Blocks.txt first, and put them inside the same folder as this program.

Code (tested with Python 2.7 and 3.6):

# -*- coding: utf-8 -*-

class UnicodeCharacter:
    def __init__(self):
        self.code = 0
        self.name = 'unnamed'
        self.category = ''
        self.combining = ''
        self.bidirectional = ''
        self.decomposition = ''
        self.asDecimal = None
        self.asDigit = None
        self.asNumeric = None
        self.mirrored = False
        self.uc1Name = None
        self.comment = ''
        self.uppercase = None
        self.lowercase = None
        self.titlecase = None
        self.block = None

    def __getitem__(self, item):
        return getattr(self, item)

    def __repr__(self):
        return '{'+self.name+'}'

class UnicodeBlock:
    def __init__(self):
        self.first = 0
        self.last = 0
        self.name = 'unnamed'

    def __repr__(self):
        return '{'+self.name+'}'

class BlockList:
    def __init__(self):
        self.blocklist = []
        with open('Blocks.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.split(';')
                    block = UnicodeBlock()
                    block.name = rawdata[1].strip()
                    rawdata = rawdata[0].split('..')
                    block.first = int(rawdata[0],16)
                    block.last = int(rawdata[1],16)
                    self.blocklist.append(block)
            # make 100% sure it's sorted, for quicker look-up later
            # (it is usually sorted in the file, but better make sure)
            self.blocklist.sort (key=lambda x: block.first)

    def lookup(self,code):
        for item in self.blocklist:
            if code >= item.first and code <= item.last:
                return item.name
        return None

class UnicodeList:
    """UnicodeList loads Unicode data from the external files
    'UnicodeData.txt' and 'Blocks.txt', both available at unicode.org

    These files must appear in the same directory as this program.

    UnicodeList is a new interpretation of the standard library
    'unicodedata'; you may first want to check if its functionality
    suffices.

    As UnicodeList loads its data from an external file, it does not depend
    on the local build from Python (in which the Unicode data gets frozen
    to the then 'current' version).

    Initialize with

        uclist = UnicodeList()
    """
    def __init__(self):

        # we need this first
        blocklist = BlockList()
        bpos = 0

        self.codelist = []
        with open('UnicodeData.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.strip().split(';')
                    parsed = UnicodeCharacter()
                    parsed.code = int(rawdata[0],16)
                    parsed.characterName = rawdata[1]
                    parsed.category = rawdata[2]
                    parsed.combining = rawdata[3]
                    parsed.bidirectional = rawdata[4]
                    parsed.decomposition = rawdata[5]
                    parsed.asDecimal = int(rawdata[6]) if rawdata[6] else None
                    parsed.asDigit = int(rawdata[7]) if rawdata[7] else None
                    # the following value may contain a slash:
                    #  ONE QUARTER ... 1/4
                    # let's make it Python 2.7 compatible :)
                    if '/' in rawdata[8]:
                        rawdata[8] = rawdata[8].replace('/','./')
                        parsed.asNumeric = eval(rawdata[8])
                    else:
                        parsed.asNumeric = int(rawdata[8]) if rawdata[8] else None
                    parsed.mirrored = rawdata[9] == 'Y'
                    parsed.uc1Name = rawdata[10]
                    parsed.comment = rawdata[11]
                    parsed.uppercase = int(rawdata[12],16) if rawdata[12] else None
                    parsed.lowercase = int(rawdata[13],16) if rawdata[13] else None
                    parsed.titlecase = int(rawdata[14],16) if rawdata[14] else None
                    while bpos < len(blocklist.blocklist) and parsed.code > blocklist.blocklist[bpos].last:
                        bpos += 1
                    parsed.block = blocklist.blocklist[bpos].name if bpos < len(blocklist.blocklist) and parsed.code >= blocklist.blocklist[bpos].first else None
                    self.codelist.append(parsed)

    def find_code(self,codepoint):
        """Find the Unicode information for a codepoint (as int).

        Returns:
            a UnicodeCharacter class object or None.
        """
        # the list is unlikely to contain duplicates but I have seen Unicode.org
        # doing that in similar situations. Again, better make sure.
        val = [x for x in self.codelist if codepoint == x.code]
        return val[0] if val else None

    def find_char(self,str):
        """Find the Unicode information for a codepoint (as character).

        Returns:
            for a single character: a UnicodeCharacter class object or
            None.
            for a multicharacter string: a list of the above, one element
            per character.
        """
        if len(str) > 1:
            result = [self.find_code(ord(x)) for x in str]
            return result
        else:
            return self.find_code(ord(str))

When loaded, you can now look up a character code with

>>> ul = UnicodeList()     # ONLY NEEDED ONCE!
>>> print (ul.find_code(0x204))
{LATIN CAPITAL LETTER E WITH DOUBLE GRAVE}

which by default is shown as the name of a character (Unicode calls this a 'code point'), but you can retrieve other properties as well:

>>> print ('%04X' % uc.find_code(0x204).lowercase)
0205
>>> print (ul.lookup(0x204).block)
Latin Extended-B

and (as long as you don't get a None) even chain them:

>>> print (ul.find_code(ul.find_code(0x204).lowercase))
{LATIN SMALL LETTER E WITH DOUBLE GRAVE}

It does not rely on your particular build of Python; you can always download an updated list from unicode.org and be assured to get the most recent information:

import unicodedata
>>> print (unicodedata.name('\U0001F903'))
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
ValueError: no such name
>>> print (uclist.find_code(0x1f903))
{LEFT HALF CIRCLE WITH FOUR DOTS}

(As tested with Python 3.5.3.)

There are currently two lookup functions defined:

  • find_code(int) looks up character information by codepoint as an integer.
  • find_char(string) looks up character information for the character(s) in string. If there is only one character, it returns a UnicodeCharacter object; if there are more, it returns a list of objects.

After import unicodelist (assuming you saved this as unicodelist.py), you can use

>>> ul = UnicodeList()
>>> hex(ul.find_char(u'Γ¨').code)
'0xe8'

to look up the hex code for any character, and a list comprehension such as

>>> l = [hex(ul.find_char(x).code) for x in 'Hello']
>>> l
['0x48', '0x65', '0x6c', '0x6c', '0x6f']

for longer strings. Note that you don't actually need all of this if all you want is a hex representation of a string! This suffices:

 l = [hex(ord(x)) for x in 'Hello']

The purpose of this module is to give easy access to other Unicode properties. A longer example:

str = 'HΓ©llo...'
dest = ''
for i in str:
    dest += chr(ul.find_char(i).uppercase) if ul.find_char(i).uppercase is not None else i
print (dest)

HÉLLO...

and showing a list of properties for a character per your example:

letter = u'Θ„'
print ('Name > '+ul.find_char(letter).name)
print ('Unicode number > U+%04x' % ul.find_char(letter).code)
print ('Bloc > '+ul.find_char(letter).block)
print ('Lowercase > %s' % chr(ul.find_char(letter).lowercase))

(I left out HTML; these names are not defined in the Unicode standard.)

Answer from Jongware on Stack Overflow
Top answer
1 of 3
5

The standard module unicodedata defines a lot of properties, but not everything. A quick peek at its source confirms this.

Fortunately unicodedata.txt, the data file where this comes from, is not hard to parse. Each line consists of exactly 15 elements, ; separated, which makes it ideal for parsing. Using the description of the elements on ftp://ftp.unicode.org/Public/3.0-Update/UnicodeData-3.0.0.html, you can create a few classes to encapsulate the data. I've taken the names of the class elements from that list; the meaning of each of the elements is explained on that same page.

Make sure to download ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt and ftp://ftp.unicode.org/Public/UNIDATA/Blocks.txt first, and put them inside the same folder as this program.

Code (tested with Python 2.7 and 3.6):

# -*- coding: utf-8 -*-

class UnicodeCharacter:
    def __init__(self):
        self.code = 0
        self.name = 'unnamed'
        self.category = ''
        self.combining = ''
        self.bidirectional = ''
        self.decomposition = ''
        self.asDecimal = None
        self.asDigit = None
        self.asNumeric = None
        self.mirrored = False
        self.uc1Name = None
        self.comment = ''
        self.uppercase = None
        self.lowercase = None
        self.titlecase = None
        self.block = None

    def __getitem__(self, item):
        return getattr(self, item)

    def __repr__(self):
        return '{'+self.name+'}'

class UnicodeBlock:
    def __init__(self):
        self.first = 0
        self.last = 0
        self.name = 'unnamed'

    def __repr__(self):
        return '{'+self.name+'}'

class BlockList:
    def __init__(self):
        self.blocklist = []
        with open('Blocks.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.split(';')
                    block = UnicodeBlock()
                    block.name = rawdata[1].strip()
                    rawdata = rawdata[0].split('..')
                    block.first = int(rawdata[0],16)
                    block.last = int(rawdata[1],16)
                    self.blocklist.append(block)
            # make 100% sure it's sorted, for quicker look-up later
            # (it is usually sorted in the file, but better make sure)
            self.blocklist.sort (key=lambda x: block.first)

    def lookup(self,code):
        for item in self.blocklist:
            if code >= item.first and code <= item.last:
                return item.name
        return None

class UnicodeList:
    """UnicodeList loads Unicode data from the external files
    'UnicodeData.txt' and 'Blocks.txt', both available at unicode.org

    These files must appear in the same directory as this program.

    UnicodeList is a new interpretation of the standard library
    'unicodedata'; you may first want to check if its functionality
    suffices.

    As UnicodeList loads its data from an external file, it does not depend
    on the local build from Python (in which the Unicode data gets frozen
    to the then 'current' version).

    Initialize with

        uclist = UnicodeList()
    """
    def __init__(self):

        # we need this first
        blocklist = BlockList()
        bpos = 0

        self.codelist = []
        with open('UnicodeData.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.strip().split(';')
                    parsed = UnicodeCharacter()
                    parsed.code = int(rawdata[0],16)
                    parsed.characterName = rawdata[1]
                    parsed.category = rawdata[2]
                    parsed.combining = rawdata[3]
                    parsed.bidirectional = rawdata[4]
                    parsed.decomposition = rawdata[5]
                    parsed.asDecimal = int(rawdata[6]) if rawdata[6] else None
                    parsed.asDigit = int(rawdata[7]) if rawdata[7] else None
                    # the following value may contain a slash:
                    #  ONE QUARTER ... 1/4
                    # let's make it Python 2.7 compatible :)
                    if '/' in rawdata[8]:
                        rawdata[8] = rawdata[8].replace('/','./')
                        parsed.asNumeric = eval(rawdata[8])
                    else:
                        parsed.asNumeric = int(rawdata[8]) if rawdata[8] else None
                    parsed.mirrored = rawdata[9] == 'Y'
                    parsed.uc1Name = rawdata[10]
                    parsed.comment = rawdata[11]
                    parsed.uppercase = int(rawdata[12],16) if rawdata[12] else None
                    parsed.lowercase = int(rawdata[13],16) if rawdata[13] else None
                    parsed.titlecase = int(rawdata[14],16) if rawdata[14] else None
                    while bpos < len(blocklist.blocklist) and parsed.code > blocklist.blocklist[bpos].last:
                        bpos += 1
                    parsed.block = blocklist.blocklist[bpos].name if bpos < len(blocklist.blocklist) and parsed.code >= blocklist.blocklist[bpos].first else None
                    self.codelist.append(parsed)

    def find_code(self,codepoint):
        """Find the Unicode information for a codepoint (as int).

        Returns:
            a UnicodeCharacter class object or None.
        """
        # the list is unlikely to contain duplicates but I have seen Unicode.org
        # doing that in similar situations. Again, better make sure.
        val = [x for x in self.codelist if codepoint == x.code]
        return val[0] if val else None

    def find_char(self,str):
        """Find the Unicode information for a codepoint (as character).

        Returns:
            for a single character: a UnicodeCharacter class object or
            None.
            for a multicharacter string: a list of the above, one element
            per character.
        """
        if len(str) > 1:
            result = [self.find_code(ord(x)) for x in str]
            return result
        else:
            return self.find_code(ord(str))

When loaded, you can now look up a character code with

>>> ul = UnicodeList()     # ONLY NEEDED ONCE!
>>> print (ul.find_code(0x204))
{LATIN CAPITAL LETTER E WITH DOUBLE GRAVE}

which by default is shown as the name of a character (Unicode calls this a 'code point'), but you can retrieve other properties as well:

>>> print ('%04X' % uc.find_code(0x204).lowercase)
0205
>>> print (ul.lookup(0x204).block)
Latin Extended-B

and (as long as you don't get a None) even chain them:

>>> print (ul.find_code(ul.find_code(0x204).lowercase))
{LATIN SMALL LETTER E WITH DOUBLE GRAVE}

It does not rely on your particular build of Python; you can always download an updated list from unicode.org and be assured to get the most recent information:

import unicodedata
>>> print (unicodedata.name('\U0001F903'))
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
ValueError: no such name
>>> print (uclist.find_code(0x1f903))
{LEFT HALF CIRCLE WITH FOUR DOTS}

(As tested with Python 3.5.3.)

There are currently two lookup functions defined:

  • find_code(int) looks up character information by codepoint as an integer.
  • find_char(string) looks up character information for the character(s) in string. If there is only one character, it returns a UnicodeCharacter object; if there are more, it returns a list of objects.

After import unicodelist (assuming you saved this as unicodelist.py), you can use

>>> ul = UnicodeList()
>>> hex(ul.find_char(u'Γ¨').code)
'0xe8'

to look up the hex code for any character, and a list comprehension such as

>>> l = [hex(ul.find_char(x).code) for x in 'Hello']
>>> l
['0x48', '0x65', '0x6c', '0x6c', '0x6f']

for longer strings. Note that you don't actually need all of this if all you want is a hex representation of a string! This suffices:

 l = [hex(ord(x)) for x in 'Hello']

The purpose of this module is to give easy access to other Unicode properties. A longer example:

str = 'HΓ©llo...'
dest = ''
for i in str:
    dest += chr(ul.find_char(i).uppercase) if ul.find_char(i).uppercase is not None else i
print (dest)

HÉLLO...

and showing a list of properties for a character per your example:

letter = u'Θ„'
print ('Name > '+ul.find_char(letter).name)
print ('Unicode number > U+%04x' % ul.find_char(letter).code)
print ('Bloc > '+ul.find_char(letter).block)
print ('Lowercase > %s' % chr(ul.find_char(letter).lowercase))

(I left out HTML; these names are not defined in the Unicode standard.)

2 of 3
3

The unicodedata documentation shows how to do most of this.

The Unicode block name is apparently not available but another Stack Overflow question has a solution of sorts and another has some additional approaches using regex.

The uppercase/lowercase mapping and character number information is not particularly Unicode-specific; just use the regular Python string functions.

So in summary

>>> import unicodedata
>>> unicodedata.name('Γ‹')
'LATIN CAPITAL LETTER E WITH DIAERESIS'
>>> 'U+%04X' % ord('Γ‹')
'U+00CB'
>>> '&#%i;' % ord('Γ‹')
'&#203;'
>>> 'Γ‹'.lower()
'Γ«'

The U+%04X formatting is sort-of correct, in that it simply avoids padding and prints the whole hex number for code points with a value higher than 65,535. Note that some other formats require the use of %08X padding in this scenario (notably \U00010000 format in Python).

🌐
python-tcod
python-tcod.readthedocs.io β€Ί en β€Ί latest β€Ί tcod β€Ί charmap-reference.html
Character Table Reference - python-tcod 21.2.1 documentation
Unicode is the Unicode code point as a hexadecimal number. You can use chr to convert these numbers into a string. Character maps such as tcod.tileset.CHARMAP_CP437 are simply a list of Unicode numbers, where the index of the list is the Tile Index. String is the Python string for that character.
Discussions

Formatting a table using unicode symbols in python - Code Review Stack Exchange
The goal of the project is to take in input for a data, perhaps in the future as a csv, then print it out using UNICODE box-drawing characters. The main function declares the necessary data and Table More on codereview.stackexchange.com
🌐 codereview.stackexchange.com
May 15, 2022
Return unicode strings from stored bytestrings
Storing strings in a table in Python 3 with something like: import tables class Tags(tables.IsDescription): tag = tables.StringCol(namelength) handle = tables.open_file('testfile.h5', '... More on github.com
🌐 github.com
16
September 9, 2015
Where can I find a full list of Python Unicode Superscript formatting?

Unicode is not specifically for Python. It is a universal standard. So any unicode table can work, like this and this.

For your request, u"y\u207B\u00B9" will do.

More on reddit.com
🌐 r/learnpython
2
2
June 20, 2018
Explain it like I'm five: Python and Unicode?
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character. More on reddit.com
🌐 r/Python
60
106
June 12, 2013
🌐
Inspired Python
inspiredpython.com β€Ί tip β€Ί python-strings-using-translation-tables-to-make-unicode-text
Python Strings: Using translation tables to make Unicode text β€’ Inspired Python
>>> import string >>> before = string.ascii_lowercase + string.digits >>> after = 'πŸ…πŸ…‘πŸ…’πŸ…“πŸ…”πŸ…•πŸ…–πŸ…—πŸ…˜πŸ…™πŸ…šπŸ…›πŸ…œπŸ…πŸ…žπŸ…ŸπŸ… πŸ…‘πŸ…’πŸ…£πŸ…€πŸ…₯πŸ…¦πŸ…§πŸ…¨πŸ…©β“ΏβžŠβž‹βžŒβžβžŽβžβžβž‘βž’' >>> translation_table = str.maketrans(before, after) >>> 'inspired python'.translate(translation_table) 'πŸ…˜πŸ…πŸ…’πŸ…ŸπŸ…˜πŸ…‘πŸ…”πŸ…“ πŸ…ŸπŸ…¨πŸ…£πŸ…—πŸ…žπŸ…'
🌐
Python documentation
docs.python.org β€Ί 3 β€Ί howto β€Ί unicode.html
Unicode HOWTO β€” Python 3.14.7 documentation
To summarize the previous section: a Unicode string is a sequence of code points, which are numbers from 0 through 0x10FFFF (1,114,111 decimal). This sequence of code points needs to be represented in memory as a set of code units, and code ...
🌐
Python Cheat Sheet
pythonsheets.com β€Ί notes β€Ί basic β€Ί python-unicode.html
Unicode β€” Python Cheat Sheet
In this case, we may acquire unexpected results when we are comparing two strings even though they look alike. Therefore, we can normalize a Unicode form to solve the issue. # python 3 >>> u1 = 'Café' # unicode string >>> u2 = 'Cafe\u0301' >>> u1, u2 ('Café', 'Café') >>> len(u1), len(u2) (4, 5) >>> u1 == u2 False >>> u1.encode('utf-8') # get u1 byte string b'Caf\xc3\xa9' >>> u2.encode('utf-8') # get u2 byte string b'Cafe\xcc\x81' >>> from unicodedata import normalize >>> s1 = normalize('NFC', u1) # get u1 NFC format >>> s2 = normalize('NFC', u2) # get u2 NFC format >>> s1 == s2 True >>> s1.encode('utf-8'), s2.encode('utf-8') (b'Caf\xc3\xa9', b'Caf\xc3\xa9') >>> s1 = normalize('NFD', u1) # get u1 NFD format >>> s2 = normalize('NFD', u2) # get u2 NFD format >>> s1, s2 ('Café', 'Café') >>> s1 == s2 True >>> s1.encode('utf-8'), s2.encode('utf-8') (b'Cafe\xcc\x81', b'Cafe\xcc\x81')
🌐
Real Python
realpython.com β€Ί python-encodings-guide
Unicode & Character Encodings in Python: A Painless Guide – Real Python
May 20, 2019 - You can copy and paste this right into a Python 3 interpreter shell: ... >>> alphabet = 'Ξ±Ξ²Ξ³Ξ΄Ξ΅ΞΆΞ·ΞΈΞΉΞΊΞ»ΞΌΞ½ΞΎΞΏΟ€ΟΟ‚ΟƒΟ„Ο…Ο†Ο‡Οˆ' >>> print(alphabet) Ξ±Ξ²Ξ³Ξ΄Ξ΅ΞΆΞ·ΞΈΞΉΞΊΞ»ΞΌΞ½ΞΎΞΏΟ€ΟΟ‚ΟƒΟ„Ο…Ο†Ο‡Οˆ Β· Besides placing the actual, unescaped Unicode characters in the console, there are other ways to type Unicode strings as well.
🌐
GitHub
gist.github.com β€Ί arrowtype β€Ί 713dad14fe9a574d58d1aab61ba9b2f0
The basics of working with unicode values in Python Β· GitHub
For example, integers will often ... them to do different types of work. To go from a string to an unicode integer, you can use ord(), like:...
Find elsewhere
Top answer
1 of 1
4

First pass

Good job in doing a first pass at type hinting! In the newest stable version of Python (3.10 as of this writing) it's no longer necessary to import List - you can hint with the built-in list. However, your head and body shouldn't really be hinted as lists (which imply mutability), but instead Sequence, which will also accept immutable tuples.

When you do simple member assignment in a constructor from parameters, the members don't also need type hints, and their type can be inferred.

Your # Table symbols can all be deleted since you don't use them. I think the code is quite legible without declaring these as constants. If you were to keep (and start using) these constants, you would want to move them out to static scope before the constructor.

Your print is not general-purpose enough. For example, if someone wants to write this table out to a file, that will be difficult. One convenient (and possibly the highest-performing) way of rewriting this is as an iterator of lines. The caller can decide to either iterate over each line and do something with it; or call into a wrapper utility method that joins the lines and returns a single string.

The first three f-strings in your print method have trivial wrappers with a very long, single field in the middle. This is difficult to read, and you're better off moving that field content to a variable. Bonus: the variable name will self-document the line after it, so you can remove your comments.

When you join on a generator, don't also wrap the generator into a list comprehension. Generators are perfectly capable of being passed bare into join.

Don't for i, _ in enumerate; this is a job for a simple i in range(len.

In your main loop, don't index + 1; pass a start parameter to enumerate.

In your body loop, rather than an enumerate, zip together your widths and elements.

find_widths is a good candidate for being re-expressed as an iterator function. Store it to a tuple and not a list in the constructor.

Your test table content in main should (almost) all be converted to tuples, with the exception of your outer body list comprehension since there's no such thing as a tuple comprehension, so a list is more convenient.

from typing import Sequence, Iterator
import random


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def lines(self) -> Iterator[str]:
        top = '┬'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"β”Œ{top}┐"

        header = 'β”‚'.join(
            title.center(width)
            for title, width in zip(self.head, self.column_widths)
        )
        yield f"β”‚{header}β”‚"

        separator = 'β”Ό'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"β”œ{separator}─"

        for row in self.body:
            row_str = "β”‚".join(
                element.ljust(width, ' ')
                for element, width in zip(row, self.column_widths)
            )
            yield f"β”‚{row_str}β”‚"

        tail = 'β”΄'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"β””{tail}β”˜"

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        "Alpha", "Bravo", "Charlie", "Delta", "Echo",
        "Foxtrot", "Golf", "Hotel", "India", "Juliet",
        "Kilo", "Lima", "Mike", "November", "Oscar",
        "Papa", "Quebec", "Romeo", "Sierra", "Tango",
        "Uniform", "Victor", "Whiskey", "X-Ray", "Yankee",
        "Zulu",
    )
    table1: Table = Table(
        head=("Id", "Name ", "Marks"),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == "__main__":
    main()

Second pass

Now, take a look at the commonalities in your formatting code. You basically only have two operations: making a hyphen-separator, and making a line with content. Put these in utility functions rather than copy-pasting the code:

from typing import Sequence, Iterator, Iterable
import random


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def make_separator(self, left: str, mid: str, right: str) -> str:
        header = mid.join(
            '─' * width for width in self.column_widths
        )
        return f'{left}{header}{right}'

    @staticmethod
    def make_line(elements: Iterable[str]) -> str:
        line = 'β”‚'.join(elements)
        return f'β”‚{line}β”‚'

    def lines(self) -> Iterator[str]:
        yield self.make_separator(left='β”Œ', mid='┬', right='┐')

        yield self.make_line(
            title.center(width)
            for title, width in zip(self.head, self.column_widths)
        )

        yield self.make_separator(left='β”œ', mid='β”Ό', right='─')

        for row in self.body:
            yield self.make_line(
                element.ljust(width, ' ')
                for element, width in zip(row, self.column_widths)
            )

        yield self.make_separator(left='β””', mid='β”΄', right='β”˜')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

Third pass

Recognise that there is still some repetition: we're zipping over widths both times that we call the line-making utility function. Move that zip to within the function, and cut away the formatting differences to their own functions, passing references to the functions into the line-making method. Also, you can splat a three-character string into your function arguments for shorter invocation.

from typing import Sequence, Iterator, Iterable, Callable
import random


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
β”‚    2502  BOX DRAWINGS LIGHT VERTICAL
          
β”Œ    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
β”œ    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
β”Ό    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
─    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
β””    2514  BOX DRAWINGS LIGHT UP AND RIGHT
β”΄    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
β”˜    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def make_separator(self, left: str, mid: str, right: str) -> str:
        line = mid.join(
            '─' * width for width in self.column_widths
        )
        return f'{left}{line}{right}'

    @staticmethod
    def format_header(title: str, width: int) -> str:
        return title.center(width)

    @staticmethod
    def format_element(element: str, width: int) -> str:
        return element.ljust(width)

    def make_line(
        self,
        elements: Iterable[str],
        format: Callable[[str, int], str],
    ) -> str:
        line = 'β”‚'.join(
            format(element, width)
            for element, width in zip(elements, self.column_widths)
        )
        return f'β”‚{line}β”‚'

    def lines(self) -> Iterator[str]:
        yield self.make_separator(*'β”Œβ”¬β”')
        yield self.make_line(self.head, self.format_header)
        yield self.make_separator(*'β”œβ”Όβ”€')

        for row in self.body:
            yield self.make_line(row, self.format_element)

        yield self.make_separator(*'β””β”΄β”˜')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

A word on Unicode

Though I'm still not convinced you should be using variable constants for your drawing characters, it would be informative and helpful to include a docstring like the one I showed at the top of the last code block.

Overall it's vaguely safe to use these characters for rendering in fixed-width terminals of modern machines. In narrow cases this might break; for interesting examples read Misalignment of Unicode block characters in preformatted text blocks.

When you carry this code and its output around editors, IDEs and source control, take care to preserve UTF-8 encoding. The output of the last sample seems byte-for-byte equivalent:

β”Œβ”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”
β”‚ Id β”‚ Name   β”‚Marksβ”‚
β”œβ”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€
β”‚ 1  β”‚Alpha   β”‚   74β”‚
β”‚ 2  β”‚Bravo   β”‚    8β”‚
β”‚ 3  β”‚Charlie β”‚   90β”‚
β”‚ 4  β”‚Delta   β”‚   75β”‚
β”‚ 5  β”‚Echo    β”‚   28β”‚
β”‚ 6  β”‚Foxtrot β”‚   84β”‚
β”‚ 7  β”‚Golf    β”‚   92β”‚
β”‚ 8  β”‚Hotel   β”‚   38β”‚
β”‚ 9  β”‚India   β”‚    6β”‚
β”‚ 10 β”‚Juliet  β”‚   59β”‚
β”‚ 11 β”‚Kilo    β”‚    0β”‚
β”‚ 12 β”‚Lima    β”‚   85β”‚
β”‚ 13 β”‚Mike    β”‚   33β”‚
β”‚ 14 β”‚Novemberβ”‚   81β”‚
β”‚ 15 β”‚Oscar   β”‚   39β”‚
β”‚ 16 β”‚Papa    β”‚   60β”‚
β”‚ 17 β”‚Quebec  β”‚   18β”‚
β”‚ 18 β”‚Romeo   β”‚   54β”‚
β”‚ 19 β”‚Sierra  β”‚   36β”‚
β”‚ 20 β”‚Tango   β”‚   50β”‚
β”‚ 21 β”‚Uniform β”‚   53β”‚
β”‚ 22 β”‚Victor  β”‚   83β”‚
β”‚ 23 β”‚Whiskey β”‚   85β”‚
β”‚ 24 β”‚X-Ray   β”‚   16β”‚
β”‚ 25 β”‚Yankee  β”‚   68β”‚
β”‚ 26 β”‚Zulu    β”‚   36β”‚
β””β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”˜

True Formatting

Your original approach (and all of the approaches above) rely on explicit string repetition through the * operator. There is a very different approach that forms true formatting strings defining alignment, width and padding. For details on these parameters read the format mini-language specification.

Once this formatting string is defined, you can pre-bind to its .format() method and hold a reference to that method. Since the formatting strings encode the column widths, you don't actually need to store the widths on the class and can drop them after the constructor.

I don't strongly consider this approach universally better or worse.

from typing import Sequence, Iterator, Iterable, Callable
import random


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
β”‚    2502  BOX DRAWINGS LIGHT VERTICAL
          
β”Œ    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
β”œ    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
β”Ό    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
─    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
β””    2514  BOX DRAWINGS LIGHT UP AND RIGHT
β”΄    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
β”˜    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        widths = tuple(self.find_widths())
        self.format_sep = self.make_separator_format(widths)
        self.format_head = self.make_content_format(widths, align='^')
        self.format_body = self.make_content_format(widths, align='<')

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    @staticmethod
    def make_separator_format(widths: Sequence[int]) -> Callable:
        fmt = (
            '{0}'
            + ''.join(
                '{1:─>%d}' % (1 + width)
                for width in widths[:-1]
            )
            + '{2:─>%d}' % (1 + widths[-1])
        )
        return fmt.format

    @staticmethod
    def make_content_format(widths: Sequence[int], align: str) -> Callable:
        fmt = (
            'β”‚'
            + 'β”‚'.join(
                '{:%s%d}' % (align, width)
                for width in widths
            )
            + 'β”‚'
        )
        return fmt.format

    def lines(self) -> Iterator[str]:
        yield self.format_sep(*'β”Œβ”¬β”')
        yield self.format_head(*self.head)
        yield self.format_sep(*'β”œβ”Όβ”€')

        for row in self.body:
            yield self.format_body(*row)

        yield self.format_sep(*'β””β”΄β”˜')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

If you debug, you can see the format strings it makes:

{0}{1:─>5}{1:─>9}{2:─>6}
β”‚{:^4}β”‚{:^8}β”‚{:^5}β”‚
β”‚{:<4}β”‚{:<8}β”‚{:<5}β”‚

Formatting responsibility

Currently your main function has stolen a little bit of the responsibility to format the cell content. This is a little awkward, and should just be transferred to the table class.

from typing import Sequence, Iterator, Callable, Optional, Any
from random import randint


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
β”‚    2502  BOX DRAWINGS LIGHT VERTICAL
          
β”Œ    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
β”œ    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
β”Ό    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
─    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
β””    2514  BOX DRAWINGS LIGHT UP AND RIGHT
β”΄    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
β”˜    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(
        self,
        head: Sequence[str],
        body: Sequence[Sequence[Any]],
        formats: Sequence[Optional[str]] = (),
    ) -> None:
        self.head = head
        self.body = tuple(self.format_body(body, formats))
        widths = tuple(self.find_widths())
        self.format_sep = self.make_separator_format(widths)
        self.format_head = self.make_content_format(widths, align='^')
        self.format_body = self.make_content_format(widths, align='<')

    @staticmethod
    def format_body(
        body: Sequence[Sequence[Any]],
        formats: Sequence[Optional[str]],
    ) -> Iterator[Sequence[str]]:
        if not formats:
            formats = (None,) * len(body[0])

        for row in body:
            yield [
                fmt.format(cell) if fmt else str(cell)
                for cell, fmt in zip(row, formats)
            ]

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    @staticmethod
    def make_separator_format(widths: Sequence[int]) -> Callable:
        fmt = (
            '{0}'
            + ''.join(
                '{1:─>%d}' % (1 + width)
                for width in widths[:-1]
            )
            + '{2:─>%d}' % (1 + widths[-1])
        )
        return fmt.format

    @staticmethod
    def make_content_format(
        widths: Sequence[int],
        align: str,
    ) -> Callable:
        fmt = (
            'β”‚'
            + 'β”‚'.join(
                '{:%s%d}' % (align, width)
                for width in widths
            )
            + 'β”‚'
        )
        return fmt.format

    def lines(self) -> Iterator[str]:
        yield self.format_sep(*'β”Œβ”¬β”')
        yield self.format_head(*self.head)
        yield self.format_sep(*'β”œβ”Όβ”€')

        for row in self.body:
            yield self.format_body(*row)

        yield self.format_sep(*'β””β”΄β”˜')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (index, name, randint(0, 100))
            for index, name in enumerate(names, 1)
        ],
        formats=('{:^4}', None, '{:>5}'),
    )
    print(table1)


if __name__ == '__main__':
    main()
🌐
Python
docs.python.org β€Ί 3 β€Ί library β€Ί unicodedata.html
unicodedata β€” Unicode Database
This module provides access to the Unicode Character Database (UCD) which defines character properties for all Unicode characters. The data contained in this database is compiled from the UCD versi...
🌐
Rmotr
learn.rmotr.com β€Ί python β€Ί understanding-unicode-in-python β€Ί strings-and-unicode β€Ί unicode-in-python
Understanding Unicode in Python > Strings and Unicode > Unicode in Python
In Python 2, a simple string literal like "hello world" will create a str, which is a byte string type. To create a unicode string in Python 2 using literals, you have to prefix your string with the lowercase letter 'u': u'Hello Unicode World'. In Python 3 a simple string literal without any ...
🌐
GitHub
github.com β€Ί PyTables β€Ί PyTables β€Ί issues β€Ί 499
Return unicode strings from stored bytestrings Β· Issue #499 Β· PyTables/PyTables
September 9, 2015 - handle = tables.open_file('tes... yields b'bark' instead of 'bark'. This is because PyTables currently stores strings as ascii byte strings instead of unicode....
Author: PyTables
🌐
Python3
python3.info β€Ί stdlib β€Ί string β€Ί unicode.html
7.1. Unicode β€” Python - from None to AI
https://symbl.cc/en/unicode-table/ >>> import string >>> >>> >>> string.ascii_lowercase 'abcdefghijklmnopqrstuvwxyz' >>> >>> string.ascii_uppercase 'ABCDEFGHIJKLMNOPQRSTUVWXYZ' >>> >>> string.ascii_letters 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ' >>> import unicodedata >>> >>> >>> unicodedata.name('a') 'LATIN SMALL LETTER A' >>> >>> unicodedata.name('Δ…') 'LATIN SMALL LETTER A WITH OGONEK' >>> >>> unicodedata.name('Ε›') 'LATIN SMALL LETTER S WITH ACUTE' >>> >>> unicodedata.name('Ε‚') 'LATIN SMALL LETTER L WITH STROKE' >>> >>> unicodedata.name('ΕΌ') 'LATIN SMALL LETTER Z WITH DOT ABOVE' >>> >>> print('\U0001F680') πŸš€ Β·
🌐
Pythonturtle
pythonturtle.academy β€Ί unicode-table
Unicode Table – Python and Turtle
February 27, 2019 - Draw a 16Γ—16 table of unicode symbols. Unicode starts from number 0x2600 (Hexadecimal). You can convert number to text with chr() function. You can start with different number to find more unicode symbols. Depending on your Operating System, symbols may look different.
🌐
Python Basics
python-basics-tutorial.readthedocs.io β€Ί en β€Ί latest β€Ί types β€Ί strings β€Ί encodings.html
Unicode and character encodings - Python Basics
>>> hepy.rstrip(string.punctuation) 'Hello Pythonistas' However, the string module works with Unicode by default, which is represented as binary data (bytes). It is obvious that the ASCII character set is not nearly large enough to cover all languages, dialects, symbols and glyphs; it is not even large enough for English. While ASCII is a complete subset of Unicode – the first 128 characters in the Unicode table correspond exactly to ASCII characters – Unicode encompasses a much larger set of characters.
🌐
Python
docs.python.org β€Ί 3 β€Ί c-api β€Ί unicode.html
Unicode Objects and Codecs β€” Python 3.14.7 documentation
Return a mapping suitable for decoding a custom single-byte encoding. Given a Unicode string string of up to 256 characters representing an encoding table, returns either a compact internal mapping object or a dictionary mapping character ordinals ...
🌐
Python GTK+ 3 Tutorial
python-gtk-3-tutorial.readthedocs.io β€Ί en β€Ί latest β€Ί unicode.html
4. How to Deal With Strings β€” Python GTK+ 3 Tutorial 3.4 documentation
The Unicode HOWTO for Python 3.x discusses Unicode support in Python 3.x. UTF-8 encoding table and Unicode characters contains a list of Unicode code points and their respective UTF-8 encoding.
🌐
YouTube
youtube.com β€Ί watch
Python and Unicode characters - YouTube
What is Unicode? How can we insert Unicode characters into strings? What's the difference between \x, \u, and \U? How can we insert characters with their nam...
Published: March 23, 2022
🌐
Asmeurer
asmeurer.com β€Ί python-unicode-variable-names
Python Unicode Variable Names | A page listing all the Unicode characters that are valid in Python variable names
You can normalize strings with Python using the unicodedata module: >>> a = 'á' >>> len(a) 2 >>> import unicodedata >>> unicodedata.normalize("NFKC", a) 'Ñ' >>> len(_) 1 · The below table lists characters that normalize to other characters, but be aware that other combinations of characters such as combining accents may not be listed below but may still normalize to a character listed below.
🌐
Reddit
reddit.com β€Ί r/learnpython β€Ί where can i find a full list of python unicode superscript formatting?
r/learnpython on Reddit: Where can I find a full list of Python Unicode Superscript formatting?
June 20, 2018 -

Hi, I am trying to format some text as a superscript but I am finding that difficult since I can't find any of the Unicode values for Python!

# For example: This would be the unicode formatting for a python superscript of x^2
print(u"x\u00B2")

I am specifically looking for y, and -1 as a superscript. Does anyone know of a direct reference or a chart or list?

Thanks!

🌐
Reddit
reddit.com β€Ί r/python β€Ί explain it like i'm five: python and unicode?
r/Python on Reddit: Explain it like I'm five: Python and Unicode?
June 12, 2013 -

I am seriously confused. And whenever I think I got it, I see some - in my opinion - inconsistent behavior. Can it be consistently explained or is it more art than science?

When do I have to encode/decode("UTF-8")? What does it do exactly? Whats so special about unicode("abc"), or is it identical to u"abc"?

Why, if I'm using a HTML-encoding of UTF8, a python-script with encoding-UTF-8 and a UTF-8 capable shell and have them all interact, do I have to randomly start adding the above functions until stuff accidentally doesn't break anymore? :)

My problem is that while I can code quite well, I have no formal computer science education and don't tend to think in bytes.

Top answer
1 of 5
83
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character.
2 of 5
22
To answer your specific questions: when you encode("UTF-8") you are converting a unicode string to a byte string. It should be called on unicode strings. When you decode("UTF-8") you are converting a byte string to a unicode string. It should be called on byte strings. unicode("abc") is the same as u"abc": they both create a unicode string with three characters. Most of the confusion comes from the fact that python 2 plays fast and loose with unicode strings. It will try and convert between them for you when you mix them together, which yields unexpected results. Python 3 has much more sane behavior: it forces you to encode or decode explicitly to convert between the two. Basically what you need to do to avoid most problems and confusion is to do your encoding/decoding at the input/output boundaries of your program. Decode as soon as you get a byte string from external sources, use unicode strings throughout the program, and encode it just before it leaves.