The standard module unicodedata defines a lot of properties, but not everything. A quick peek at its source confirms this.

Fortunately unicodedata.txt, the data file where this comes from, is not hard to parse. Each line consists of exactly 15 elements, ; separated, which makes it ideal for parsing. Using the description of the elements on ftp://ftp.unicode.org/Public/3.0-Update/UnicodeData-3.0.0.html, you can create a few classes to encapsulate the data. I've taken the names of the class elements from that list; the meaning of each of the elements is explained on that same page.

Make sure to download ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt and ftp://ftp.unicode.org/Public/UNIDATA/Blocks.txt first, and put them inside the same folder as this program.

Code (tested with Python 2.7 and 3.6):

# -*- coding: utf-8 -*-

class UnicodeCharacter:
    def __init__(self):
        self.code = 0
        self.name = 'unnamed'
        self.category = ''
        self.combining = ''
        self.bidirectional = ''
        self.decomposition = ''
        self.asDecimal = None
        self.asDigit = None
        self.asNumeric = None
        self.mirrored = False
        self.uc1Name = None
        self.comment = ''
        self.uppercase = None
        self.lowercase = None
        self.titlecase = None
        self.block = None

    def __getitem__(self, item):
        return getattr(self, item)

    def __repr__(self):
        return '{'+self.name+'}'

class UnicodeBlock:
    def __init__(self):
        self.first = 0
        self.last = 0
        self.name = 'unnamed'

    def __repr__(self):
        return '{'+self.name+'}'

class BlockList:
    def __init__(self):
        self.blocklist = []
        with open('Blocks.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.split(';')
                    block = UnicodeBlock()
                    block.name = rawdata[1].strip()
                    rawdata = rawdata[0].split('..')
                    block.first = int(rawdata[0],16)
                    block.last = int(rawdata[1],16)
                    self.blocklist.append(block)
            # make 100% sure it's sorted, for quicker look-up later
            # (it is usually sorted in the file, but better make sure)
            self.blocklist.sort (key=lambda x: block.first)

    def lookup(self,code):
        for item in self.blocklist:
            if code >= item.first and code <= item.last:
                return item.name
        return None

class UnicodeList:
    """UnicodeList loads Unicode data from the external files
    'UnicodeData.txt' and 'Blocks.txt', both available at unicode.org

    These files must appear in the same directory as this program.

    UnicodeList is a new interpretation of the standard library
    'unicodedata'; you may first want to check if its functionality
    suffices.

    As UnicodeList loads its data from an external file, it does not depend
    on the local build from Python (in which the Unicode data gets frozen
    to the then 'current' version).

    Initialize with

        uclist = UnicodeList()
    """
    def __init__(self):

        # we need this first
        blocklist = BlockList()
        bpos = 0

        self.codelist = []
        with open('UnicodeData.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.strip().split(';')
                    parsed = UnicodeCharacter()
                    parsed.code = int(rawdata[0],16)
                    parsed.characterName = rawdata[1]
                    parsed.category = rawdata[2]
                    parsed.combining = rawdata[3]
                    parsed.bidirectional = rawdata[4]
                    parsed.decomposition = rawdata[5]
                    parsed.asDecimal = int(rawdata[6]) if rawdata[6] else None
                    parsed.asDigit = int(rawdata[7]) if rawdata[7] else None
                    # the following value may contain a slash:
                    #  ONE QUARTER ... 1/4
                    # let's make it Python 2.7 compatible :)
                    if '/' in rawdata[8]:
                        rawdata[8] = rawdata[8].replace('/','./')
                        parsed.asNumeric = eval(rawdata[8])
                    else:
                        parsed.asNumeric = int(rawdata[8]) if rawdata[8] else None
                    parsed.mirrored = rawdata[9] == 'Y'
                    parsed.uc1Name = rawdata[10]
                    parsed.comment = rawdata[11]
                    parsed.uppercase = int(rawdata[12],16) if rawdata[12] else None
                    parsed.lowercase = int(rawdata[13],16) if rawdata[13] else None
                    parsed.titlecase = int(rawdata[14],16) if rawdata[14] else None
                    while bpos < len(blocklist.blocklist) and parsed.code > blocklist.blocklist[bpos].last:
                        bpos += 1
                    parsed.block = blocklist.blocklist[bpos].name if bpos < len(blocklist.blocklist) and parsed.code >= blocklist.blocklist[bpos].first else None
                    self.codelist.append(parsed)

    def find_code(self,codepoint):
        """Find the Unicode information for a codepoint (as int).

        Returns:
            a UnicodeCharacter class object or None.
        """
        # the list is unlikely to contain duplicates but I have seen Unicode.org
        # doing that in similar situations. Again, better make sure.
        val = [x for x in self.codelist if codepoint == x.code]
        return val[0] if val else None

    def find_char(self,str):
        """Find the Unicode information for a codepoint (as character).

        Returns:
            for a single character: a UnicodeCharacter class object or
            None.
            for a multicharacter string: a list of the above, one element
            per character.
        """
        if len(str) > 1:
            result = [self.find_code(ord(x)) for x in str]
            return result
        else:
            return self.find_code(ord(str))

When loaded, you can now look up a character code with

>>> ul = UnicodeList()     # ONLY NEEDED ONCE!
>>> print (ul.find_code(0x204))
{LATIN CAPITAL LETTER E WITH DOUBLE GRAVE}

which by default is shown as the name of a character (Unicode calls this a 'code point'), but you can retrieve other properties as well:

>>> print ('%04X' % uc.find_code(0x204).lowercase)
0205
>>> print (ul.lookup(0x204).block)
Latin Extended-B

and (as long as you don't get a None) even chain them:

>>> print (ul.find_code(ul.find_code(0x204).lowercase))
{LATIN SMALL LETTER E WITH DOUBLE GRAVE}

It does not rely on your particular build of Python; you can always download an updated list from unicode.org and be assured to get the most recent information:

import unicodedata
>>> print (unicodedata.name('\U0001F903'))
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
ValueError: no such name
>>> print (uclist.find_code(0x1f903))
{LEFT HALF CIRCLE WITH FOUR DOTS}

(As tested with Python 3.5.3.)

There are currently two lookup functions defined:

  • find_code(int) looks up character information by codepoint as an integer.
  • find_char(string) looks up character information for the character(s) in string. If there is only one character, it returns a UnicodeCharacter object; if there are more, it returns a list of objects.

After import unicodelist (assuming you saved this as unicodelist.py), you can use

>>> ul = UnicodeList()
>>> hex(ul.find_char(u'è').code)
'0xe8'

to look up the hex code for any character, and a list comprehension such as

>>> l = [hex(ul.find_char(x).code) for x in 'Hello']
>>> l
['0x48', '0x65', '0x6c', '0x6c', '0x6f']

for longer strings. Note that you don't actually need all of this if all you want is a hex representation of a string! This suffices:

 l = [hex(ord(x)) for x in 'Hello']

The purpose of this module is to give easy access to other Unicode properties. A longer example:

str = 'Héllo...'
dest = ''
for i in str:
    dest += chr(ul.find_char(i).uppercase) if ul.find_char(i).uppercase is not None else i
print (dest)

HÉLLO...

and showing a list of properties for a character per your example:

letter = u'Ȅ'
print ('Name > '+ul.find_char(letter).name)
print ('Unicode number > U+%04x' % ul.find_char(letter).code)
print ('Bloc > '+ul.find_char(letter).block)
print ('Lowercase > %s' % chr(ul.find_char(letter).lowercase))

(I left out HTML; these names are not defined in the Unicode standard.)

Answer from Jongware on Stack Overflow
Top answer
1 of 3
5

The standard module unicodedata defines a lot of properties, but not everything. A quick peek at its source confirms this.

Fortunately unicodedata.txt, the data file where this comes from, is not hard to parse. Each line consists of exactly 15 elements, ; separated, which makes it ideal for parsing. Using the description of the elements on ftp://ftp.unicode.org/Public/3.0-Update/UnicodeData-3.0.0.html, you can create a few classes to encapsulate the data. I've taken the names of the class elements from that list; the meaning of each of the elements is explained on that same page.

Make sure to download ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt and ftp://ftp.unicode.org/Public/UNIDATA/Blocks.txt first, and put them inside the same folder as this program.

Code (tested with Python 2.7 and 3.6):

# -*- coding: utf-8 -*-

class UnicodeCharacter:
    def __init__(self):
        self.code = 0
        self.name = 'unnamed'
        self.category = ''
        self.combining = ''
        self.bidirectional = ''
        self.decomposition = ''
        self.asDecimal = None
        self.asDigit = None
        self.asNumeric = None
        self.mirrored = False
        self.uc1Name = None
        self.comment = ''
        self.uppercase = None
        self.lowercase = None
        self.titlecase = None
        self.block = None

    def __getitem__(self, item):
        return getattr(self, item)

    def __repr__(self):
        return '{'+self.name+'}'

class UnicodeBlock:
    def __init__(self):
        self.first = 0
        self.last = 0
        self.name = 'unnamed'

    def __repr__(self):
        return '{'+self.name+'}'

class BlockList:
    def __init__(self):
        self.blocklist = []
        with open('Blocks.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.split(';')
                    block = UnicodeBlock()
                    block.name = rawdata[1].strip()
                    rawdata = rawdata[0].split('..')
                    block.first = int(rawdata[0],16)
                    block.last = int(rawdata[1],16)
                    self.blocklist.append(block)
            # make 100% sure it's sorted, for quicker look-up later
            # (it is usually sorted in the file, but better make sure)
            self.blocklist.sort (key=lambda x: block.first)

    def lookup(self,code):
        for item in self.blocklist:
            if code >= item.first and code <= item.last:
                return item.name
        return None

class UnicodeList:
    """UnicodeList loads Unicode data from the external files
    'UnicodeData.txt' and 'Blocks.txt', both available at unicode.org

    These files must appear in the same directory as this program.

    UnicodeList is a new interpretation of the standard library
    'unicodedata'; you may first want to check if its functionality
    suffices.

    As UnicodeList loads its data from an external file, it does not depend
    on the local build from Python (in which the Unicode data gets frozen
    to the then 'current' version).

    Initialize with

        uclist = UnicodeList()
    """
    def __init__(self):

        # we need this first
        blocklist = BlockList()
        bpos = 0

        self.codelist = []
        with open('UnicodeData.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.strip().split(';')
                    parsed = UnicodeCharacter()
                    parsed.code = int(rawdata[0],16)
                    parsed.characterName = rawdata[1]
                    parsed.category = rawdata[2]
                    parsed.combining = rawdata[3]
                    parsed.bidirectional = rawdata[4]
                    parsed.decomposition = rawdata[5]
                    parsed.asDecimal = int(rawdata[6]) if rawdata[6] else None
                    parsed.asDigit = int(rawdata[7]) if rawdata[7] else None
                    # the following value may contain a slash:
                    #  ONE QUARTER ... 1/4
                    # let's make it Python 2.7 compatible :)
                    if '/' in rawdata[8]:
                        rawdata[8] = rawdata[8].replace('/','./')
                        parsed.asNumeric = eval(rawdata[8])
                    else:
                        parsed.asNumeric = int(rawdata[8]) if rawdata[8] else None
                    parsed.mirrored = rawdata[9] == 'Y'
                    parsed.uc1Name = rawdata[10]
                    parsed.comment = rawdata[11]
                    parsed.uppercase = int(rawdata[12],16) if rawdata[12] else None
                    parsed.lowercase = int(rawdata[13],16) if rawdata[13] else None
                    parsed.titlecase = int(rawdata[14],16) if rawdata[14] else None
                    while bpos < len(blocklist.blocklist) and parsed.code > blocklist.blocklist[bpos].last:
                        bpos += 1
                    parsed.block = blocklist.blocklist[bpos].name if bpos < len(blocklist.blocklist) and parsed.code >= blocklist.blocklist[bpos].first else None
                    self.codelist.append(parsed)

    def find_code(self,codepoint):
        """Find the Unicode information for a codepoint (as int).

        Returns:
            a UnicodeCharacter class object or None.
        """
        # the list is unlikely to contain duplicates but I have seen Unicode.org
        # doing that in similar situations. Again, better make sure.
        val = [x for x in self.codelist if codepoint == x.code]
        return val[0] if val else None

    def find_char(self,str):
        """Find the Unicode information for a codepoint (as character).

        Returns:
            for a single character: a UnicodeCharacter class object or
            None.
            for a multicharacter string: a list of the above, one element
            per character.
        """
        if len(str) > 1:
            result = [self.find_code(ord(x)) for x in str]
            return result
        else:
            return self.find_code(ord(str))

When loaded, you can now look up a character code with

>>> ul = UnicodeList()     # ONLY NEEDED ONCE!
>>> print (ul.find_code(0x204))
{LATIN CAPITAL LETTER E WITH DOUBLE GRAVE}

which by default is shown as the name of a character (Unicode calls this a 'code point'), but you can retrieve other properties as well:

>>> print ('%04X' % uc.find_code(0x204).lowercase)
0205
>>> print (ul.lookup(0x204).block)
Latin Extended-B

and (as long as you don't get a None) even chain them:

>>> print (ul.find_code(ul.find_code(0x204).lowercase))
{LATIN SMALL LETTER E WITH DOUBLE GRAVE}

It does not rely on your particular build of Python; you can always download an updated list from unicode.org and be assured to get the most recent information:

import unicodedata
>>> print (unicodedata.name('\U0001F903'))
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
ValueError: no such name
>>> print (uclist.find_code(0x1f903))
{LEFT HALF CIRCLE WITH FOUR DOTS}

(As tested with Python 3.5.3.)

There are currently two lookup functions defined:

  • find_code(int) looks up character information by codepoint as an integer.
  • find_char(string) looks up character information for the character(s) in string. If there is only one character, it returns a UnicodeCharacter object; if there are more, it returns a list of objects.

After import unicodelist (assuming you saved this as unicodelist.py), you can use

>>> ul = UnicodeList()
>>> hex(ul.find_char(u'è').code)
'0xe8'

to look up the hex code for any character, and a list comprehension such as

>>> l = [hex(ul.find_char(x).code) for x in 'Hello']
>>> l
['0x48', '0x65', '0x6c', '0x6c', '0x6f']

for longer strings. Note that you don't actually need all of this if all you want is a hex representation of a string! This suffices:

 l = [hex(ord(x)) for x in 'Hello']

The purpose of this module is to give easy access to other Unicode properties. A longer example:

str = 'Héllo...'
dest = ''
for i in str:
    dest += chr(ul.find_char(i).uppercase) if ul.find_char(i).uppercase is not None else i
print (dest)

HÉLLO...

and showing a list of properties for a character per your example:

letter = u'Ȅ'
print ('Name > '+ul.find_char(letter).name)
print ('Unicode number > U+%04x' % ul.find_char(letter).code)
print ('Bloc > '+ul.find_char(letter).block)
print ('Lowercase > %s' % chr(ul.find_char(letter).lowercase))

(I left out HTML; these names are not defined in the Unicode standard.)

2 of 3
3

The unicodedata documentation shows how to do most of this.

The Unicode block name is apparently not available but another Stack Overflow question has a solution of sorts and another has some additional approaches using regex.

The uppercase/lowercase mapping and character number information is not particularly Unicode-specific; just use the regular Python string functions.

So in summary

>>> import unicodedata
>>> unicodedata.name('Ë')
'LATIN CAPITAL LETTER E WITH DIAERESIS'
>>> 'U+%04X' % ord('Ë')
'U+00CB'
>>> '&#%i;' % ord('Ë')
'&#203;'
>>> 'Ë'.lower()
'ë'

The U+%04X formatting is sort-of correct, in that it simply avoids padding and prints the whole hex number for code points with a value higher than 65,535. Note that some other formats require the use of %08X padding in this scenario (notably \U00010000 format in Python).

🌐
python-tcod
python-tcod.readthedocs.io › en › latest › tcod › charmap-reference.html
Character Table Reference - python-tcod 21.2.1 documentation
Unicode is the Unicode code point as a hexadecimal number. You can use chr to convert these numbers into a string. Character maps such as tcod.tileset.CHARMAP_CP437 are simply a list of Unicode numbers, where the index of the list is the Tile Index. String is the Python string for that character.
Discussions

Formatting a table using unicode symbols in python - Code Review Stack Exchange
The goal of the project is to take in input for a data, perhaps in the future as a csv, then print it out using UNICODE box-drawing characters. The main function declares the necessary data and Table More on codereview.stackexchange.com
🌐 codereview.stackexchange.com
May 15, 2022
Where can I find a full list of Python Unicode Superscript formatting?

Unicode is not specifically for Python. It is a universal standard. So any unicode table can work, like this and this.

For your request, u"y\u207B\u00B9" will do.

More on reddit.com
🌐 r/learnpython
2
2
June 20, 2018
Return unicode strings from stored bytestrings
There was an error while loading. Please reload this page More on github.com
🌐 github.com
16
September 9, 2015
Explain it like I'm five: Python and Unicode?
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character. More on reddit.com
🌐 r/Python
60
106
June 12, 2013
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
The Unicode standard contains a lot of tables listing characters and their corresponding code points:
Top answer
1 of 1
4

First pass

Good job in doing a first pass at type hinting! In the newest stable version of Python (3.10 as of this writing) it's no longer necessary to import List - you can hint with the built-in list. However, your head and body shouldn't really be hinted as lists (which imply mutability), but instead Sequence, which will also accept immutable tuples.

When you do simple member assignment in a constructor from parameters, the members don't also need type hints, and their type can be inferred.

Your # Table symbols can all be deleted since you don't use them. I think the code is quite legible without declaring these as constants. If you were to keep (and start using) these constants, you would want to move them out to static scope before the constructor.

Your print is not general-purpose enough. For example, if someone wants to write this table out to a file, that will be difficult. One convenient (and possibly the highest-performing) way of rewriting this is as an iterator of lines. The caller can decide to either iterate over each line and do something with it; or call into a wrapper utility method that joins the lines and returns a single string.

The first three f-strings in your print method have trivial wrappers with a very long, single field in the middle. This is difficult to read, and you're better off moving that field content to a variable. Bonus: the variable name will self-document the line after it, so you can remove your comments.

When you join on a generator, don't also wrap the generator into a list comprehension. Generators are perfectly capable of being passed bare into join.

Don't for i, _ in enumerate; this is a job for a simple i in range(len.

In your main loop, don't index + 1; pass a start parameter to enumerate.

In your body loop, rather than an enumerate, zip together your widths and elements.

find_widths is a good candidate for being re-expressed as an iterator function. Store it to a tuple and not a list in the constructor.

Your test table content in main should (almost) all be converted to tuples, with the exception of your outer body list comprehension since there's no such thing as a tuple comprehension, so a list is more convenient.

from typing import Sequence, Iterator
import random


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def lines(self) -> Iterator[str]:
        top = '┬'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"┌{top}┐"

        header = '│'.join(
            title.center(width)
            for title, width in zip(self.head, self.column_widths)
        )
        yield f"│{header}│"

        separator = '┼'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"├{separator}┤"

        for row in self.body:
            row_str = "│".join(
                element.ljust(width, ' ')
                for element, width in zip(row, self.column_widths)
            )
            yield f"│{row_str}│"

        tail = '┴'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"└{tail}┘"

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        "Alpha", "Bravo", "Charlie", "Delta", "Echo",
        "Foxtrot", "Golf", "Hotel", "India", "Juliet",
        "Kilo", "Lima", "Mike", "November", "Oscar",
        "Papa", "Quebec", "Romeo", "Sierra", "Tango",
        "Uniform", "Victor", "Whiskey", "X-Ray", "Yankee",
        "Zulu",
    )
    table1: Table = Table(
        head=("Id", "Name ", "Marks"),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == "__main__":
    main()

Second pass

Now, take a look at the commonalities in your formatting code. You basically only have two operations: making a hyphen-separator, and making a line with content. Put these in utility functions rather than copy-pasting the code:

from typing import Sequence, Iterator, Iterable
import random


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def make_separator(self, left: str, mid: str, right: str) -> str:
        header = mid.join(
            '─' * width for width in self.column_widths
        )
        return f'{left}{header}{right}'

    @staticmethod
    def make_line(elements: Iterable[str]) -> str:
        line = '│'.join(elements)
        return f'│{line}│'

    def lines(self) -> Iterator[str]:
        yield self.make_separator(left='┌', mid='┬', right='┐')

        yield self.make_line(
            title.center(width)
            for title, width in zip(self.head, self.column_widths)
        )

        yield self.make_separator(left='├', mid='┼', right='┤')

        for row in self.body:
            yield self.make_line(
                element.ljust(width, ' ')
                for element, width in zip(row, self.column_widths)
            )

        yield self.make_separator(left='└', mid='┴', right='┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

Third pass

Recognise that there is still some repetition: we're zipping over widths both times that we call the line-making utility function. Move that zip to within the function, and cut away the formatting differences to their own functions, passing references to the functions into the line-making method. Also, you can splat a three-character string into your function arguments for shorter invocation.

from typing import Sequence, Iterator, Iterable, Callable
import random


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
│    2502  BOX DRAWINGS LIGHT VERTICAL
          
┌    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
├    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
┼    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
┤    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
└    2514  BOX DRAWINGS LIGHT UP AND RIGHT
┴    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
┘    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def make_separator(self, left: str, mid: str, right: str) -> str:
        line = mid.join(
            '─' * width for width in self.column_widths
        )
        return f'{left}{line}{right}'

    @staticmethod
    def format_header(title: str, width: int) -> str:
        return title.center(width)

    @staticmethod
    def format_element(element: str, width: int) -> str:
        return element.ljust(width)

    def make_line(
        self,
        elements: Iterable[str],
        format: Callable[[str, int], str],
    ) -> str:
        line = '│'.join(
            format(element, width)
            for element, width in zip(elements, self.column_widths)
        )
        return f'│{line}│'

    def lines(self) -> Iterator[str]:
        yield self.make_separator(*'┌┬┐')
        yield self.make_line(self.head, self.format_header)
        yield self.make_separator(*'├┼┤')

        for row in self.body:
            yield self.make_line(row, self.format_element)

        yield self.make_separator(*'└┴┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

A word on Unicode

Though I'm still not convinced you should be using variable constants for your drawing characters, it would be informative and helpful to include a docstring like the one I showed at the top of the last code block.

Overall it's vaguely safe to use these characters for rendering in fixed-width terminals of modern machines. In narrow cases this might break; for interesting examples read Misalignment of Unicode block characters in preformatted text blocks.

When you carry this code and its output around editors, IDEs and source control, take care to preserve UTF-8 encoding. The output of the last sample seems byte-for-byte equivalent:

┌────┬────────┬─────┐
│ Id │ Name   │Marks│
├────┼────────┼─────┤
│ 1  │Alpha   │   74│
│ 2  │Bravo   │    8│
│ 3  │Charlie │   90│
│ 4  │Delta   │   75│
│ 5  │Echo    │   28│
│ 6  │Foxtrot │   84│
│ 7  │Golf    │   92│
│ 8  │Hotel   │   38│
│ 9  │India   │    6│
│ 10 │Juliet  │   59│
│ 11 │Kilo    │    0│
│ 12 │Lima    │   85│
│ 13 │Mike    │   33│
│ 14 │November│   81│
│ 15 │Oscar   │   39│
│ 16 │Papa    │   60│
│ 17 │Quebec  │   18│
│ 18 │Romeo   │   54│
│ 19 │Sierra  │   36│
│ 20 │Tango   │   50│
│ 21 │Uniform │   53│
│ 22 │Victor  │   83│
│ 23 │Whiskey │   85│
│ 24 │X-Ray   │   16│
│ 25 │Yankee  │   68│
│ 26 │Zulu    │   36│
└────┴────────┴─────┘

True Formatting

Your original approach (and all of the approaches above) rely on explicit string repetition through the * operator. There is a very different approach that forms true formatting strings defining alignment, width and padding. For details on these parameters read the format mini-language specification.

Once this formatting string is defined, you can pre-bind to its .format() method and hold a reference to that method. Since the formatting strings encode the column widths, you don't actually need to store the widths on the class and can drop them after the constructor.

I don't strongly consider this approach universally better or worse.

from typing import Sequence, Iterator, Iterable, Callable
import random


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
│    2502  BOX DRAWINGS LIGHT VERTICAL
          
┌    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
├    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
┼    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
┤    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
└    2514  BOX DRAWINGS LIGHT UP AND RIGHT
┴    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
┘    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        widths = tuple(self.find_widths())
        self.format_sep = self.make_separator_format(widths)
        self.format_head = self.make_content_format(widths, align='^')
        self.format_body = self.make_content_format(widths, align='<')

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    @staticmethod
    def make_separator_format(widths: Sequence[int]) -> Callable:
        fmt = (
            '{0}'
            + ''.join(
                '{1:─>%d}' % (1 + width)
                for width in widths[:-1]
            )
            + '{2:─>%d}' % (1 + widths[-1])
        )
        return fmt.format

    @staticmethod
    def make_content_format(widths: Sequence[int], align: str) -> Callable:
        fmt = (
            '│'
            + '│'.join(
                '{:%s%d}' % (align, width)
                for width in widths
            )
            + '│'
        )
        return fmt.format

    def lines(self) -> Iterator[str]:
        yield self.format_sep(*'┌┬┐')
        yield self.format_head(*self.head)
        yield self.format_sep(*'├┼┤')

        for row in self.body:
            yield self.format_body(*row)

        yield self.format_sep(*'└┴┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

If you debug, you can see the format strings it makes:

{0}{1:─>5}{1:─>9}{2:─>6}
│{:^4}│{:^8}│{:^5}│
│{:<4}│{:<8}│{:<5}│

Formatting responsibility

Currently your main function has stolen a little bit of the responsibility to format the cell content. This is a little awkward, and should just be transferred to the table class.

from typing import Sequence, Iterator, Callable, Optional, Any
from random import randint


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
│    2502  BOX DRAWINGS LIGHT VERTICAL
          
┌    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
├    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
┼    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
┤    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
└    2514  BOX DRAWINGS LIGHT UP AND RIGHT
┴    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
┘    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(
        self,
        head: Sequence[str],
        body: Sequence[Sequence[Any]],
        formats: Sequence[Optional[str]] = (),
    ) -> None:
        self.head = head
        self.body = tuple(self.format_body(body, formats))
        widths = tuple(self.find_widths())
        self.format_sep = self.make_separator_format(widths)
        self.format_head = self.make_content_format(widths, align='^')
        self.format_body = self.make_content_format(widths, align='<')

    @staticmethod
    def format_body(
        body: Sequence[Sequence[Any]],
        formats: Sequence[Optional[str]],
    ) -> Iterator[Sequence[str]]:
        if not formats:
            formats = (None,) * len(body[0])

        for row in body:
            yield [
                fmt.format(cell) if fmt else str(cell)
                for cell, fmt in zip(row, formats)
            ]

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    @staticmethod
    def make_separator_format(widths: Sequence[int]) -> Callable:
        fmt = (
            '{0}'
            + ''.join(
                '{1:─>%d}' % (1 + width)
                for width in widths[:-1]
            )
            + '{2:─>%d}' % (1 + widths[-1])
        )
        return fmt.format

    @staticmethod
    def make_content_format(
        widths: Sequence[int],
        align: str,
    ) -> Callable:
        fmt = (
            '│'
            + '│'.join(
                '{:%s%d}' % (align, width)
                for width in widths
            )
            + '│'
        )
        return fmt.format

    def lines(self) -> Iterator[str]:
        yield self.format_sep(*'┌┬┐')
        yield self.format_head(*self.head)
        yield self.format_sep(*'├┼┤')

        for row in self.body:
            yield self.format_body(*row)

        yield self.format_sep(*'└┴┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (index, name, randint(0, 100))
            for index, name in enumerate(names, 1)
        ],
        formats=('{:^4}', None, '{:>5}'),
    )
    print(table1)


if __name__ == '__main__':
    main()
🌐
Real Python
realpython.com › python-encodings-guide
Unicode & Character Encodings in Python: A Painless Guide – Real Python
May 20, 2019 - Rather, Unicode is implemented by different character encodings, which you’ll see soon. Unicode is better thought of as a map (something like a dict) or a 2-column database table.
🌐
Inspired Python
inspiredpython.com › tip › python-strings-using-translation-tables-to-make-unicode-text
Python Strings: Using translation tables to make Unicode text • Inspired Python
>>> import string >>> before = string.ascii_lowercase + string.digits >>> after = '🅐🅑🅒🅓🅔🅕🅖🅗🅘🅙🅚🅛🅜🅝🅞🅟🅠🅡🅢🅣🅤🅥🅦🅧🅨🅩⓿➊➋➌➍➎➏➐➑➒' >>> translation_table = str.maketrans(before, after) >>> 'inspired python'.translate(translation_table) '🅘🅝🅢🅟🅘🅡🅔🅓 🅟🅨🅣🅗🅞🅝'
🌐
Python Cheat Sheet
pythonsheets.com › notes › basic › python-unicode.html
Unicode — Python Cheat Sheet
The main goal of this cheat sheet is to collect some common snippets which are related to Unicode. In Python 3, strings are represented by Unicode instead of bytes.
Find elsewhere
🌐
Pythonturtle
pythonturtle.academy › unicode-table
Unicode Table – Python and Turtle
February 27, 2019 - Draw a 16×16 table of unicode symbols. Unicode starts from number 0x2600 (Hexadecimal). You can convert number to text with chr() function. You can start with different number to find more unicode symbols.
🌐
GitHub
gist.github.com › arrowtype › 713dad14fe9a574d58d1aab61ba9b2f0
The basics of working with unicode values in Python · GitHub
Unicodes can either be integers (“A” is 65, “B” is 66, etc) or hex (“A” is 0x41, “B” is 0x42, etc). When scripting with RoboFont or FontTools, a hard thing at first is that different styles come up in different contexts. For example, integers will often be used in scripts, but hex values are shown in UIs and in the TTX output of cmap (the table that maps unicode values to glyphs).
🌐
GeeksforGeeks
geeksforgeeks.org › python › unicode_literals-in-python
unicode_literals in Python - GeeksforGeeks
June 21, 2021 - The issue with the ASCII is that it can only support the English language but what if we want to use another language like Hindi, Russian, Chinese, etc. We didn't have enough space in ASCII to covers up all these languages and emojis. This is where Unicode comes, Unicode provides us a huge table to which can store ASCII table and also the extent to store other languages, symbols, and emojis.
🌐
Reddit
reddit.com › r/learnpython › where can i find a full list of python unicode superscript formatting?
r/learnpython on Reddit: Where can I find a full list of Python Unicode Superscript formatting?
June 20, 2018 -

Hi, I am trying to format some text as a superscript but I am finding that difficult since I can't find any of the Unicode values for Python!

# For example: This would be the unicode formatting for a python superscript of x^2
print(u"x\u00B2")

I am specifically looking for y, and -1 as a superscript. Does anyone know of a direct reference or a chart or list?

Thanks!

🌐
Asmeurer
asmeurer.com › python-unicode-variable-names
Python Unicode Variable Names | A page listing all the Unicode characters that are valid in Python variable names
You can normalize strings with Python using the unicodedata module: >>> a = 'á' >>> len(a) 2 >>> import unicodedata >>> unicodedata.normalize("NFKC", a) 'á' >>> len(_) 1 · The below table lists characters that normalize to other characters, but be aware that other combinations of characters such as combining accents may not be listed below but may still normalize to a character listed below.
🌐
GitHub
github.com › PyTables › PyTables › issues › 499
Return unicode strings from stored bytestrings · Issue #499 · PyTables/PyTables
September 9, 2015 - handle = tables.open_file('tes... yields b'bark' instead of 'bark'. This is because PyTables currently stores strings as ascii byte strings instead of unicode....
Author: PyTables
🌐
Python
docs.python.org › 3.0 › howto › unicode.html
Unicode HOWTO — Python v3.0.1 documentation
To help understand the standard, Jukka Korpela has written an introductory guide to reading the Unicode character tables, available at <http://www.cs.tut.fi/~jkorpela/unicode/guide.html>.
🌐
Python
docs.python.org › 3 › library › unicodedata.html
unicodedata — Unicode Database
This module provides access to the Unicode Character Database (UCD) which defines character properties for all Unicode characters. The data contained in this database is compiled from the UCD versi...
🌐
Reddit
reddit.com › r/python › explain it like i'm five: python and unicode?
r/Python on Reddit: Explain it like I'm five: Python and Unicode?
June 12, 2013 -

I am seriously confused. And whenever I think I got it, I see some - in my opinion - inconsistent behavior. Can it be consistently explained or is it more art than science?

When do I have to encode/decode("UTF-8")? What does it do exactly? Whats so special about unicode("abc"), or is it identical to u"abc"?

Why, if I'm using a HTML-encoding of UTF8, a python-script with encoding-UTF-8 and a UTF-8 capable shell and have them all interact, do I have to randomly start adding the above functions until stuff accidentally doesn't break anymore? :)

My problem is that while I can code quite well, I have no formal computer science education and don't tend to think in bytes.

Top answer
1 of 5
83
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character.
2 of 5
22
To answer your specific questions: when you encode("UTF-8") you are converting a unicode string to a byte string. It should be called on unicode strings. When you decode("UTF-8") you are converting a byte string to a unicode string. It should be called on byte strings. unicode("abc") is the same as u"abc": they both create a unicode string with three characters. Most of the confusion comes from the fact that python 2 plays fast and loose with unicode strings. It will try and convert between them for you when you mix them together, which yields unexpected results. Python 3 has much more sane behavior: it forces you to encode or decode explicitly to convert between the two. Basically what you need to do to avoid most problems and confusion is to do your encoding/decoding at the input/output boundaries of your program. Decode as soon as you get a byte string from external sources, use unicode strings throughout the program, and encode it just before it leaves.
🌐
GitHub
github.com › PyTables › PyTables › issues › 268
Unicode, python 3 · Issue #268 · PyTables/PyTables
July 18, 2013 - It seems that it isn't possible to create a table to store numpy.str dtypes, where this worked before in Python 2.x. Guessing this is because the default python string is now unicode and numpy.str types have a 'kind' of U.
Author: PyTables
🌐
Python
docs.python.org › 3 › c-api › unicode.html
Unicode Objects and Codecs — Python 3.14.7 documentation
Translate a string by applying a character mapping table to it and return the resulting Unicode object.