The standard module unicodedata defines a lot of properties, but not everything. A quick peek at its source confirms this.

Fortunately unicodedata.txt, the data file where this comes from, is not hard to parse. Each line consists of exactly 15 elements, ; separated, which makes it ideal for parsing. Using the description of the elements on ftp://ftp.unicode.org/Public/3.0-Update/UnicodeData-3.0.0.html, you can create a few classes to encapsulate the data. I've taken the names of the class elements from that list; the meaning of each of the elements is explained on that same page.

Make sure to download ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt and ftp://ftp.unicode.org/Public/UNIDATA/Blocks.txt first, and put them inside the same folder as this program.

Code (tested with Python 2.7 and 3.6):

# -*- coding: utf-8 -*-

class UnicodeCharacter:
    def __init__(self):
        self.code = 0
        self.name = 'unnamed'
        self.category = ''
        self.combining = ''
        self.bidirectional = ''
        self.decomposition = ''
        self.asDecimal = None
        self.asDigit = None
        self.asNumeric = None
        self.mirrored = False
        self.uc1Name = None
        self.comment = ''
        self.uppercase = None
        self.lowercase = None
        self.titlecase = None
        self.block = None

    def __getitem__(self, item):
        return getattr(self, item)

    def __repr__(self):
        return '{'+self.name+'}'

class UnicodeBlock:
    def __init__(self):
        self.first = 0
        self.last = 0
        self.name = 'unnamed'

    def __repr__(self):
        return '{'+self.name+'}'

class BlockList:
    def __init__(self):
        self.blocklist = []
        with open('Blocks.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.split(';')
                    block = UnicodeBlock()
                    block.name = rawdata[1].strip()
                    rawdata = rawdata[0].split('..')
                    block.first = int(rawdata[0],16)
                    block.last = int(rawdata[1],16)
                    self.blocklist.append(block)
            # make 100% sure it's sorted, for quicker look-up later
            # (it is usually sorted in the file, but better make sure)
            self.blocklist.sort (key=lambda x: block.first)

    def lookup(self,code):
        for item in self.blocklist:
            if code >= item.first and code <= item.last:
                return item.name
        return None

class UnicodeList:
    """UnicodeList loads Unicode data from the external files
    'UnicodeData.txt' and 'Blocks.txt', both available at unicode.org

    These files must appear in the same directory as this program.

    UnicodeList is a new interpretation of the standard library
    'unicodedata'; you may first want to check if its functionality
    suffices.

    As UnicodeList loads its data from an external file, it does not depend
    on the local build from Python (in which the Unicode data gets frozen
    to the then 'current' version).

    Initialize with

        uclist = UnicodeList()
    """
    def __init__(self):

        # we need this first
        blocklist = BlockList()
        bpos = 0

        self.codelist = []
        with open('UnicodeData.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.strip().split(';')
                    parsed = UnicodeCharacter()
                    parsed.code = int(rawdata[0],16)
                    parsed.characterName = rawdata[1]
                    parsed.category = rawdata[2]
                    parsed.combining = rawdata[3]
                    parsed.bidirectional = rawdata[4]
                    parsed.decomposition = rawdata[5]
                    parsed.asDecimal = int(rawdata[6]) if rawdata[6] else None
                    parsed.asDigit = int(rawdata[7]) if rawdata[7] else None
                    # the following value may contain a slash:
                    #  ONE QUARTER ... 1/4
                    # let's make it Python 2.7 compatible :)
                    if '/' in rawdata[8]:
                        rawdata[8] = rawdata[8].replace('/','./')
                        parsed.asNumeric = eval(rawdata[8])
                    else:
                        parsed.asNumeric = int(rawdata[8]) if rawdata[8] else None
                    parsed.mirrored = rawdata[9] == 'Y'
                    parsed.uc1Name = rawdata[10]
                    parsed.comment = rawdata[11]
                    parsed.uppercase = int(rawdata[12],16) if rawdata[12] else None
                    parsed.lowercase = int(rawdata[13],16) if rawdata[13] else None
                    parsed.titlecase = int(rawdata[14],16) if rawdata[14] else None
                    while bpos < len(blocklist.blocklist) and parsed.code > blocklist.blocklist[bpos].last:
                        bpos += 1
                    parsed.block = blocklist.blocklist[bpos].name if bpos < len(blocklist.blocklist) and parsed.code >= blocklist.blocklist[bpos].first else None
                    self.codelist.append(parsed)

    def find_code(self,codepoint):
        """Find the Unicode information for a codepoint (as int).

        Returns:
            a UnicodeCharacter class object or None.
        """
        # the list is unlikely to contain duplicates but I have seen Unicode.org
        # doing that in similar situations. Again, better make sure.
        val = [x for x in self.codelist if codepoint == x.code]
        return val[0] if val else None

    def find_char(self,str):
        """Find the Unicode information for a codepoint (as character).

        Returns:
            for a single character: a UnicodeCharacter class object or
            None.
            for a multicharacter string: a list of the above, one element
            per character.
        """
        if len(str) > 1:
            result = [self.find_code(ord(x)) for x in str]
            return result
        else:
            return self.find_code(ord(str))

When loaded, you can now look up a character code with

>>> ul = UnicodeList()     # ONLY NEEDED ONCE!
>>> print (ul.find_code(0x204))
{LATIN CAPITAL LETTER E WITH DOUBLE GRAVE}

which by default is shown as the name of a character (Unicode calls this a 'code point'), but you can retrieve other properties as well:

>>> print ('%04X' % uc.find_code(0x204).lowercase)
0205
>>> print (ul.lookup(0x204).block)
Latin Extended-B

and (as long as you don't get a None) even chain them:

>>> print (ul.find_code(ul.find_code(0x204).lowercase))
{LATIN SMALL LETTER E WITH DOUBLE GRAVE}

It does not rely on your particular build of Python; you can always download an updated list from unicode.org and be assured to get the most recent information:

import unicodedata
>>> print (unicodedata.name('\U0001F903'))
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
ValueError: no such name
>>> print (uclist.find_code(0x1f903))
{LEFT HALF CIRCLE WITH FOUR DOTS}

(As tested with Python 3.5.3.)

There are currently two lookup functions defined:

  • find_code(int) looks up character information by codepoint as an integer.
  • find_char(string) looks up character information for the character(s) in string. If there is only one character, it returns a UnicodeCharacter object; if there are more, it returns a list of objects.

After import unicodelist (assuming you saved this as unicodelist.py), you can use

>>> ul = UnicodeList()
>>> hex(ul.find_char(u'è').code)
'0xe8'

to look up the hex code for any character, and a list comprehension such as

>>> l = [hex(ul.find_char(x).code) for x in 'Hello']
>>> l
['0x48', '0x65', '0x6c', '0x6c', '0x6f']

for longer strings. Note that you don't actually need all of this if all you want is a hex representation of a string! This suffices:

 l = [hex(ord(x)) for x in 'Hello']

The purpose of this module is to give easy access to other Unicode properties. A longer example:

str = 'Héllo...'
dest = ''
for i in str:
    dest += chr(ul.find_char(i).uppercase) if ul.find_char(i).uppercase is not None else i
print (dest)

HÉLLO...

and showing a list of properties for a character per your example:

letter = u'Ȅ'
print ('Name > '+ul.find_char(letter).name)
print ('Unicode number > U+%04x' % ul.find_char(letter).code)
print ('Bloc > '+ul.find_char(letter).block)
print ('Lowercase > %s' % chr(ul.find_char(letter).lowercase))

(I left out HTML; these names are not defined in the Unicode standard.)

Answer from Jongware on Stack Overflow
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
Release, 1.12,. This HOWTO discusses Python’s support for the Unicode specification for representing textual data, and explains various problems that people commonly encounter when trying to work w...
Top answer
1 of 3
5

The standard module unicodedata defines a lot of properties, but not everything. A quick peek at its source confirms this.

Fortunately unicodedata.txt, the data file where this comes from, is not hard to parse. Each line consists of exactly 15 elements, ; separated, which makes it ideal for parsing. Using the description of the elements on ftp://ftp.unicode.org/Public/3.0-Update/UnicodeData-3.0.0.html, you can create a few classes to encapsulate the data. I've taken the names of the class elements from that list; the meaning of each of the elements is explained on that same page.

Make sure to download ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt and ftp://ftp.unicode.org/Public/UNIDATA/Blocks.txt first, and put them inside the same folder as this program.

Code (tested with Python 2.7 and 3.6):

# -*- coding: utf-8 -*-

class UnicodeCharacter:
    def __init__(self):
        self.code = 0
        self.name = 'unnamed'
        self.category = ''
        self.combining = ''
        self.bidirectional = ''
        self.decomposition = ''
        self.asDecimal = None
        self.asDigit = None
        self.asNumeric = None
        self.mirrored = False
        self.uc1Name = None
        self.comment = ''
        self.uppercase = None
        self.lowercase = None
        self.titlecase = None
        self.block = None

    def __getitem__(self, item):
        return getattr(self, item)

    def __repr__(self):
        return '{'+self.name+'}'

class UnicodeBlock:
    def __init__(self):
        self.first = 0
        self.last = 0
        self.name = 'unnamed'

    def __repr__(self):
        return '{'+self.name+'}'

class BlockList:
    def __init__(self):
        self.blocklist = []
        with open('Blocks.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.split(';')
                    block = UnicodeBlock()
                    block.name = rawdata[1].strip()
                    rawdata = rawdata[0].split('..')
                    block.first = int(rawdata[0],16)
                    block.last = int(rawdata[1],16)
                    self.blocklist.append(block)
            # make 100% sure it's sorted, for quicker look-up later
            # (it is usually sorted in the file, but better make sure)
            self.blocklist.sort (key=lambda x: block.first)

    def lookup(self,code):
        for item in self.blocklist:
            if code >= item.first and code <= item.last:
                return item.name
        return None

class UnicodeList:
    """UnicodeList loads Unicode data from the external files
    'UnicodeData.txt' and 'Blocks.txt', both available at unicode.org

    These files must appear in the same directory as this program.

    UnicodeList is a new interpretation of the standard library
    'unicodedata'; you may first want to check if its functionality
    suffices.

    As UnicodeList loads its data from an external file, it does not depend
    on the local build from Python (in which the Unicode data gets frozen
    to the then 'current' version).

    Initialize with

        uclist = UnicodeList()
    """
    def __init__(self):

        # we need this first
        blocklist = BlockList()
        bpos = 0

        self.codelist = []
        with open('UnicodeData.txt','r') as uc_f:
            for line in uc_f:
                line = line.strip(' \r\n')
                if '#' in line:
                    line = line.split('#')[0].strip()
                if line != '':
                    rawdata = line.strip().split(';')
                    parsed = UnicodeCharacter()
                    parsed.code = int(rawdata[0],16)
                    parsed.characterName = rawdata[1]
                    parsed.category = rawdata[2]
                    parsed.combining = rawdata[3]
                    parsed.bidirectional = rawdata[4]
                    parsed.decomposition = rawdata[5]
                    parsed.asDecimal = int(rawdata[6]) if rawdata[6] else None
                    parsed.asDigit = int(rawdata[7]) if rawdata[7] else None
                    # the following value may contain a slash:
                    #  ONE QUARTER ... 1/4
                    # let's make it Python 2.7 compatible :)
                    if '/' in rawdata[8]:
                        rawdata[8] = rawdata[8].replace('/','./')
                        parsed.asNumeric = eval(rawdata[8])
                    else:
                        parsed.asNumeric = int(rawdata[8]) if rawdata[8] else None
                    parsed.mirrored = rawdata[9] == 'Y'
                    parsed.uc1Name = rawdata[10]
                    parsed.comment = rawdata[11]
                    parsed.uppercase = int(rawdata[12],16) if rawdata[12] else None
                    parsed.lowercase = int(rawdata[13],16) if rawdata[13] else None
                    parsed.titlecase = int(rawdata[14],16) if rawdata[14] else None
                    while bpos < len(blocklist.blocklist) and parsed.code > blocklist.blocklist[bpos].last:
                        bpos += 1
                    parsed.block = blocklist.blocklist[bpos].name if bpos < len(blocklist.blocklist) and parsed.code >= blocklist.blocklist[bpos].first else None
                    self.codelist.append(parsed)

    def find_code(self,codepoint):
        """Find the Unicode information for a codepoint (as int).

        Returns:
            a UnicodeCharacter class object or None.
        """
        # the list is unlikely to contain duplicates but I have seen Unicode.org
        # doing that in similar situations. Again, better make sure.
        val = [x for x in self.codelist if codepoint == x.code]
        return val[0] if val else None

    def find_char(self,str):
        """Find the Unicode information for a codepoint (as character).

        Returns:
            for a single character: a UnicodeCharacter class object or
            None.
            for a multicharacter string: a list of the above, one element
            per character.
        """
        if len(str) > 1:
            result = [self.find_code(ord(x)) for x in str]
            return result
        else:
            return self.find_code(ord(str))

When loaded, you can now look up a character code with

>>> ul = UnicodeList()     # ONLY NEEDED ONCE!
>>> print (ul.find_code(0x204))
{LATIN CAPITAL LETTER E WITH DOUBLE GRAVE}

which by default is shown as the name of a character (Unicode calls this a 'code point'), but you can retrieve other properties as well:

>>> print ('%04X' % uc.find_code(0x204).lowercase)
0205
>>> print (ul.lookup(0x204).block)
Latin Extended-B

and (as long as you don't get a None) even chain them:

>>> print (ul.find_code(ul.find_code(0x204).lowercase))
{LATIN SMALL LETTER E WITH DOUBLE GRAVE}

It does not rely on your particular build of Python; you can always download an updated list from unicode.org and be assured to get the most recent information:

import unicodedata
>>> print (unicodedata.name('\U0001F903'))
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
ValueError: no such name
>>> print (uclist.find_code(0x1f903))
{LEFT HALF CIRCLE WITH FOUR DOTS}

(As tested with Python 3.5.3.)

There are currently two lookup functions defined:

  • find_code(int) looks up character information by codepoint as an integer.
  • find_char(string) looks up character information for the character(s) in string. If there is only one character, it returns a UnicodeCharacter object; if there are more, it returns a list of objects.

After import unicodelist (assuming you saved this as unicodelist.py), you can use

>>> ul = UnicodeList()
>>> hex(ul.find_char(u'è').code)
'0xe8'

to look up the hex code for any character, and a list comprehension such as

>>> l = [hex(ul.find_char(x).code) for x in 'Hello']
>>> l
['0x48', '0x65', '0x6c', '0x6c', '0x6f']

for longer strings. Note that you don't actually need all of this if all you want is a hex representation of a string! This suffices:

 l = [hex(ord(x)) for x in 'Hello']

The purpose of this module is to give easy access to other Unicode properties. A longer example:

str = 'Héllo...'
dest = ''
for i in str:
    dest += chr(ul.find_char(i).uppercase) if ul.find_char(i).uppercase is not None else i
print (dest)

HÉLLO...

and showing a list of properties for a character per your example:

letter = u'Ȅ'
print ('Name > '+ul.find_char(letter).name)
print ('Unicode number > U+%04x' % ul.find_char(letter).code)
print ('Bloc > '+ul.find_char(letter).block)
print ('Lowercase > %s' % chr(ul.find_char(letter).lowercase))

(I left out HTML; these names are not defined in the Unicode standard.)

2 of 3
3

The unicodedata documentation shows how to do most of this.

The Unicode block name is apparently not available but another Stack Overflow question has a solution of sorts and another has some additional approaches using regex.

The uppercase/lowercase mapping and character number information is not particularly Unicode-specific; just use the regular Python string functions.

So in summary

>>> import unicodedata
>>> unicodedata.name('Ë')
'LATIN CAPITAL LETTER E WITH DIAERESIS'
>>> 'U+%04X' % ord('Ë')
'U+00CB'
>>> '&#%i;' % ord('Ë')
'&#203;'
>>> 'Ë'.lower()
'ë'

The U+%04X formatting is sort-of correct, in that it simply avoids padding and prints the whole hex number for code points with a value higher than 65,535. Note that some other formats require the use of %08X padding in this scenario (notably \U00010000 format in Python).

Discussions

Help me understand - printing Unicode characters to draw a letter
Looks overly complicated to me, but I am probably missing something. For printing, I'd go with, table = pd.read_html(url, header =0, flavor ='bs4') rows = table[0].sort_values(['y-coordinate','x-coordinate'], ascending=[False, True]) nl = "" # newline when x returns to 0 for i, row in rows.iterrows(): if row['x-coordinate'] == 0: print(end=nl) nl = '\n' print(row['Character'], end="") print() More on reddit.com
🌐 r/learnpython
15
3
September 21, 2024
Return unicode strings from stored bytestrings
There was an error while loading. Please reload this page More on github.com
🌐 github.com
16
September 9, 2015
Explain it like I'm five: Python and Unicode?
There are two types of strings in python: byte strings and unicode strings. Each element in a byte string is a byte. There are only 256 possible bytes. Each element in a unicode string is a character (also called a unicode code point). There are a little over a million characters defined in unicode. Meaning each element/character in a unicode string can be one of those million characters. Byte strings are useful because you can write them to files, transmit them over the network, etc. Unicode strings are useful because you can store pretty much any character that exists. So people usually like to manipulate unicode strings in their programs. But how do you convert a unicode string to a byte string? You encode it. An encoding is a representation of a unicode string. It defines a byte or byte sequence for every* unicode code point; essentially a translation table. For every unicode code point, there is a byte or sequence of bytes. There's more to it than that, but those are the essential bits you need to know. What this means when you're writing a program is that you want to manipulate unicode strings throughout, and when you want to output a string (to a file, or over the network), you encode it. When you read in a byte string from external sources, you decode it. Does that make sense? *some encodings may not support every unicode character; they may only support some subset of unicode. UTF-8 is nice because it supports everything. It defines a sequence of bytes for every unicode character. More on reddit.com
🌐 r/Python
60
106
June 12, 2013
Top answer
1 of 1
4

First pass

Good job in doing a first pass at type hinting! In the newest stable version of Python (3.10 as of this writing) it's no longer necessary to import List - you can hint with the built-in list. However, your head and body shouldn't really be hinted as lists (which imply mutability), but instead Sequence, which will also accept immutable tuples.

When you do simple member assignment in a constructor from parameters, the members don't also need type hints, and their type can be inferred.

Your # Table symbols can all be deleted since you don't use them. I think the code is quite legible without declaring these as constants. If you were to keep (and start using) these constants, you would want to move them out to static scope before the constructor.

Your print is not general-purpose enough. For example, if someone wants to write this table out to a file, that will be difficult. One convenient (and possibly the highest-performing) way of rewriting this is as an iterator of lines. The caller can decide to either iterate over each line and do something with it; or call into a wrapper utility method that joins the lines and returns a single string.

The first three f-strings in your print method have trivial wrappers with a very long, single field in the middle. This is difficult to read, and you're better off moving that field content to a variable. Bonus: the variable name will self-document the line after it, so you can remove your comments.

When you join on a generator, don't also wrap the generator into a list comprehension. Generators are perfectly capable of being passed bare into join.

Don't for i, _ in enumerate; this is a job for a simple i in range(len.

In your main loop, don't index + 1; pass a start parameter to enumerate.

In your body loop, rather than an enumerate, zip together your widths and elements.

find_widths is a good candidate for being re-expressed as an iterator function. Store it to a tuple and not a list in the constructor.

Your test table content in main should (almost) all be converted to tuples, with the exception of your outer body list comprehension since there's no such thing as a tuple comprehension, so a list is more convenient.

from typing import Sequence, Iterator
import random


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def lines(self) -> Iterator[str]:
        top = '┬'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"┌{top}┐"

        header = '│'.join(
            title.center(width)
            for title, width in zip(self.head, self.column_widths)
        )
        yield f"│{header}│"

        separator = '┼'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"├{separator}┤"

        for row in self.body:
            row_str = "│".join(
                element.ljust(width, ' ')
                for element, width in zip(row, self.column_widths)
            )
            yield f"│{row_str}│"

        tail = '┴'.join(
            '─' * self.column_widths[i]
            for i in range(len(self.head))
        )
        yield f"└{tail}┘"

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        "Alpha", "Bravo", "Charlie", "Delta", "Echo",
        "Foxtrot", "Golf", "Hotel", "India", "Juliet",
        "Kilo", "Lima", "Mike", "November", "Oscar",
        "Papa", "Quebec", "Romeo", "Sierra", "Tango",
        "Uniform", "Victor", "Whiskey", "X-Ray", "Yankee",
        "Zulu",
    )
    table1: Table = Table(
        head=("Id", "Name ", "Marks"),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == "__main__":
    main()

Second pass

Now, take a look at the commonalities in your formatting code. You basically only have two operations: making a hyphen-separator, and making a line with content. Put these in utility functions rather than copy-pasting the code:

from typing import Sequence, Iterator, Iterable
import random


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def make_separator(self, left: str, mid: str, right: str) -> str:
        header = mid.join(
            '─' * width for width in self.column_widths
        )
        return f'{left}{header}{right}'

    @staticmethod
    def make_line(elements: Iterable[str]) -> str:
        line = '│'.join(elements)
        return f'│{line}│'

    def lines(self) -> Iterator[str]:
        yield self.make_separator(left='┌', mid='┬', right='┐')

        yield self.make_line(
            title.center(width)
            for title, width in zip(self.head, self.column_widths)
        )

        yield self.make_separator(left='├', mid='┼', right='┤')

        for row in self.body:
            yield self.make_line(
                element.ljust(width, ' ')
                for element, width in zip(row, self.column_widths)
            )

        yield self.make_separator(left='└', mid='┴', right='┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

Third pass

Recognise that there is still some repetition: we're zipping over widths both times that we call the line-making utility function. Move that zip to within the function, and cut away the formatting differences to their own functions, passing references to the functions into the line-making method. Also, you can splat a three-character string into your function arguments for shorter invocation.

from typing import Sequence, Iterator, Iterable, Callable
import random


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
│    2502  BOX DRAWINGS LIGHT VERTICAL
          
┌    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
├    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
┼    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
┤    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
└    2514  BOX DRAWINGS LIGHT UP AND RIGHT
┴    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
┘    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        self.column_widths: tuple[int] = tuple(self.find_widths())

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    def make_separator(self, left: str, mid: str, right: str) -> str:
        line = mid.join(
            '─' * width for width in self.column_widths
        )
        return f'{left}{line}{right}'

    @staticmethod
    def format_header(title: str, width: int) -> str:
        return title.center(width)

    @staticmethod
    def format_element(element: str, width: int) -> str:
        return element.ljust(width)

    def make_line(
        self,
        elements: Iterable[str],
        format: Callable[[str, int], str],
    ) -> str:
        line = '│'.join(
            format(element, width)
            for element, width in zip(elements, self.column_widths)
        )
        return f'│{line}│'

    def lines(self) -> Iterator[str]:
        yield self.make_separator(*'┌┬┐')
        yield self.make_line(self.head, self.format_header)
        yield self.make_separator(*'├┼┤')

        for row in self.body:
            yield self.make_line(row, self.format_element)

        yield self.make_separator(*'└┴┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

A word on Unicode

Though I'm still not convinced you should be using variable constants for your drawing characters, it would be informative and helpful to include a docstring like the one I showed at the top of the last code block.

Overall it's vaguely safe to use these characters for rendering in fixed-width terminals of modern machines. In narrow cases this might break; for interesting examples read Misalignment of Unicode block characters in preformatted text blocks.

When you carry this code and its output around editors, IDEs and source control, take care to preserve UTF-8 encoding. The output of the last sample seems byte-for-byte equivalent:

┌────┬────────┬─────┐
│ Id │ Name   │Marks│
├────┼────────┼─────┤
│ 1  │Alpha   │   74│
│ 2  │Bravo   │    8│
│ 3  │Charlie │   90│
│ 4  │Delta   │   75│
│ 5  │Echo    │   28│
│ 6  │Foxtrot │   84│
│ 7  │Golf    │   92│
│ 8  │Hotel   │   38│
│ 9  │India   │    6│
│ 10 │Juliet  │   59│
│ 11 │Kilo    │    0│
│ 12 │Lima    │   85│
│ 13 │Mike    │   33│
│ 14 │November│   81│
│ 15 │Oscar   │   39│
│ 16 │Papa    │   60│
│ 17 │Quebec  │   18│
│ 18 │Romeo   │   54│
│ 19 │Sierra  │   36│
│ 20 │Tango   │   50│
│ 21 │Uniform │   53│
│ 22 │Victor  │   83│
│ 23 │Whiskey │   85│
│ 24 │X-Ray   │   16│
│ 25 │Yankee  │   68│
│ 26 │Zulu    │   36│
└────┴────────┴─────┘

True Formatting

Your original approach (and all of the approaches above) rely on explicit string repetition through the * operator. There is a very different approach that forms true formatting strings defining alignment, width and padding. For details on these parameters read the format mini-language specification.

Once this formatting string is defined, you can pre-bind to its .format() method and hold a reference to that method. Since the formatting strings encode the column widths, you don't actually need to store the widths on the class and can drop them after the constructor.

I don't strongly consider this approach universally better or worse.

from typing import Sequence, Iterator, Iterable, Callable
import random


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
│    2502  BOX DRAWINGS LIGHT VERTICAL
          
┌    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
├    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
┼    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
┤    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
└    2514  BOX DRAWINGS LIGHT UP AND RIGHT
┴    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
┘    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(self, head: Sequence[str], body: Sequence[Sequence[str]]) -> None:
        self.head = head
        self.body = body
        widths = tuple(self.find_widths())
        self.format_sep = self.make_separator_format(widths)
        self.format_head = self.make_content_format(widths, align='^')
        self.format_body = self.make_content_format(widths, align='<')

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    @staticmethod
    def make_separator_format(widths: Sequence[int]) -> Callable:
        fmt = (
            '{0}'
            + ''.join(
                '{1:─>%d}' % (1 + width)
                for width in widths[:-1]
            )
            + '{2:─>%d}' % (1 + widths[-1])
        )
        return fmt.format

    @staticmethod
    def make_content_format(widths: Sequence[int], align: str) -> Callable:
        fmt = (
            '│'
            + '│'.join(
                '{:%s%d}' % (align, width)
                for width in widths
            )
            + '│'
        )
        return fmt.format

    def lines(self) -> Iterator[str]:
        yield self.format_sep(*'┌┬┐')
        yield self.format_head(*self.head)
        yield self.format_sep(*'├┼┤')

        for row in self.body:
            yield self.format_body(*row)

        yield self.format_sep(*'└┴┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (
                str(index).center(4),
                name,
                str(random.randint(0, 100)).rjust(5, ' ')
            ) for index, name in enumerate(names, 1)
        ],
    )
    print(table1)


if __name__ == '__main__':
    main()

If you debug, you can see the format strings it makes:

{0}{1:─>5}{1:─>9}{2:─>6}
│{:^4}│{:^8}│{:^5}│
│{:<4}│{:<8}│{:<5}│

Formatting responsibility

Currently your main function has stolen a little bit of the responsibility to format the cell content. This is a little awkward, and should just be transferred to the table class.

from typing import Sequence, Iterator, Callable, Optional, Any
from random import randint


"""
Unicode characters used:

Char Code  Name
─    2500  BOX DRAWINGS LIGHT HORIZONTAL
│    2502  BOX DRAWINGS LIGHT VERTICAL
          
┌    250C  BOX DRAWINGS LIGHT DOWN AND RIGHT
┬    252C  BOX DRAWINGS LIGHT DOWN AND HORIZONTAL
┐    2510  BOX DRAWINGS LIGHT DOWN AND LEFT 
          
├    251C  BOX DRAWINGS LIGHT VERTICAL AND RIGHT
┼    253C  BOX DRAWINGS LIGHT VERTICAL AND HORIZONTAL
┤    2524  BOX DRAWINGS LIGHT VERTICAL AND LEFT
          
└    2514  BOX DRAWINGS LIGHT UP AND RIGHT
┴    2534  BOX DRAWINGS LIGHT UP AND HORIZONTAL
┘    2518  BOX DRAWINGS LIGHT UP AND LEFT
"""


class Table:
    def __init__(
        self,
        head: Sequence[str],
        body: Sequence[Sequence[Any]],
        formats: Sequence[Optional[str]] = (),
    ) -> None:
        self.head = head
        self.body = tuple(self.format_body(body, formats))
        widths = tuple(self.find_widths())
        self.format_sep = self.make_separator_format(widths)
        self.format_head = self.make_content_format(widths, align='^')
        self.format_body = self.make_content_format(widths, align='<')

    @staticmethod
    def format_body(
        body: Sequence[Sequence[Any]],
        formats: Sequence[Optional[str]],
    ) -> Iterator[Sequence[str]]:
        if not formats:
            formats = (None,) * len(body[0])

        for row in body:
            yield [
                fmt.format(cell) if fmt else str(cell)
                for cell, fmt in zip(row, formats)
            ]

    def __str__(self) -> str:
        return '\n'.join(self.lines())

    @staticmethod
    def make_separator_format(widths: Sequence[int]) -> Callable:
        fmt = (
            '{0}'
            + ''.join(
                '{1:─>%d}' % (1 + width)
                for width in widths[:-1]
            )
            + '{2:─>%d}' % (1 + widths[-1])
        )
        return fmt.format

    @staticmethod
    def make_content_format(
        widths: Sequence[int],
        align: str,
    ) -> Callable:
        fmt = (
            '│'
            + '│'.join(
                '{:%s%d}' % (align, width)
                for width in widths
            )
            + '│'
        )
        return fmt.format

    def lines(self) -> Iterator[str]:
        yield self.format_sep(*'┌┬┐')
        yield self.format_head(*self.head)
        yield self.format_sep(*'├┼┤')

        for row in self.body:
            yield self.format_body(*row)

        yield self.format_sep(*'└┴┘')

    def find_widths(self) -> Iterator[int]:
        table = (*self.body, self.head)
        for index in range(len(self.head)):
            yield max(len(row[index]) for row in table)


def main() -> None:
    names = (
        'Alpha', 'Bravo', 'Charlie', 'Delta', 'Echo',
        'Foxtrot', 'Golf', 'Hotel', 'India', 'Juliet',
        'Kilo', 'Lima', 'Mike', 'November', 'Oscar',
        'Papa', 'Quebec', 'Romeo', 'Sierra', 'Tango',
        'Uniform', 'Victor', 'Whiskey', 'X-Ray', 'Yankee',
        'Zulu',
    )
    table1: Table = Table(
        head=('Id', 'Name ', 'Marks'),
        body=[
            (index, name, randint(0, 100))
            for index, name in enumerate(names, 1)
        ],
        formats=('{:^4}', None, '{:>5}'),
    )
    print(table1)


if __name__ == '__main__':
    main()
🌐
GitHub
github.com › MerlijnWajer › unitable
GitHub - MerlijnWajer/unitable: Python code to transform Python lists into a Unicode Table
Simple program to turn python lists into unicode tables. Planning to add support for Restructured Text tables later. Usage in python file. Example output: ┌──────┬────────┬──────────────────────────┐ │ Foo │ Bar │ Quux │ ├──────┼────────┼──────────────────────────┤ │ Baz │ Wubble │ Wobble │ │ Baz │ Wubble │ Wobble │ │ Baz │ Wubble │ Wobble │ │ Baz │ Wubble │ Wobble │ │ Baz │ Wubble │ WobbleWobbleWobbleWobble │ └──────┴────────┴──────────────────────────┘
Author: MerlijnWajer
🌐
Inspired Python
inspiredpython.com › tip › python-strings-using-translation-tables-to-make-unicode-text
Python Strings: Using translation tables to make Unicode text • Inspired Python
>>> import string >>> before = string.ascii_lowercase + string.digits >>> after = '🅐🅑🅒🅓🅔🅕🅖🅗🅘🅙🅚🅛🅜🅝🅞🅟🅠🅡🅢🅣🅤🅥🅦🅧🅨🅩⓿➊➋➌➍➎➏➐➑➒' >>> translation_table = str.maketrans(before, after) >>> 'inspired python'.translate(translation_table) '🅘🅝🅢🅟🅘🅡🅔🅓 🅟🅨🅣🅗🅞🅝'
🌐
python-tcod
python-tcod.readthedocs.io › en › latest › tcod › charmap-reference.html
Character Table Reference - python-tcod 21.2.1 documentation
Unicode is the Unicode code point as a hexadecimal number. You can use chr to convert these numbers into a string. Character maps such as tcod.tileset.CHARMAP_CP437 are simply a list of Unicode numbers, where the index of the list is the Tile Index. String is the Python string for that character.
🌐
Real Python
realpython.com › python-encodings-guide
Unicode & Character Encodings in Python: A Painless Guide – Real Python
May 20, 2019 - For \xhh, \uxxxx, and \Uxxxxxxxx, exactly as many digits are required as are shown in these examples. This can throw you for a loop because of the way that Unicode tables conventionally display the codes for characters, with a leading U+ and variable number of hex characters.
Find elsewhere
🌐
Pythonturtle
pythonturtle.academy › unicode-table
Unicode Table – Python and Turtle
February 27, 2019 - Draw a 16×16 table of unicode symbols. Unicode starts from number 0x2600 (Hexadecimal). You can convert number to text with chr() function. You can start with different number to find more unicode symbols. Depending on your Operating System, symbols may look different.
🌐
Python Cheat Sheet
pythonsheets.com › notes › basic › python-unicode.html
Unicode — Python Cheat Sheet
For example, the character, é can be written as e ́ (Canonical Decomposition) or é (Canonical Composition). In this case, we may acquire unexpected results when we are comparing two strings even though they look alike.
🌐
GitHub
gist.github.com › arrowtype › 713dad14fe9a574d58d1aab61ba9b2f0
The basics of working with unicode values in Python · GitHub
# and import unicodedata x = unicodedata.category('A') print(x) # Ll -- lowercase # Lu -- uppercase # Lt -- titlecase # Lm -- modifier # Lo -- other
🌐
Rmotr
learn.rmotr.com › python › understanding-unicode-in-python › strings-and-unicode › unicode-in-python
Understanding Unicode in Python > Strings and Unicode > Unicode in Python
Here is a stub of a table from unicode-table.com that shows the first 128 symbols: For example, the character A has the code point 0041. Usually, when specifying unicode code points we use the U+ prefix, and we say that the code point is U+0041. We can also see this with Python:
🌐
Python
docs.python.org › 3.3 › howto › unicode.html
Unicode HOWTO — Python 3.3.7 documentation
September 19, 2017 - A code point is an integer value, usually denoted in base 16. In the standard, a code point is written using the notation U+12CA to mean the character with value 0x12ca (4,810 decimal). The Unicode standard contains a lot of tables listing characters and their corresponding code points:
🌐
Python
docs.python.org › 3.8 › library › unicodedata.html
unicodedata — Unicode Database — Python 3.8.20 documentation
>>> import unicodedata >>> unicodedata.lookup('LEFT CURLY BRACKET') '{' >>> unicodedata.name('/') 'SOLIDUS' >>> unicodedata.decimal('9') 9 >>> unicodedata.decimal('a') Traceback (most recent call last): File "<stdin>", line 1, in <module> ValueError: not a decimal >>> unicodedata.category('A') # 'L'etter, 'u'ppercase 'Lu' >>> unicodedata.bidirectional('\u0660') # 'A'rabic, 'N'umber 'AN' ... © Copyright 2001-2025, Python Software Foundation. This page is licensed under the Python Software Foundation License Version 2. Examples, recipes, and other code in the documentation are additionally licensed under the Zero Clause BSD License.
🌐
Python
docs.python.org › 3.11 › library › unicodedata.html
unicodedata — Unicode Database
>>> import unicodedata >>> unicodedata.lookup('LEFT CURLY BRACKET') '{' >>> unicodedata.name('/') 'SOLIDUS' >>> unicodedata.decimal('9') 9 >>> unicodedata.decimal('a') Traceback (most recent call last): File "<stdin>", line 1, in <module> ValueError: not a decimal >>> unicodedata.category('A') # 'L'etter, 'u'ppercase 'Lu' >>> unicodedata.bidirectional('\u0660') # 'A'rabic, 'N'umber 'AN' ... © Copyright 2001-2025, Python Software Foundation. This page is licensed under the Python Software Foundation License Version 2. Examples, recipes, and other code in the documentation are additionally licensed under the Zero Clause BSD License.
🌐
Pythonlang
docs.pythonlang.net › 3 › library › unicodedata.html
unicodedata — Unicode Database — Python 3.14.0 documentation
The Unicode HOWTO for more information about Unicode and how to use this module. ... Look up character by name. If a character with the given name is found, return the corresponding character. If not found, KeyError is raised. For example:
🌐
Python
docs.python.org › 3.0 › howto › unicode.html
Unicode HOWTO — Python v3.0.1 documentation
A code point is an integer value, usually denoted in base 16. In the standard, a code point is written using the notation U+12ca to mean the character with value 0x12ca (4810 decimal). The Unicode standard contains a lot of tables listing characters and their corresponding code points:
🌐
Python
docs.python.org › 3.9 › library › unicodedata.html
unicodedata — Unicode Database — Python 3.9.23 documentation
>>> import unicodedata >>> unicodedata.lookup('LEFT CURLY BRACKET') '{' >>> unicodedata.name('/') 'SOLIDUS' >>> unicodedata.decimal('9') 9 >>> unicodedata.decimal('a') Traceback (most recent call last): File "<stdin>", line 1, in <module> ValueError: not a decimal >>> unicodedata.category('A') # 'L'etter, 'u'ppercase 'Lu' >>> unicodedata.bidirectional('\u0660') # 'A'rabic, 'N'umber 'AN' ... © Copyright 2001-2025, Python Software Foundation. This page is licensed under the Python Software Foundation License Version 2. Examples, recipes, and other code in the documentation are additionally licensed under the Zero Clause BSD License.
🌐
Reddit
reddit.com › r/learnpython › help me understand - printing unicode characters to draw a letter
r/learnpython on Reddit: Help me understand - printing Unicode characters to draw a letter
September 21, 2024 -

Hi all,

I am relatively new to programming, mostly just learning by examples I find online and looking up concepts. I have been working on this project for a little bit and I am really struggling to understand how to loop properly to print out what I need in the right way.

The task is this:

write a function that takes in the URL for such a Google Doc as an argument, retrieves and parses the data in the document, and prints the grid of characters.
When printed in a fixed-width font, the characters in the grid will form a graphic showing a sequence of uppercase letters, which is the secret message.

The document specifies the Unicode characters in the grid, along with the x- and y-coordinates of each character.

The minimum possible value of these coordinates is 0. There is no maximum possible value, so the grid can be arbitrarily large.

Any positions in the grid that do not have a specified character should be filled with a space character.

You can assume the document will always have the same format as the example document linked above.

For example, the simplified example document linked above draws out the letter 'F':

█▀▀▀
█▀▀
█  

Note that the coordinates (0, 0) will always correspond to the same corner of the grid as in this example, so make sure to understand in which directions the x- and y-coordinates
increase.

Specifications

Your code must be written in Python (preferred) or JavaScript.

You may use external libraries.

You may write helper functions, but there should be one function that:

  1. Takes in one argument, which is a string containing the URL for the Google Doc with the input data, AND

  2. When called, prints the grid of characters specified by the input data, displaying a graphic of correctly oriented uppercase letters.'''

The URL i am using is this https://docs.google.com/document/d/e/2PACX-1vRMx5YQlZNa3ra8dYYxmv-QIQ3YJe8tbI3kqcuC7lQiZm-CSEznKfN_HYNSpoXcZIV3Y_O3YoUB1ecq/pub

and this is what I've done so far:

import pandas as pd
from bs4 import BeautifulSoup

#this function uses the pandas library to take a url and read the table.
def pull_link(url):
    table = pd.read_html(url, header =0, flavor ='bs4')
    datatable = table[0].sort_values(by=['x-coordinate','y-coordinate'], ignore_index=True)
    xcoords = datatable['x-coordinate']
    character = datatable['Character']
    ycoords = datatable['y-coordinate']


    for i in range(0,len(character)):
        if (xcoords[i] == 0):
            print(character[i])
        #if (xcoords[i] == xcoords[i-1]) and (ycoords[i] - ycoords[i-1] == 1):
         #   print(character[i], end='')
        #if ycoords[i] > 0:
            #if (xcoords[i] == 0 and (ycoords[i] - ycoords[i-1] == 1)):
                #print(character[i])
        #if xcoords[i] > 0:
         #   if xcoords[i] - xcoords[i-1] != 1:
          #      print(character[i], end='')
            
        #print(character[i],end='')

    #return xcoords,character,ycoords
     
print(pull_link('https://docs.google.com/document/d/e/2PACX-1vRMx5YQlZNa3ra8dYYxmv-QIQ3YJe8tbI3kqcuC7lQiZm-CSEznKfN_HYNSpoXcZIV3Y_O3YoUB1ecq/pub'))

I'm able to parse the table into variables, but i cannot understand how to loop through this, I've put a bunch of code into comments cause i've been messing around with different ideas but nothing is working for me. I appreciate any tips or if anyone knows what type of topics this exercise covers i'd be interested in learning more about it.

🌐
GitHub
github.com › PyTables › PyTables › issues › 499
Return unicode strings from stored bytestrings · Issue #499 · PyTables/PyTables
September 9, 2015 - Return unicode strings from stored bytestrings#499 · Copy link · Labels · python3strings · dotsdl · opened · on Sep 8, 2015 · Issue body actions · Storing strings in a table in Python 3 with something like: import tables class Tags(tables.IsDescription): tag = tables.StringCol(namelength) handle = tables.open_file('testfile.h5', 'a') table = handle.create_table('/', 'tags', Tags, 'tags') newtags = ['bark', 'lark', 'snark'] for tag in newtags: table.row['tag'] = tag table.row.append() handle.close() and then reading them back with: handle = tables.open_file('testfile.h5', 'r') table = handle.get_node('/', 'tags') tags = [x['tag'] for x in table.read()] print(tags[0]) yields b'bark' instead of 'bark'.
Author: PyTables
🌐
UW PCE
uwpce-pythoncert.github.io › SystemDevelopment › unicode.html
Unicode in Python 2 — System Development With Python 2.0 documentation
Lots of tables of code points online: One example: http://inamidst.com/stuff/unidata/ hello_unicode.py. Use unicode objects in all your code · Decode on input · Encode on output · Many packages do this for you: XML processing, databases, ... Gotcha: Python has a default encoding (usually ...