In Python 2

>>> plain_string = "Hi!"
>>> unicode_string = u"Hi!"
>>> type(plain_string), type(unicode_string)
(<type 'str'>, <type 'unicode'>)

^ This is the difference between a byte string (plain_string) and a unicode string.

>>> s = "Hello!"
>>> u = unicode(s, "utf-8")

^ Converting to unicode and specifying the encoding.

In Python 3

All strings are unicode. The unicode function does not exist anymore. See answer from @Noumenon

Answer from user225312 on Stack Overflow
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-a-string-to-utf-8-in-python
Convert a String to Utf-8 in Python - GeeksforGeeks
July 23, 2025 - Converting a string to UTF-8 in Python is a simple task with multiple methods at your disposal. Whether you choose the encode method, the bytes constructor, or the str.encode method, the key is to specify the UTF-8 encoding.
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.8 documentation
Emacs supports many different variables, but Python only supports ‘coding’. The -*- symbols indicate to Emacs that the comment is special; they have no significance to Python but are a convention. Python looks for coding: name or coding=name in the comment. If you don’t include such a comment, the default encoding used will be UTF-8 as already mentioned. See also PEP 263 for more information. The Unicode specification includes a database of information about code points.
🌐
Finxter
blog.finxter.com › home › learn python blog › python convert unicode to bytes, ascii, utf-8, raw string
Python Convert Unicode to Bytes, ASCII, UTF-8, Raw String - Be on the Right Side of Change
June 30, 2021 - Due to the fact that UTF-8 encoding ... For this task, both of the above methods are applicable. With encode(), we first get a byte string by applying UTF-8 encoding to the input Unicode string, and then use decode(), which will give us a UTF-8 encoded Unicode string that is already ...
Find elsewhere
🌐
Medium
medium.com › @nawazmohtashim › method-to-encode-a-string-to-utf-8-in-python-b287027b7be9
Method to Encode a String to UTF-8 in Python | by Mohd Mohtashim Nawaz | Medium
February 2, 2024 - While encoding a string to UTF-8, it’s essential to consider potential errors that may arise. The encode() method allows you to handle errors by specifying the errors parameter. Common error handling options include: 'strict': Raises a UnicodeEncodeError if an error occurs (default behavior).
Top answer
1 of 2
2
  • The L postfixes signify long integers. They are the same thing as (short) integers really; there really is no need to convert these. It is only their repr() output that includes the L; print the value directly or write it to a file and the L postfix is not included.

  • Unicode values can be encoded to UTF-8 with the unicode.encode() method:

    encoded = unicodestr.encode('utf8')
    

Your beef is with the list representation here; you logged all rows, and Python containers represent their contents by calling repr() on each value. These representations are great for debugging as their types are made obvious.

It depends on what you do with these values next. It is generally a good idea to use Unicode throughout your code, and only encode at the last moment (when writing to a file, or printing or sending over the network). A good many methods handle this for you. Printing will encode to your terminal codec automatically, for example. When adding to an XML file, most XML libraries handle Unicode for you. Etc.

2 of 2
0

Your problem is that you are trying to display data, INSTEAD you are displaying python representation if this object.

So it contains meta-data like u, L, etc. If you want to display data the way you want, you should write a code to deal with it.

For example:

for row in cur.fetchall():
    print u"'{row[0]}', '{row[1]}', '{row[2]}', '{row[3]}', '{row[4]}'".format(row=row)

So it will look like

'1', '2', '3', '4'
'1', '2', '3', '4'
'1', '2', '3', '4'

But... as I can see, you make structure look like CSV-file(comma-separated values), do you? So, maybe, you should read about csv python module?

🌐
Medium
medium.com › @agustinb › introduction-to-unicode-and-utf-8-in-python-9e7a844edddd
Introduction to Unicode and UTF-8 in Python 2 | by agustinb | Medium
December 1, 2018 - Python 2 does implicit decoding in order to make a single Unicode object, but its default codec is ASCII (you can run sys.getdefaultencoding() to check). So, in the first example, ‘world’ wasn’t a problem but ‘π’ cannot be decoded using ASCII. Codec is a shortcut for Encoder/Decoder. >>> content = '\xcf\x80-zza'.decode('utf-8') # π-zza>>> type(content) <type 'unicode'>>>> print content π-zza>>> output_string = content.encode('utf-8')>>> type(output_string) <type 'str'>>>> output_string '\xcf\x80-zza' # bytes!
🌐
Java2Blog
java2blog.com › home › python › python string › encode string to utf-8 in python
Encode String to UTF-8 in Python [2 ways] - Java2Blog
December 25, 2022 - The same is not the case for Python 2. In this version, bytes and string are basically the same thing. So this, function is redundant since the string is already encoded. We can observe the same in the code below. ... To encode string to UTF-8 in Python, use the codecs.encode() function.
🌐
Python Forum
python-forum.io › thread-11676.html
unicode to utf-8
July 21, 2018 - if i have a list with a series of Unicode code point values as ints and want to convert to a list of utf-8 code values as ints, how could i achieve this in Python? the reverse would be nice, too.
🌐
O'Reilly
oreilly.com › library › view › python-cookbook › 0596001673 › ch03s18.html
Converting Between Unicode and Plain Strings - Python Cookbook [Book]
July 19, 2002 - # Convert Unicode to plain Python string: "encode" unicodestring = u"Hello world" utf8string = unicodestring.encode("utf-8") asciistring = unicodestring.encode("ascii") isostring = unicodestring.encode("ISO-8859-1") utf16string = unicodestring.encode("utf-16") # Convert plain Python string to Unicode: "decode" plainstring1 = unicode(utf8string, "utf-8") plainstring2 = unicode(asciistring, "ascii") plainstring3 = unicode(isostring, "ISO-8859-1") plainstring4 = unicode(utf16string, "utf-16") assert plainstring1==plainstring2==plainstring3==plainstring4 ·
Authors: Alex MartelliDavid Ascher
Published: 2002
Pages: 608
Top answer
1 of 2
65

You don't need to encode data that is already encoded. When you try to do that, Python will first try to decode it to unicode before it can encode it back to UTF-8. That is what is failing here:

>>> data = u'\u00c3'            # Unicode data
>>> data = data.encode('utf8')  # encoded to UTF-8
>>> data
'\xc3\x83'
>>> data.encode('utf8')         # Try to *re*-encode it
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0: ordinal not in range(128)

Just write your data directly to the file, there is no need to encode already-encoded data.

If you instead build up unicode values instead, you would indeed have to encode those to be writable to a file. You'd want to use codecs.open() instead, which returns a file object that will encode unicode values to UTF-8 for you.

You also really don't want to write out the UTF-8 BOM, unless you have to support Microsoft tools that cannot read UTF-8 otherwise (such as MS Notepad).

For your MySQL insert problem, you need to do two things:

  • Add charset='utf8' to your MySQLdb.connect() call.

  • Use unicode objects, not str objects when querying or inserting, but use sql parameters so the MySQL connector can do the right thing for you:

    artiste = artiste.decode('utf8')  # it is already UTF8, decode to unicode
    
    c.execute('SELECT COUNT(id) AS nbr FROM artistes WHERE nom=%s', (artiste,))
    
    # ...
    
    c.execute('INSERT INTO artistes(nom,status,path) VALUES(%s, 99, %s)', (artiste, artiste + u'/'))
    

It may actually work better if you used codecs.open() to decode the contents automatically instead:

import codecs

sql = mdb.connect('localhost','admin','ugo&(-@F','music_vibration', charset='utf8')

with codecs.open('config/index/'+index, 'r', 'utf8') as findex:
    for line in findex:
        if u'#artiste' not in line:
            continue

        artiste=line.split(u'[:::]')[1].strip()

    cursor = sql.cursor()
    cursor.execute('SELECT COUNT(id) AS nbr FROM artistes WHERE nom=%s', (artiste,))
    if not cursor.fetchone()[0]:
        cursor = sql.cursor()
        cursor.execute('INSERT INTO artistes(nom,status,path) VALUES(%s, 99, %s)', (artiste, artiste + u'/'))
        artists_inserted += 1

You may want to brush up on Unicode and UTF-8 and encodings. I can recommend the following articles:

  • The Python Unicode HOWTO

  • Pragmatic Unicode by Ned Batchelder

  • The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!) by Joel Spolsky

2 of 2
2

Unfortunately, the string.encode() method is not always reliable. Check out this thread for more information: What is the fool proof way to convert some string (utf-8 or else) to a simple ASCII string in python

🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-unicode-string-to-a-string-in-python
Convert Unicode String to a Byte String in Python - GeeksforGeeks
July 23, 2025 - In this example, the Unicode string is transformed into a byte string using the str.encode() method with UTF-8 encoding.
🌐
Programiz
programiz.com › python-programming › methods › string › encode
Python String encode()
# unicode string string = 'pythön!' # print string print('The string is:', string) # default encoding to utf-8 · string_utf = string.encode() # print result print('The encoded version is:', string_utf) ... The string is: pythön! The encoded version (with ignore) is: b'pythn!' The encoded version (with replace) is: b'pyth?n!' Note: Try different encoding and error parameters as well. Since Python 3.0, strings are stored as Unicode, i.e.