In Python3 all strings are unicode, so the problem you're having is likely due to your locale settings not being correct. The Python3 interpreter looks to use the locale environment variables and if it cannot find them it emulates basic ASCII

From locale.py:

except ImportError:

    # Locale emulation

    CHAR_MAX = 127
    LC_ALL = 6
    LC_COLLATE = 3
    LC_CTYPE = 0
    LC_MESSAGES = 5
    LC_MONETARY = 4
    LC_NUMERIC = 1
    LC_TIME = 2
    Error = ValueError

Double check the locale on your shell from which you are executing. Here are a few work arounds you can try to see if they get you working before you go through the task of getting your env setup correctly.

1) Validate UTF-8 locale or language files are installed (see link above)

2) Try adding this to the top of your script

#!/usr/bin/env LC_ALL=en_US.UTF-8 /usr/local/bin/python3
print('カタカナ')

or

#!/usr/bin/env LANG=en_US.UTF-8 /usr/local/bin/python3
print('カタカナ')

Or export shell variables before executing the Python interpreter

export LANG=en_US.UTF-8
export LC_ALL=en_US.UTF-8
python3
>>> print('カタカナ')

Sorry I cannot be more specific, as these settings are platform and OS specific. You can forcefully attempt to set the locale in Python directly using the locale module, but I don't recommend that, and it won't help if they are not installed.

Hope that helps.

Answer from sehafoc on Stack Overflow
Top answer
1 of 2
5

In Python3 all strings are unicode, so the problem you're having is likely due to your locale settings not being correct. The Python3 interpreter looks to use the locale environment variables and if it cannot find them it emulates basic ASCII

From locale.py:

except ImportError:

    # Locale emulation

    CHAR_MAX = 127
    LC_ALL = 6
    LC_COLLATE = 3
    LC_CTYPE = 0
    LC_MESSAGES = 5
    LC_MONETARY = 4
    LC_NUMERIC = 1
    LC_TIME = 2
    Error = ValueError

Double check the locale on your shell from which you are executing. Here are a few work arounds you can try to see if they get you working before you go through the task of getting your env setup correctly.

1) Validate UTF-8 locale or language files are installed (see link above)

2) Try adding this to the top of your script

#!/usr/bin/env LC_ALL=en_US.UTF-8 /usr/local/bin/python3
print('カタカナ')

or

#!/usr/bin/env LANG=en_US.UTF-8 /usr/local/bin/python3
print('カタカナ')

Or export shell variables before executing the Python interpreter

export LANG=en_US.UTF-8
export LC_ALL=en_US.UTF-8
python3
>>> print('カタカナ')

Sorry I cannot be more specific, as these settings are platform and OS specific. You can forcefully attempt to set the locale in Python directly using the locale module, but I don't recommend that, and it won't help if they are not installed.

Hope that helps.

2 of 2
0

What's new in Python 3.0 says:

All text is Unicode; however encoded Unicode is represented as binary data

If you want to try outputting utf-8, here's an example:

b'\x41'.decode("utf-8", "strict")

If you'd like to use unicode in a string literal, use the unicode escape and its coded representation. For your example:

print("\u24B6")
🌐
Reddit
reddit.com › r/python › debate: enforcing python source encoding as utf-8
r/Python on Reddit: Debate: Enforcing Python source encoding as UTF-8
July 31, 2014 -

I currently put this at the top of all my .py files:

# -*- coding: utf-8 -*-

I've been taught this for years as best practice. To me the idea of enforcing UTF-8 by default makes sense, especially with my tests containing a lot of Unicode characters. It allows me to write Unicode literals in my code directly.

However, I recently was told that forcing the source encoding to UTF-8 can be bad for cross-platform compatibility, since Windows doesn't default to UTF-8.

Both approaches seem to have strong arguments. In more detail, what are the benefits of enforcing/not enforcing a source encoding? What are the problems?

What do you consider the pros and cons of either approach?

Discussions

python - Force UTF-8 while writing to file - Stack Overflow
How do I enforce UTF-8 encoding when writing a string to a file in Python? I need it in a larger toolchain, but cannot get it running reliably. Following different, failing, approaches from Stack More on stackoverflow.com
🌐 stackoverflow.com
Working with UTF-8 encoding in Python source - Stack Overflow
You need not use unicode(), simply write string in UTF-8 encoding. 2011-06-09T08:03:22.567Z+00:00 ... In Python versions older than 3, you also need to prefix unicode string literals with "u": some_string = u'idzie wąż wąską dróżką'. 2011-06-09T08:06:28.08Z+00:00 More on stackoverflow.com
🌐 stackoverflow.com
unicode - python encoding utf-8 - Stack Overflow
But I dont understand why, but I got problem with encoding. My MySQL database is in utf8, or seems to be SQL query SHOW variables LIKE 'char%' returns me only utf8 or binary. ... Copy#!/usr/bin/python # -*- coding: utf-8 -*- def saveIndex(index,date): import MySQLdb as mdb import codecs sql ... More on stackoverflow.com
🌐 stackoverflow.com
How to convert a string to utf-8 in Python - Stack Overflow
I have a browser which sends utf-8 characters to my Python server, but when I retrieve it from the query string, the encoding that Python returns is ASCII. How can I convert the plain string to utf... More on stackoverflow.com
🌐 stackoverflow.com
🌐
Python
peps.python.org › pep-0686
PEP 686 – Make UTF-8 mode default | peps.python.org
March 18, 2022 - Inconsistent default encoding causes many bugs. Python will enable UTF-8 mode by default from Python 3.15.
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.7 documentation
Ideally, you’d want to be able to write literals in your language’s natural encoding. You could then edit Python source code with your favorite editor which would display the accented characters naturally, and have the right characters used at runtime. Python supports writing source code in UTF-8 by default, but you can use almost any encoding if you declare the encoding being used.
🌐
GitHub
gist.github.com › embayer › 2774afb51b188dc53ed8
set python default encoding to utf-8 · GitHub
cd ~/.virtualenvs/myvirtualenv/lib/python2.x/site-packages echo 'import sys;sys.setdefaultencoding("utf-8")' > sitecustomize.py · to automatically create a sitecustomize.py every time you create a virtualenv, edit your · #!/bin/bash # This hook is run after a new virtualenv is activated. PY_VERSION=`ls $VIRTUAL_ENV/lib/` echo 'import sys;sys.setdefaultencoding("utf-8")' > $VIRTUAL_ENV/lib/$PY_VERSION/site-packages/sitecustomize.py ... Even after setting all these encoding formats in the file.
🌐
DEV Community
dev.to › methane › python-use-utf-8-mode-on-windows-212i
Python: Use the UTF-8 mode on Windows! - DEV Community
January 10, 2020 - Summary: Set the PYTHONUTF8=1 environment variable. On macOS and Linux, UTF-8 is the standard encoding already. But Windows still uses legacy encoding (e.g.
Find elsewhere
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-a-string-to-utf-8-in-python
Convert a String to Utf-8 in Python - GeeksforGeeks
July 23, 2025 - UTF-8 String (Using str.encode method): b'Hello, World!' Converting a string to UTF-8 in Python is a simple task with multiple methods at your disposal. Whether you choose the encode method, the bytes constructor, or the str.encode method, the key is to specify the UTF-8 encoding.
Top answer
1 of 2
65

You don't need to encode data that is already encoded. When you try to do that, Python will first try to decode it to unicode before it can encode it back to UTF-8. That is what is failing here:

>>> data = u'\u00c3'            # Unicode data
>>> data = data.encode('utf8')  # encoded to UTF-8
>>> data
'\xc3\x83'
>>> data.encode('utf8')         # Try to *re*-encode it
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0: ordinal not in range(128)

Just write your data directly to the file, there is no need to encode already-encoded data.

If you instead build up unicode values instead, you would indeed have to encode those to be writable to a file. You'd want to use codecs.open() instead, which returns a file object that will encode unicode values to UTF-8 for you.

You also really don't want to write out the UTF-8 BOM, unless you have to support Microsoft tools that cannot read UTF-8 otherwise (such as MS Notepad).

For your MySQL insert problem, you need to do two things:

  • Add charset='utf8' to your MySQLdb.connect() call.

  • Use unicode objects, not str objects when querying or inserting, but use sql parameters so the MySQL connector can do the right thing for you:

    artiste = artiste.decode('utf8')  # it is already UTF8, decode to unicode
    
    c.execute('SELECT COUNT(id) AS nbr FROM artistes WHERE nom=%s', (artiste,))
    
    # ...
    
    c.execute('INSERT INTO artistes(nom,status,path) VALUES(%s, 99, %s)', (artiste, artiste + u'/'))
    

It may actually work better if you used codecs.open() to decode the contents automatically instead:

import codecs

sql = mdb.connect('localhost','admin','ugo&(-@F','music_vibration', charset='utf8')

with codecs.open('config/index/'+index, 'r', 'utf8') as findex:
    for line in findex:
        if u'#artiste' not in line:
            continue

        artiste=line.split(u'[:::]')[1].strip()

    cursor = sql.cursor()
    cursor.execute('SELECT COUNT(id) AS nbr FROM artistes WHERE nom=%s', (artiste,))
    if not cursor.fetchone()[0]:
        cursor = sql.cursor()
        cursor.execute('INSERT INTO artistes(nom,status,path) VALUES(%s, 99, %s)', (artiste, artiste + u'/'))
        artists_inserted += 1

You may want to brush up on Unicode and UTF-8 and encodings. I can recommend the following articles:

  • The Python Unicode HOWTO

  • Pragmatic Unicode by Ned Batchelder

  • The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!) by Joel Spolsky

2 of 2
2

Unfortunately, the string.encode() method is not always reliable. Check out this thread for more information: What is the fool proof way to convert some string (utf-8 or else) to a simple ASCII string in python

🌐
Medium
medium.com › @nawazmohtashim › method-to-encode-a-string-to-utf-8-in-python-b287027b7be9
Method to Encode a String to UTF-8 in Python | by Mohd Mohtashim Nawaz | Medium
February 2, 2024 - However, when it comes to storing or transmitting these strings, they need to be encoded into a specific format, such as UTF-8. Python provides a built-in method called encode() that allows you to encode a string into a specified encoding format, ...
🌐
Python Morsels
pythonmorsels.com › unicode-character-encodings-in-python
Unicode character encodings - Python Morsels
May 2, 2022 - Like the string encode method, ... utf-8 by default: ... But if you have bytes that represent data in a different character encoding, you'll need to specify that character encoding instead: >>> data = b"H\x00e\x00l\x00l\x00o\x00 \x00t\x00h\x00e\x00r\x00e\x00!\x00 \x00('" >>> data.decode("utf-16le") 'Hello there! ✨' When you open a file in Python, whether for ...
🌐
Vstinner
vstinner.github.io › python37-new-utf8-mode.html
Python 3.7 UTF-8 Mode — Victor Stinner blog 3
March 27, 2018 - The locale encoding remains the best default filesystem encoding for Python. I would say that the locale encoding is the least bad filesystem encoding. This article tells the story of my PEP 540: Add a new UTF-8 Mode which adds an opt-in option to "use UTF-8" everywhere".
🌐
Honeybadger
honeybadger.io › blog › python-character-encoding
Python developer's guide to character encoding - Honeybadger Developer Blog
March 6, 2023 - Python uses UTF-8 by default, which means it does not need to be specified in every Python file. To encode a string into bytes, add the encode method, which will return the binary representation of the string.
🌐
Python
peps.python.org › pep-0528
PEP 528 – Change Windows console encoding to UTF-8 | peps.python.org
Since the readline interface is required to return an 8-bit encoded string with no embedded nulls, the _PyOS_WindowsConsoleReadline function transcodes from utf-16-le as read from the operating system into utf-8.
🌐
MojoAuth
mojoauth.com › character-encoding-decoding › utf-8-encoding--python
UTF-8 Encoding : Python | Encoding Solutions Across Programming Languages
When encoded, the UTF-8 representation is a bytes sequence that can be stored or transmitted efficiently. Decoding is the reverse process of encoding, and it is equally easy in Python. The decode() method converts a bytes object back into a string.
🌐
Evanjones
evanjones.ca › python-utf8.html
How to Use UTF-8 with Python (evanjones.ca)
You can do this in one of two ways. First, you can place a UTF-8 byte-order marker at the beginning of your file, if your editor supports it. Secondly, you can place the following special comment in the first or second lines of your script: ... Any ASCII-compatible encoding is permitted.