>>> u'aあä'.encode('ascii', 'ignore')
'a'

Decode the string you get back, using either the charset in the the appropriate meta tag in the response or in the Content-Type header, then encode.

The method encode(encoding, errors) accepts custom handlers for errors. The default values, besides ignore, are:

>>> u'aあä'.encode('ascii', 'replace')
b'a??'
>>> u'aあä'.encode('ascii', 'xmlcharrefreplace')
b'aあä'
>>> u'aあä'.encode('ascii', 'backslashreplace')
b'a\\u3042\\xe4'

See https://docs.python.org/3/library/stdtypes.html#str.encode

Answer from Ignacio Vazquez-Abrams on Stack Overflow
Top answer
1 of 12
248
>>> u'aあä'.encode('ascii', 'ignore')
'a'

Decode the string you get back, using either the charset in the the appropriate meta tag in the response or in the Content-Type header, then encode.

The method encode(encoding, errors) accepts custom handlers for errors. The default values, besides ignore, are:

>>> u'aあä'.encode('ascii', 'replace')
b'a??'
>>> u'aあä'.encode('ascii', 'xmlcharrefreplace')
b'aあä'
>>> u'aあä'.encode('ascii', 'backslashreplace')
b'a\\u3042\\xe4'

See https://docs.python.org/3/library/stdtypes.html#str.encode

2 of 12
141

As an extension to Ignacio Vazquez-Abrams' answer

>>> u'aあä'.encode('ascii', 'ignore')
'a'

It is sometimes desirable to remove accents from characters and print the base form. This can be accomplished with

>>> import unicodedata
>>> unicodedata.normalize('NFKD', u'aあä').encode('ascii', 'ignore')
'aa'

You may also want to translate other characters (such as punctuation) to their nearest equivalents, for instance the RIGHT SINGLE QUOTATION MARK unicode character does not get converted to an ascii APOSTROPHE when encoding.

>>> print u'\u2019'
’
>>> unicodedata.name(u'\u2019')
'RIGHT SINGLE QUOTATION MARK'
>>> u'\u2019'.encode('ascii', 'ignore')
''
# Note we get an empty string back
>>> u'\u2019'.replace(u'\u2019', u'\'').encode('ascii', 'ignore')
"'"

Although there are more efficient ways to accomplish this. See this question for more details Where is Python's "best ASCII for this Unicode" database?

🌐
Delft Stack
delftstack.com › "delft stack" › "howto" › "python how-to's" › "how to convert unicode characters to ascii string in python"
How to Convert Unicode Characters to ASCII String in Python | Delft Stack
February 2, 2024 - In this code, we import the unidecode function from the unidecode library. Then, we pass your Unicode string, stringVal, to the unidecode function, which will return an ASCII representation of the string.
🌐
GeeksforGeeks
geeksforgeeks.org › python › convert-unicode-to-ascii-in-python
Convert Unicode to ASCII in Python - GeeksforGeeks
July 23, 2025 - from anyascii import anyascii# working with emoji example emoji_uni = anyascii('???? ???? ????')print("The ASCII from emojis : " + str(emoji_uni))# checking for Symbols sym_uni = anyascii('➕ ☆ ℳ')print("The ASCII from Symbols : " + str(sym_uni)) ... The iconv utility is a system command-line tool that can convert text from one character encoding to another. You can use the subprocess module to call the iconv utility from Python. ... import subprocess unicode_string = "Héllo, Wörld!" process = subprocess.Popen(['iconv', '-f', 'utf-8', '-t', 'ascii//TRANSLIT'], stdin=subprocess.PIPE, stdout=subprocess.PIPE) output, error = process.communicate(input=unicode_string.encode()) ascii_string = output.decode() print(ascii_string)
🌐
Delft Stack
delftstack.com › "delft stack" › "howto" › "python how-to's" › "how to convert unicode to ascii in python"
How to Convert Unicode to ASCII in Python | Delft Stack
February 15, 2024 - If you are a beginner and do not want to go into detail, install a Python package called unidecode using the following command. It will convert Unicode to ASCII directly; it will be helpful when working with an application where you need to convert Unicode to ASCII.
🌐
Finxter
blog.finxter.com › home › learn python blog › python convert unicode to bytes, ascii, utf-8, raw string
Python Convert Unicode to Bytes, ASCII, UTF-8, Raw String - Be on the Right Side of Change
June 30, 2021 - PyPi has a unidecode module, it exports a function that takes a Unicode string and returns a string that can be encoded into ASCII bytes in Python 3.x:
🌐
GitHub
gist.github.com › BlaayLock › e496d4341a820cb28629a7d38757a3f3
Python: Convert Unicode to ASCII without errors, utf8 -> cp1251 · GitHub
Python: Convert Unicode to ASCII without errors, utf8 -> cp1251 · Raw · Convert Unicode to ASCII without errors, utf8 -> cp1251 · This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than ...
Find elsewhere
🌐
O'Reilly
oreilly.com › library › view › python-cookbook › 0596001673 › ch03s18.html
Converting Between Unicode and Plain Strings - Python Cookbook [Book]
July 19, 2002 - You need to deal with data that doesn’t fit in the ASCII character set. Unicode strings can be encoded in plain strings in a variety of ways, according to whichever encoding you choose: # Convert Unicode to plain Python string: "encode" unicodestring = u"Hello world" utf8string = unicodestring.encode("utf-8") asciistring = unicodestring.encode("ascii") isostring = unicodestring.encode("ISO-8859-1") utf16string = unicodestring.encode("utf-16") # Convert plain Python string to Unicode: "decode" plainstring1 = unicode(utf8string, "utf-8") plainstring2 = unicode(asciistring, "ascii") plainstring3 = unicode(isostring, "ISO-8859-1") plainstring4 = unicode(utf16string, "utf-16") assert plainstring1==plainstring2==plainstring3==plainstring4
Authors: Alex MartelliDavid Ascher
Published: 2002
Pages: 608
🌐
GitHub
github.com › ajanin › uni2ascii
GitHub - ajanin/uni2ascii: Convert unicode to closest ASCII equivalent. · GitHub
To convert to ASCII: > echo і lоѵе üńісοdе | uni2ascii i love unicode · The default action is to leave untouched any non-ascii that uni2ascii.py doesn't know about. This can be overridden with command line arguments. Call uni2ascii -h for help. ... It's quite easy to add new transliterations by just copying and pasting offending strings into the code.
Author: ajanin
Top answer
1 of 3
40

The Unicode characters u'\xce0' and u'\xc9' do not have any corresponding ASCII values. So, if you don't want to lose data, you have to encode that data in some way that's valid as ASCII. Options include:

>>> print s.encode('ascii', errors='backslashreplace')
ABRA\xc3O JOS\xc9
>>> print s.encode('ascii', errors='xmlcharrefreplace')
ABRAÃO JOSÉ
>>> print s.encode('unicode-escape')
ABRA\xc3O JOS\xc9
>>> print s.encode('punycode')
ABRAO JOS-jta5e

All of these are ASCII strings, and contain all of the information from your original Unicode string (so they can all be reversed without loss of data), but none of them are all that pretty for an end-user (and none of them can be reversed just by decode('ascii')).

See str.encode, Python Specific Encodings, and Unicode HOWTO for more info.


As a side note, when some people say "ASCII", they really don't mean "ASCII" but rather "any 8-bit character set that's a superset of ASCII" or "some particular 8-bit character set that I have in mind". If that's what you meant, the solution is to encode to the right 8-bit character set:

>>> s.encode('utf-8')
'ABRA\xc3\x83O JOS\xc3\x89'
>>> s.encode('cp1252')
'ABRA\xc3O JOS\xc9'
>>> s.encode('iso-8859-15')
'ABRA\xc3O JOS\xc9'

The hard part is knowing which character set you meant. If you're writing both the code that produces the 8-bit strings and the code that consumes it, and you don't know any better, you meant UTF-8. If the code that consumes the 8-bit strings is, say, the open function or a web browser that you're serving a page to or something else, things are more complicated, and there's no easy answer without a lot more information.

2 of 3
1

I found https://pypi.org/project/Unidecode/ this library very useful

>>> from unidecode import unidecode
>>> unidecode('ko\u017eu\u0161\u010dek')
'kozuscek'
>>> unidecode('30 \U0001d5c4\U0001d5c6/\U0001d5c1')
'30 km/h'
>>> unidecode('\u5317\u4EB0')
'Bei Jing '
🌐
Python
docs.python.org › 3.0 › howto › unicode.html
Unicode HOWTO — Python v3.0.1 documentation
Usually this is implemented by converting the Unicode string into some encoding that varies depending on the system. For example, Mac OS X uses UTF-8 while Windows uses a configurable encoding; on Windows, Python uses the name “mbcs” to refer to whatever the currently configured encoding is. On Unix systems, there will only be a filesystem encoding if you’ve set the LANG or LC_CTYPE environment variables; if you haven’t, the default encoding is ASCII.
🌐
Peterbe.com
peterbe.com › plog › unicode-to-ascii
Unicode strings to ASCII ...nicely - Peterbe.com
Hi, I wrote a script based on your idea. It transforms number, str and unicode to ASCII: http://www.haypocalc.com/perso/prog/python/any2ascii.py It takes care of some caracters like "ßøł" (just fill smart_unicode dictionnary ;-)).
Top answer
1 of 8
44

You can convert the file easily enough just using the unicode function, but you'll run into problems with Unicode characters without a straight ASCII equivalent.

This blog recommends the unicodedata module, which seems to take care of roughly converting characters without direct corresponding ASCII values, e.g.

>>> title = u"Klüft skräms inför på fédéral électoral große"

is typically converted to

Klft skrms infr p fdral lectoral groe

which is pretty wrong. However, using the unicodedata module, the result can be much closer to the original text:

>>> import unicodedata
>>> unicodedata.normalize('NFKD', title).encode('ascii','ignore')
'Kluft skrams infor pa federal electoral groe'
2 of 8
11

I think this is a deeper issue than you realize. Simply changing the file from Unicode into ASCII is easy, however, getting all of the Unicode characters to translate into reasonable ASCII counterparts (many letters are not available in both encodings) is another.

This Python Unicode tutorial may give you a better idea of what happens to Unicode strings that are translated to ASCII: http://www.reportlab.com/i18n/python_unicode_tutorial.html

Here's a useful quote from the site:

Python 1.6 also gets a "unicode" built-in function, to which you can specify the encoding:

> >>> unicode('hello') u'hello'
> >>> unicode('hello', 'ascii') u'hello'
> >>> unicode('hello', 'iso-8859-1') u'hello'
> >>>

All three of these return the same thing, since the characters in 'Hello' are common to all three encodings.

Now let's encode something with a European accent, which is outside of ASCII. What you see at a console may depend on your operating system locale; Windows lets me type in ISO-Latin-1.

> >>> a = unicode('André','latin-1')
> >>> a u'Andr\202'

If you can't type an acute letter e, you can enter the string 'Andr\202', which is unambiguous.

Unicode supports all the common operations such as iteration and splitting. We won't run over them here.

🌐
Python
docs.python.org › 3.3 › howto › unicode.html
Unicode HOWTO — Python 3.3.7 documentation
September 19, 2017 - Encodings don’t have to handle every possible Unicode character, and most encodings don’t. The rules for converting a Unicode string into the ASCII encoding, for example, are simple; for each code point:
🌐
GeeksforGeeks
geeksforgeeks.org › python › change-unicode-to-ascii-character-using-unihandecode
Change Unicode to ASCII Character using Unihandecode - GeeksforGeeks
July 23, 2025 - In simple language we can say that it is a transliteration to convert all character in Unicode to ASCII alphabet. ... This module does not come built-in with Python. To install this type the below command in the terminal. ... from unihandecode import Unihandecoder data1 = Unihandecoder(lang='zh') print(data1.decode("\u660e\u5929\u7684\u98ce\u5439")) ... The first line argument takes the name of the decoder you want to use. Then the decoder takes a string as argument an returns the transliterated string.
🌐
Codemia
codemia.io › home › knowledge hub › convert unicode to ascii without errors in python
Convert Unicode to ASCII without errors in Python | Codemia
September 24, 2025 - That choice matters because ASCII is tiny compared with Unicode. Characters such as é, —, 你好, and emoji do not have direct ASCII equivalents, so Python needs instructions about what to do with them. If you only care about avoiding exceptions, encode to ASCII with an error handler:
🌐
Python documentation
docs.python.org › 3 › howto › unicode.html
Unicode HOWTO — Python 3.14.8 documentation
A Unicode string is turned into a sequence of bytes that contains embedded zero bytes only where they represent the null character (U+0000). This means that UTF-8 strings can be processed by C functions such as strcpy() and sent through protocols that can’t handle zero bytes for anything other than end-of-string markers. A string of ASCII text is also valid UTF-8 text.