As mentioned in the comments, your question isn't very specific, so I'll try to give you some hints about character encodings, see if you can apply those to your specific case!

Unicode and Encoding

Here's a small primer about encoding. Basically, there are two ways to represent text in Python:

  • unicode. You can consider that unicode is the ultimate encoding, you should strive to use it everywhere. In Python 2.x source files, unicode strings look like u'some unicode'.
  • str. This is encoded text - to be able to read it, you need to know the encoding (or guess it). In Python 2.x, those strings look like 'some str'.

This changed in Python 3 (unicode is now str and str is now bytes).

How does that play out?

Usually, it's pretty straightforward to ensure that you code uses unicode for its execution, and uses str for I/O:

  • Everything you receive is encoded, so you do input_string.decode('encoding') to convert it to unicode.
  • Everything you need to output is unicode but needs to be encoded, so you do output_string.encode('encoding').

The most common encodings are cp-1252 on Windows (on US or EU systems), and utf-8 on Linux.

Applying this to your case

I DO have to write äöü in a path, or it will not work

Windows natively uses unicode for file paths and names, so you should actually always use unicode for those.

It DOES have to be an ANSI-"encoded" file, or it will not work

When you write to the file, be sure to always run your output through output.encode('cp1252') (or whatever encoding ANSI would be on your system).

Things like line.write(str.decode('utf-8')) break the funktion of the file

By now you probably realized that:

  • If str as indeed an str instance, Python will try to convert it to unicode using the utf-8 encoding, but then try to encode it again (likely in ascii) to write it to the file
  • If str is actually an unicode instance, Python will first encode it (likely in ascii, and that will probably crash) to then be able to decode it.

Bottom line is, you need to know if str is unicode, you should encode it. If it's already encoded, don't touch it (or decode it then encode it if the encoding is not the one you want!).

A magical comment at the beginning of the script like # -- coding: iso-8859-1 -- does nothing here (though it is helpful when it comes to the mentioned Metadata and allowed characters in it...)

Not a surprise, this only tells Python what encoding should be used to read your source file so that non-ascii characters are properly recognized.

Oh, and i'm using Python 2.7.3. Third-Party modules dependencies, you know...

Python 3 probably is a big update in terms of unicode and encoding, but that doesn't mean Python 2.x can't make it work!

Will that solve your issue?

You can't be sure, it's possible that the problem lies in the player you're using, not in your code.

Once you output it, you should make sure that your script's output is readable using reference tools (such as Windows Explorer). If it is, but the player still can't open it, you should consider updating to a newer version.

Answer from Thomas Orozco on Stack Overflow
Top answer
1 of 3
30

As mentioned in the comments, your question isn't very specific, so I'll try to give you some hints about character encodings, see if you can apply those to your specific case!

Unicode and Encoding

Here's a small primer about encoding. Basically, there are two ways to represent text in Python:

  • unicode. You can consider that unicode is the ultimate encoding, you should strive to use it everywhere. In Python 2.x source files, unicode strings look like u'some unicode'.
  • str. This is encoded text - to be able to read it, you need to know the encoding (or guess it). In Python 2.x, those strings look like 'some str'.

This changed in Python 3 (unicode is now str and str is now bytes).

How does that play out?

Usually, it's pretty straightforward to ensure that you code uses unicode for its execution, and uses str for I/O:

  • Everything you receive is encoded, so you do input_string.decode('encoding') to convert it to unicode.
  • Everything you need to output is unicode but needs to be encoded, so you do output_string.encode('encoding').

The most common encodings are cp-1252 on Windows (on US or EU systems), and utf-8 on Linux.

Applying this to your case

I DO have to write äöü in a path, or it will not work

Windows natively uses unicode for file paths and names, so you should actually always use unicode for those.

It DOES have to be an ANSI-"encoded" file, or it will not work

When you write to the file, be sure to always run your output through output.encode('cp1252') (or whatever encoding ANSI would be on your system).

Things like line.write(str.decode('utf-8')) break the funktion of the file

By now you probably realized that:

  • If str as indeed an str instance, Python will try to convert it to unicode using the utf-8 encoding, but then try to encode it again (likely in ascii) to write it to the file
  • If str is actually an unicode instance, Python will first encode it (likely in ascii, and that will probably crash) to then be able to decode it.

Bottom line is, you need to know if str is unicode, you should encode it. If it's already encoded, don't touch it (or decode it then encode it if the encoding is not the one you want!).

A magical comment at the beginning of the script like # -- coding: iso-8859-1 -- does nothing here (though it is helpful when it comes to the mentioned Metadata and allowed characters in it...)

Not a surprise, this only tells Python what encoding should be used to read your source file so that non-ascii characters are properly recognized.

Oh, and i'm using Python 2.7.3. Third-Party modules dependencies, you know...

Python 3 probably is a big update in terms of unicode and encoding, but that doesn't mean Python 2.x can't make it work!

Will that solve your issue?

You can't be sure, it's possible that the problem lies in the player you're using, not in your code.

Once you output it, you should make sure that your script's output is readable using reference tools (such as Windows Explorer). If it is, but the player still can't open it, you should consider updating to a newer version.

2 of 3
6

On Windows there is special encoding available called mbcs, it converts between current default ANSI codepage and UNICODE. For example on a Spanish Language PC:

u'ñ'.encode('mbcs') -> '\xf1'
'\xf1'.decode('mbcs') -> u'ñ'

On Windows ANSI means current default multi-byte code page. For western European languages Windows ISO-8859-1, for eastern European languages windows ISO-8859-2) encoded byte string and other encodings for other languages as appropriate.

More info available at:

https://docs.python.org/2.4/lib/standard-encodings.html

See also:

https://docs.python.org/2/library/sys.html#sys.getfilesystemencoding

🌐
Python
docs.python.org › 3 › library › codecs.html
codecs — Codec registry and base classes
This module implements the ANSI codepage (CP_ACP). Availability: Windows. Changed in version 3.2: Before 3.2, the errors argument was ignored; 'replace' was always used to encode, and 'ignore' to decode.
Discussions

encode - ANSI encoding in python - Stack Overflow
On Windows i can do like this: >>> "\x98".encode("ANSI") b'\x98' On Linux it throws an error: >>> "\x98".encode("ANSI") Traceback (most re... More on stackoverflow.com
🌐 stackoverflow.com
April 1, 2021
UTF-8 and ANSI encoding issue
Hi all, i have a code that writes to a file with utf-8 encoding. But when i try to open the same file i created in the same script i get an error message saying it can’t read it because the file is in ANSI. Here is the code of creating the file: with open(new_file, 'w', encoding='utf-8') ... More on discuss.python.org
🌐 discuss.python.org
10
0
November 23, 2023
Python: Multiple files ANSI to utf-8 converter | Notepad++ Community
hello, I want to use “Python Script” Plugin as to convert multiple files to UTF-8 (not UTF-8-BOM), on a particular folder. Can be this done? More on community.notepad-plus-plus.org
🌐 community.notepad-plus-plus.org
8
0
March 7, 2023
How to set encoding as 'ANSI' using Python? - Stack Overflow
I am using Python 3.7.4 version. I want to set the encoding as 'ANSI' at the time of reading a text file and also writing a text file. I another case I read a file by providing 'utf-8' ( please f... More on stackoverflow.com
🌐 stackoverflow.com
🌐
PyPI
pypi.org › project › ansipants
ansipants · PyPI
A Python module and command-line utility for converting .ANS format ANSI art to HTML.
      » pip install ansipants
    
Published: Dec 25, 2021
Version: 0.2
🌐
Example Code
example-code.com › python › charset_convert_file_from_utf8_to_ansi.asp
CkPython Convert a File from utf-8 to ANSI (such as Windows-1252)
Chilkat • HOME • Android™ • AutoIt • C • C# • C++ • Chilkat2-Python • CkPython • Classic ASP • DataFlex • Delphi DLL • Go • Java • Node.js • Objective-C • PHP Extension • Perl • PowerBuilder • PowerShell • PureBasic • Ruby • SQL Server • Swift • Tcl • Unicode C • Unicode C++ • VB.NET • VBScript • Visual Basic 6.0 • Visual FoxPro • Xojo Plugin
🌐
Python.org
discuss.python.org › python help
UTF-8 and ANSI encoding issue - Python Help - Discussions on Python.org
November 23, 2023 - Hi all, i have a code that writes to a file with utf-8 encoding. But when i try to open the same file i created in the same script i get an error message saying it can’t read it because the file is in ANSI. Here is the code of creating the file: with open(new_file, 'w', encoding='utf-8') as f: for item in items: f.write('%s\n' % item) with open(new_file, 'a', encoding='utf-8') as f: for line in lines: f.write('%s\n' % line) This is the code which supposed to open the ...
🌐
GitHub
gist.github.com › 3e7e43ab85717e81925656f70f5bae8d
A guide to character encoding aware development · GitHub
December 15, 2021 - In general the name is just cp followed by the code page number (e.g. "cp1251"). The codecs module documentation has a complete list14. If running on a Windows system, Python aliases "mbcs" to the system ANSI code page for convenience.
🌐
Notepad++ Community
community.notepad-plus-plus.org › topic › 24214 › python-multiple-files-ansi-to-utf-8-converter
Python: Multiple files ANSI to utf-8 converter | Notepad++ Community
March 7, 2023 - You could probably modify this script to try to guess the encoding of files, but I’ve tried using automatic encoding detection in Python and it’s pretty hit-or-miss. If you’re really determined to try guessing encoding, try looking at codecs. ... Ok, I change the lines. If I understand well enough: help='d:\\2022_12_02\\word 2\\1') # name of directory in which you want to change file encodings · help='ANSI') # the previous encoding of files found
Find elsewhere
🌐
Esri Community
community.esri.com › t5 › python-questions › python-script-file-not-encoded-correctly-with-ansi › td-p › 705853
Solved: Python script file not encoded correctly with ANSI... - Esri Community
December 12, 2021 - File "<ipython-input-1-d259788efdf0>", line 1 "C:\Users\Is\For\Losers" ^ SyntaxError: (unicode error) 'unicodeescape' codec can't decode bytes in position 2-3: truncated \UXXXXXXXX escape
🌐
YouTube
youtube.com › watch
python encode ansi
Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.
🌐
Medium
medium.com › towardsdev › mastering-ansi-escape-codes-in-python-parsing-and-processing-text-4d7fc9645bf5
Mastering ANSI Escape Codes in Python: Parsing and Processing Text | by Py-Core Python Programming | Towards Dev
January 15, 2025 - Parsing and processing text files containing these ANSI escape characters using Python can help automate tasks, analyze terminal outputs, and capture data.
🌐
PyPI
pypi.org › project › ansiparser
ansiparser · PyPI
Parse ANSI escape sequences into screen outputs.
      » pip install ansiparser
    
Published: Nov 05, 2024
Version: 1.3.2
Top answer
1 of 2
16

MS Notepad gives the user a choice of 4 encodings, expressed in clumsy confusing terminology:

"Unicode" is UTF-16, written little-endian. "Unicode big endian" is UTF-16, written big-endian. In both UTF-16 cases, this means that the appropriate BOM will be written. Use utf-16 to decode such a file.

"UTF-8" is UTF-8; Notepad explicitly writes a "UTF-8 BOM". Use utf-8-sig to decode such a file.

"ANSI" is a shocker. This is MS terminology for "whatever the default legacy encoding is on this computer".

Here is a list of Windows encodings that I know of and the languages/scripts that they are used for:

cp874  Thai
cp932  Japanese 
cp936  Unified Chinese (P.R. China, Singapore)
cp949  Korean 
cp950  Traditional Chinese (Taiwan, Hong Kong, Macao(?))
cp1250 Central and Eastern Europe 
cp1251 Cyrillic ( Belarusian, Bulgarian, Macedonian, Russian, Serbian, Ukrainian)
cp1252 Western European languages
cp1253 Greek 
cp1254 Turkish 
cp1255 Hebrew 
cp1256 Arabic script
cp1257 Baltic languages 
cp1258 Vietnamese
cp???? languages/scripts of India  

If the file has been created on the computer where it is being read, then you can obtain the "ANSI" encoding by locale.getpreferredencoding(). Otherwise if you know where it came from, you can specify what encoding to use if it's not UTF-16. Failing that, guess.

Be careful using codecs.open() to read files on Windows. The docs say: """Note Files are always opened in binary mode, even if no binary mode was specified. This is done to avoid data loss due to encodings using 8-bit values. This means that no automatic conversion of '\n' is done on reading and writing.""" This means that your lines will end in \r\n and you will need/want to strip those off.

Putting it all together:

Sample text file, saved with all 4 encoding choices, looks like this in Notepad:

The quick brown fox jumped over the lazy dogs.
àáâãäå

Here is some demo code:

import locale

def guess_notepad_encoding(filepath, default_ansi_encoding=None):
    with open(filepath, 'rb') as f:
        data = f.read(3)
    if data[:2] in ('\xff\xfe', '\xfe\xff'):
        return 'utf-16'
    if data == u''.encode('utf-8-sig'):
        return 'utf-8-sig'
    # presumably "ANSI"
    return default_ansi_encoding or locale.getpreferredencoding()

if __name__ == "__main__":
    import sys, glob, codecs
    defenc = sys.argv[1]
    for fpath in glob.glob(sys.argv[2]):
        print
        print (fpath, defenc)
        with open(fpath, 'rb') as f:
            print "raw:", repr(f.read())
        enc = guess_notepad_encoding(fpath, defenc)
        print "guessed encoding:", enc
        with codecs.open(fpath, 'r', enc) as f:
            for lino, line in enumerate(f, 1):
                print lino, repr(line)
                print lino, repr(line.rstrip('\r\n'))

and here is the output when run in a Windows "Command Prompt" window using the command \python27\python read_notepad.py "" t1-*.txt

('t1-ansi.txt', '')
raw: 'The quick brown fox jumped over the lazy dogs.\r\n\xe0\xe1\xe2\xe3\xe4\xe5
\r\n'
guessed encoding: cp1252
1 u'The quick brown fox jumped over the lazy dogs.\r\n'
1 u'The quick brown fox jumped over the lazy dogs.'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5\r\n'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5'

('t1-u8.txt', '')
raw: '\xef\xbb\xbfThe quick brown fox jumped over the lazy dogs.\r\n\xc3\xa0\xc3
\xa1\xc3\xa2\xc3\xa3\xc3\xa4\xc3\xa5\r\n'
guessed encoding: utf-8-sig
1 u'The quick brown fox jumped over the lazy dogs.\r\n'
1 u'The quick brown fox jumped over the lazy dogs.'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5\r\n'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5'

('t1-uc.txt', '')
raw: '\xff\xfeT\x00h\x00e\x00 \x00q\x00u\x00i\x00c\x00k\x00 \x00b\x00r\x00o\x00w
\x00n\x00 \x00f\x00o\x00x\x00 \x00j\x00u\x00m\x00p\x00e\x00d\x00 \x00o\x00v\x00e
\x00r\x00 \x00t\x00h\x00e\x00 \x00l\x00a\x00z\x00y\x00 \x00d\x00o\x00g\x00s\x00.
\x00\r\x00\n\x00\xe0\x00\xe1\x00\xe2\x00\xe3\x00\xe4\x00\xe5\x00\r\x00\n\x00'
guessed encoding: utf-16
1 u'The quick brown fox jumped over the lazy dogs.\r\n'
1 u'The quick brown fox jumped over the lazy dogs.'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5\r\n'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5'

('t1-ucb.txt', '')
raw: '\xfe\xff\x00T\x00h\x00e\x00 \x00q\x00u\x00i\x00c\x00k\x00 \x00b\x00r\x00o\
x00w\x00n\x00 \x00f\x00o\x00x\x00 \x00j\x00u\x00m\x00p\x00e\x00d\x00 \x00o\x00v\
x00e\x00r\x00 \x00t\x00h\x00e\x00 \x00l\x00a\x00z\x00y\x00 \x00d\x00o\x00g\x00s\
x00.\x00\r\x00\n\x00\xe0\x00\xe1\x00\xe2\x00\xe3\x00\xe4\x00\xe5\x00\r\x00\n'
guessed encoding: utf-16
1 u'The quick brown fox jumped over the lazy dogs.\r\n'
1 u'The quick brown fox jumped over the lazy dogs.'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5\r\n'
2 u'\xe0\xe1\xe2\xe3\xe4\xe5'

Things to be aware of:

(1) "mbcs" is a file-system pseudo-encoding which has no relevance at all to decoding the contents of files. On a system where the default encoding is cp1252, it makes like latin1 (aarrgghh!!); see below

>>> all_bytes = "".join(map(chr, range(256)))
>>> u1 = all_bytes.decode('cp1252', 'replace')
>>> u2 = all_bytes.decode('mbcs', 'replace')
>>> u1 == u2
False
>>> [(i, u1[i], u2[i]) for i in xrange(256) if u1[i] != u2[i]]
[(129, u'\ufffd', u'\x81'), (141, u'\ufffd', u'\x8d'), (143, u'\ufffd', u'\x8f')
, (144, u'\ufffd', u'\x90'), (157, u'\ufffd', u'\x9d')]
>>>

(2) chardet is very good at detecting encodings based on non-Latin scripts (Chinese/Japanese/Korean, Cyrillic, Hebrew, Greek) but not much good at Latin-based encodings (Western/Central/Eastern Europe, Turkish, Vietnamese) and doesn't grok Arabic at all.

2 of 2
3

Notepad saves Unicode files with a byte order mark. This means that the first bytes of the file will be:

  • EF BB BF -- UTF-8
  • FF FE -- "Unicode" (actually UTF-16 little-endian, looks like)
  • FE FF -- "Unicode big-endian" (looks like UTF-16 big-endian)

Other text editors may or may not have the same behavior, but if you know for sure Notepad is being used, this will give you a decent heuristic for auto-selecting the encoding. All these sequences are valid in the ANSI encoding as well, however, so it is possible for this heuristic to make mistakes. It is not possible to guarantee that the correct encoding is used.

🌐
Programmersought
programmersought.com › article › 11051642134
UTF8 to ANSI encoding using Python - Programmer Sought
Windows uses Notepad to save files with an ANSI encoding, and Python with mbcs encoding (Windows only) for ANSI: with open('hello.txt', 'w') as f: F.write(u'hello'.encode('mbcs')) Execute the above code to create an ANSI encoded file. ANSI == Windows native encoding In Simplified Chinese Windows: ansi == gbk : >>> u'hello'.encode('mbcs') '\xc4\xe3\xba\xc3' >>> u'Hello'.encode('mbcs').decode('gbk') u'\u4f60\u597d'
🌐
Originlab
my.originlab.com › forum › topic.asp
The Origin Forum - ANSI encoding
Over 1 Million registered users across corporations, universities and government research labs worldwide, rely on Origin to import, graph, explore, analyze and interpret their data. With a point-and-click interface and tools for batch operations, Origin helps them optimize their daily workflow.