ANSI encoding is a slightly generic term used to refer to the standard code page on a system, usually Windows. It is more properly referred to as Windows-1252 on Western/U.S. systems. (It can represent certain other Windows code pages on other systems.) This is essentially an extension of the ASCII character set in that it includes all the ASCII characters with an additional 128 character codes. This difference is due to the fact that "ANSI" encoding is 8-bit rather than 7-bit as ASCII is (ASCII is almost always encoded nowadays as 8-bit bytes with the MSB set to 0). See the article for an explanation of why this encoding is usually referred to as ANSI.

The name "ANSI" is a misnomer, since it doesn't correspond to any actual ANSI standard, but the name has stuck. ANSI is not the same as UTF-8.

Answer from Noldorin on Stack Overflow
Top answer
1 of 11
334

ANSI encoding is a slightly generic term used to refer to the standard code page on a system, usually Windows. It is more properly referred to as Windows-1252 on Western/U.S. systems. (It can represent certain other Windows code pages on other systems.) This is essentially an extension of the ASCII character set in that it includes all the ASCII characters with an additional 128 character codes. This difference is due to the fact that "ANSI" encoding is 8-bit rather than 7-bit as ASCII is (ASCII is almost always encoded nowadays as 8-bit bytes with the MSB set to 0). See the article for an explanation of why this encoding is usually referred to as ANSI.

The name "ANSI" is a misnomer, since it doesn't correspond to any actual ANSI standard, but the name has stuck. ANSI is not the same as UTF-8.

2 of 11
70

Technically, ANSI should be the same as US-ASCII. It refers to the ANSI X3.4 standard, which is simply the ANSI organisation's ratified version of ASCII. Use of the top-bit-set characters is not defined in ASCII/ANSI as it is a 7-bit character set.

However years of misuse of the term by the DOS and subsequently Windows community has left its practical meaning as “the system codepage of whatever machine is being used”. The system codepage is also sometimes known as ‘mbcs’, since on East Asian systems that can be a multiple-byte-per-character encoding. Some code pages can even use top-bit-clear bytes as trailing bytes in a multibyte sequence, so it's not even strict compatible with plain ASCII... but even then, it's still called “ANSI”.

On US and Western European default settings, “ANSI” maps to Windows code page 1252. This is not the same as ISO-8859-1 (although it is quite similar). On other machines it could be anything else at all. This makes “ANSI” utterly useless as an external encoding identifier.

🌐
Gaijin
gaijin.at › en › infos › ascii-ansi-character-table
ASCII and ANSI Character Table
ASCII (American Standard Code for Information Interchange) is a 7-bit character set that contains characters from 0 to 127. The generic term ANSI (American National Standards Institute) is used for 8-bit character sets. These character sets contain the unchanged ASCII character set.
🌐
Wikipedia
en.wikipedia.org › wiki › ANSI_character_set
ANSI character set - Wikipedia
August 13, 2026 - Windows code pages, a collection of 8-bit character sets compatible with ASCII but incompatible with each other, especially those code pages that are partly compatible with ISO-8859, most commonly Windows Latin 1 · Windows-1252 is referred to as "ANSI" especially often
🌐
University of Edinburgh
dcs.ed.ac.uk › home › SUNWspro › 3.0 › c-compiler › transition › tguide.doc.html
Making the Transition to ANSI C: 1 - Making the Transition to ANSI C
The only requirement placed on the encoding is that no multibyte character can use a null character as part of its encoding. ANSI C specifies that program comments, string literals, character constants, and header names are all sequences of multibyte characters.
🌐
EDI Academy Blog
ediacademy.com › blog › ansi-encoding
ANSI Encoding | EDI Blog
March 24, 2024 - The first 128 characters of ANSI encoding are identical to those of ASCII, ensuring backward compatibility with older systems and software.
🌐
Vovsoft
vovsoft.com › blog › difference-between-ansi-and-utf-8
Difference between ANSI and UTF-8 - Vovsoft
September 19, 2022 - If you don't want emojis and characters from other languages ​​to be corrupted, you should use UTF-8. ANSI encoding is a generic term used to refer to the standard code page on a system.
🌐
Post.Byes
post.bytes.com › home › forum › topic
ansi c compiler character encoding - Post.Byes
August 18, 2008 - Re: ansi c compiler character encoding Many inputs and some disagreement. A simple example may be the letter Ö that in ASCII is represented by the number 153, but in ISO-8859-1 and Unicode is represented by the number 214. From what I have read out, I have to specify to customers that a specific method has an input of a city name _coded with ISO-8859-1_ in a char pointer.
Top answer
1 of 2
5

What is the difference between Windows-1252 and ANSI encoding?

See below. In practice it probably won't make much difference to your conversion.

If you keep a copy of the original file then you can always apply a different conversion if necessary.

Having said that there are ways of converting UTF-8 to ANSI.


Windows-1252

This character encoding is a superset of ISO 8859-1 in terms of printable characters, but differs from the IANA's ISO-8859-1 by using displayable characters rather than control characters in the 80 to 9F (hex) range. Notable additional characters include curly quotation marks and all the printable characters that are in ISO 8859-15. It is known to Windows by the code page number 1252, and by the IANA-approved name "windows-1252".

...

Historically, the phrase "ANSI Code Page" (ACP) is used in Windows to refer to various code pages considered as native. The intention was that most of these would be ANSI standards such as ISO-8859-1. Even though Windows-1252 was the first and by far most popular code page named so in Microsoft Windows parlance, the code page has never been an ANSI standard. Microsoft explains, "The term ANSI as used to signify Windows code pages is a historical reference, but is nowadays a misnomer that continues to persist in the Windows community.

Source Windows-1252

Note that is spite of the above statement by Microsoft they still call Windows 1252 "ANSI":

Source Code Page 1252 Windows Latin 1 (ANSI)

2 of 2
4

Are Windows-1252 and ANSI the same thing?

– Yes, for Western European languages, they are identically the same encodings. 1

For other natural languages, see my table at the end of this answer.

Unfortunately, the Wikipedia pages on the topic, this and this,
are riddled with confusing statements and unreferenced claims.

You are much better off going directly to one of the sources that Wikipedia does reference.
It was written in May 2002 and says :

The term “ANSI” as used to signify Windows code pages is a historical reference, but is nowadays a misnomer that continues to persist in the Windows community. The source of this comes from the fact that the Windows code page 1252 was originally based on an ANSI draft, which became ISO Standard 8859-1. However, in adding code points to the range reserved for control codes in the ISO standard, the Windows code page 1252 and subsequent Windows code pages originally based on the ISO 8859-x series deviated from ISO. To this day [May 2002], it is not uncommon to have the development community, both within and outside of Microsoft, confuse the 8859-1 code page with Windows 1252, as well as see “ANSI” or “A” used to signify Windows code page support.

References

  • Charts showing the Windows-1252 character set | Section 4
  • Post containing a table of all ten Windows code pages
  • Windows-1252 | Wikipedia
  • ISO/IEC 8859-1 | Wikipedia
  • Unicode and Windows XP | Cathy Wissink, Microsoft

1 For charts displaying the Windows-1252 character set, see this post, Section 4.

Find elsewhere
🌐
W3Schools
w3schools.com › charsets › ref_html_ansi.asp
HTML Windows-1252 - ANSI Reference
The intention was that these character sets would be an ANSI standard like ISO-8859-1.
🌐
Alan Wood
alanwood.net › demos › ansi.html
ANSI character set and equivalent Unicode and HTML characters
ANSI characters 32 to 127 correspond to those in the 7-bit ASCII character set, which forms the Basic Latin Unicode character range. Characters 160–255 correspond to those in the Latin-1 Supplement Unicode character range. Positions 128–159 in Latin-1 Supplement are reserved for controls, but most of them are used for printable characters in ANSI; the Unicode equivalents are noted in the table below.
🌐
Wikipedia
en.wikipedia.org › wiki › Windows-1252
Windows-1252 - Wikipedia
3 weeks ago - Windows-1252 or CP-1252 (Windows code page 1252) is a legacy single-byte character encoding that is used by default (as the "ANSI code page") in Microsoft Windows throughout the Americas, Western Europe, Oceania, and much of Africa.
🌐
TutorialsPoint
tutorialspoint.com › difference-between-ansi-and-utf-8
Difference between ANSI and Unicode
May 15, 2023 - In conclusion, ANSI is a set of character encodings with restricted language coverage that is primarily used in older systems, whereas Unicode is a comprehensive character encoding standard that supports all languages and symbols, making it ...
🌐
URL Decode
urldecoder.org › dec › ansi
URL Decoding of "ansi" - Online
Although it is known as URL encoding it is, in fact, used more generally within the main Uniform Resource Identifier (URI) set, which includes both Uniform Resource Locator (URL) and Uniform Resource Name (URN). As such it is also used in the preparation of data of the "application/x-www-form-urlencoded" media type, as is often employed in the submission of HTML form data in HTTP requests. Character set: In case of textual data, the encoding scheme does not contain the character set, so you have to specify which character set was used during the encoding process.
🌐
MedCalc
medcalc.org › home › manual › appendices › miscellaneous tables
ANSI character set | MedCalc software
1 month ago - Note that the term ANSI actually refers to Windows code page 1252.
🌐
Medium
medium.com › @jimmy760205 › ascii-ansi-and-unicode-41e0241b75d4
ASCII, ANSI, Unicode, and UTF-8 | Medium
March 28, 2020 - After converted Unicode to ANSI, 人 is stored as 0xA4, 0x48 in std::string, which is encoded by Big-5.
🌐
STT Media
sttmedia.com › unicode-ansi
ASCII and ANSI
Even with these extensions, the first 128 characters are the common ASCII characters, while the other 128 characters are characters, which are required for the appropriate language or corresponding character set. Although the ANSI encoding requires only one byte per character and thus is the most effective encoding, there are disadvantages, because this efficiency results from the impossibility to store different character systems or other special characters in one file.
🌐
Community
community.safe.com › home › forums › fme form › transformers › character encoding - ansi = iso-8859-1?
Character encoding - ANSI = iso-8859-1? - FME Community
May 27, 2020 - So it seems that ANSI means "system encoding". You would set it to the system encoding you want (like cp-1252) or leave it unset if you wanted the same encoding as the source.
🌐
Blue Prism Community
community.blueprism.com › t5 › Product-Forum › Saving-csv-file-with-ANSI-encoding › td-p › 102039
Saving csv file with ANSI encoding - SS&C Blue Prism Community
October 19, 2022 - We tried writing a code where we are able to define encoding type but it doesn't take "ANSI" to be a valid encoding type.
🌐
EDI Academy Blog
ediacademy.com › blog › ansi-format
ANSI Format (8-bit Encoding) | EDI Blog
September 15, 2021 - So, the ASCII standard set has 128 characters. Also it includes letters as well as symbols (exclamation marks, commas etc.). Sometimes text documents may contain non-ASCII symbols. In order to display a text document properly, the ANSI encodes an extended set of symbols.