You can use Notepad++ to evaluate a file's encoding without needing to write code. The evaluated encoding of the open file will display on the bottom bar, far right side. The encodings supported can be seen by going to Settings -> Preferences -> New Document/Default Directory and looking in the drop down.
You can use Notepad++ to evaluate a file's encoding without needing to write code. The evaluated encoding of the open file will display on the bottom bar, far right side. The encodings supported can be seen by going to Settings -> Preferences -> New Document/Default Directory and looking in the drop down.
In Linux systems, you can use file command. It will give the correct encoding
Sample:
file blah.csv
Output:
blah.csv: ISO-8859 text, with very long lines
Why should I encode CSV files to UTF-8?
Is this CSV encoder private?
What encodings can be detected?
Hi!
I am going a bit crazy, I am sure I am getting several things fundamentally wrong, but I cannot find any information online about it.
We have an online database in which our clients can upload .csv files with lines of text to upload them in bulk, the database then processes the lines . The system only accepts .csv files in UTF 8 encoding.
Since some weeks or maybe months, the system is extremely finicky, and doesn't accept the vast majority of csv files on account that they are not "UTF 8 encoded". Please note absolutely nothing was changed in the system (I wish I had the budget to change it, but...). Sometimes manually saving the csv files in "CSV UTF 8" format solves the issue, but sometimes after doing that the same error happens or a different one in which some lines are not recognized. This has led me to suspect that the issue is not with the files, which are saved as "CSV UTF 8", but with the text itself, as the issues tend to come from staff that operates in countries with slightly different "alphabets" (cyrillic or central/eastern european latin aphabet, i.e, special letters and accents), so my suspicion is that the text is actually in the encoding of their local language, even if they are writing in English.
To validate this theory, I searched everywhere for a way to find out the individual encoding of each cell, but I cannot seem to find even a mention of this, so I guess it is not possible?
In any case importing the data into a new file selecting UTF 8 creates other issues, but that's something I can troubleshoot after validating or discarding this theory of mine.... any help would be super welcome!
TLDR: Is there a way to identify the specific encoding of the text of a worksheet cell?