You can use Notepad++ to evaluate a file's encoding without needing to write code. The evaluated encoding of the open file will display on the bottom bar, far right side. The encodings supported can be seen by going to Settings -> Preferences -> New Document/Default Directory and looking in the drop down.
You can use Notepad++ to evaluate a file's encoding without needing to write code. The evaluated encoding of the open file will display on the bottom bar, far right side. The encodings supported can be seen by going to Settings -> Preferences -> New Document/Default Directory and looking in the drop down.
In Linux systems, you can use file command. It will give the correct encoding
Sample:
file blah.csv
Output:
blah.csv: ISO-8859 text, with very long lines
This error because of without specifying encoding. Add this line at the beginning your python script
# -*- coding: utf-8 -*-
I was able to figure this out. It's not the most eligant solution, but it works. I made a method that finds all csv files in the current working directory if any of the filenames contain a "µ" character replace with an "_". Return a list of all csv file names. I understand that this could potentially create naming conflicts, but since I'm the end user I'll be careful.
# -*- coding: Latin-1 -*-
import os
import pandas as pd
filenames = os.listdir(path_to_dir)
filenames_fixed = []
for filename in filenames:
if filename.endswith(suffix) and 'µ' in filename:
new_filename = filename.replace('µ', '_')
os.rename(os.path.join(path_to_dir, filename),
os.path.join(path_to_dir, new_filename))
filenames_fixed.append(new_filename)
elif filename.endswith(suffix):
filenames_fixed.append(filename)
return filenames_fixed
csv_list_cwd = find_csv_filenames_remove_nonASCII(os.getcwd())
for csv_file in csv_list_cwd:
df_cwd = pd.read_csv(csv_file, encoding="Latin-1")