I was able to read ISO-8859-1 using spark but when I store the same data to S3/hdfs back and read it, the format is converting to UTF-8.
ex: é to é
val df = spark.read.format("csv").option("delimiter", ",").option("ESCAPE quote", '"'). option("header",true).option("encoding", "ISO-8859-1").load("s3://bucket/folder")
Answer from Saida on Stack OverflowMedium
medium.com › @mike_82447 › pyspark-character-encoding-fccfad3989bd
PySpark — character encoding
October 31, 2019 - So — its obviously a text encoding\decoding thing, turns out the answer is to give spark a few clues about what it is dealing with by adding an “Encoding” option: raw_notes_df2 = spark.read.options(header=”True”).option(“encoding”, “ISO-8859–1”).csv(notes_path2) BOOM!
Author: AbsaOSS
Namastespark
namastespark.com › docs › spark › csv › handle-csv-files-of-different-encodings
Handling Encodings | Namaste Spark
May 30, 2025 - val encodeDf = spark.read .option("header", "true") .option("inferSchema", "true") .option("encoding","Windows-1252") .csv("csvFiles/empSalary.csv") encodeDf.show() encodeDf.printSchema()
Dbmstutorials
dbmstutorials.com › pyspark › spark-read-write-dataframe-options.html
PySpark: Dataframe Options
February 27, 2022 - df=spark.read.options(header=True, delimiter="|").csv("file:///path_to_file/tutorial_file_with_header.txt") ➠ lineSep example: This attribute can be used to specify single as a separator for each row while reading or writing a file using either option() or options() functions.
SQLRelease
sqlrelease.com › home › spark read file with special characters using pyspark
Spark read file with special characters using PySpark - SQLRelease
May 24, 2021 - In other words, the Spanish characters are not being replaced with the junk characters. In conclusion, we are able to read this file correctly into a Spark data frame by adding option(“encoding”, “windows-1252”) in the PySpark code.
Top answer 1 of 3
20
I was able to read ISO-8859-1 using spark but when I store the same data to S3/hdfs back and read it, the format is converting to UTF-8.
ex: é to é
val df = spark.read.format("csv").option("delimiter", ",").option("ESCAPE quote", '"'). option("header",true).option("encoding", "ISO-8859-1").load("s3://bucket/folder")
2 of 3
4
My guess is that the input file is not in UTF-8 and hence you get the incorrect characters.
My recommendation would be to write a pure Java application (with no Spark at all) and see if reading and writing gives the same results with UTF-8 encoding.
Databricks
community.databricks.com › s › question › 0D53f00001HKHWdCAP › how-to-import-data-and-apply-multiline-and-charset-utf8-at-the-same-time
How to import data and apply multiline and charset... - Databricks Community - 29116
September 25, 2021 - T_new_exp = spark.read\ .option("charset", "ISO-8859-1")\ .option("parserLib", "univocity")\ .option("multiLine", "true")\ .schema(schema)\ .csv(file)
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-load-save-functions.html
Generic Load/Save Functions - Spark 4.2.0 Documentation
The following ORC example will create bloom filter and use dictionary encoding only for favorite_color. For Parquet, there exists parquet.bloom.filter.enabled and parquet.enable.dictionary, too. To find more detailed information about the extra ORC/Parquet options, visit the official Apache ORC / Parquet websites. ... users_df = spark.read.orc("examples/src/main/resources/users.orc") (users_df.write.format("orc") .option("orc.bloom.filter.columns", "favorite_color") .option("orc.dictionary.key.threshold", "1.0") .option("orc.column.encoding.direct", "name") .save("users_with_options.orc"))
Apache Spark
spark.apache.org › docs › latest › api › java › org › apache › spark › sql › DataFrameReader.html
DataFrameReader (Spark 4.1.1 JavaDoc)
The text files must be encoded as UTF-8. If the directory structure of the text files contains partitioning information, those are ignored in the resulting Dataset. To include partitioning information as columns, use text. By default, each line in the text files is a new row in the resulting ...
The Internals of Spark SQL
jaceklaskowski.gitbooks.io › mastering-spark-sql › content › spark-sql-CSVFileFormat.html
CSVFileFormat · The Internals of Spark SQL
spark.read.format("csv").load("csv-datasets") // or the same as above using a shortcut spark.read.csv("csv-datasets") CSVFileFormat uses CSV options (that in turn are used to configure the underlying CSV parser from uniVocity-parsers project).
Apache JIRA
issues.apache.org › jira › browse › SPARK-20336
[SPARK-20336] spark.read.csv() with wholeFile=True option fails to read non ASCII unicode characters - ASF Jira
April 26, 2017 - When it is loaded without wholeFile=True option, non-ASCII characters are shown correctly although multi-line records are parsed incorrectly as follows: testdf_default = spark.read.csv("test.encoding.csv", header=True) testdf_default.show() +----+----+----+ |col1|col2|col3| +----+----+----+ | 1| a|text| | 2| b|テキスト| | 3| c| 텍스트| | 4| d|text| |テキスト|null|null| | 텍스트"|null|null| | 5| e|last| +----+----+----+ When wholeFile=True option is used, non-ASCII characters are broken as follows: testdf_wholefile = spark.read.csv("test.encoding.csv", header=True, wholeFile=True)
Stack Overflow
stackoverflow.com › questions › 76620125 › how-to-use-the-encoding-as-per-the-file-while-reading-from-spark
python - How to use the encoding as per the file, while reading from spark - Stack Overflow
import chardet.universaldetector as chardet_detector def get_dataframe(spark, path, schema, delimiter): """Return a dataframe setup to read a CSV file""" s3_file_content = spark.sparkContext.binaryFiles(path).flatMap(lambda x: x[1].splitlines()).collect() file_encoding = detect_file_encoding(s3_file_content) options = {"encoding": file_encoding} data_frame = spark.read.schema(schema).options(**options).csv(path).toDF( *[f.metadata.get("alias") for f in schema]) data_frame.show() return data_frame def detect_file_encoding(content): detector = chardet_detector.UniversalDetector() for line in content: detector.feed(line) if detector.done: break detector.close() return detector.result['encoding']
Apache JIRA
issues.apache.org › jira › browse › SPARK-13108
[SPARK-13108] Encoding not working with non-ascii compatible encodings (UTF-16/32 etc.) - ASF Jira
January 24, 2018 - According to MAPREDUCE-232#comment-13183601, it still looks fine with most encodings though but without UTF-16/32. ... I tested this in Max OS. I converted `cars_iso-8859-1.csv` into `cars_utf-16.csv` as below: iconv -f iso-8859-1 -t utf-16 < cars_iso-8859-1.csv > cars_utf-16.csv ... val cars = "cars_utf-16.csv" sqlContext.read .format("csv") .option("charset", "utf-16") .option("delimiter", 'þ') .load(cars) .show()
Stack Overflow
stackoverflow.com › questions › 58262846 › reading-csv-with-multiline-option-and-encoding-option
python - Reading CSV with multiLine option and encoding option - Stack Overflow
Hello, Unfortunately, you cannot use "multiline" and "charset" together, if you use together encoding will be set to default. For more details, refer “Spark – Known issue": mail-archives.apache.org/mod_mbox/spark-commits/201804.mbox/… 2019-10-14T06:42:28.753Z+00:00
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-csv.html
CSV Files - Spark 4.2.0 Documentation
Spark SQL provides spark.read().csv("file_name") to read a file or directory of files in CSV format into Spark DataFrame, and dataframe.write().csv("path") to write to a CSV file. Function option() can be used to customize the behavior of reading or writing, such as controlling behavior of the header, delimiter character, character set, and so on.