I was able to read ISO-8859-1 using spark but when I store the same data to S3/hdfs back and read it, the format is converting to UTF-8.

ex: é to é

val df = spark.read.format("csv").option("delimiter", ",").option("ESCAPE quote", '"'). option("header",true).option("encoding", "ISO-8859-1").load("s3://bucket/folder")
Answer from Saida on Stack Overflow
🌐
Medium
medium.com › @mike_82447 › pyspark-character-encoding-fccfad3989bd
PySpark — character encoding
October 31, 2019 - So — its obviously a text encoding\decoding thing, turns out the answer is to give spark a few clues about what it is dealing with by adding an “Encoding” option: raw_notes_df2 = spark.read.options(header=”True”).option(“encoding”, “ISO-8859–1”).csv(notes_path2) BOOM!
🌐
Namastespark
namastespark.com › docs › spark › csv › handle-csv-files-of-different-encodings
Handling Encodings | Namaste Spark
May 30, 2025 - val encodeDf = spark.read .option("header", "true") .option("inferSchema", "true") .option("encoding","Windows-1252") .csv("csvFiles/empSalary.csv") encodeDf.show() encodeDf.printSchema()
🌐
Dbmstutorials
dbmstutorials.com › pyspark › spark-read-write-dataframe-options.html
PySpark: Dataframe Options
February 27, 2022 - df=spark.read.options(header=True, delimiter="|").csv("file:///path_to_file/tutorial_file_with_header.txt") ➠ lineSep example: This attribute can be used to specify single as a separator for each row while reading or writing a file using either option() or options() functions.
🌐
SQLRelease
sqlrelease.com › home › spark read file with special characters using pyspark
Spark read file with special characters using PySpark - SQLRelease
May 24, 2021 - In other words, the Spanish characters are not being replaced with the junk characters. In conclusion, we are able to read this file correctly into a Spark data frame by adding option(“encoding”, “windows-1252”) in the PySpark code.
🌐
Spark By {Examples}
sparkbyexamples.com › home › apache spark › spark read() options
Spark Read() options - Spark By {Examples}
May 6, 2026 - Spark provides several read options that help you to read files. The spark.read() is a method used to read data from various data sources such as CSV,
🌐
Databricks
community.databricks.com › s › question › 0D53f00001HKHWdCAP › how-to-import-data-and-apply-multiline-and-charset-utf8-at-the-same-time
How to import data and apply multiline and charset... - Databricks Community - 29116
September 25, 2021 - T_new_exp = spark.read\ .option("charset", "ISO-8859-1")\ .option("parserLib", "univocity")\ .option("multiLine", "true")\ .schema(schema)\ .csv(file)
Find elsewhere
🌐
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-load-save-functions.html
Generic Load/Save Functions - Spark 4.2.0 Documentation
The following ORC example will create bloom filter and use dictionary encoding only for favorite_color. For Parquet, there exists parquet.bloom.filter.enabled and parquet.enable.dictionary, too. To find more detailed information about the extra ORC/Parquet options, visit the official Apache ORC / Parquet websites. ... users_df = spark.read.orc("examples/src/main/resources/users.orc") (users_df.write.format("orc") .option("orc.bloom.filter.columns", "favorite_color") .option("orc.dictionary.key.threshold", "1.0") .option("orc.column.encoding.direct", "name") .save("users_with_options.orc"))
🌐
Apache Spark
spark.apache.org › docs › latest › api › java › org › apache › spark › sql › DataFrameReader.html
DataFrameReader (Spark 4.1.1 JavaDoc)
The text files must be encoded as UTF-8. If the directory structure of the text files contains partitioning information, those are ignored in the resulting Dataset. To include partitioning information as columns, use text. By default, each line in the text files is a new row in the resulting ...
🌐
The Internals of Spark SQL
jaceklaskowski.gitbooks.io › mastering-spark-sql › content › spark-sql-CSVFileFormat.html
CSVFileFormat · The Internals of Spark SQL
spark.read.format("csv").load("csv-datasets") // or the same as above using a shortcut spark.read.csv("csv-datasets") CSVFileFormat uses CSV options (that in turn are used to configure the underlying CSV parser from uniVocity-parsers project).
🌐
Apache JIRA
issues.apache.org › jira › browse › SPARK-20336
[SPARK-20336] spark.read.csv() with wholeFile=True option fails to read non ASCII unicode characters - ASF Jira
April 26, 2017 - When it is loaded without wholeFile=True option, non-ASCII characters are shown correctly although multi-line records are parsed incorrectly as follows: testdf_default = spark.read.csv("test.encoding.csv", header=True) testdf_default.show() +----+----+----+ |col1|col2|col3| +----+----+----+ | 1| a|text| | 2| b|テキスト| | 3| c| 텍스트| | 4| d|text| |テキスト|null|null| | 텍스트"|null|null| | 5| e|last| +----+----+----+ When wholeFile=True option is used, non-ASCII characters are broken as follows: testdf_wholefile = spark.read.csv("test.encoding.csv", header=True, wholeFile=True)
🌐
Stack Overflow
stackoverflow.com › questions › 76620125 › how-to-use-the-encoding-as-per-the-file-while-reading-from-spark
python - How to use the encoding as per the file, while reading from spark - Stack Overflow
import chardet.universaldetector as chardet_detector def get_dataframe(spark, path, schema, delimiter): """Return a dataframe setup to read a CSV file""" s3_file_content = spark.sparkContext.binaryFiles(path).flatMap(lambda x: x[1].splitlines()).collect() file_encoding = detect_file_encoding(s3_file_content) options = {"encoding": file_encoding} data_frame = spark.read.schema(schema).options(**options).csv(path).toDF( *[f.metadata.get("alias") for f in schema]) data_frame.show() return data_frame def detect_file_encoding(content): detector = chardet_detector.UniversalDetector() for line in content: detector.feed(line) if detector.done: break detector.close() return detector.result['encoding']
🌐
GitHub
github.com › apache › spark › pull › 20937 › files › f2a259ffb0a322daf4aca5a127ec990071d5f935
[SPARK-23094][SPARK-23723][SPARK-23724][SQL] Support custom encoding for json files by MaxGekk · Pull Request #20937 · apache/spark
Here is an example of using of the option: spark.read.schema(schema) .option("multiline", "true") .option("encoding", "UTF-16LE") .json(fileName) If the option is not specified, charset auto-detection mechanism is used by default.
Author: apache
🌐
Apache JIRA
issues.apache.org › jira › browse › SPARK-13108
[SPARK-13108] Encoding not working with non-ascii compatible encodings (UTF-16/32 etc.) - ASF Jira
January 24, 2018 - According to MAPREDUCE-232#comment-13183601, it still looks fine with most encodings though but without UTF-16/32. ... I tested this in Max OS. I converted `cars_iso-8859-1.csv` into `cars_utf-16.csv` as below: iconv -f iso-8859-1 -t utf-16 < cars_iso-8859-1.csv > cars_utf-16.csv ... val cars = "cars_utf-16.csv" sqlContext.read .format("csv") .option("charset", "utf-16") .option("delimiter", 'þ') .load(cars) .show()
🌐
Stack Overflow
stackoverflow.com › questions › 58262846 › reading-csv-with-multiline-option-and-encoding-option
python - Reading CSV with multiLine option and encoding option - Stack Overflow
Hello, Unfortunately, you cannot use "multiline" and "charset" together, if you use together encoding will be set to default. For more details, refer “Spark – Known issue": mail-archives.apache.org/mod_mbox/spark-commits/201804.mbox/… 2019-10-14T06:42:28.753Z+00:00
🌐
Spark Code Hub
sparkcodehub.com › dataframe › read text
Read Text | SPARK Tutorial | Spark Code Hub
encoding, allowing developers to customize how text data is interpreted. These options make it versatile for various text formats, from single-line logs to multi-line documents with custom separators.
🌐
Scribd
scribd.com › document › 925459117 › Spark-Read-Options-CheatSheet
Spark .read.option() Overview Guide | PDF
Save Spark Read Options CheatSheet ... automatically .option('delimiter', "," or ";" or "\t") # Field delimiter .option('encoding', "UTF-8") # Character encoding .option('multiLine', true/false) # Allow multi-line fields (es...
🌐
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-csv.html
CSV Files - Spark 4.2.0 Documentation
Spark SQL provides spark.read().csv("file_name") to read a file or directory of files in CSV format into Spark DataFrame, and dataframe.write().csv("path") to write to a CSV file. Function option() can be used to customize the behavior of reading or writing, such as controlling behavior of the header, delimiter character, character set, and so on.