I was able to read ISO-8859-1 using spark but when I store the same data to S3/hdfs back and read it, the format is converting to UTF-8.

ex: é to é

val df = spark.read.format("csv").option("delimiter", ",").option("ESCAPE quote", '"'). option("header",true).option("encoding", "ISO-8859-1").load("s3://bucket/folder")
Answer from Saida on Stack Overflow
🌐
Medium
medium.com › @mike_82447 › pyspark-character-encoding-fccfad3989bd
PySpark — character encoding
October 31, 2019 - So — its obviously a text encoding\decoding thing, turns out the answer is to give spark a few clues about what it is dealing with by adding an “Encoding” option: raw_notes_df2 = spark.read.options(header=”True”).option(“encoding”, “ISO-8859–1”).csv(notes_path2) BOOM!
🌐
Namastespark
namastespark.com › docs › spark › csv › handle-csv-files-of-different-encodings
Handling Encodings | Namaste Spark
May 30, 2025 - val encodeDf = spark.read .option("header", "true") .option("inferSchema", "true") .option("encoding","Windows-1252") .csv("csvFiles/empSalary.csv") encodeDf.show() encodeDf.printSchema()
🌐
Dbmstutorials
dbmstutorials.com › pyspark › spark-read-write-dataframe-options.html
PySpark: Dataframe Options
February 27, 2022 - Spark can read gzip without specifying codec but for writing gzip codec must be specified. compression is the synonym for codec. ... df.write.options(codec= "gzip").json("file:///path_to_file/data_files") ➠ quoteAll example: This attribute ...
🌐
SQLRelease
sqlrelease.com › home › spark read file with special characters using pyspark
Spark read file with special characters using PySpark - SQLRelease
May 24, 2021 - To read the above file correctly in PySpark, we need to add the file encoding option in the Spark read method. That is to say, we need to add an extra option in the previous read method.
🌐
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-load-save-functions.html
Generic Load/Save Functions - Spark 4.2.0 Documentation
The following ORC example will create bloom filter and use dictionary encoding only for favorite_color. For Parquet, there exists parquet.bloom.filter.enabled and parquet.enable.dictionary, too. To find more detailed information about the extra ORC/Parquet options, visit the official Apache ORC / Parquet websites. ORC data source: users_df = spark.read.orc("examples/src/main/resources/users.orc") (users_df.write.format("orc") .option("orc.bloom.filter.columns", "favorite_color") .option("orc.dictionary.key.threshold", "1.0") .option("orc.column.encoding.direct", "name") .save("users_with_options.orc")) Find full example code at "examples/src/main/python/sql/datasource.py" in the Spark repo.
🌐
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-csv.html
CSV Files - Spark 4.2.0 Documentation
Spark SQL provides spark.read().csv("file_name") to read a file or directory of files in CSV format into Spark DataFrame, and dataframe.write().csv("path") to write to a CSV file. Function option() can be used to customize the behavior of reading or writing, such as controlling behavior of the header, delimiter character, character set, and so on.
Find elsewhere
🌐
Databricks
community.databricks.com › s › question › 0D53f00001HKHWdCAP › how-to-import-data-and-apply-multiline-and-charset-utf8-at-the-same-time
How to import data and apply multiline and charset... - Databricks Community - 29116
September 25, 2021 - T_new_exp = spark.read\ .option("charset", "ISO-8859-1")\ .option("parserLib", "univocity")\ .option("multiLine", "true")\ .schema(schema)\ .csv(file)
🌐
Spark By {Examples}
sparkbyexamples.com › home › apache spark › spark read() options
Spark Read() options - Spark By {Examples}
May 6, 2026 - Spark provides several read options that help you to read files. The spark.read() is a method used to read data from various data sources such as CSV,
🌐
Apache JIRA
issues.apache.org › jira › browse › SPARK-20336
[SPARK-20336] spark.read.csv() with wholeFile=True option fails to read non ASCII unicode characters - ASF Jira
April 26, 2017 - When it is loaded without wholeFile=True option, non-ASCII characters are shown correctly although multi-line records are parsed incorrectly as follows: testdf_default = spark.read.csv("test.encoding.csv", header=True) testdf_default.show() +----+----+----+ |col1|col2|col3| +----+----+----+ ...
🌐
Apache Spark
spark.apache.org › docs › latest › api › java › org › apache › spark › sql › DataFrameReader.html
DataFrameReader (Spark 4.1.1 JavaDoc)
The text files must be encoded as UTF-8. If the directory structure of the text files contains partitioning information, those are ignored in the resulting Dataset. To include partitioning information as columns, use text. By default, each line in the text files is a new row in the resulting ...
🌐
The Internals of Spark SQL
jaceklaskowski.gitbooks.io › mastering-spark-sql › content › spark-sql-CSVFileFormat.html
CSVFileFormat · The Internals of Spark SQL
spark.read.format("csv").load("csv-datasets") // or the same as above using a shortcut spark.read.csv("csv-datasets") CSVFileFormat uses CSV options (that in turn are used to configure the underlying CSV parser from uniVocity-parsers project).
🌐
Apache JIRA
issues.apache.org › jira › browse › SPARK-13108
[SPARK-13108] Encoding not working with non-ascii compatible encodings (UTF-16/32 etc.) - ASF Jira
January 24, 2018 - According to MAPREDUCE-232#comment-13183601, it still looks fine with most encodings though but without UTF-16/32. ... I tested this in Max OS. I converted `cars_iso-8859-1.csv` into `cars_utf-16.csv` as below: iconv -f iso-8859-1 -t utf-16 < cars_iso-8859-1.csv > cars_utf-16.csv ... val cars = "cars_utf-16.csv" sqlContext.read .format("csv") .option("charset", "utf-16") .option("delimiter", 'þ') .load(cars) .show()
🌐
Stack Overflow
stackoverflow.com › questions › 76620125 › how-to-use-the-encoding-as-per-the-file-while-reading-from-spark
python - How to use the encoding as per the file, while reading from spark - Stack Overflow
import chardet.universaldetector as chardet_detector def get_dataframe(spark, path, schema, delimiter): """Return a dataframe setup to read a CSV file""" s3_file_content = spark.sparkContext.binaryFiles(path).flatMap(lambda x: x[1].splitlines()).collect() file_encoding = detect_file_encoding(s3_file_content) options = {"encoding": file_encoding} data_frame = spark.read.schema(schema).options(**options).csv(path).toDF( *[f.metadata.get("alias") for f in schema]) data_frame.show() return data_frame def detect_file_encoding(content): detector = chardet_detector.UniversalDetector() for line in content: detector.feed(line) if detector.done: break detector.close() return detector.result['encoding']
🌐
GitHub
github.com › apache › spark › pull › 20937 › files › f2a259ffb0a322daf4aca5a127ec990071d5f935
[SPARK-23094][SPARK-23723][SPARK-23724][SQL] Support custom encoding for json files by MaxGekk · Pull Request #20937 · apache/spark
Here is an example of using of the option: spark.read.schema(schema) .option("multiline", "true") .option("encoding", "UTF-16LE") .json(fileName) If the option is not specified, charset auto-detection mechanism is used by default.
Author: apache
🌐
Stack Overflow
stackoverflow.com › questions › 58262846 › reading-csv-with-multiline-option-and-encoding-option
python - Reading CSV with multiLine option and encoding option - Stack Overflow
df= sqlContext.read.format('csv').options(header='true',inferSchema='false',delimiter='\t',encoding='SJIS',multiline='true').load('/mnt/Data/Data.tsv') ... Please try to update encoding with charset. For more details, please refer to github.com/databricks/spark-csv.
🌐
Stack Overflow
stackoverflow.com › questions › 66842070 › encoding-option-in-spark-jdbc
oracle database - Encoding option in Spark JDBC - Stack Overflow
March 28, 2021 - I want to read data from an Oracle DB using Spark JDBC in a specific charset encoding like us-ascii but I am unable to. ... val res=spark.read.format("jdbc") .option("url", url) .option("user", "userid") .option("password", "pwd") .option("driver","oracle.jdbc.OracleDriver") .option("encoding", "us-ascii") .option("characterEncoding", "us-ascii") .option("query", tableQuery).option("fetchsize","10000") .load()
🌐
Spark Code Hub
sparkcodehub.com › dataframe › read text
Read Text | SPARK Tutorial | Spark Code Hub
spark.read.text method is the primary entry point for loading text files into DataFrames in Scala Spark. This section details its syntax, core options, and basic usage, with examples demonstrating how to read our sample text file. ... UTF-8). Defines the file encoding (e.g.,