I was able to read ISO-8859-1 using spark but when I store the same data to S3/hdfs back and read it, the format is converting to UTF-8.
ex: é to é
val df = spark.read.format("csv").option("delimiter", ",").option("ESCAPE quote", '"'). option("header",true).option("encoding", "ISO-8859-1").load("s3://bucket/folder")
Answer from Saida on Stack OverflowMedium
medium.com › @mike_82447 › pyspark-character-encoding-fccfad3989bd
PySpark — character encoding
October 31, 2019 - So — its obviously a text encoding\decoding thing, turns out the answer is to give spark a few clues about what it is dealing with by adding an “Encoding” option: raw_notes_df2 = spark.read.options(header=”True”).option(“encoding”, “ISO-8859–1”).csv(notes_path2) BOOM!
Author: AbsaOSS
Namastespark
namastespark.com › docs › spark › csv › handle-csv-files-of-different-encodings
Handling Encodings | Namaste Spark
May 30, 2025 - val encodeDf = spark.read .option("header", "true") .option("inferSchema", "true") .option("encoding","Windows-1252") .csv("csvFiles/empSalary.csv") encodeDf.show() encodeDf.printSchema()
Top answer 1 of 3
20
I was able to read ISO-8859-1 using spark but when I store the same data to S3/hdfs back and read it, the format is converting to UTF-8.
ex: é to é
val df = spark.read.format("csv").option("delimiter", ",").option("ESCAPE quote", '"'). option("header",true).option("encoding", "ISO-8859-1").load("s3://bucket/folder")
2 of 3
4
My guess is that the input file is not in UTF-8 and hence you get the incorrect characters.
My recommendation would be to write a pure Java application (with no Spark at all) and see if reading and writing gives the same results with UTF-8 encoding.
Dbmstutorials
dbmstutorials.com › pyspark › spark-read-write-dataframe-options.html
PySpark: Dataframe Options
February 27, 2022 - Spark can read gzip without specifying codec but for writing gzip codec must be specified. compression is the synonym for codec. ... df.write.options(codec= "gzip").json("file:///path_to_file/data_files") ➠ quoteAll example: This attribute ...
SQLRelease
sqlrelease.com › home › spark read file with special characters using pyspark
Spark read file with special characters using PySpark - SQLRelease
May 24, 2021 - To read the above file correctly in PySpark, we need to add the file encoding option in the Spark read method. That is to say, we need to add an extra option in the previous read method.
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-load-save-functions.html
Generic Load/Save Functions - Spark 4.2.0 Documentation
The following ORC example will create bloom filter and use dictionary encoding only for favorite_color. For Parquet, there exists parquet.bloom.filter.enabled and parquet.enable.dictionary, too. To find more detailed information about the extra ORC/Parquet options, visit the official Apache ORC / Parquet websites. ORC data source: users_df = spark.read.orc("examples/src/main/resources/users.orc") (users_df.write.format("orc") .option("orc.bloom.filter.columns", "favorite_color") .option("orc.dictionary.key.threshold", "1.0") .option("orc.column.encoding.direct", "name") .save("users_with_options.orc")) Find full example code at "examples/src/main/python/sql/datasource.py" in the Spark repo.
Apache Spark
spark.apache.org › docs › latest › sql-data-sources-csv.html
CSV Files - Spark 4.2.0 Documentation
Spark SQL provides spark.read().csv("file_name") to read a file or directory of files in CSV format into Spark DataFrame, and dataframe.write().csv("path") to write to a CSV file. Function option() can be used to customize the behavior of reading or writing, such as controlling behavior of the header, delimiter character, character set, and so on.
Databricks
community.databricks.com › s › question › 0D53f00001HKHWdCAP › how-to-import-data-and-apply-multiline-and-charset-utf8-at-the-same-time
How to import data and apply multiline and charset... - Databricks Community - 29116
September 25, 2021 - T_new_exp = spark.read\ .option("charset", "ISO-8859-1")\ .option("parserLib", "univocity")\ .option("multiLine", "true")\ .schema(schema)\ .csv(file)
Apache JIRA
issues.apache.org › jira › browse › SPARK-20336
[SPARK-20336] spark.read.csv() with wholeFile=True option fails to read non ASCII unicode characters - ASF Jira
April 26, 2017 - When it is loaded without wholeFile=True option, non-ASCII characters are shown correctly although multi-line records are parsed incorrectly as follows: testdf_default = spark.read.csv("test.encoding.csv", header=True) testdf_default.show() +----+----+----+ |col1|col2|col3| +----+----+----+ ...
Apache Spark
spark.apache.org › docs › latest › api › java › org › apache › spark › sql › DataFrameReader.html
DataFrameReader (Spark 4.1.1 JavaDoc)
The text files must be encoded as UTF-8. If the directory structure of the text files contains partitioning information, those are ignored in the resulting Dataset. To include partitioning information as columns, use text. By default, each line in the text files is a new row in the resulting ...
The Internals of Spark SQL
jaceklaskowski.gitbooks.io › mastering-spark-sql › content › spark-sql-CSVFileFormat.html
CSVFileFormat · The Internals of Spark SQL
spark.read.format("csv").load("csv-datasets") // or the same as above using a shortcut spark.read.csv("csv-datasets") CSVFileFormat uses CSV options (that in turn are used to configure the underlying CSV parser from uniVocity-parsers project).
Apache JIRA
issues.apache.org › jira › browse › SPARK-13108
[SPARK-13108] Encoding not working with non-ascii compatible encodings (UTF-16/32 etc.) - ASF Jira
January 24, 2018 - According to MAPREDUCE-232#comment-13183601, it still looks fine with most encodings though but without UTF-16/32. ... I tested this in Max OS. I converted `cars_iso-8859-1.csv` into `cars_utf-16.csv` as below: iconv -f iso-8859-1 -t utf-16 < cars_iso-8859-1.csv > cars_utf-16.csv ... val cars = "cars_utf-16.csv" sqlContext.read .format("csv") .option("charset", "utf-16") .option("delimiter", 'þ') .load(cars) .show()
Stack Overflow
stackoverflow.com › questions › 76620125 › how-to-use-the-encoding-as-per-the-file-while-reading-from-spark
python - How to use the encoding as per the file, while reading from spark - Stack Overflow
import chardet.universaldetector as chardet_detector def get_dataframe(spark, path, schema, delimiter): """Return a dataframe setup to read a CSV file""" s3_file_content = spark.sparkContext.binaryFiles(path).flatMap(lambda x: x[1].splitlines()).collect() file_encoding = detect_file_encoding(s3_file_content) options = {"encoding": file_encoding} data_frame = spark.read.schema(schema).options(**options).csv(path).toDF( *[f.metadata.get("alias") for f in schema]) data_frame.show() return data_frame def detect_file_encoding(content): detector = chardet_detector.UniversalDetector() for line in content: detector.feed(line) if detector.done: break detector.close() return detector.result['encoding']
Stack Overflow
stackoverflow.com › questions › 58262846 › reading-csv-with-multiline-option-and-encoding-option
python - Reading CSV with multiLine option and encoding option - Stack Overflow
df= sqlContext.read.format('csv').options(header='true',inferSchema='false',delimiter='\t',encoding='SJIS',multiline='true').load('/mnt/Data/Data.tsv') ... Please try to update encoding with charset. For more details, please refer to github.com/databricks/spark-csv.
Stack Overflow
stackoverflow.com › questions › 66842070 › encoding-option-in-spark-jdbc
oracle database - Encoding option in Spark JDBC - Stack Overflow
March 28, 2021 - I want to read data from an Oracle DB using Spark JDBC in a specific charset encoding like us-ascii but I am unable to. ... val res=spark.read.format("jdbc") .option("url", url) .option("user", "userid") .option("password", "pwd") .option("driver","oracle.jdbc.OracleDriver") .option("encoding", "us-ascii") .option("characterEncoding", "us-ascii") .option("query", tableQuery).option("fetchsize","10000") .load()