Pyspark will not decode correctly if the hex vales are preceded by double backslashes (ex: \\xBA instead of \xBA).

Using "take(3)" instead of "show()" showed that in fact there was a second backslash:

[Row(item_name='Jogador n\\xBA 10'),
 Row(item_name='Camisa N\\xB0 9'),
 Row(item_name='Uniforme M\\xE9dio')]

To solve this I created a UDF to decode using "unicode-escape" method:

import pyspark.sql.functions as F
import pyspark.sql.types as T
my_udf = F.udf(lambda x: x.encode().decode('unicode-escape'),T.StringType())
df.withColumn('test', my_udf('item_name')).show()
+------------------+---------------+
|         item_name|           test|
+------------------+---------------+
|  Jogador n\xBA 10|  Jogador nº 10|
|    Camisa N\xB0 9|    Camisa N° 9|
| Uniforme M\xE9dio| Uniforme Médio|
+------------------+---------------+
Answer from Maviles on Stack Overflow
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › pyspark.sql › api › pyspark.sql.functions.encode.html
pyspark.sql.functions.encode — PySpark 4.2.0 documentation
Computes the first argument into a binary from a string using the provided character set (one of ‘US-ASCII’, ‘ISO-8859-1’, ‘UTF-8’, ‘UTF-16BE’, ‘UTF-16LE’, ‘UTF-16’, ‘UTF-32’). New in version 1.5.0. Changed in version 3.4.0: Supports Spark Connect.
🌐
Medium
medium.com › @diangermishuizen › convert-the-character-set-encoding-of-a-string-field-in-a-pyspark-dataframe-on-databricks-3c841531e5a8
Convert the Character Set/Encoding of a String field in a PySpark DataFrame on Databricks | by Dian Germishuizen | Medium
May 21, 2022 - We are using the .withColumns() function on the dataFrame to overwrite the contents of the field. The new contents is retrieved from the output of the sql.functions.encode() function. In that function, we are converting the data to the utf-8 ...
🌐
GitHub
github.com › databricks › spark-xml › issues › 229
How to Encode UTF-8 for Dataframe · Issue #229 · databricks/spark-xml
January 23, 2017 - Hi! please, maybe you can help me. How to encoding output to UTF-8, there are wrong values. coDf=Df.selectExpr("explode(copy) as e").select("e._sourceOwner", "e._targetOwner","e._sourceKeySet", "e._targetKeySet", "e._sourceMetadata", "e....
Author: databricks
🌐
Medium
medium.com › @mike_82447 › pyspark-character-encoding-fccfad3989bd
PySpark — character encoding
October 31, 2019 - raw_notes_df2 = spark.read.options(header=”True”).option(“encoding”, “ISO-8859–1”).csv(notes_path2)
🌐
Databricks
community.databricks.com › s › question › 0D53f00001HKHWdCAP › how-to-import-data-and-apply-multiline-and-charset-utf8-at-the-same-time
How to import data and apply multiline and charset... - Databricks Community - 29116
September 25, 2021 - Please make sure you are using or enforcing python 3. python 2 is default and it will have issues with encoding ... You could also potentially use the .withColumns() function on the data frame, and use the pyspark.sql.functions.encode function to convert the characterset to the one you need.
🌐
Stack Overflow
stackoverflow.com › questions › 47525063 › write-spark-dataframe-in-postgresql-with-utf-8-encoding
python - Write Spark Dataframe in PostgreSQL with UTF-8 encoding - Stack Overflow
November 28, 2017 - It seems by default Pyspark is trying to encode the dataframe as ASCII, thus I should specify the correct encoding (UTF-8).
Find elsewhere
🌐
Databricks
community.databricks.com › s › feed › 0D53f00001HKHfWCAX
Removing non-ascii and special character in pyspark
November 1, 2022 - from pyspark.sql import SparkSession spark = SparkSession.builder.appName("Python Spark").getOrCreate() df = spark.read.csv("filepath\Customers_v01.csv",header=True,sep=","); myres = df.rdd.map(lambda x: x[1].encode().decode('utf-8')) print(myres.collect())
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › pyspark.sql › api › pyspark.sql.functions.decode.html
pyspark.sql.functions.decode — PySpark 4.2.0 documentation
Computes the first argument into a string from a binary using the provided character set (one of ‘US-ASCII’, ‘ISO-8859-1’, ‘UTF-8’, ‘UTF-16BE’, ‘UTF-16LE’, ‘UTF-16’, ‘UTF-32’). New in version 1.5.0. Changed in version 3.4.0: Supports Spark Connect.
🌐
The Internals of Spark SQL
jaceklaskowski.gitbooks.io › mastering-spark-sql › content › spark-sql-CSVFileFormat.html
CSVFileFormat · The Internals of Spark SQL
CSVFileFormat is a TextBasedFileFormat for csv format (i.e. registers itself to handle files in csv format and converts them to Spark SQL rows) · CSVFileFormat uses CSV options (that in turn are used to configure the underlying CSV parser from uniVocity-parsers project)
🌐
Stack Overflow
stackoverflow.com › questions › 42819708 › string-encoding-issue-in-spark-sql-dataframe
python - String encoding issue in Spark SQL/DataFrame - Stack Overflow
When I read the file into pyspark throught the following code: schema = StructType([ StructField("id", IntegerType(), True), StructField("name", StringType(), True)]) df = sqlContext.read.csv("file.csv", header=False, schema = schema) On executing df.first() I get the following output: Row(artistid=1240105, artistname=u'Andr\xe9 Visior') ... Sign up to request clarification or add additional context in comments. ... no open file in eexcel and save it back as CSV(utf-8) and then apply the same program.
🌐
Stack Overflow
stackoverflow.com › questions › 59386160 › unable-to-include-utf-8-character-as-value-for-a-column-in-spark-dataframe
python - Unable to include UTF-8 character as value for a column in Spark Dataframe - Stack Overflow
December 18, 2019 - I am trying to convert a list of key value pairs of "ID" and "Symbols" in to a Dataframe of two columns ["id","value"].Sample list is as given below with 4 values.(Spark session is defined as "spark") L=[(1,"@"),(2,"#"),(3,"$"),(4,"£")] df=spark.sparkContext.parallelize(L).toDF(["ID","VALUE"]) df.show() ... "UnicodeEncodeError: 'ascii' codec can not encode characters in position 66-67:ordinal not in range(128) I know it's happening since £ is not a ascii character. I have tried converting the characters in list into utf-8 / unicode , but still I am getting above error.
🌐
Stack Overflow
stackoverflow.com › questions › 43744181 › pyspark-writing-accented-character-in-a-dataframe
encoding - pyspark writing accented character in a dataframe - Stack Overflow
May 2, 2017 - How do I make pyspark write accented character in the csv file ? I have tried using log['a_key'].encode('UTF-8') but it was the same result
🌐
UnoGeeks
unogeeks.com › home › blog › databricks utf-8
Databricks UTF-8
May 30, 2024 - Double-check your input data encoding, connection settings, and output configurations. Community Resources: The Databricks community forums and documentation often have discussions and solutions for handling UTF-8 related issues. ... from pyspark.sql.functions import encode, decode # Read a CSV file with UTF-8 encoding df = spark.read.csv("path/to/file.csv", header=True, inferSchema=True, encoding='utf-8') # Convert a column to UTF-8 encoded bytes df = df.withColumn("encoded_column", encode(df["column_name"], "utf-8")) # Convert UTF-8 encoded bytes back to a string df = df.withColumn("decoded_column", decode(df["encoded_column"], "utf-8