Pyspark will not decode correctly if the hex vales are preceded by double backslashes (ex: \\xBA instead of \xBA).
Using "take(3)" instead of "show()" showed that in fact there was a second backslash:
[Row(item_name='Jogador n\\xBA 10'),
Row(item_name='Camisa N\\xB0 9'),
Row(item_name='Uniforme M\\xE9dio')]
To solve this I created a UDF to decode using "unicode-escape" method:
import pyspark.sql.functions as F
import pyspark.sql.types as T
my_udf = F.udf(lambda x: x.encode().decode('unicode-escape'),T.StringType())
df.withColumn('test', my_udf('item_name')).show()
+------------------+---------------+
| item_name| test|
+------------------+---------------+
| Jogador n\xBA 10| Jogador nº 10|
| Camisa N\xB0 9| Camisa N° 9|
| Uniforme M\xE9dio| Uniforme Médio|
+------------------+---------------+
Answer from Maviles on Stack OverflowSo this workaround helped to solve this, By changing the default encoding for the session
import sys
reload(sys)
sys.setdefaultencoding('UTF-8')
and then
df_pandas_df = df_pandas_df.astype(str)
converts whole dataframe as string df.
Instead of directly casting it to string try to infer types of pandas DataFrame using following statement:
df_pandas_df .apply(lambda x: pd.lib.infer_dtype(x.values))
UPD:
try to perform mapping without .str invocation.
Maybe something like below:
for cols in df_pandas_df.columns:
df_pandas_df[cols] = df_pandas_df[cols].apply(lambda x: unicode(x, errors='ignore'))
https://issues.apache.org/jira/browse/SPARK-11772 talks about this issue and gives a solution that runs:
export PYTHONIOENCODING=utf8
before running pyspark. I wonder why above works, because sys.getdefaultencoding() returned utf-8 for me even without it.
How to set sys.stdout encoding in Python 3? also talks about this and gives the following solution for Python 3:
import sys
sys.stdout = open(sys.stdout.fileno(), mode='w', encoding='utf8', buffering=1)
import sys
reload(sys)
sys.setdefaultencoding('utf-8')
This works for me, I am setting the encoding upfront and it is valid throughout the script.