IIUC, you want to return the rows in which column_a is "like" (in the SQL sense) any of the values in list_a.

One way is to use functools.reduce:

from functools import reduce

list_a = ['string', 'third']

df1 = df.where(
    reduce(lambda a, b: a|b, (df['column_a'].like('%'+pat+"%") for pat in list_a))
)
df1.show()
#+------------+-----+
#|    column_a|count|
#+------------+-----+
#| some_string|   10|
#|third_string|   30|
#+------------+-----+

Essentially you loop over all of the possible strings in list_a to compare in like and "OR" the results. Here is the execution plan:

df1.explain()
#== Physical Plan ==
#*(1) Filter (Contains(column_a#0, string) || Contains(column_a#0, third))
#+- Scan ExistingRDD[column_a#0,count#1]

Another option is to use pyspark.sql.Column.rlike instead of like.

df2 = df.where(
    df['column_a'].rlike("|".join(["(" + pat + ")" for pat in list_a]))
)

df2.show()
#+------------+-----+
#|    column_a|count|
#+------------+-----+
#| some_string|   10|
#|third_string|   30|
#+------------+-----+

Which has the corresponding execution plan:

df2.explain()
#== Physical Plan ==
#*(1) Filter (isnotnull(column_a#0) && column_a#0 RLIKE (string)|(third))
#+- Scan ExistingRDD[column_a#0,count#1]
Answer from pault on Stack Overflow
🌐
GeeksforGeeks
geeksforgeeks.org › python › filtering-a-row-in-pyspark-dataframe-based-on-matching-values-from-a-list
Filtering a row in PySpark DataFrame based on matching values from a list - GeeksforGeeks
July 28, 2021 - In this article, we are going to filter the rows in the dataframe based on matching values in the list by using isin in Pyspark dataframe · isin(): This is used to find the elements contains in a given dataframe, it will take the elements and get the elements to match to the data
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › pyspark.sql › api › pyspark.sql.Column.contains.html
pyspark.sql.Column.contains — PySpark 4.1.2 documentation
>>> df = spark.createDataFrame( ... [(2, "Alice"), (5, "Bob")], ["age", "name"]) >>> df.filter(df.name.contains('o')).collect() [Row(age=5, name='Bob')]
🌐
Spark By {Examples}
sparkbyexamples.com › home › apache spark › spark filter using contains() examples
Spark Filter Using contains() Examples - Spark By {Examples}
May 6, 2026 - In Spark & PySpark, contains() function is used to match a column value contains in a literal string (matches on part of the string), this is mostly
🌐
Stack Overflow
stackoverflow.com › questions › 69728936 › loop-through-pyspark-dataframe-once-to-find-rows-that-contains-list-of-values
Loop through pyspark dataframe once to find rows that contains list of values - Stack Overflow
Put country_list in a dataframe with one column country. Then join/merge/union with df, which will all scale where a for loop does not. ... from pyspark.sql import functions as F (df .withColumn('ExtractColumn', F .when(F.col('projectname').startswith('TBD'), F.regexp_replace('projectname', 'TBD ', '')) ) .show(10, False) ) # Output # +------------------+--------------+--------------+ # |projectname |projectnumber |ExtractColumn | # +------------------+--------------+--------------+ # |TBD Canada |10000000000029|Canada | # |TBD China |10000000000033|China | # |TBD United Kingdom|10000000000033|United Kingdom| # |US Diver |10000000000234|null | # |Flower 2.0 |10000000023947|null | # +------------------+--------------+--------------+
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › pyspark.sql › api › pyspark.sql.functions.contains.html
pyspark.sql.functions.contains — PySpark 4.2.0 documentation
>>> df = spark.createDataFrame([("Spark SQL", "Spark")], ['a', 'b']) >>> df.select(contains(df.a, df.b).alias('r')).collect() [Row(r=True)]
Find elsewhere
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › pyspark.sql › api › pyspark.sql.functions.array_contains.html
pyspark.sql.functions.array_contains — PySpark 4.2.0 documentation
Example 1: Basic usage of array_contains function. >>> from pyspark.sql import functions as sf >>> df = spark.createDataFrame([(["a", "b", "c"],), ([],)], ['data']) >>> df.select(sf.array_contains(df.data, "a")).show() +-----------------------+ |array_contains(data, a)| +-----------------------+ ...
🌐
Spark By {Examples}
sparkbyexamples.com › home › pyspark › pyspark isin() & sql in operator
PySpark isin() & SQL IN Operator - Spark By {Examples}
May 6, 2026 - In PySpark, the isin() function, or the IN operator is used to check DataFrame values and see if they're present in a given list of values. This function
🌐
Spark By {Examples}
sparkbyexamples.com › home › pyspark › pyspark filter using contains() examples
PySpark Filter using contains() Examples - Spark By {Examples}
May 6, 2026 - PySpark SQL contains() function is used to match a column value contains in a literal string (matches on part of the string), this is mostly used to
🌐
Statology
statology.org › home › pyspark: how to filter using “contains”
PySpark: How to Filter Using "Contains"
October 12, 2023 - #filter DataFrame where team column contains 'avs' df.filter(df.team.contains('avs')).show() The following example shows how to use this syntax in practice. Suppose we have the following PySpark DataFrame that contains information about points scored by various basketball players:
🌐
MungingData
mungingdata.com › pyspark › filter-array
Filtering PySpark Arrays and DataFrame Array Columns - MungingData
Use filter to append an arr_evens column that only contains the even numbers from some_arr: from pyspark.sql.functions import * is_even = lambda x: x % 2 == 0 res = df.withColumn("arr_evens", filter(col("some_arr"), is_even)) res.show() +---------------+---------+ | some_arr|arr_evens| +---------------+---------+ |[1, 2, 3, 5, 7]| [2]| | [2, 4, 9]| [2, 4]| +---------------+---------+ The vanilla filter method in Python works similarly: list(filter(is_even, [2, 4, 9])) # [2, 4] The Spark filter function takes is_even as the second argument and the Python filter function takes is_even as the first argument.
🌐
Stack Overflow
stackoverflow.com › questions › 54759084 › pyspark-how-do-we-check-if-a-column-value-is-contained-in-a-list
apache spark - pyspark how do we check if a column value is contained in a list - Stack Overflow
Try to extract all of the values in the list l and concatenate the results. If the resulting concatenated string is an empty string, that means none of the values matched. ... from pyspark.sql.functions import concat, regexp_extract records = df.where(concat(*[regexp_extract("score", str(val), 0) for val in l]) != "") records.show() #+---+-----+ #| id|score| #+---+-----+ #| 0| 100| #| 0| 1| #| 1| 10| #| 3| 18| #| 3| 18| #| 3| 18| #+---+-----+
🌐
Stack Overflow
stackoverflow.com › questions › 59158666 › pyspark-how-to-loop-through-a-dataframe-column-which-contains-list-elements
Pyspark: how to loop through a dataframe column which contains list elements? - Stack Overflow
December 3, 2019 - How do I loop through the all_posts['tagged_persons'] column to see if an element of the list AND the corresponding year equal a row of the headliners dataframe? ... Save this answer. ... Show activity on this post. You can explode the tagged_person column first refer. After that join it with headliners based on tagged_person, Artist and year column and then filter the rows where Artist is null will give you the resultant data. Copyfrom pyspark.sql.functions import explode all_posts = all_posts.select(all_posts.Artist, explode(all_posts.tagged_persons)) cond = [all_posts.tagged_persons == headliners.Artist, all_posts.year == headliners.Year] join_df = all_posts.join(headliners, cond, 'left') filter_df = join_df.filter(col("Artist").isNotNull())
🌐
Spark Code Hub
sparkcodehub.com › dataframe › check value in list
Checking if a Value Exists in a List in Spark DataFrames
They integrate with other operations, such as string manipulation (Spark How to Do String Manipulation), regex (Spark DataFrame Regex Expressions), or conditional logic (Spark How to Use Case Statement), making them versatile for data cleaning (Spark How to Cleaning and Preprocessing Data in Spark DataFrame) and analytics. For Python-based filtering, see PySpark DataFrame Filter. Spark provides several functions to check if a value exists in a list, primarily ... array_contains, along with SQL expressions and custom approaches.
🌐
Statology
statology.org › home › pyspark: how to filter rows based on values in a list
PySpark: How to Filter Rows Based on Values in a List
October 30, 2023 - This tutorial explains how to filter a PySpark DataFrame for rows that contain a value from a list, including an example.
🌐
Statology
statology.org › home › pyspark: filter for rows that contain one of multiple values
PySpark: Filter for Rows that Contain One of Multiple Values
October 12, 2023 - #define array of substrings to search for my_values = ['ets', 'urs'] regex_values = "|".join(my_values) filter DataFrame where team column contains any substring from array df.filter(df.team.rlike(regex_values)).show() The following example shows how to use this syntax in practice. Suppose we have the following PySpark DataFrame that contains information about points scored by various basketball players: