🌐
Statology
statology.org › home › pyspark: how to filter for “not contains”
PySpark: How to Filter for "Not Contains"
October 12, 2023 - We can use the following syntax to filter the DataFrame to only contain rows where the team column does not contain “avs” anywhere in the string:
Discussions

Filtering rows that does not contain a string
Both methods fail due to syntax error could you please help me filter rows that does not contain a certain string in pyspark. More on community.databricks.com
🌐 community.databricks.com
0
August 6, 2020
Is there a way to filter a field not containing something in a spark dataframe using scala? - Stack Overflow
Hopefully I'm stupid and this will be easy. I have a dataframe containing the columns 'url' and 'referrer'. I want to extract all the referrers that contain the top level domain 'www.mydomain.c... More on stackoverflow.com
🌐 stackoverflow.com
PySpark: Filtering for a value not working?
It shouldn't make a difference, but does it work if you do df.filter(df.something_id == 'CyVmB5kL')? More on reddit.com
🌐 r/learnpython
2
1
April 3, 2023
Pyspark dataframe operator "IS NOT IN" - Stack Overflow
I would like to rewrite this from R to Pyspark, any nice looking suggestions? array More on stackoverflow.com
🌐 stackoverflow.com
🌐
Spark By {Examples}
sparkbyexamples.com › home › pyspark › pyspark not isin() or is not in operator
PySpark NOT isin() or IS NOT IN Operator - Spark By {Examples}
May 6, 2026 - The NOT isin() operation in PySpark is used to filter rows in a DataFrame where the column's value is not present in a specified list of values. This is
🌐
Arab Psychology
scales.arabpsychology.com › home › pyspark: filter for “not contains”
PySpark: Filter For “Not Contains”
November 17, 2025 - The mechanism for filtering a DataFrame using a “Not Contains” condition relies on applying a logical negation operator to the output of the standard string search function. In Python and, by extension, PySpark, the tilde symbol (~) serves ...
🌐
Statology
statology.org › home › how to use “is not in” in pyspark (with example)
How to Use "IS NOT IN" in PySpark (With Example)
October 10, 2023 - By using this operator along with the isin function, we are able to filter the DataFrame to only contain rows where the value in a particular column is not in a list of values. The following tutorials explain how to perform other common tasks in PySpark:
🌐
Sqlandhadoop
sqlandhadoop.com › pyspark-filter-25-examples-to-teach-you-everything
PySpark Filter – 25 examples to teach you everything – SQL & Hadoop
Refer to below diagram for easy reference to the multiple options available in PySpark Filter conditions. ... NOT : NOT is special conditional check which can reverse the obvious output. It basically negates the actual output by applying opposite FILTER condition on the input dataset to yield ...
Find elsewhere
🌐
MungingData
mungingdata.com › pyspark › filter-array
Filtering PySpark Arrays and DataFrame Array Columns - MungingData
Filter out all the rows that don't contain a word that starts with the letter a. starts_with_a = lambda s: s.startswith("a") res = df.filter(exists(col("some_words"), starts_with_a)) res.show() +-------------+ | some_words| +-------------+ |[apple, pear]| | [cat, ant]| +-------------+ ... See the PySpark exists and forall post for a detailed discussion of exists and the other method we'll talk about next, forall.
🌐
Statology
statology.org › home › pyspark: how to filter rows using not like
PySpark: How to Filter Rows Using NOT LIKE
October 30, 2023 - Notice that each of the rows in the resulting DataFrame do not contain a pattern like “avs” in the team column. Note that we used the like function to find all strings in the team column that had a pattern like “avs” and then we used the ~ symbol to negate this function. The end result is that we’re able to filter for only the rows in the DataFrame that do not have a pattern like “avs” in the team column. Note: You can find the complete documentation for the PySpark like function here.
🌐
Spark By {Examples}
sparkbyexamples.com › home › pyspark › pyspark where() & filter() for efficient data filtering
PySpark where() & filter() for efficient data filtering - Spark By {Examples}
May 6, 2026 - In this PySpark article, you will learn how to apply a filter on DataFrame columns of string, arrays, and struct types by using single and multiple
🌐
Reddit
reddit.com › r/learnpython › pyspark: filtering for a value not working?
r/learnpython on Reddit: PySpark: Filtering for a value not working?
April 3, 2023 -

Hi all, thanks for taking the time try and help me.

So this seems pretty basic but I'm really struggling. I have a pyspark dataframe that has a column called something_id and the ids in the column are a mixture of capital letters, lowercase letters and numbers. For example CyVmB5kL. I can see when I do df.show() that the something_id as a value called CyVmB5kL, but when I try df.filter("something_id = 'CyVmB5kL'").show() I get an empty dataframe. There's no whitespace, I've literally copied and pasted different values from df.show(). I've created my own df from stratch and tried the exact same steps and it worked as expected. So I'm lost and stackoverflow.com, chat-gpt and google have all failed me. You're my only hope. Any ideas?

Thanks in advanced

  • data_scallion

🌐
Dbmstutorials
dbmstutorials.com › pyspark › spark-dataframe-filters.html
PySpark: Dataframe Filters
This tutorial will explain how filters can be used on dataframes in Pyspark. where() function is an alias for filter() function. Following topics will be covered on this page: Basic filters · Filter using IN clause · Filter using not IN clause · Filter using List · Filter Null Values · Filter not Null Values · Filter using LIKE operator · Filter using not LIKE operator · Filter using Contains ·
Top answer
1 of 1
1

Split text into tokens/words and use arrays_overlap function to check if wanted or unwanted token is present:

df = df.filter(
    (
      F.arrays_overlap(
          F.split(F.regexp_replace(F.lower("message"), r"[^a-zA-Z0-9\s]+", ""), "\s+"),
          F.array([F.lit(c) for c in wanted_words])
          )
    )
    & 
    (
      ~F.arrays_overlap(
          F.split(F.regexp_replace(F.lower("message"), r"[^a-zA-Z0-9\s]+", ""), "\s+"),
          F.array([F.lit(c) for c in unwanted_words])
          )
    )
)

Full example:

columns = ["id","message"]
data = [["ab123","Hello my name is Chris"],["cd345","The room should be 2301"],["ef567","Welcome! What is your name?"],["gh873","That way please"],["kj893","The current year is 2022"]]
df = spark.createDataFrame(data).toDF(*columns)

wanted_words = ['name','room']
unwanted_words = ['welcome','year']

df = df.filter(
    (
      F.arrays_overlap(
          F.split(F.regexp_replace(F.lower("message"), r"[^a-zA-Z0-9\s]+", ""), "\s+"),
          F.array([F.lit(c) for c in wanted_words])
          )
    )
    & 
    (
      ~F.arrays_overlap(
          F.split(F.regexp_replace(F.lower("message"), r"[^a-zA-Z0-9\s]+", ""), "\s+"),
          F.array([F.lit(c) for c in unwanted_words])
          )
    )
)

[Out]:
+-----+------------------------+
|id   |message                 |
+-----+------------------------+
|ab123|Hello my name is Chris  |
|cd345|The room should be 2301 |
+-----+------------------------+

You can also pre-compute the tokens at once for efficiency:

df = df.withColumn("tokens", F.split(F.regexp_replace(F.lower("message"), r"[^a-zA-Z0-9\s]+", ""), "\s+"))

and use in "arrays_overlap":

F.arrays_overlap(F.col("tokens"), F.array([F.lit(c) for c in wanted_words]))
🌐
Palantir
palantir.com › docs › foundry › transforms-python-spark › pyspark-filtering
Python (Spark) • PySpark reference • Filtering • Palantir
Returns a boolean column expression indicating whether the column's string value contains the string (literal, or other column) provided in the parameter.
🌐
Apache
lists.apache.org › thread › k52smc1nlw4rll5b4sdm8x7ct84hhbyh
Not contains on Spark dataframe-Apache Mail Archives
Email display mode: · Modern rendering · Legacy rendering · This site requires JavaScript enabled. Please enable it