To select rows whose column value equals a scalar, some_value, use ==:

df.loc[df['column_name'] == some_value]

To select rows whose column value is in an iterable, some_values, use isin:

df.loc[df['column_name'].isin(some_values)]

Combine multiple conditions with &:

df.loc[(df['column_name'] >= A) & (df['column_name'] <= B)]

Note the parentheses. Due to Python's operator precedence rules, & binds more tightly than <= and >=. Thus, the parentheses in the last example are necessary. Without the parentheses

df['column_name'] >= A & df['column_name'] <= B

is parsed as

df['column_name'] >= (A & df['column_name']) <= B

which results in a Truth value of a Series is ambiguous error.


To select rows whose column value does not equal some_value, use !=:

df.loc[df['column_name'] != some_value]

The isin returns a boolean Series, so to select rows whose value is not in some_values, negate the boolean Series using ~:

df = df.loc[~df['column_name'].isin(some_values)] # .loc is not in-place replacement

For example,

import pandas as pd
import numpy as np
df = pd.DataFrame({'A': 'foo bar foo bar foo bar foo foo'.split(),
                   'B': 'one one two three two two one three'.split(),
                   'C': np.arange(8), 'D': np.arange(8) * 2})
print(df)
#      A      B  C   D
# 0  foo    one  0   0
# 1  bar    one  1   2
# 2  foo    two  2   4
# 3  bar  three  3   6
# 4  foo    two  4   8
# 5  bar    two  5  10
# 6  foo    one  6  12
# 7  foo  three  7  14

print(df.loc[df['A'] == 'foo'])

yields

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

If you have multiple values you want to include, put them in a list (or more generally, any iterable) and use isin:

print(df.loc[df['B'].isin(['one','three'])])

yields

     A      B  C   D
0  foo    one  0   0
1  bar    one  1   2
3  bar  three  3   6
6  foo    one  6  12
7  foo  three  7  14

Note, however, that if you wish to do this many times, it is more efficient to make an index first, and then use df.loc:

df = df.set_index(['B'])
print(df.loc['one'])

yields

       A  C   D
B              
one  foo  0   0
one  bar  1   2
one  foo  6  12

or, to include multiple values from the index use df.index.isin:

df.loc[df.index.isin(['one','two'])]

yields

       A  C   D
B              
one  foo  0   0
one  bar  1   2
two  foo  2   4
two  foo  4   8
two  bar  5  10
one  foo  6  12
Answer from unutbu on Stack Overflow
🌐
Statology
statology.org › home › pandas: how to select rows based on column values
Pandas: How to Select Rows Based on Column Values
March 25, 2025 - You’re right that this is a common need when working with pandas. To select rows where a string column equals one of multiple values (like team “B” OR “C”), the `.isin()` method actually works perfectly.
Top answer
1 of 16
6655

To select rows whose column value equals a scalar, some_value, use ==:

df.loc[df['column_name'] == some_value]

To select rows whose column value is in an iterable, some_values, use isin:

df.loc[df['column_name'].isin(some_values)]

Combine multiple conditions with &:

df.loc[(df['column_name'] >= A) & (df['column_name'] <= B)]

Note the parentheses. Due to Python's operator precedence rules, & binds more tightly than <= and >=. Thus, the parentheses in the last example are necessary. Without the parentheses

df['column_name'] >= A & df['column_name'] <= B

is parsed as

df['column_name'] >= (A & df['column_name']) <= B

which results in a Truth value of a Series is ambiguous error.


To select rows whose column value does not equal some_value, use !=:

df.loc[df['column_name'] != some_value]

The isin returns a boolean Series, so to select rows whose value is not in some_values, negate the boolean Series using ~:

df = df.loc[~df['column_name'].isin(some_values)] # .loc is not in-place replacement

For example,

import pandas as pd
import numpy as np
df = pd.DataFrame({'A': 'foo bar foo bar foo bar foo foo'.split(),
                   'B': 'one one two three two two one three'.split(),
                   'C': np.arange(8), 'D': np.arange(8) * 2})
print(df)
#      A      B  C   D
# 0  foo    one  0   0
# 1  bar    one  1   2
# 2  foo    two  2   4
# 3  bar  three  3   6
# 4  foo    two  4   8
# 5  bar    two  5  10
# 6  foo    one  6  12
# 7  foo  three  7  14

print(df.loc[df['A'] == 'foo'])

yields

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

If you have multiple values you want to include, put them in a list (or more generally, any iterable) and use isin:

print(df.loc[df['B'].isin(['one','three'])])

yields

     A      B  C   D
0  foo    one  0   0
1  bar    one  1   2
3  bar  three  3   6
6  foo    one  6  12
7  foo  three  7  14

Note, however, that if you wish to do this many times, it is more efficient to make an index first, and then use df.loc:

df = df.set_index(['B'])
print(df.loc['one'])

yields

       A  C   D
B              
one  foo  0   0
one  bar  1   2
one  foo  6  12

or, to include multiple values from the index use df.index.isin:

df.loc[df.index.isin(['one','two'])]

yields

       A  C   D
B              
one  foo  0   0
one  bar  1   2
two  foo  2   4
two  foo  4   8
two  bar  5  10
one  foo  6  12
2 of 16
854

There are several ways to select rows from a Pandas dataframe:

  1. Boolean indexing (df[df['col'] == value] )
  2. Positional indexing (df.iloc[...])
  3. Label indexing (df.xs(...))
  4. df.query(...) API

Below I show you examples of each, with advice when to use certain techniques. Assume our criterion is column 'A' == 'foo'

(Note on performance: For each base type, we can keep things simple by using the Pandas API or we can venture outside the API, usually into NumPy, and speed things up.)


Setup

The first thing we'll need is to identify a condition that will act as our criterion for selecting rows. We'll start with the OP's case column_name == some_value, and include some other common use cases.

Borrowing from @unutbu:

import pandas as pd, numpy as np

df = pd.DataFrame({'A': 'foo bar foo bar foo bar foo foo'.split(),
                   'B': 'one one two three two two one three'.split(),
                   'C': np.arange(8), 'D': np.arange(8) * 2})

1. Boolean indexing

... Boolean indexing requires finding the true value of each row's 'A' column being equal to 'foo', then using those truth values to identify which rows to keep. Typically, we'd name this series, an array of truth values, mask. We'll do so here as well.

mask = df['A'] == 'foo'

We can then use this mask to slice or index the data frame

df[mask]

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

This is one of the simplest ways to accomplish this task and if performance or intuitiveness isn't an issue, this should be your chosen method. However, if performance is a concern, then you might want to consider an alternative way of creating the mask.


2. Positional indexing

Positional indexing (df.iloc[...]) has its use cases, but this isn't one of them. In order to identify where to slice, we first need to perform the same boolean analysis we did above. This leaves us performing one extra step to accomplish the same task.

mask = df['A'] == 'foo'
pos = np.flatnonzero(mask)
df.iloc[pos]

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

3. Label indexing

Label indexing can be very handy, but in this case, we are again doing more work for no benefit

df.set_index('A', append=True, drop=False).xs('foo', level=1)

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

4. df.query() API

pd.DataFrame.query is a very elegant/intuitive way to perform this task, but is often slower. However, if you pay attention to the timings below, for large data, the query is very efficient. More so than the standard approach and of similar magnitude as my best suggestion.

df.query('A == "foo"')

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

My preference is to use the Boolean mask

Actual improvements can be made by modifying how we create our Boolean mask.

mask alternative 1 Use the underlying NumPy array and forgo the overhead of creating another pd.Series

mask = df['A'].values == 'foo'

I'll show more complete time tests at the end, but just take a look at the performance gains we get using the sample data frame. First, we look at the difference in creating the mask

%timeit mask = df['A'].values == 'foo'
%timeit mask = df['A'] == 'foo'

5.84 µs ± 195 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)
166 µs ± 4.45 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

Evaluating the mask with the NumPy array is ~ 30 times faster. This is partly due to NumPy evaluation often being faster. It is also partly due to the lack of overhead necessary to build an index and a corresponding pd.Series object.

Next, we'll look at the timing for slicing with one mask versus the other.

mask = df['A'].values == 'foo'
%timeit df[mask]
mask = df['A'] == 'foo'
%timeit df[mask]

219 µs ± 12.3 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
239 µs ± 7.03 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)

The performance gains aren't as pronounced. We'll see if this holds up over more robust testing.


mask alternative 2 We could have reconstructed the data frame as well. There is a big caveat when reconstructing a dataframe—you must take care of the dtypes when doing so!

Instead of df[mask] we will do this

pd.DataFrame(df.values[mask], df.index[mask], df.columns).astype(df.dtypes)

If the data frame is of mixed type, which our example is, then when we get df.values the resulting array is of dtype object and consequently, all columns of the new data frame will be of dtype object. Thus requiring the astype(df.dtypes) and killing any potential performance gains.

%timeit df[m]
%timeit pd.DataFrame(df.values[mask], df.index[mask], df.columns).astype(df.dtypes)

216 µs ± 10.4 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
1.43 ms ± 39.6 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)

However, if the data frame is not of mixed type, this is a very useful way to do it.

Given

np.random.seed([3,1415])
d1 = pd.DataFrame(np.random.randint(10, size=(10, 5)), columns=list('ABCDE'))

d1

   A  B  C  D  E
0  0  2  7  3  8
1  7  0  6  8  6
2  0  2  0  4  9
3  7  3  2  4  3
4  3  6  7  7  4
5  5  3  7  5  9
6  8  7  6  4  7
7  6  2  6  6  5
8  2  8  7  5  8
9  4  7  6  1  5

%%timeit
mask = d1['A'].values == 7
d1[mask]

179 µs ± 8.73 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

Versus

%%timeit
mask = d1['A'].values == 7
pd.DataFrame(d1.values[mask], d1.index[mask], d1.columns)

87 µs ± 5.12 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

We cut the time in half.


mask alternative 3

@unutbu also shows us how to use pd.Series.isin to account for each element of df['A'] being in a set of values. This evaluates to the same thing if our set of values is a set of one value, namely 'foo'. But it also generalizes to include larger sets of values if needed. Turns out, this is still pretty fast even though it is a more general solution. The only real loss is in intuitiveness for those not familiar with the concept.

mask = df['A'].isin(['foo'])
df[mask]

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

However, as before, we can utilize NumPy to improve performance while sacrificing virtually nothing. We'll use np.in1d

mask = np.in1d(df['A'].values, ['foo'])
df[mask]

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

Timing

I'll include other concepts mentioned in other posts as well for reference.

Code Below

Each column in this table represents a different length data frame over which we test each function. Each column shows relative time taken, with the fastest function given a base index of 1.0.

res.div(res.min())

                         10        30        100       300       1000      3000      10000     30000
mask_standard         2.156872  1.850663  2.034149  2.166312  2.164541  3.090372  2.981326  3.131151
mask_standard_loc     1.879035  1.782366  1.988823  2.338112  2.361391  3.036131  2.998112  2.990103
mask_with_values      1.010166  1.000000  1.005113  1.026363  1.028698  1.293741  1.007824  1.016919
mask_with_values_loc  1.196843  1.300228  1.000000  1.000000  1.038989  1.219233  1.037020  1.000000
query                 4.997304  4.765554  5.934096  4.500559  2.997924  2.397013  1.680447  1.398190
xs_label              4.124597  4.272363  5.596152  4.295331  4.676591  5.710680  6.032809  8.950255
mask_with_isin        1.674055  1.679935  1.847972  1.724183  1.345111  1.405231  1.253554  1.264760
mask_with_in1d        1.000000  1.083807  
Discussions

python - Select columns in pandas dataframe by value in rows - Stack Overflow
I have a pandas.DataFrame with too much columns. I want to select all columns with values in rows equals to 0 and 1. Type of all columns is int64 and I can't select they by object or other type. Ho... More on stackoverflow.com
🌐 stackoverflow.com
June 7, 2017
Filtering a pandas float column by “less than”
Can also do df = df.query(“column_name < 100.0”) IMO this is never a bad option since it’s extremely concise and clear. Anyone familiar with SQL, excel, etc will immediately understand what they’re looking at. More on reddit.com
🌐 r/learnpython
5
1
March 26, 2020
how to update a pandas dataframe column value, when a specific string appears in another column?
It's not something you'd really use .apply for. You would use boolean indexing, e.g. df['A'].str.contains('foo') would give you a Series of True/False values. You can then use .loc to set column(s) to a particular value for the True rows: df.loc[df['A'].str.contains('foo'), 'B'] = 'bar' More on reddit.com
🌐 r/learnpython
7
3
July 24, 2024
Filter pandas columns with count of non-null value less than 7
df.dropna(axis=1, thresh=7) try it out and see In Pandas there is always another method to do the same thing, just some ways are simpler than others More on reddit.com
🌐 r/learnpython
5
3
April 14, 2021
🌐
Spark By {Examples}
sparkbyexamples.com › home › pandas › pandas select rows based on column values
Pandas Select Rows Based on Column Values - Spark By {Examples}
June 12, 2025 - In pandas, you can select rows based on column values using boolean indexing or using methods like DataFrame.loc[] attribute, DataFrame.query(), or
🌐
Pandas
pandas.pydata.org › docs › getting_started › intro_tutorials › 03_subset_data.html
How do I select a subset of a DataFrame? — pandas 3.0.6 documentation
Use loc for label-based selection (using row/column names). Use iloc for position-based selection (using table positions). You can assign new values to a selection based on loc/iloc.
🌐
Saturn Cloud
saturncloud.io › blog › pandas-tips-select-rows-by-column-value
How to select rows by column value in Pandas | Saturn Cloud Blog
September 10, 2023 - import pandas as pd data = pd.DataFrame({'Color': 'Tabby Black Calico Tabby Tabby Black'.split(), 'Name': 'Maxine Angel Delilah Tom Jeff Fluffy'.split(), 'Age': [2, 5, 17, 10, 7, 2]}) #select by scalar value data.loc[data['Color'] == 'Tabby'] #select by iterable value data.loc[data['Age'].isin([2, 5])] Boolean indexing also allows for selection by negation, or by multiple conditions (with &, |): #select rows where column value does NOT equal some scalar value data.loc[data['Color'] != 'Calico'] #select rows where column value is NOT in some iterable value data.loc[~data['Name'].isin(['Tom', 'Fluffy'])] #select rows by multiple conditions data.loc[(data['Color'] == 'Tabby') & (data['Age'] <= 7)]
🌐
Pandas
pandas.pydata.org › docs › user_guide › indexing.html
Indexing and selecting data — pandas 3.0.6 documentation
A boolean array (any NA values will be treated as False). A callable function with one argument (the calling Series or DataFrame) and that returns valid output for indexing (one of the above). A tuple of row (and column) indices whose elements are one of the above inputs. See more at Selection by Position, Advanced Indexing and Advanced Hierarchical.
🌐
InterviewQs
interviewqs.com › ddi-code-snippets › rows-cols-python
Select rows from a Pandas DataFrame based on values in a column - InterviewQs
#To select rows whose column value is in an iterable array, which we'll define as array, you can use isin:array = ['yellow', 'green']df.loc[df['favorite_color'].isin(array)]
Find elsewhere
🌐
Medium
medium.com › @akaivdo › pandas-select-rows-from-a-dataframe-based-on-column-values-29aef08388ec
Pandas >> Select Rows From a DataFrame Based on Column Values | by NextGenTechDawn | Medium
May 6, 2023 - import pandas as pd # Create a sample DataFrame df = pd.DataFrame({ 'Name': ['Alice', 'Bob', 'Charlie', 'Dave', 'Eva'], 'Age': [25, 30, 35, 40, 45], 'Gender': ['F', 'M', 'M', 'M', 'F'] }) # Select rows where Age is greater than or equal to 35 result = df[df['Age'] >= 35] # Display the result print(result) In this example, we first create a sample DataFrame with columns ‘Name’, ‘Age’, and ‘Gender’. We then use boolean indexing to select rows where the…
🌐
Towards Data Science
towardsdatascience.com › home › latest › how to select rows from pandas dataframe based on column values
How To Select Rows From Pandas DataFrame Based on Column Values | Towards Data Science
January 20, 2025 - Now let's assume we want to select only those rows whose values in a specific column are included in a list. To do so, we simply need to use isin() as shown below. ... Now if you want to select rows whose value is not in an iterable (e.g. a ...
🌐
GeeksforGeeks
geeksforgeeks.org › pandas › how-to-select-rows-from-a-dataframe-based-on-column-values
How to Select Rows from a Dataframe based on Column Values ? - GeeksforGeeks
July 15, 2025 - This simple operation showcases power of pandas in filtering data efficiently. The loc method is significant because it allows you to select rows based on labels and conditions. It is particularly useful when you need to filter data using specific criteria, such as selecting rows where a column value meets a certain condition.
🌐
Sentry
sentry.io › sentry answers › python › select rows from a python pandas dataframe based on column values
Select rows from a Python Pandas DataFrame based on column values | Sentry
February 15, 2023 - In other words, what is the DataFrame equivalent of a SELECT WHERE statement in SQL? This can be achieved using the DataFrame’s loc property. ... lower_limit = 1 upper_limit = 3 my_dataframe.loc[(my_dataframe["column_name"] >= lower_limit) & (my_dataframe["column_name"] <= upper_limit)] If you’re looking to get a deeper understanding of how Python application monitoring works, take a look at the following articles: ... Tasty treats for web developers brought to you by Sentry.
🌐
PythonHow
pythonhow.com › how › select-rows-from-a-dataframe-based-on-column-values-with-pandas
Here is how to select rows from a DataFrame based on column values with Pandas in Python
For example, you can use it to select rows based on multiple column values, combine multiple boolean expressions, or even use custom functions to evaluate each row. Overall, the DataFrame.loc method is the recommended way to select rows from a Pandas DataFrame based on column values.
🌐
Towards Data Science
towardsdatascience.com › home › latest › interesting ways to select pandas dataframe columns
Interesting Ways to Select Pandas DataFrame Columns | Towards Data Science
April 16, 2021 - Data types include 'float64' and ... the same data type, you'll get a series of True/False. Use the values method to get just the True/False values and not the index....
🌐
Statology
statology.org › home › pandas: how to select columns based on condition
Pandas: How to Select Columns Based on Condition
November 4, 2022 - You can use the following methods to select columns in a pandas DataFrame by condition: Method 1: Select Columns Where At Least One Row Meets Condition · #select columns where at least one row has a value greater than 2 df.loc[:, (df > 2).any()]
🌐
KDnuggets
kdnuggets.com › 2019 › 06 › select-rows-columns-pandas.html
How to Select Rows and Columns in Pandas Using [ ], .loc, iloc, .at and .iat - KDnuggets
In this example, there are 11 columns that are float and one column that is an integer. To select only the float columns, use wine_df.select_dtypes(include = ['float']). The select_dtypes method takes in a list of datatypes in its include parameter. The list values can be a string or a Python object.
🌐
Spark By {Examples}
sparkbyexamples.com › home › pandas › pandas select columns by name or index
Pandas Select Columns by Name or Index - Spark By {Examples}
June 4, 2025 - In Pandas, selecting columns by name or index allows you to access specific columns in a DataFrame based on their labels (names) or positions (indices).
🌐
Stack Overflow
stackoverflow.com › questions › 32350114 › select-columns-in-pandas-dataframe-by-value-in-rows
python - Select columns in pandas dataframe by value in rows - Stack Overflow
June 7, 2017 - I have a pandas.DataFrame with too much columns. I want to select all columns with values in rows equals to 0 and 1. Type of all columns is int64 and I can't select they by object or other type. Ho...
🌐
Shane Lynn
shanelynn.ie › home › pandas iloc and loc – quickly select rows and columns in dataframes
Pandas iloc and loc – quickly select data in DataFrames
October 16, 2021 - For example, the statement data[‘first_name’] == ‘Antonio’] produces a Pandas Series with a True/False value for every row in the ‘data’ DataFrame, where there are “True” values for the rows where the first_name is “Antonio”. These type of boolean arrays can be passed directly to the .loc indexer as so: Using a boolean True/False series to select rows in a pandas data frame – all rows with first name of “Antonio” are selected. As before, a second argument can be passed to .loc to select particular columns out of the data frame. Again, columns are referred to by name for the loc indexer and can be a single string, a list of columns, or a slice “:” operation.
🌐
Note.nkmk.me
note.nkmk.me › home › python › pandas
pandas: Select rows/columns by index (numbers and names) | note.nkmk.me
August 8, 2023 - Using a Boolean Series, you can select rows by conditions. Refer to the following article for details. ... Consider the following Series as an example. s = df['col_0'] print(s) # row_0 00 # row_1 10 # row_2 20 # row_3 30 # row_4 40 # Name: col_0, dtype: object ... You can get the value of the element by specifying the numbers (positions) or names (labels).
🌐
Apps Developer Blog
appsdeveloperblog.com › home › python › pandas: how to select rows based on column values
Pandas: How to Select Rows Based on Column Values - Apps Developer Blog
February 27, 2023 - This will create a new DataFrame selected_rows that contains only the rows where the name column is either Alice or Charlie. Note that the isin method takes a list of values to match against and returns a boolean Series that can be used to select the desired rows from the DataFrame using the indexing operator []. Another way to filter DataFrame in Pandas is to use the where method.