Given a sample dataframe df as:
a b
0 1 2
1 2 3
2 3 4
3 4 5
what you want is:
df['a'] = df['a'].apply(lambda x: x + 1)
that returns:
a b
0 2 2
1 3 3
2 4 4
3 5 5
Answer from Fabio Lamanna on Stack OverflowGiven a sample dataframe df as:
a b
0 1 2
1 2 3
2 3 4
3 4 5
what you want is:
df['a'] = df['a'].apply(lambda x: x + 1)
that returns:
a b
0 2 2
1 3 3
2 4 4
3 5 5
For a single column better to use map(), like this:
df = pd.DataFrame([{'a': 15, 'b': 15, 'c': 5}, {'a': 20, 'b': 10, 'c': 7}, {'a': 25, 'b': 30, 'c': 9}])
a b c
0 15 15 5
1 20 10 7
2 25 30 9
df['a'] = df['a'].map(lambda a: a / 2.)
a b c
0 7.5 15 5
1 10.0 10 7
2 12.5 30 9
How to properly apply a lambda function into a pandas data frame column - Stack Overflow
How to apply a lambda on pandas DataFrame column headers
python - How to apply lambda function to specific column based on the values in the adjacent column - Stack Overflow
Optimal way to apply functions for different columns in pandas
You need mask:
sample['PR'] = sample['PR'].mask(sample['PR'] < 90, np.nan)
Another solution with loc and boolean indexing:
sample.loc[sample['PR'] < 90, 'PR'] = np.nan
Sample:
import pandas as pd
import numpy as np
sample = pd.DataFrame({'PR':[10,100,40] })
print (sample)
PR
0 10
1 100
2 40
sample['PR'] = sample['PR'].mask(sample['PR'] < 90, np.nan)
print (sample)
PR
0 NaN
1 100.0
2 NaN
sample.loc[sample['PR'] < 90, 'PR'] = np.nan
print (sample)
PR
0 NaN
1 100.0
2 NaN
EDIT:
Solution with apply:
sample['PR'] = sample['PR'].apply(lambda x: np.nan if x < 90 else x)
Timings len(df)=300k:
sample = pd.concat([sample]*100000).reset_index(drop=True)
In [853]: %timeit sample['PR'].apply(lambda x: np.nan if x < 90 else x)
10 loops, best of 3: 102 ms per loop
In [854]: %timeit sample['PR'].mask(sample['PR'] < 90, np.nan)
The slowest run took 4.28 times longer than the fastest. This could mean that an intermediate result is being cached.
100 loops, best of 3: 3.71 ms per loop
You need to add else in your lambda function. Because you are telling what to do in case your condition(here x < 90) is met, but you are not telling what to do in case the condition is not met.
sample['PR'] = sample['PR'].apply(lambda x: 'NaN' if x < 90 else x)
Hi experts,
Basically what I want to do is to allow the users to get whatever columns they want with a lambda function. I tried a lot of methods but none of them worked.
e.g. Let's say I have the following code:
dfa = pd.DataFrame({"BZ_Impressions": [1000, 2000, 3000],
"BZ_Clicks": [40, 50, 60],
"MF_Impressions": [500, 600, 1200],
"MZ_Clicks": [50, 30, 120],
"BZ_PX_Joins": [10,30,25],
"MF_PX_Joins": [5, 12, 25]})I could, for example, only want columns that start with "BZ", and I could also maybe only want columns that has "Impressions" in it. I could also want columns that have a length smaller than 10 characters. The possibilities are too many so a lambda function is the only solution. But I'm willing to focus only on string functions on headers only (i.e. user cannot select columns based on actual data, only on headers)
I have tried to convert the column headers into a list and then apply a lambda function on it, but it says "list has no attribution 'apply'". I also tried numpy arrays but it didn't work. I'm wondering if I can directly apply something on DataFrame.columns so the users may get a sub-DataFrame easily?
Thank you!