If you have to use the "apply" variant, the code should be:
df['product_AH'] = df.apply(lambda row: row.Age * row.Height, axis=1)
The parameter to the function applied is the whole row.
But much quicker solution is:
df['product_AH'] = df.Age * df.Height
(1.43 ms, compared to 5.08 ms for the "apply" variant).
This way computation is performed using vectorization, whereas apply refers to each row separately, applies the function to it, then assembles all results and saves them in the target column, which is considerably slower.
Answer from Valdi_Bo on Stack Overflowpython - Using lambda functions with apply for Pandas DataFrame - Stack Overflow
How to properly apply a lambda function into a pandas data frame column - Stack Overflow
Lambda function in pandas
I'm slightly addicted to lambda functions on Pandas. Is it bad practice?
You need mask:
sample['PR'] = sample['PR'].mask(sample['PR'] < 90, np.nan)
Another solution with loc and boolean indexing:
sample.loc[sample['PR'] < 90, 'PR'] = np.nan
Sample:
import pandas as pd
import numpy as np
sample = pd.DataFrame({'PR':[10,100,40] })
print (sample)
PR
0 10
1 100
2 40
sample['PR'] = sample['PR'].mask(sample['PR'] < 90, np.nan)
print (sample)
PR
0 NaN
1 100.0
2 NaN
sample.loc[sample['PR'] < 90, 'PR'] = np.nan
print (sample)
PR
0 NaN
1 100.0
2 NaN
EDIT:
Solution with apply:
sample['PR'] = sample['PR'].apply(lambda x: np.nan if x < 90 else x)
Timings len(df)=300k:
sample = pd.concat([sample]*100000).reset_index(drop=True)
In [853]: %timeit sample['PR'].apply(lambda x: np.nan if x < 90 else x)
10 loops, best of 3: 102 ms per loop
In [854]: %timeit sample['PR'].mask(sample['PR'] < 90, np.nan)
The slowest run took 4.28 times longer than the fastest. This could mean that an intermediate result is being cached.
100 loops, best of 3: 3.71 ms per loop
You need to add else in your lambda function. Because you are telling what to do in case your condition(here x < 90) is met, but you are not telling what to do in case the condition is not met.
sample['PR'] = sample['PR'].apply(lambda x: 'NaN' if x < 90 else x)
Hi! So I'm trying to do a course on kaggle and I got stuck in using the lambda function - I'm not even sure to start in the first place.
For starters, kaggle introduces the lambda function with this line of code here, which makes me think that it's just like a normal function in math:
reviews.points.map(lambda p: p - review_points_mean)
However, as I looked through the tutorial, I didn't understand what was going on with the lambda function as it's being used to source out words from the data set as well. For example:
n_trop = reviews.description.map(lambda desc: "tropical" in desc).sum()
This was used as a way to sum up the number of times that the word "tropical" was used in the dataset. this also confused me with how you would use the lambda function itself as mentioned above that i thought it would be used just like a normal function in math.
My question is - how exactly is the lambda function from pandas being used in different ways? Because I don't understand what is going on in here. I've tried googling for answers online, but the pandas documentation does not answer my question. Thank you!