You can apply an arbitrary function across a dataframe row using DataFrame.apply.
In your case, you could define a function like:
def conditions(s):
if (s['discount'] > 20) or (s['tax'] == 0) or (s['total'] > 100):
return 1
else:
return 0
And use it to add a new column to your data:
df_full['Class'] = df_full.apply(conditions, axis=1)
Answer from Gustavo Bezerra on Stack OverflowAdd new column to Python Pandas DataFrame based on multiple conditions - Stack Overflow
pandas - Create new columns based on multiple conditions in Python - Stack Overflow
Adding a new column based on conditions (PANDAS)
dataset - Creating new column in dataframe based on conditions in 2 other columns - Data Science Stack Exchange
You can apply an arbitrary function across a dataframe row using DataFrame.apply.
In your case, you could define a function like:
def conditions(s):
if (s['discount'] > 20) or (s['tax'] == 0) or (s['total'] > 100):
return 1
else:
return 0
And use it to add a new column to your data:
df_full['Class'] = df_full.apply(conditions, axis=1)
Judging by the image of your data is rather unclear what you mean by a discount 20%.
However, you can likely do something like this.
df['class'] = 0 # add a class column with 0 as default value
# find all rows that fulfills your conditions and set class to 1
df.loc[(df['discount'] / df['total'] > .2) & # if discount is more than .2 of total
(df['tax'] == 0) & # if tax is 0
(df['total'] > 100), # if total is > 100
'class'] = 1 # then set class to 1
Note that & means and here, if you want or instead use |.
import numpy as np
import pandas as pd
data = [(27450, 27450, 29420,"10/10/2016"),
(29420 , 36142, 29420, "10/10/2016"),
(11 , 11, 27450, "10/10/2016")]
df = pd.DataFrame(data, columns=("User_id","Actor1","Actor2", "Time"))
mask = (df['User_id'] == df['Actor1'])
df['first actor'] = mask.astype(int)
df['other actor'] = np.where(mask, df['Actor2'], df['Actor1'])
print(df)
yields
User_id Actor1 Actor2 Time first actor other actor
0 27450 27450 29420 10/10/2016 1 29420
1 29420 36142 29420 10/10/2016 0 36142
2 11 11 27450 10/10/2016 1 27450
First create a boolean mask which is True when User_id equals Actor1:
In [51]: mask = (df['User_id'] == df['Actor1']); mask
Out[51]:
0 True
1 False
2 True
dtype: bool
Converting mask to ints creates the first column:
In [52]: mask.astype(int)
Out[52]:
0 1
1 0
2 1
dtype: int64
Then use np.where to select between two values. np.where(mask, A, B) returns an array whose ith value is A[i] if mask[i] is True, and B[i] otherwise. Thus,
np.where(mask, df['Actor2'], df['Actor1']) takes the value from Actor2 where mask is True, and the value from Actor1 otherwise:
In [53]: np.where(mask, df['Actor2'], df['Actor1'])
Out[53]: array([29420, 36142, 27450])
Heres my solution - I have assumed that if userid appears in actor1 column its not necessary it'll be in the same row...
df["Col1"] = [1 if i in df["Actor1"].values else 0 for i in df["User_id"].values]
df["Col2"] = [df.iloc[i]["Actor2"] if j == 1 else df.iloc[i]["Actor1"] for i, j in enumerate(df["Col1"].values)]
Output -
User_id Actor1 Actor2 Time Col1 Col2
0 27450 27450 29420 10/10/2016 1 29420
1 29420 36142 29420 10/10/2016 0 36142
2 11 11 27450 10/10/2016 1 27450
I have a column named destination that has over 50 countries. I am trying to add a new column called continents that has each country categorized based on the location. I have used the loc method, but I was only able to add a county one by one.
ex.loc[ex['Destination'] =='MALAYSIA','Continent']='Asia'
Any ideas on this? Thanks