You can do this using interpolate:
df['Price'].interpolate(method='linear', inplace=True)
Result:
Price Date
0 NaN 1
1 NaN 2
2 1800.000000 3
3 1900.000000 4
4 1933.333333 5
5 1966.666667 6
6 2000.000000 7
7 2200.000000 8
As you can see, this only fills the missing values in a forward direction. If you want to fill the first two values as well, use the parameter limit_direction="both":
df['Price'].interpolate(method='linear', inplace=True, limit_direction="both")
There are different interpolation methods, e.g. quadratic or spline, for more info see the docs: https://pandas.pydata.org/pandas-docs/stable/generated/pandas.Series.interpolate.html
Answer from rje on Stack OverflowYou can do this using interpolate:
df['Price'].interpolate(method='linear', inplace=True)
Result:
Price Date
0 NaN 1
1 NaN 2
2 1800.000000 3
3 1900.000000 4
4 1933.333333 5
5 1966.666667 6
6 2000.000000 7
7 2200.000000 8
As you can see, this only fills the missing values in a forward direction. If you want to fill the first two values as well, use the parameter limit_direction="both":
df['Price'].interpolate(method='linear', inplace=True, limit_direction="both")
There are different interpolation methods, e.g. quadratic or spline, for more info see the docs: https://pandas.pydata.org/pandas-docs/stable/generated/pandas.Series.interpolate.html
Lets build a MCVE for Linear Regression Imputation. First we load libraries and your dataset:
import io
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression, HuberRegressor
data = pd.read_csv(io.StringIO("""Day,Price
1,NaN
2,NaN
3,1800
4,1900
5,NaN
6,NaN
7,2000
8,2200"""))
We query missing values:
q = data["Price"].isnull()
We build a Linear Regression model and fit it with available data:
model = LinearRegression(fit_intercept=True)
model.fit(data.loc[~q, ["Day"]], data.loc[~q, "Price"])
Then we can predict missing values (imputation):
data.loc[q, "Price"] = model.predict(data.loc[q, ["Day"]])
Result looks like:

If we want the outliers having less impact on the regression, we can choose a robust regressor such as Huber Regressor. Simply change the model for:
model = HuberRegressor()
It renders as follow:

You can use apply and lambda for this:
missing_data_df['horsepower']= missing_data_df.apply(
lambda row:
0.25743277 * row.displacement + 0.00958711 * row.weight + 25.874947903262651
if np.isnan(row.horsepower) else row.horsepower, axis=1)
Several things
- missing_data_df.horsepower has no missing values
- missing_data_df.weight, a variable in your formula, does have missing values
- if hp = 0.25743277 * disp + 0.00958711 * weight + 25.874947903262651
then weight = (0.25743277 * disp + 25.874947903262651 - hp) / -0.00958711
To calculate weight try
for idx in missing_data_df.index:
if pd.isnull(missing_data_df.loc[idx,"weight"]):
disp = missing_data_df.loc[idx,"displacement"]
hp = missing_data_df.loc[idx,"horsepower"]
missing_data_df.loc[idx,"weight"] = (0.25743277 * disp + 25.874947903262651 - hp) / -0.00958711
In general, .loc[] and .iloc[] are a better way to go when finding or setting values
