You can do this using interpolate:

df['Price'].interpolate(method='linear', inplace=True)

Result:

    Price   Date
0   NaN     1
1   NaN     2
2   1800.000000     3
3   1900.000000     4
4   1933.333333     5
5   1966.666667     6
6   2000.000000     7
7   2200.000000     8

As you can see, this only fills the missing values in a forward direction. If you want to fill the first two values as well, use the parameter limit_direction="both":

df['Price'].interpolate(method='linear', inplace=True, limit_direction="both")

There are different interpolation methods, e.g. quadratic or spline, for more info see the docs: https://pandas.pydata.org/pandas-docs/stable/generated/pandas.Series.interpolate.html

Answer from rje on Stack Overflow
🌐
Kaggle
kaggle.com › code › shashankasubrahmanya › missing-data-imputation-using-regression
Missing Data Imputation using Regression
June 7, 2018 - Python · Missing Data Imputation using RegressionDetermining missing valuesUsing Regression to impute missing dataReferences · This Notebook has been released under the Apache 2.0 open source license. Input1 file · arrow_right_alt · Output0 files · arrow_right_alt ·
🌐
HH.IO
henrikhain.io › post › stochastic-regression-imputation
Stochastic regression imputation | HH.IO
June 3, 2021 - This blog post introduced stochastic regression imputation in comparison to deterministic regression imputation together with some code examples in Python.
Top answer
1 of 2
8

You can do this using interpolate:

df['Price'].interpolate(method='linear', inplace=True)

Result:

    Price   Date
0   NaN     1
1   NaN     2
2   1800.000000     3
3   1900.000000     4
4   1933.333333     5
5   1966.666667     6
6   2000.000000     7
7   2200.000000     8

As you can see, this only fills the missing values in a forward direction. If you want to fill the first two values as well, use the parameter limit_direction="both":

df['Price'].interpolate(method='linear', inplace=True, limit_direction="both")

There are different interpolation methods, e.g. quadratic or spline, for more info see the docs: https://pandas.pydata.org/pandas-docs/stable/generated/pandas.Series.interpolate.html

2 of 2
0

Lets build a MCVE for Linear Regression Imputation. First we load libraries and your dataset:

import io
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression, HuberRegressor

data = pd.read_csv(io.StringIO("""Day,Price
1,NaN
2,NaN
3,1800
4,1900
5,NaN
6,NaN
7,2000
8,2200"""))

We query missing values:

q = data["Price"].isnull()

We build a Linear Regression model and fit it with available data:

model = LinearRegression(fit_intercept=True)
model.fit(data.loc[~q, ["Day"]], data.loc[~q, "Price"])

Then we can predict missing values (imputation):

data.loc[q, "Price"] = model.predict(data.loc[q, ["Day"]])

Result looks like:

If we want the outliers having less impact on the regression, we can choose a robust regressor such as Huber Regressor. Simply change the model for:

model = HuberRegressor()

It renders as follow:

🌐
scikit-learn
scikit-learn.org › stable › modules › impute.html
8.4. Imputation of missing values — scikit-learn 1.9.1 documentation
Then, the regressor is used to predict the missing values of y. This is done for each feature in an iterative fashion, and then is repeated for max_iter imputation rounds.
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to handle missing data in python? [explained in 5 easy steps]
How to Handle Missing Data in Python? [Explained in 5 Easy Steps]
May 1, 2025 - See that this model produces more accuracy than the previous model as we are using a specific regression model for filling in the missing values. We can also use models KNN for filling in the missing values. But sometimes, using models for imputation can result in overfitting the data.
🌐
PyPI
pypi.org › project › autoimpute
Client Challenge
JavaScript is disabled in your browser · Please enable JavaScript to proceed · A required part of this site couldn’t load. This may be due to a browser extension, network issues, or browser settings. Please check your connection, disable any ad blockers, or try using a different browser
Find elsewhere
🌐
Bookdown
bookdown.org › mwheymans › bookmi › single-missing-data-imputation.html
Chapter3 Single Missing data imputation | Book_MI.knit
You can extract the mean imputed dataset by using the complete function as follows: complete(imp_mean) With regression imputation the information of other variables is used to predict the missing values in a variable by using a regression model.
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › iterative imputation for missing values in machine learning
Iterative Imputation for Missing Values in Machine Learning - MachineLearningMastery.com
August 18, 2020 - All input variables in machine learning are float, even integers – at least in Python/sklearn this is the case. Converting data types into your preferred type after data preparation is a good idea if you’re not modeling. ... Thank you for the amazing information. According to the sklearn.impute.IterativeImputer user guide, the estimator parameter can be a different kind of regression algorithm such as BayesianRidge ,DecisionTreeRegressor, and so on.
🌐
Medium
medium.com › @Cambridge_Spark › tutorial-introduction-to-missing-data-imputation-4912b51c34eb
Tutorial: Introduction to Missing Data Imputation | by Cambridge Spark | Medium
September 3, 2019 - Thus, we can use a simple linear model regressing total_bill on tip to fill the missing values in total_bill. ... As we can see, the imputed total_bill from a simple linear model from tips does not exactly recover the truth but capture the general trend (and is better a single value imputation such as mean imputation).
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › data-imputation-techniques-in-ml
Data Imputation Techniques in ML - GeeksforGeeks
October 30, 2025 - Generate 𝑚 imputed datasets. Analyze each dataset. ... Predict missing values based on other variables using regression.
🌐
GitHub
github.com › kearnz › autoimpute
GitHub - kearnz/autoimpute: Python package for Imputation Methods
Autoimpute also extends supervised machine learning methods from scikit-learn and statsmodels to apply them to multiply imputed datasets (using the MiceImputer under the hood). Right now, Autoimpute supports linear regression and binary logistic regression.
Starred by 251 users
Forked by 18 users
Languages: Python 100.0% | Python 100.0%
🌐
DataCamp
campus.datacamp.com › courses › dealing-with-missing-data-in-python › advanced-imputation-techniques
Evaluation of different imputation techniques | Python
Another way to analyze the performance of different imputations is to observe their density plots and see which one most resembles the shape of the original data. To perform linear regression, we will use the statsmodels package as it produces various statistical summaries.
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › impute › plot_iterative_imputer_variants_comparison.html
Imputing missing values with variants of IterativeImputer — scikit-learn 1.9.1 documentation
make_pipeline (Nystroem, Ridge): a pipeline with the expansion of a degree 2 polynomial kernel and regularized linear regression · KNeighborsRegressor: comparable to other KNN imputation approaches
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › impute › plot_missing_values.html
Imputing missing values before building an estimator — scikit-learn 1.9.0 documentation
We use the class’s default choice of the regressor model (BayesianRidge) to predict missing feature values. The performance of the predictor may be negatively affected by vastly different scales of the features, so we re-scale the features in the California housing dataset. imputer = IterativeImputer(add_indicator=True) mses_diabetes[4], stds_diabetes[4] = get_score( X_miss_diabetes, y_miss_diabetes, imputer ) mses_california[4], stds_california[4] = get_score( X_miss_california, y_miss_california, make_pipeline(RobustScaler(), imputer) ) x_labels.append("Iterative Imputation") mses_diabetes = mses_diabetes * -1 mses_california = mses_california * -1
🌐
GeeksforGeeks
geeksforgeeks.org › imputing-missing-values-before-building-an-estimator-in-scikit-learn
Imputing Missing Values Before Building an Estimator in Scikit Learn - GeeksforGeeks
April 28, 2025 - In the above example, we first loaded a dataset which containing missing values. We then identified missing values in the following dataset using the NumPy library. We then used Scikit Learn's SimpleImputer class to impute missing values in the dataset. Finally, we built a linear regression estimator using the imputed dataset.
🌐
Datasciencestunt
datasciencestunt.com › regression-imputation
Regression Imputation: A Technique for Dealing with Missing ...
April 30, 2023 - Aviso legal: las referencias a una empresa, producto o servicios específicos en este Sitio no están controladas por GoDaddy.com LLC y no constituyen ni implican la asociación ni respaldo a anunciantes externos