You can do the following, let's say df is your dataset:
from sklearn.preprocessing import Imputer
imputer = Imputer(missing_values ='NaN', strategy = 'mean', axis = 0)
df[['Age','Salary']]=imputer.fit_transform(df[['Age','Salary']])
print(df)
Country Age Salary Purchased
0 France 44.000000 72000.000000 No
1 Spain 27.000000 48000.000000 Yes
2 Germany 30.000000 54000.000000 No
3 Spain 38.000000 61000.000000 No
4 Germany 40.000000 63777.777778 Yes
5 France 35.000000 58000.000000 Yes
6 Spain 38.777778 52000.000000 No
7 France 48.000000 79000.000000 Yes
8 Germany 50.000000 83000.000000 No
9 France 37.000000 67000.000000 Yes
Answer from YOLO on Stack OverflowAnalytics Vidhya
analyticsvidhya.com › home › how to handle missing data in python? [explained in 5 easy steps]
How to Handle Missing Data in Python? [Explained in 5 Easy Steps]
May 1, 2025 - This article taught us about the different ways of handling missing values in our dataset. If there are way too many missing values in a column then you can drop that column. Otherwise we can impute missing values with mean, median and mode.
Medium
medium.com › @hassankhan2608 › missing-value-imputation-methods-using-python-f1b8796901ba
Missing Value Imputation Methods using Python | by Mohd Hassan Khan | Medium
January 15, 2024 - K-Nearest Neighbors (KNN) Imputation: Identifies ‘k’ samples in the dataset that are similar to the observation with missing data and imputes values based on the average (or majority) of these ‘k’ neighbours. For predictive imputation, let’s use k-nearest neighbours (KNN). We’ll use the KNNImputer from the scikit-learn library using Python:
python - How Do I impute missing values using pandas? - Stack Overflow
I am trying to impute missing values as the mean of other values in the column; however, my code is having no effect. Does anyone know what I may be doing wrong? Thanks! My code: from sklearn. More on stackoverflow.com
pandas - how to impute missing value in python using some condition - Stack Overflow
so this code is replacing the value of date with last update if 'yyy' is present in visited column. but the task is if we have date missing for rajesh now we will check does rajesh has any entry in visited column with name yyy if yes then his missing date will get replace with latest update 2022-04-25T15:45:14.47Z+00:00 ... Find the answer to your question by asking. Ask question ... See similar questions with these tags. ... 2 I need to impute ... More on stackoverflow.com
How to deal with missing values?
Generally you have 3 options: remove the row entirely, fill the missing value, or some machine learning algorithms can work even when you have missing entries. When you have more than enough data, just remove the rows with empties. Like if I had 2 million rows of data with a random 10% of them empty, I generally wouldn't break a sweat over removing them; it's still a ton of data. 900 is still a pretty good amount of data, so you might consider this. There are a few ways to fill the empties, depending on the data you've got. You could use an average of some kind: mean, median, or mode. The right approach depends on the distribution of the data, so you should plot the data to take a look. For instance, if the attribute has a lot of skewness or otherwise isn't normally distributed, a mean could be the wrong choice. Generally, the idea behind filling the empties with an average is to replace them with benign data close to/at the center of the distribution, so ML algorithms won't be pulled one way or the other. I personally use this approach when I don't have much data, both in terms of total rows as well as columns - too few rows to simply remove them, too few columns to use the next approach, which is imputation. You could impute the missing data, like with k-Nearest Neighbors imputation, through which you give the data to KNN to predict what the missing values ought to be. This approach can provide higher accuracy in certain cases, but I'd only advise it if you have enough columns in your data such that you'd be able to arrive at accurate predictions. For instance, if you have a weather dataset with many columns of pertinent info to predicting temperatures like latitude/longitude, dates, dew point, pressure, etc. etc., and some of your rows have missing data, the imputation would probably work well, because the other attributes in the data are sufficient to predict the empty values. Finally, some algorithms are totally cool with missing values or you can tell them to ignore them. Hope this helps EDIT: I forgot to mention...the cool thing about learning data science is that you get to try all the approaches and compare results. Give them all a shot and see what works best. Record what did and the features of the data. Then when you do another analysis with a different dataset see if a different approach works out better. Compare and contrast depending on the data you have and you'll get the hang of it. More on reddit.com
Best way of handling missing values?
The way I’ve done it thus far is impute a column based on its distribution (median if the data is skewed) or fill it with the mean. I never drop missing values. This is nowhere near an ideal approach. Depending on the exact nature of your data you'll want to use different techniques to impute missing information, but at the absolute very least you should be using kNN as baseline. Mean and median imputation consistently perform the worst out of common techniques. The only method I'm aware of that has lower performance than mean and median is simply randomly taking another datapoint and filling that in. On the other hand, kNN has been demonstrated in multiple research papers to be a top performer in improving the precision of the final result when used with missing data, and to provide a low degree of distortion. The best technique isn't always kNN, you'll see good performance with predictive mean matching, 1NN, fuzzy mean matching, bayesian PCA, etc depending on the circumstances. However, kNN works as a very decent and more importantly consistently high performing starting point for filling missing values. This hold true when applied across a wide range of datasets both in terms of content and size. One approach worth considering when evaluating which technique to use is to simply drop data on purpose, and then apply a variety of filling techniques as mentioned above. You can then compare the distortion, precision, and errors to the known value and see which method gives the best results on your particular data. This is what I recommend personally. More on reddit.com
11:03
A Better Approach to Categorical Data Imputation in Python - YouTube
14:32
Handling Missing Data in Python: Simple Imputer in Python for Machine ...
08:41
How to impute missing data in categorical features (using MICE) ...
17:23
Missing value imputation application In Python | Python missing ...
09:16
Python Missing Data Filling Techniques - Simple Methods To Handle ...
05:50
Impute missing values using KNNImputer or IterativeImputer - YouTube
scikit-learn
scikit-learn.org › stable › modules › impute.html
8.4. Imputation of missing values — scikit-learn 1.9.1 documentation
Missing values can be imputed with a provided constant value, or using the statistics (mean, median or most frequent) of each column in which the missing values are located. This class also allows for different missing values encodings.
MachineLearningMastery
machinelearningmastery.com › home › blog › how to handle missing data with python
How to Handle Missing Data with Python - MachineLearningMastery.com
November 27, 2023 - In this tutorial, you will learn how to handle missing data for machine learning with Python. Specifically, after completing this tutorial you will know: How to mark invalid or corrupt values as missing in your dataset. How to remove rows with missing data from your dataset. How to impute missing values with mean values in your dataset.
Kaggle
kaggle.com › code › parulpandey › a-guide-to-handling-missing-values-in-python
A Guide to Handling Missing values in Python
July 11, 2020 - Handling Missing Values in PythonTable of ContentsObjectiveDataLoading necessary libraries and datasetsReading in the datasetExamining the Target columnDetecting Missing valuesDetecting missing values numericallyDetecting missing data visually using Missingno libraryReasons for Missing ValuesFinding reason for missing data using matrix plotFinding reason for missing data using a HeatmapFinding reason for missing data using DendrogramTreating Missing valuesDeletionsImputations Techniques for non Time Series ProblemsImputations Techniques for Time Series ProblemsAdvanced Imputation TechniquesAlgorithms which handle missing valuesConclusionReferences and good resources
ProjectPro
projectpro.io › recipes › impute-missing-values-with-means-in-python
How to Impute Missing Values with Mean in Python? -
April 12, 2023 - Mean imputation is a simple method of imputation where missing values are replaced with the mean of the available data. This method assumes that the missing values are missing at random and that the data is normally distributed. You don't have to remember all the machine learning algorithms by heart because of amazing libraries in Python.
Top answer 1 of 2
4
You can do the following, let's say df is your dataset:
from sklearn.preprocessing import Imputer
imputer = Imputer(missing_values ='NaN', strategy = 'mean', axis = 0)
df[['Age','Salary']]=imputer.fit_transform(df[['Age','Salary']])
print(df)
Country Age Salary Purchased
0 France 44.000000 72000.000000 No
1 Spain 27.000000 48000.000000 Yes
2 Germany 30.000000 54000.000000 No
3 Spain 38.000000 61000.000000 No
4 Germany 40.000000 63777.777778 Yes
5 France 35.000000 58000.000000 Yes
6 Spain 38.777778 52000.000000 No
7 France 48.000000 79000.000000 Yes
8 Germany 50.000000 83000.000000 No
9 France 37.000000 67000.000000 Yes
2 of 2
1
You're assigning an Imputer object to the variable imputer:
imputer = Imputer(missing_values ='NaN', strategy = 'mean', axis = 0)
You then call the fit() function on your Imputer object, and then the transform() function.
Then you print the dataset variable, which I'm not sure where it comes from. Did you mean to print the Imputer object, or the result of one of those calls instead?
MachineLearningMastery
machinelearningmastery.com › home › blog › dealing with missing data strategically: advanced imputation techniques in pandas and scikit-learn
Dealing with Missing Data Strategically: Advanced Imputation Techniques in Pandas and Scikit-learn - MachineLearningMastery.com
June 7, 2025 - While there exist basic strategies to deal with instances or attributes containing missing values, — like removing rows or columns entirely, or imputing missing values with a default value (typically the mean or median of the attribute) — these strategies are sometimes not sufficient. This article presents some advanced strategies to handle missing data, namely, imputation techniques made possible through a combined use of Pandas and Scikit-learn libraries in Python.
DataCamp
datacamp.com › tutorial › techniques-to-handle-missing-data-values
Top Techniques to Handle Missing Values Every Data Scientist Should Know | DataCamp
January 31, 2023 - Mean and median imputations are respectively used to replace missing values of a given column with the mean and median of the non-missing values in that column. Normal distribution is the ideal scenario. Unfortunately, it is not always the case. This is where the median imputation can be helpful because it is not sensitive to outliers. In Python, the fillna() function from pandas can be used to make these replacements.
James LeDoux's Blog
jamesrledoux.com › code › imputation
Impute Missing Values - James LeDoux’s Blog
June 1, 2019 - One approach to imputing categorical features is to replace missing values with the most common class. You can do with by taking the index of the most common feature given in Pandas’ value_counts function.
Train in Data
blog.trainindata.com › your-guide-to-missing-values-imputation
Your Guide to Missing Values Imputation | Train in Data Blog
July 22, 2024 - In the following output, we see the fraction of missing data per variable: LotFrontage 0.184932 MasVnrArea 0.004892 BsmtQual 0.023483 FireplaceQu 0.467710 GarageYrBlt 0.052838 dtype: float64 · Now, we’ll use different univariate imputation methods from the Feature-engine open source Python library to impute the null values ...
DataCamp
campus.datacamp.com › courses › exploratory-data-analysis-in-python › data-cleaning-and-imputation
Addressing missing data | Python
We then filter for the remaining columns with missing values, giving us four columns. To impute the mode for the first three columns, we loop through them and call the dot-fillna method, passing the respective column's mode and indexing the first item, which contains the mode, in square brackets.
Stack Overflow
stackoverflow.com › questions › 71988275 › how-to-impute-missing-value-in-python-using-some-condition
pandas - how to impute missing value in python using some condition - Stack Overflow
Here i have to replace missing value in date column with New date and if we have date missing for eg: Rajesh, we will check does Rajesh have any entry in visited column with name yyy if yes then his missing date will get replace with last_update.
Medium
medium.com › @vayakakshay08 › missing-value-imputation-with-pandas-4efd5f68da23
Missing Value Imputation with Pandas | by Akshay V. | Medium
January 25, 2024 - Impute the categorical values by the most relevant value of the column. import pandas as pd # Create a sample DataFrame with missing values in categorical columns data = {'Category': ['A', 'B', 'A', None, 'B', 'A'], 'Status': ['Active', 'Inactive', 'Active', None, 'Inactive', 'Active'], 'Value': [10, 20, 15, None, 25, 18]} df = pd.DataFrame(data) # Display the original DataFrame print("Original DataFrame:") print(df) # Impute missing categorical values with the most frequent value (mode) df_categorical_imputed = df.apply(lambda x: x.fillna(x.mode()[0]) if x.dtype == 'O' else x) # Display the DataFrame after imputation print("\nDataFrame after imputing categorical values:") print(df_categorical_imputed)
CodeSignal
codesignal.com › learn › courses › data-preprocessing-for-machine-learning › lessons › handling-missing-values
Handling Missing Values | CodeSignal Learn
One of the easiest ways to handle missing values in Python is by using the SimpleImputer class from the sklearn.impute module.
Medium
medium.com › @cyberdud3 › mastering-missing-data-in-python-tips-for-data-scientists-8662d93945a1
Mastering Missing Data in Python: Tips for Data Scientists | by ABHIJITH SUDHAKAR | Medium
March 14, 2023 - Scikit-learn also provides several functions for handling missing values in machine learning pipelines, including make_column_transformer() for applying different imputation strategies to different columns. Missingno: Missingno is a Python library for visualizing missing data in data science projects.