There is no need to use the SimpleImputer.
DataFrame.fillna() can do the work as well

  • For the second column, use

    column.fillna(column.mean(), inplace=True)

  • For the third column, use

    column.fillna(constant, inplace=True)

Of course, you will need to replace column with your DataFrame's column you want to change and constant with your desired constant.


Edit
Since the use of inplace is discouraged and will be deprecated, the syntax should be

column = column.fillna(column.mean())
Answer from liakoyras on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.9.1 documentation
Added in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed. ... The placeholder for the missing values. All occurrences of missing_values will be imputed. For pandas’ dataframes with nullable integer dtypes with missing values, missing_values can be set to either np.nan or pd.NA. ... The imputation strategy. If “mean”, then replace missing values using the mean along each column.
🌐
DevSkrol
devskrol.com › home › how to impute missing values using simpleimputer and columntransformer?
How to Impute missing values using SimpleImputer and ColumnTransformer? » DevSkrol
January 14, 2022 - SimpleImputer returns the result in array format instead of DataFrame. ColumnTransformer allows transformation in different columns with different imputations and applies at the same time.
🌐
Towards Data Science
towardsdatascience.com › home › latest › imputing missing values using the simpleimputer class in sklearn
Imputing Missing Values using the SimpleImputer Class in sklearn | Towards Data Science
September 19, 2021 - In my example, I assign the value back to the column B: ... To replace the missing values for multiple columns in your dataframe, you just need to pass in a dataframe containing the relevant columns:
🌐
scikit-learn
scikit-learn.org › stable › modules › impute.html
8.4. Imputation of missing values — scikit-learn 1.9.1 documentation
The SimpleImputer class provides basic strategies for imputing missing values. Missing values can be imputed with a provided constant value, or using the statistics (mean, median or most frequent) of each column in which the missing values are located.
🌐
Kaggle
kaggle.com › learn-forum › 62035
Checking your browser - reCAPTCHA
July 26, 2018 - Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
Top answer
1 of 1
17

If you want to impute different features with different arbitrary values, or the median, you need to set up several SimpleImputer steps within a pipeline and then join them with the ColumnTransformer:

import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer

# first we need to make lists, indicating which features
# will be imputed with each method

features_numeric = ['LotFrontage', 'MasVnrArea', 'GarageYrBlt']
features_categoric = ['BsmtQual', 'FireplaceQu']

# then we instantiate the imputers, within a pipeline
# we create one imputer for numerical and one imputer
# for categorical

# this imputer imputes with the mean
imputer_numeric = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='mean')),
])

# this imputer imputes with an arbitrary value
imputer_categoric = Pipeline(
    steps=[('imputer',
            SimpleImputer(strategy='constant', fill_value='Missing'))])

# then we put the features list and the transformers together
# using the column transformer

preprocessor = ColumnTransformer(transformers=[('imputer_numeric',
                                                imputer_numeric,
                                                features_numeric),
                                               ('imputer_categoric',
                                                imputer_categoric,
                                                features_categoric)])

# now we fit the preprocessor
preprocessor.fit(X_train)

# and now we can impute the data
# remember it returs a numpy array

X_train = preprocessor.transform(X_train)
X_test = preprocessor.transform(X_test)

Alternatively, you can use the package Feature-Engine which transformers allow you to specify the features:

from feature_engine import imputation as msi
from sklearn.pipeline import Pipeline as pipe

pipe = pipe([
    # add a binary variable to indicate missing information for the 2 variables below
    ('continuous_var_imputer', msi.AddMissingIndicator(variables = ['LotFrontage', 'GarageYrBlt'])),
     
    # replace NA by the median in the 3 variables below, they are numerical
    ('continuous_var_median_imputer', msi.MeanMedianImputer(imputation_method='median', variables = ['LotFrontage', 'GarageYrBlt', 'MasVnrArea'])),
     
    # replace NA by adding the label "Missing" in categorical variables (transformer will skip those variables where there is no NA)
    ('categorical_imputer', msi.CategoricalImputer(variables = ['var1', 'var2'])),
     
    # median imputer
    # to handle those, I will add an additional step here
    ('additional_median_imputer', msi.MeanMedianImputer(imputation_method='median', variables = ['var4', 'var5'])),
     ])

pipe.fit(X_train)
X_train_t = pipe.transform(X_train)

Feature-engine returns dataframes. More info in this link.

To install Feature-Engine do:

pip install feature-engine

Hope that helps

Find elsewhere
Top answer
1 of 1
5

Why

It's not the SimpleImputer exactly; it's the ColumnTransformer itself. ColumnTransformer applies its transformers in parallel, not sequentially (see also [1], [2]), so if a column gets passed to multiple transformers, you'll end up with that column multiple times in the output. In your case, output column 4 comes from the "adder" on total_bedrooms (which has done nothing, so has missing values still), and output column 10 comes from the "imputer" (and so will have no missing values).

Fixes

In this particular case, two approaches seem the easiest.

Impute everything

Any of your numeric features that don't have missings won't be affected. However, if you want the pipeline to error out on future data that has missing values, then don't do this.

num_pipe = Pipeline([
    ("add_feat", AttributeAdder()),
    ("impute", SimpleImputer(strategy="median")),
])
transformer = ColumnTransformer(
    [
        ('num', num_pipe, num_att),
        ('cat', OneHotEncoder(), category_att),
    ],
    remainder = 'passthrough',
)

Smaller transformer column sets

Since you don't actually need your total_bedrooms column for your AttributeAdder, you don't need to pass it into that transformer. The specifics of this will depend on how you're using t_rooms, t_households, etc., but generally:

transformer = ColumnTransformer(
    [
        ('adder', AttributeAdder(), [["total_rooms", "households", "population"]]),
        ('imputer', SimpleImputer(strategy='median'), ['total_bedrooms']),
        ('ohe', OneHotEncoder(), category_att)
    ],
    remainder = 'passthrough'  # now you're relying on this one much more
)

In a related approach, you have more flexibility in how the added features are computed. Change your AttributeAdder to return just the new features (don't concatenate to X in the last step of transform), and rely on the ColumnTransformer to pass those features along. (Note that we can't rely on remainder for those, but we can use "passthrough" as one of the transformers.)

class AttributeAdder(BaseEstimator, TransformerMixin):
    ...
    def transform(self,X,y=None):
        ...
        return np.c_[room_per_household,population_per_household]


transformer = ColumnTransformer(
    [
        ('adder', AttributeAdder(), num_att),
        ('num', "passthrough", num_att.drop(['total_bedrooms'])),
        ('imputer', SimpleImputer(strategy='median'), ['total_bedrooms']),
        ('ohe', OneHotEncoder(), category_att)
    ],
    passthrough=True,  # if you have columns in neither of num_att and category_att that you want kept
)
🌐
Machinelearningknowledge
machinelearningknowledge.ai › how-to-use-sklearn-simple-imputer-simpleimputer-for-filling-missing-values-in-dataset
How To Use Sklearn Simple Imputer (SimpleImputer) for ...
MLK is a knowledge sharing community platform for machine learning enthusiasts, beginners & experts. Let us create a powerful hub together to Make AI Simple
🌐
Towards Data Science
towardsdatascience.com › home › latest › how to use the simpleimputer class in machine learning with python
How to use the SimpleImputer Class in Machine Learning with Python | Towards Data Science
March 5, 2025 - My main purpose here is not to go into this dataset in depth, but rather demonstrate a use case for the utilisation of the SimpleImputer class for predictive modelling. First import the libraries required and perform some exploratory data analysis (EDA). The head of the dataframe and columns attribute called on the dataframe reveal all 15 data attributes.
🌐
Stack Overflow
stackoverflow.com › questions › 62683385 › simpleimputer-using-both-columns-to-calculate-average
python - SimpleImputer Using both columns to calculate average - Stack Overflow
July 1, 2020 - Not if you pass the full matrix (df). If you want to go on a column-by-column case, you'd need to loop, or use .apply() let me work a quick example so I can post an answer
🌐
Haifengl
haifengl.github.io › api › java › smile › feature › imputation › SimpleImputer.html
SimpleImputer
Fits the missing value imputation values. Impute all the numeric columns with median, boolean/nominal columns with mode, and text columns with empty string.
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.5.2 documentation
Added in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed. ... The placeholder for the missing values. All occurrences of missing_values will be imputed. For pandas’ dataframes with nullable integer dtypes with missing values, missing_values can be set to either np.nan or pd.NA. ... The imputation strategy. If “mean”, then replace missing values using the mean along each column.
🌐
Ray
docs.ray.io › en › latest › data › api › doc › ray.data.preprocessors.SimpleImputer.html
ray.data.preprocessors.SimpleImputer — Ray 2.49.2 - Ray Docs
class ray.data.preprocessors.SimpleImputer(columns: List[str], strategy: str = 'mean', fill_value: str | Number | None = None, *, output_columns: List[str] | None = None)[source]#
🌐
Code Underscored
codeunderscored.com › how-to-handle-missing-data-using-simpleimputer-of-scikit-learn
How to handle missing data using SimpleImputer of Scikit-learn
February 22, 2022 - The discarding of entire rows and columns having missing values is the primary method for using incomplete datasets. However, attaining the latter is at the cost of potentially valuable data being lost (even though incomplete). Instead, imputing the missing values, i.e., inferring them from the known part of the data, is preferable. In this article, we’ve shown you how to use sklearn’s SimpleImputer ...