🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.9.1 documentation
Added in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed.
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › ml-handle-missing-data-with-simple-imputer
ML | Handle Missing Data with Simple Imputer - GeeksforGeeks
September 28, 2021 - SimpleImputer is a scikit-learn class which is helpful in handling the missing data in the predictive model dataset. It replaces the NaN values with a specified placeholder.
Discussions

python - How to use SimpleImputer class to impute missing values in different columns with different constant values? - Stack Overflow
If you want to impute different features with different arbitrary values, or the median, you need to set up several SimpleImputer steps within a pipeline and then join them with the ColumnTransformer: More on stackoverflow.com
🌐 stackoverflow.com
I'm having a hard time grasping a behavior I observed in the SimpleImputer in scikit-learn where it seems to update derived values correctly. Can someone explain this to me?
You need to use the SimpleImputer from scikit to fill in the median for total_bedrooms for some rows that have it missing. More on reddit.com
🌐 r/MLQuestions
1
1
May 1, 2024
I am using SimpleImputer in a columntransformer + pipeline and I continue to receive message that my input…

I think what your code does is pass cols_ordinal to an imputer, return it as colums, and in parallel pass cols_ordinal to ordinal encoder and return those as even more columns. So ordinal encoder does not get the imputed columns! For that you need to pipeline imputer end encoder, and pass them to columntransformer as one pipeline.

More on reddit.com
🌐 r/scikit_learn
1
1
March 20, 2020
PROJECT STRUGGLES : r/learnpython
Subreddit for posting questions and asking for general advice about all topics related to learning python. More on reddit.com
🌐 r/learnpython
🌐
Towards Data Science
towardsdatascience.com › home › latest › imputing missing values using the simpleimputer class in sklearn
Imputing Missing Values using the SimpleImputer Class in sklearn | Towards Data Science
September 19, 2021 - An alternative to using the fillna() method is to use the SimpleImputer class from sklearn. You can find the SimpleImputer class from the sklearn.impute package.
🌐
scikit-learn
scikit-learn.org › 1.6 › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.6.1 documentation
Added in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed.
🌐
Analytics Vidhya
analyticsvidhya.com › home › handling missing data with simpleimputer
Handling Missing Data with SimpleImputer - Analytics Vidhya
November 9, 2022 - In simple words, SimpleImputer is a sci-kit library used to fill in the missing values in the datasets.
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.5.2 documentation
Added in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed.
🌐
Train in Data
blog.trainindata.com › imputing-missing-data-with-scikit-learns-simple-imputer
Imputing missing data with Scikit-learn’s simple imputer | Train in Data Blog
April 9, 2024 - Here, we added missing indicators by using SimpleImputer. Scikit-learn has the MissingIndicator() transformer that just adds missing indicators. So if you want to do multivariate imputation while adding placeholders for the nan values, use this class instead.
Find elsewhere
🌐
scikit-learn
scikit-learn.org › 0.20 › modules › generated › sklearn.impute.SimpleImputer.html
sklearn.impute.SimpleImputer — scikit-learn 0.20.4 documentation
sklearn.impute.SimpleImputer · Examples using sklearn.impute.SimpleImputer · class sklearn.impute.SimpleImputer(missing_values=nan, strategy='mean', fill_value=None, verbose=0, copy=True)[source]¶ · Imputation transformer for completing missing values. Read more in the User Guide.
Top answer
1 of 1
17

If you want to impute different features with different arbitrary values, or the median, you need to set up several SimpleImputer steps within a pipeline and then join them with the ColumnTransformer:

import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer

# first we need to make lists, indicating which features
# will be imputed with each method

features_numeric = ['LotFrontage', 'MasVnrArea', 'GarageYrBlt']
features_categoric = ['BsmtQual', 'FireplaceQu']

# then we instantiate the imputers, within a pipeline
# we create one imputer for numerical and one imputer
# for categorical

# this imputer imputes with the mean
imputer_numeric = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='mean')),
])

# this imputer imputes with an arbitrary value
imputer_categoric = Pipeline(
    steps=[('imputer',
            SimpleImputer(strategy='constant', fill_value='Missing'))])

# then we put the features list and the transformers together
# using the column transformer

preprocessor = ColumnTransformer(transformers=[('imputer_numeric',
                                                imputer_numeric,
                                                features_numeric),
                                               ('imputer_categoric',
                                                imputer_categoric,
                                                features_categoric)])

# now we fit the preprocessor
preprocessor.fit(X_train)

# and now we can impute the data
# remember it returs a numpy array

X_train = preprocessor.transform(X_train)
X_test = preprocessor.transform(X_test)

Alternatively, you can use the package Feature-Engine which transformers allow you to specify the features:

from feature_engine import imputation as msi
from sklearn.pipeline import Pipeline as pipe

pipe = pipe([
    # add a binary variable to indicate missing information for the 2 variables below
    ('continuous_var_imputer', msi.AddMissingIndicator(variables = ['LotFrontage', 'GarageYrBlt'])),
     
    # replace NA by the median in the 3 variables below, they are numerical
    ('continuous_var_median_imputer', msi.MeanMedianImputer(imputation_method='median', variables = ['LotFrontage', 'GarageYrBlt', 'MasVnrArea'])),
     
    # replace NA by adding the label "Missing" in categorical variables (transformer will skip those variables where there is no NA)
    ('categorical_imputer', msi.CategoricalImputer(variables = ['var1', 'var2'])),
     
    # median imputer
    # to handle those, I will add an additional step here
    ('additional_median_imputer', msi.MeanMedianImputer(imputation_method='median', variables = ['var4', 'var5'])),
     ])

pipe.fit(X_train)
X_train_t = pipe.transform(X_train)

Feature-engine returns dataframes. More info in this link.

To install Feature-Engine do:

pip install feature-engine

Hope that helps

🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.10.dev0 documentation
Added in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed.
🌐
scikit-learn
scikit-learn.org › 0.23 › modules › generated › sklearn.impute.SimpleImputer.html
sklearn.impute.SimpleImputer — scikit-learn 0.23.2 documentation
New in version 0.20: SimpleImputer replaces the previous sklearn.preprocessing.Imputer estimator which is now removed.
🌐
Snowflake Documentation
docs.snowflake.com › en › developer-guide › snowpark-ml › reference › 1.2.0 › api › modeling › snowflake.ml.modeling.impute.SimpleImputer
snowflake.ml.modeling.impute.SimpleImputer | Snowflake Documentation
Univariate imputer for completing missing values with simple strategies. Note that the add_indicator parameter is not implemented. For more details on this class, see sklearn.impute.SimpleImputer.
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › main › sklearn › impute › _base.py
scikit-learn/sklearn/impute/_base.py at main · scikit-learn/scikit-learn
class SimpleImputer(_BaseImputer): """Univariate imputer for completing missing values with simple strategies. · Replace missing values using a descriptive statistic (e.g. mean, median, or · most frequent) along each column, or ...
Author: scikit-learn
🌐
Sklearner
sklearner.com › scikit-learn-simpleimputer
Scikit-Learn SimpleImputer for Data Imputation | SKLearner
SimpleImputer is used for handling missing values in a dataset by replacing them with a specified strategy such as the mean, median, or most frequent value.