To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.

import pandas as pd
import numpy as np

from sklearn.base import TransformerMixin

class DataFrameImputer(TransformerMixin):

    def __init__(self):
        """Impute missing values.

        Columns of dtype object are imputed with the most frequent value 
        in column.

        Columns of other types are imputed with mean of column.

        """
    def fit(self, X, y=None):

        self.fill = pd.Series([X[c].value_counts().index[0]
            if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
            index=X.columns)

        return self

    def transform(self, X, y=None):
        return X.fillna(self.fill)

data = [
    ['a', 1, 2],
    ['b', 1, 1],
    ['b', 2, 2],
    [np.nan, np.nan, np.nan]
]

X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)

print('before...')
print(X)
print('after...')
print(xt)

which prints,

before...
     0   1   2
0    a   1   2
1    b   1   1
2    b   2   2
3  NaN NaN NaN
after...
   0         1         2
0  a  1.000000  2.000000
1  b  1.000000  1.000000
2  b  2.000000  2.000000
3  b  1.333333  1.666667
Answer from sveitser on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.9.1 documentation
The imputation strategy. If “mean”, then replace missing values using the mean along each column. Can only be used with numeric data.
🌐
DZone
dzone.com › data engineering › data › imputing missing data using sklearn simpleimputer
Imputing Missing Data Using Sklearn SimpleImputer
August 18, 2020 - It is used to impute / replace the numerical or categorical missing data related to one or more features with appropriate values such as following: Each of the above type represents strategy when creating an instance of SimpleImputer.
Discussions

python - Impute categorical missing values in scikit-learn - Stack Overflow
Python generates an error: 'could ... data. ... Imputer works on numbers, not strings. Convert to numbers, then impute, then convert back. ... Why would it not allow categorical vars for most_frequent strategy? strange. ... You can now use from sklearn.impute import SimpleImputer and then imp ... More on stackoverflow.com
🌐 stackoverflow.com
python - Why is SimpleImputer returning categorical data? - Stack Overflow
I'm imputing values into a dataframe using fillna for the numerical columns and SimpleImputer for the categorical columns. The problem is that when I ran the following code, I noticed that all of my More on stackoverflow.com
🌐 stackoverflow.com
Understanding behavior of Simple Imputer with categorical values
Trying to understand the behaviour of an already fitted simple imputer on a dataframe that have only missing values. Adding a sample code for better understanding. import pandas as pd import numpy ... More on github.com
🌐 github.com
2
1
python - sklearn SimpleImputer too slow for categorical data represented as string values - Data Science Stack Exchange
I have a data set with categorical features represented as string values and I want to fill-in missing values in it. I’ve tried to use sklearn’s SimpleImputer but it takes too much time to fulfill ... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
January 7, 2020
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-4-data-preprocessing-feature-engineering › using-imputers
Using Imputers Scikit-learn (SimpleImputer)
The choice of strategy (mean, median, most_frequent, constant) should be guided by the nature of your data and the algorithm you intend to use. For instance, median is often preferred over mean for numerical data with outliers. most_frequent or constant (with a meaningful fill_value) are necessary ...
Top answer
1 of 12
112

To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.

import pandas as pd
import numpy as np

from sklearn.base import TransformerMixin

class DataFrameImputer(TransformerMixin):

    def __init__(self):
        """Impute missing values.

        Columns of dtype object are imputed with the most frequent value 
        in column.

        Columns of other types are imputed with mean of column.

        """
    def fit(self, X, y=None):

        self.fill = pd.Series([X[c].value_counts().index[0]
            if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
            index=X.columns)

        return self

    def transform(self, X, y=None):
        return X.fillna(self.fill)

data = [
    ['a', 1, 2],
    ['b', 1, 1],
    ['b', 2, 2],
    [np.nan, np.nan, np.nan]
]

X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)

print('before...')
print(X)
print('after...')
print(xt)

which prints,

before...
     0   1   2
0    a   1   2
1    b   1   1
2    b   2   2
3  NaN NaN NaN
after...
   0         1         2
0  a  1.000000  2.000000
1  b  1.000000  1.000000
2  b  2.000000  2.000000
3  b  1.333333  1.666667
2 of 12
16

You can use sklearn_pandas.CategoricalImputer for the categorical columns. Details:

First, (from the book Hands-On Machine Learning with Scikit-Learn and TensorFlow) you can have subpipelines for numerical and string/categorical features, where each subpipeline's first transformer is a selector that takes a list of column names (and the full_pipeline.fit_transform() takes a pandas DataFrame):

class DataFrameSelector(BaseEstimator, TransformerMixin):
    def __init__(self, attribute_names):
        self.attribute_names = attribute_names
    def fit(self, X, y=None):
        return self
    def transform(self, X):
        return X[self.attribute_names].values

You can then combine these sub pipelines with sklearn.pipeline.FeatureUnion, for example:

full_pipeline = FeatureUnion(transformer_list=[
    ("num_pipeline", num_pipeline),
    ("cat_pipeline", cat_pipeline)
])

Now, in the num_pipeline you can simply use sklearn.preprocessing.Imputer(), but in the cat_pipline, you can use CategoricalImputer() from the sklearn_pandas package.

note: sklearn-pandas package can be installed with pip install sklearn-pandas, but it is imported as import sklearn_pandas

🌐
GitHub
github.com › scikit-learn › scikit-learn › discussions › 19445
Understanding behavior of Simple Imputer with categorical values · scikit-learn/scikit-learn · Discussion #19445
import pandas as pd import numpy as np from sklearn.impute import SimpleImputer df_2 = pd.DataFrame({"col_1": [1, 2, np.nan]}, dtype="category") imputer = SimpleImputer(strategy="constant", fill_value="missing") # This will error imputer.fit(df_2)
Author: scikit-learn
Find elsewhere
🌐
scikit-learn
scikit-learn.org › stable › modules › impute.html
8.4. Imputation of missing values — scikit-learn 1.9.1 documentation
A better strategy is to impute the missing values, i.e., to infer them from the known part of the data. See the glossary entry on imputation. One type of imputation algorithm is univariate, which imputes values in the i-th feature dimension using only non-missing values in that feature dimension (e.g. SimpleImputer).
🌐
Substack
substack.com › shravankumar’s substack › handling missing data in pandas using simpleimputer
Handling Missing Data in Pandas Using SimpleImputer
June 4, 2024 - SimpleImputer is a class in scikit-learn that provides basic strategies for imputing missing values. It can be used to fill missing values with a specified strategy for both numeric and categorical data.
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 14270
SimpleImputer strategy "most_frequent": infinite loop for categorical variables as strings · Issue #14270 · scikit-learn/scikit-learn
July 5, 2019 - I tried to reset the index of the X_train data frame but it still did not change anything. for cat in df_cat_attributes: imp = SimpleImputer(missing_values=np.nan, strategy='most_frequent',verbose=10) X_train[cat] = imp.fit_transform(X_train[[cat]]).ravel() Edit: This code works with numerical categories (e.g.
Author: scikit-learn
🌐
KDnuggets
kdnuggets.com › 2022 › 07 › scikitlearn-imputer.html
Using Scikit-learn’s Imputer - KDnuggets
To create numerical and categorical transformation data pipelines, we will use sklearn’s Pipeline function. ... In the second step, we have used OrdinalEncoder, to convert categories into numbers. numeric_transformer = Pipeline(steps=[ ('imputer', KNNImputer(n_neighbors=2, weights="uniform")), ('scaler', StandardScaler())]) categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='most_frequent')), ('onehot', OrdinalEncoder())])
🌐
Analytics Vidhya
analyticsvidhya.com › home › sklearn impute for effective missing data handling in machine learning
Sklearn Impute for Effective Missing Data Handling in Machine Learning
May 29, 2025 - We can use the ‘most_frequent’ strategy to impute missing values with the most frequent category in the feature column for categorical data. ... For numerical data, we can use the ‘mean’, ‘median’, or ‘constant’ strategy to impute missing values. The ‘constant’ strategy requires specifying the constant value to replace missing values. imputer = SimpleImputer(strategy='mean') imputer = SimpleImputer(strategy='median') imputer = SimpleImputer(strategy='constant', fill_value=0)
🌐
GeeksforGeeks
geeksforgeeks.org › ml-handle-missing-data-with-simple-imputer
ML | Handle Missing Data with Simple Imputer - GeeksforGeeks
September 28, 2021 - SimpleImputer is a scikit-learn class which is helpful in handling the missing data in the predictive model dataset. It replaces the NaN values with a specified placeholder. It is implemented by the use of the SimpleImputer() method which takes ...
🌐
Medium
medium.com › @noorfatimaafzalbutt › handling-missing-categorical-data-in-pandas-and-sklearn-0740b8ea64c8
Handling Missing Categorical Data in Pandas and sklearn | by Noor Fatima | Medium
July 1, 2024 - from sklearn.model_selection import train_test_split from sklearn.impute import SimpleImputer X_train, X_test, y_train, y_test = train_test_split(df.drop(columns=['SalePrice']), df['SalePrice'], test_size=0.2) imputer = SimpleImputer(strategy='most_frequent') X_train = imputer.fit_transform(X_train) X_test = imputer.transform(X_test) print(imputer.statistics_) This confirms that the imputer has learned the most frequent values ('Gd' for FireplaceQu and 'TA' for GarageQual) and used them for imputation. Handling missing categorical data is crucial for the performance of machine learning models.
🌐
Medium
medium.com › @paghadalsneh › handling-missing-values-in-categorical-data-12b8afaa1be4
Handling Missing Values in Categorical Data | by Sneh Paghdal | Medium
February 5, 2025 - Over-represents the mode category. import pandas as pd from sklearn.impute import SimpleImputer · # Load dataset data = pd.read_csv('housing_data.csv') # Initialize imputer with 'most_frequent' strategy imputer = SimpleImputer(strategy='most_frequent') # Apply to 'Garage_Quality' column data['Garage_Quality'] = imputer.fit_transform(data[['Garage_Quality']])
🌐
The Neural Base
theneuralbase.com › home › scikit learn › beginner course › simpleimputer: missing values
SimpleImputer: missing values | Scikit Learn Beginner Course | The Neural Base
Your data contains string values but SimpleImputer expects numeric. Use strategy='most_frequent' for categorical data, or encode strings to numbers first with LabelEncoder.
🌐
Machinelearningknowledge
machinelearningknowledge.ai › how-to-use-sklearn-simple-imputer-simpleimputer-for-filling-missing-values-in-dataset
How To Use Sklearn Simple Imputer (SimpleImputer) for ...
MLK is a knowledge sharing community platform for machine learning enthusiasts, beginners & experts. Let us create a powerful hub together to Make AI Simple
🌐
Amueller
amueller.github.io › aml › 01-ml-workflow › 08-imputation.html
Missing Values — Applied Machine Learning in Python
Usually, these are unsupervised, so they only make use of the information of features on the training data. The simplest strategy is to fill in a feature with the mean or median of that features over the non-missing samples. That is implemented in the SimpleImputer in ...
🌐
Ray
docs.ray.io › en › latest › data › api › doc › ray.data.preprocessors.SimpleImputer.html
ray.data.preprocessors.SimpleImputer — Ray 2.49.2 - Ray Docs
>>> import pandas as pd >>> import ray >>> from ray.data.preprocessors import SimpleImputer >>> df = pd.DataFrame({"X": [0, None, 3, 3], "Y": [None, "b", "c", "c"]}) >>> ds = ray.data.from_pandas(df) >>> ds.to_pandas() X Y 0 0.0 None 1 NaN b 2 3.0 c 3 3.0 c · The "mean" strategy imputes missing values with the mean of non-missing values. This strategy doesn’t work with categorical data.