Searching the source code of Sklearn for SimpleImputer (with strategy= "most_frequent"), the most frequent value is calculated within a loop in python, therefore that is the part of code that is so slow. In the source code of SimpleImputer there is also the comment that explains why they do not use the scipy.stats.mstats.mode, which is mush faster:

scipy.stats.mstats.mode cannot be used because it will no work properly if the first element is masked and if its frequency is equal to the frequency of the most frequent valid element. See https://github.com/scipy/scipy/issues/2636

So if you want to use the SimpleImputer with this strategy, a faster way would be to use the "constant" strategy and pass the most frequent value by yourself (ser.mode()[0]) then the time is almost the same:

t0 = time.time()
imp = SimpleImputer(missing_values=np.nan, strategy='most_frequent')
imp.fit_transform(arr)
print('Simple Imputer (Most Frequent) Time Elapsed:', time.time()-t0)

t0 = time.time()
imp = SimpleImputer(missing_values=np.nan, strategy='constant', fill_value=ser.mode()[0])
imp.fit_transform(arr)
print('Simple Imputer (Constant) Time Elapsed:', time.time()-t0)

t0 = time.time()
ser.fillna(value=ser.mode()[0])
print('Pandas Time Elapsed:', time.time()-t0)

And the time elapsed for each strategy:

Simple Imputer (Most Frequent) Time Elapsed: 14.320188045501709
Simple Imputer (Constant) Time Elapsed: 0.052472829818725586
Pandas Time Elapsed: 0.04726815223693848
Answer from Giannis Krilis on Stack Exchange
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.9.1 documentation
If “most_frequent”, then replace missing using the most frequent value along each column. Can be used with strings or numeric data.
🌐
Analytics Vidhya
analyticsvidhya.com › home › handling missing data with simpleimputer
Handling Missing Data with SimpleImputer - Analytics Vidhya
November 9, 2022 - Using a most frequent imputation technique on the particular categorical column will allow us to fill the missing values bu the most frequent value from the column occurring in the dataset.
Discussions

python - sklearn SimpleImputer too slow for categorical data represented as string values - Data Science Stack Exchange
Simple Imputer (Most Frequent) Time Elapsed: 14.320188045501709 Simple Imputer (Constant) Time Elapsed: 0.052472829818725586 Pandas Time Elapsed: 0.04726815223693848 ... $\begingroup$ Thanks for the answer! I hoped to use SimpleImputer in a sklearn pipeline. Now I'm not sure how to deal with ... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
January 7, 2020
python - sklearn.impute.SimpleImputer: Unable to fill in the most common value for a list of dataframe columns - Stack Overflow
I have a list of columns of a dataframe that have NA's in them (below). The dtype of all these columns is str. X_train_objects = ['HomePlanet', 'Destination', 'Name', 'Cabin_letter', 'Cabin_num... More on stackoverflow.com
🌐 stackoverflow.com
SimpleImputer strategy "most_frequent" returning ValueError: could not convert string to float when imputing strings
Describe the bug When using SimpleImputer with strategy most_frequent to impute string variables, no error is raised on fitting but when transforming the missing data, a value error is raised. Step... More on github.com
🌐 github.com
4
February 26, 2021
SimpleImputer strategy "most_frequent": infinite loop for categorical variables as strings
Description Infinite loop when applying SimpleImputer with strategy "most_frequent" to data frame or pandas Series. Right from the first column, Spyder (the IDE I use) gets stuck and noth... More on github.com
🌐 github.com
8
July 5, 2019
🌐
scikit-learn
scikit-learn.org › stable › modules › impute.html
8.4. Imputation of missing values — scikit-learn 1.9.1 documentation
The SimpleImputer class also supports categorical data represented as string values or pandas categoricals when using the 'most_frequent' or 'constant' strategy:
🌐
Towards Data Science
towardsdatascience.com › home › latest › imputing missing values using the simpleimputer class in sklearn
Imputing Missing Values using the SimpleImputer Class in sklearn | Towards Data Science
September 19, 2021 - In the above example, the "most_frequent" strategy is applied to the entire dataframe. If you use the median or mean strategies, you will get an error as column D is not a numerical column. In this article, I discuss how to replace missing values in your dataframe using sklearn's SimpleImputer class.
🌐
DZone
dzone.com › data engineering › data › imputing missing data using sklearn simpleimputer
Imputing Missing Data Using Sklearn SimpleImputer
August 18, 2020 - The code example below represents the instantiation of SimpleImputer with appropriate strategies for imputing numerical missing data ... For handling categorical missing values, you could use one of the following strategies. However, it is the "most_frequent" strategy which is preferably used.
Find elsewhere
🌐
Ray
docs.ray.io › en › latest › data › api › doc › ray.data.preprocessors.SimpleImputer.html
ray.data.preprocessors.SimpleImputer — Ray 2.49.2 - Ray Docs
>>> preprocessor = SimpleImputer(columns=["X"], strategy="mean") >>> preprocessor.fit_transform(ds).to_pandas() X Y 0 0.0 None 1 2.0 b 2 3.0 c 3 3.0 c · The "most_frequent" strategy imputes missing values with the most frequent value in each column.
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-4-data-preprocessing-feature-engineering › using-imputers
Using Imputers Scikit-learn (SimpleImputer)
Can be sensitive to outliers. median: Replaces missing values using the median along each column. Also only for numerical data. Generally more effective against outliers than the mean. most_frequent: Replaces missing values using the mode (the most frequent value) along each column.
🌐
Train in Data
blog.trainindata.com › imputing-missing-data-with-scikit-learns-simple-imputer
Imputing missing data with Scikit-learn’s simple imputer | Train in Data Blog
April 9, 2024 - Implement the most common missing value imputation methods, like mean, median, and most frequent imputation with sklearn’s simple imputer.
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.5.2 documentation
If “most_frequent”, then replace missing using the most frequent value along each column. Can be used with strings or numeric data.
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 19572
SimpleImputer strategy "most_frequent" returning ValueError: could not convert string to float when imputing strings · Issue #19572 · scikit-learn/scikit-learn
February 26, 2021 - from sklearn.impute import SimpleImputer imp_frequent = SimpleImputer(strategy="most_frequent").fit([['a', 'b', 'c'], ['c', 'b','d']]) imp_frequent.transform([[np.nan, np.nan, np.nan],[np.nan, np.nan, np.nan]])
Author: scikit-learn
🌐
Quatrope
scikit-criteria.quatrope.org › en › 0.8.2 › api › preprocessing › impute.html
skcriteria.preprocessing.impute module — scikit-criteria 0.8.2 documentation
If “most_frequent”, then replace missing using the most frequent value along each column. Can be used with strings or numeric data.
🌐
GeeksforGeeks
geeksforgeeks.org › ml-handle-missing-data-with-simple-imputer
ML | Handle Missing Data with Simple Imputer - GeeksforGeeks
September 28, 2021 - It is implemented by the use of the SimpleImputer() method which takes the following arguments : missing_values : The missing_values placeholder which has to be imputed. By default is NaN strategy : The data which will replace the NaN values from the dataset. The strategy argument can take the values - 'mean'(default), 'median', 'most_frequent' and 'constant'. fill_value : The constant value to be given to the NaN data using the constant strategy.
🌐
Sklearner
sklearner.com › scikit-learn-simpleimputer
Scikit-Learn SimpleImputer for Data Imputation | SKLearner
The key hyperparameters of SimpleImputer include strategy (which determines the replacement value). Common strategies include ‘mean’, ‘median’, ‘most_frequent’, and ‘constant’.
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 14270
SimpleImputer strategy "most_frequent": infinite loop for categorical variables as strings · Issue #14270 · scikit-learn/scikit-learn
July 5, 2019 - for cat in df_cat_attributes: imp = SimpleImputer(missing_values=np.nan, strategy='most_frequent',verbose=10) X_train[cat] = imp.fit_transform(X_train[[cat]]).ravel()
Author: scikit-learn
Top answer
1 of 12
112

To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.

import pandas as pd
import numpy as np

from sklearn.base import TransformerMixin

class DataFrameImputer(TransformerMixin):

    def __init__(self):
        """Impute missing values.

        Columns of dtype object are imputed with the most frequent value 
        in column.

        Columns of other types are imputed with mean of column.

        """
    def fit(self, X, y=None):

        self.fill = pd.Series([X[c].value_counts().index[0]
            if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
            index=X.columns)

        return self

    def transform(self, X, y=None):
        return X.fillna(self.fill)

data = [
    ['a', 1, 2],
    ['b', 1, 1],
    ['b', 2, 2],
    [np.nan, np.nan, np.nan]
]

X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)

print('before...')
print(X)
print('after...')
print(xt)

which prints,

before...
     0   1   2
0    a   1   2
1    b   1   1
2    b   2   2
3  NaN NaN NaN
after...
   0         1         2
0  a  1.000000  2.000000
1  b  1.000000  1.000000
2  b  2.000000  2.000000
3  b  1.333333  1.666667
2 of 12
16

You can use sklearn_pandas.CategoricalImputer for the categorical columns. Details:

First, (from the book Hands-On Machine Learning with Scikit-Learn and TensorFlow) you can have subpipelines for numerical and string/categorical features, where each subpipeline's first transformer is a selector that takes a list of column names (and the full_pipeline.fit_transform() takes a pandas DataFrame):

class DataFrameSelector(BaseEstimator, TransformerMixin):
    def __init__(self, attribute_names):
        self.attribute_names = attribute_names
    def fit(self, X, y=None):
        return self
    def transform(self, X):
        return X[self.attribute_names].values

You can then combine these sub pipelines with sklearn.pipeline.FeatureUnion, for example:

full_pipeline = FeatureUnion(transformer_list=[
    ("num_pipeline", num_pipeline),
    ("cat_pipeline", cat_pipeline)
])

Now, in the num_pipeline you can simply use sklearn.preprocessing.Imputer(), but in the cat_pipline, you can use CategoricalImputer() from the sklearn_pandas package.

note: sklearn-pandas package can be installed with pip install sklearn-pandas, but it is imported as import sklearn_pandas

🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.10.dev0 documentation
If “most_frequent”, then replace missing using the most frequent value along each column. Can be used with strings or numeric data.
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 2888
Improve Imputer 'most_frequent' strategy · Issue #2888 · scikit-learn/scikit-learn
February 24, 2014 - Imputer 'most_frequent' strategy should support any type of value such as String or Boolean and not just numeric value
Author: scikit-learn
🌐
Snowflake Documentation
docs.snowflake.com › en › developer-guide › snowpark-ml › reference › 1.5.4 › api › modeling › snowflake.ml.modeling.impute.SimpleImputer
snowflake.ml.modeling.impute.SimpleImputer | Snowflake Documentation
... If “mean”, replace missing ... median along each column. Can only be used with numeric data. If “most_frequent”, replace missing using the most frequent value along each column....