Before I say anything else, note that you are fitting your imputer on X and then on X_test. You should never do this. Instead, you should always fit your imputer on the training data and then use that instance to transform both datasets (training and testing data).

Having said that, your problem is that you are fitting and transforming all columns. As a consequence, the imputer is transforming all columns to type object.

I believe this will solve your problem:

# Impute NaNs of numeric columns
X = X.fillna(X.mean())
X_test = X_test.fillna(X_test.mean())

# Subset of categorical columns
cat = ['Loan_ID','Gender','Married','Dependents','Education','Self_Employed',
       'Credit_History','Loan_Status']
# Fit Imputer on traing data and ONLY on categorical columns
object_imputer = SimpleImputer(strategy='most_frequent').fit(X[cat])
# Transform ONLY categorical columns
X[cat] = object_imputer.transform(X[cat])
X_test[cat] = object_imputer.transform(X_test[cat])

As you can see, all columns have the correct data type now.

X.dtypes
Loan_ID               object
Gender                object
Married               object
Dependents            object
Education             object
Self_Employed         object
ApplicantIncome        int64
CoapplicantIncome    float64
LoanAmount           float64
Loan_Amount_Term     float64
Credit_History       float64
Property_Area         object
Loan_Status           object
dtype: object
Answer from Arturo Sbr on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › impute.html
8.4. Imputation of missing values — scikit-learn 1.9.1 documentation
The SimpleImputer class also supports categorical data represented as string values or pandas categoricals when using the 'most_frequent' or 'constant' strategy:
Discussions

python - Impute categorical missing values in scikit-learn - Stack Overflow
Python generates an error: 'could ... data. ... Imputer works on numbers, not strings. Convert to numbers, then impute, then convert back. ... Why would it not allow categorical vars for most_frequent strategy? strange. ... You can now use from sklearn.impute import SimpleImputer and then imp ... More on stackoverflow.com
🌐 stackoverflow.com
python - sklearn SimpleImputer too slow for categorical data represented as string values - Data Science Stack Exchange
I have a data set with categorical features represented as string values and I want to fill-in missing values in it. I’ve tried to use sklearn’s SimpleImputer but it takes too much time to fulfill ... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
January 7, 2020
Understanding behavior of Simple Imputer with categorical values
Trying to understand the behaviour of an already fitted simple imputer on a dataframe that have only missing values. Adding a sample code for better understanding. import pandas as pd import numpy ... More on github.com
🌐 github.com
2
1
SimpleImputer strategy "most_frequent": infinite loop for categorical variables as strings
Infinite loop when applying SimpleImputer with strategy "most_frequent" to data frame or pandas Series. Right from the first column, Spyder (the IDE I use) gets stuck and nothing happens. Each column of the data frame contains about 900,000 rows. With the code below, I am basically looping through the categorical ... More on github.com
🌐 github.com
8
July 5, 2019
🌐
DZone
dzone.com › data engineering › data › imputing missing data using sklearn simpleimputer
Imputing Missing Data Using Sklearn SimpleImputer
August 18, 2020 - It is used to impute / replace the numerical or categorical missing data related to one or more features with appropriate values such as following: Each of the above type represents strategy when creating an instance of SimpleImputer.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.9.1 documentation
>>> import numpy as np >>> from sklearn.impute import SimpleImputer >>> imp_mean = SimpleImputer(missing_values=np.nan, strategy='mean') >>> imp_mean.fit([[7, 2, 3], [4, np.nan, 6], [10, 5, 9]]) SimpleImputer() >>> X = [[np.nan, 2, 3], [4, np.nan, 6], [10, np.nan, 9]] >>> print(imp_mean.transform(X)) [[ 7. 2. 3. ] [ 4. 3.5 6. ] [10. 3.5 9. ]] For a more detailed example see Imputing missing values before building an estimator. ... Fit the imputer on X. ... Input data, where n_samples is the number of samples and n_features is the number of features.
Top answer
1 of 12
112

To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.

import pandas as pd
import numpy as np

from sklearn.base import TransformerMixin

class DataFrameImputer(TransformerMixin):

    def __init__(self):
        """Impute missing values.

        Columns of dtype object are imputed with the most frequent value 
        in column.

        Columns of other types are imputed with mean of column.

        """
    def fit(self, X, y=None):

        self.fill = pd.Series([X[c].value_counts().index[0]
            if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
            index=X.columns)

        return self

    def transform(self, X, y=None):
        return X.fillna(self.fill)

data = [
    ['a', 1, 2],
    ['b', 1, 1],
    ['b', 2, 2],
    [np.nan, np.nan, np.nan]
]

X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)

print('before...')
print(X)
print('after...')
print(xt)

which prints,

before...
     0   1   2
0    a   1   2
1    b   1   1
2    b   2   2
3  NaN NaN NaN
after...
   0         1         2
0  a  1.000000  2.000000
1  b  1.000000  1.000000
2  b  2.000000  2.000000
3  b  1.333333  1.666667
2 of 12
16

You can use sklearn_pandas.CategoricalImputer for the categorical columns. Details:

First, (from the book Hands-On Machine Learning with Scikit-Learn and TensorFlow) you can have subpipelines for numerical and string/categorical features, where each subpipeline's first transformer is a selector that takes a list of column names (and the full_pipeline.fit_transform() takes a pandas DataFrame):

class DataFrameSelector(BaseEstimator, TransformerMixin):
    def __init__(self, attribute_names):
        self.attribute_names = attribute_names
    def fit(self, X, y=None):
        return self
    def transform(self, X):
        return X[self.attribute_names].values

You can then combine these sub pipelines with sklearn.pipeline.FeatureUnion, for example:

full_pipeline = FeatureUnion(transformer_list=[
    ("num_pipeline", num_pipeline),
    ("cat_pipeline", cat_pipeline)
])

Now, in the num_pipeline you can simply use sklearn.preprocessing.Imputer(), but in the cat_pipline, you can use CategoricalImputer() from the sklearn_pandas package.

note: sklearn-pandas package can be installed with pip install sklearn-pandas, but it is imported as import sklearn_pandas

🌐
GitHub
github.com › scikit-learn › scikit-learn › discussions › 19445
Understanding behavior of Simple Imputer with categorical values · scikit-learn/scikit-learn · Discussion #19445
We officially support numpy arrays, not pandas dataframe. For now, it's impossible to get a df as output of our transformers for example, so you might prefer having homogeneous objects right from the start. Most of the time, passing a df as input goes smoothly, as we still try to be as compatible as possible. But in some edge-cases like this one (categorical dtype is specific to pandas, not to numpy, and we don't account for it everywhere), it can lead to surprising results.
Author: scikit-learn
Find elsewhere
🌐
Stackademic
blog.stackademic.com › using-simpleimputer-to-impute-categorical-columns-in-python-2f9c65605646
Using SimpleImputer to impute Categorical Columns in Python | by Shashanka Shekhar | Stackademic
May 15, 2024 - The placeholder can be a constant value or derived from statistics (mean, median, or most frequent) of each column in which the missing values are located. It can handle both numerical and categorical data.
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-4-data-preprocessing-feature-engineering › using-imputers
Using Imputers Scikit-learn (SimpleImputer)
The choice of strategy (mean, median, ... median is often preferred over mean for numerical data with outliers. most_frequent or constant (with a meaningful fill_value) are necessary for categorical features if using SimpleImputer....
🌐
Medium
medium.com › @noorfatimaafzalbutt › handling-missing-categorical-data-in-pandas-and-sklearn-0740b8ea64c8
Handling Missing Categorical Data in Pandas and sklearn | by Noor Fatima | Medium
July 1, 2024 - from sklearn.model_selection import train_test_split from sklearn.impute import SimpleImputer X_train, X_test, y_train, y_test = train_test_split(df.drop(columns=['SalePrice']), df['SalePrice'], test_size=0.2) imputer = SimpleImputer(strategy='most_frequent') X_train = imputer.fit_transform(X_train) X_test = imputer.transform(X_test) print(imputer.statistics_) This confirms that the imputer has learned the most frequent values ('Gd' for FireplaceQu and 'TA' for GarageQual) and used them for imputation. Handling missing categorical data is crucial for the performance of machine learning models.
🌐
Amueller
amueller.github.io › aml › 01-ml-workflow › 08-imputation.html
Missing Values — Applied Machine Learning in Python
This is known as imputation, and, as mentioned above, usually only applied to continuous features, while for categorical features a new category is created. ... Imputation means filling in missing values, and there is a wide variety of methods. Usually, these are unsupervised, so they only ...
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 14270
SimpleImputer strategy "most_frequent": infinite loop for categorical variables as strings · Issue #14270 · scikit-learn/scikit-learn
July 5, 2019 - I tried to reset the index of the X_train data frame but it still did not change anything. for cat in df_cat_attributes: imp = SimpleImputer(missing_values=np.nan, strategy='most_frequent',verbose=10) X_train[cat] = imp.fit_transform(X_train[[cat]]).ravel() Edit: This code works with numerical categories (e.g.
Author: scikit-learn
🌐
Medium
medium.com › @paulamaranon › simultaneous-imputation-of-categorical-and-numerical-data-using-sklearn-imputer-857496fcbce
Simultaneous imputation of categorical and numerical data using Sklearn Imputer | by Paula Maranon | Medium
December 29, 2023 - Although it allows replacing missing values using different strategies (like median for numerical and mode for categorical features), it would often require creating two different data frames.
🌐
Kaggle
kaggle.com › c › home-data-for-ml-course › discussion › 133633
Checking your browser - reCAPTCHA
March 3, 2020 - Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
Kanaries
docs.kanaries.net › topics › Python › how-to-use-scikit-learn-imputer
The Ultimate Guide: How to Use Scikit-learn Imputer – Kanaries
To impute categorical values, we can use the 'most_frequent' strategy provided by the SimpleImputer.
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › impute.html
6.4. Imputation of missing values — scikit-learn 1.5.2 documentation
The SimpleImputer class also supports categorical data represented as string values or pandas categoricals when using the 'most_frequent' or 'constant' strategy:
🌐
Analytics Vidhya
analyticsvidhya.com › home › handling missing data with simpleimputer
Handling Missing Data with SimpleImputer - Analytics Vidhya
November 9, 2022 - Let’s suppose we have a numerical column named “Age” in our data set in which some of the values are missing. Then using the Mean strategy will allow us to fill in the missing values in the column by the mean of all age values. ... imputer = SimpleImputer(missing_values=np.nan, strategy='mean') imputer.fit(df['age']) df['age']= imputer.fit_transform(df['age'])
🌐
KDnuggets
kdnuggets.com › 2022 › 07 › scikitlearn-imputer.html
Using Scikit-learn’s Imputer - KDnuggets
For numerical values, it uses mean, median, and constant. For categorical values, it uses the most frequently used and constant value. You can also train your model to predict the missing labels.