To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.

import pandas as pd
import numpy as np

from sklearn.base import TransformerMixin

class DataFrameImputer(TransformerMixin):

    def __init__(self):
        """Impute missing values.

        Columns of dtype object are imputed with the most frequent value 
        in column.

        Columns of other types are imputed with mean of column.

        """
    def fit(self, X, y=None):

        self.fill = pd.Series([X[c].value_counts().index[0]
            if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
            index=X.columns)

        return self

    def transform(self, X, y=None):
        return X.fillna(self.fill)

data = [
    ['a', 1, 2],
    ['b', 1, 1],
    ['b', 2, 2],
    [np.nan, np.nan, np.nan]
]

X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)

print('before...')
print(X)
print('after...')
print(xt)

which prints,

before...
     0   1   2
0    a   1   2
1    b   1   1
2    b   2   2
3  NaN NaN NaN
after...
   0         1         2
0  a  1.000000  2.000000
1  b  1.000000  1.000000
2  b  2.000000  2.000000
3  b  1.333333  1.666667
Answer from sveitser on Stack Overflow
Top answer
1 of 12
112

To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.

import pandas as pd
import numpy as np

from sklearn.base import TransformerMixin

class DataFrameImputer(TransformerMixin):

    def __init__(self):
        """Impute missing values.

        Columns of dtype object are imputed with the most frequent value 
        in column.

        Columns of other types are imputed with mean of column.

        """
    def fit(self, X, y=None):

        self.fill = pd.Series([X[c].value_counts().index[0]
            if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
            index=X.columns)

        return self

    def transform(self, X, y=None):
        return X.fillna(self.fill)

data = [
    ['a', 1, 2],
    ['b', 1, 1],
    ['b', 2, 2],
    [np.nan, np.nan, np.nan]
]

X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)

print('before...')
print(X)
print('after...')
print(xt)

which prints,

before...
     0   1   2
0    a   1   2
1    b   1   1
2    b   2   2
3  NaN NaN NaN
after...
   0         1         2
0  a  1.000000  2.000000
1  b  1.000000  1.000000
2  b  2.000000  2.000000
3  b  1.333333  1.666667
2 of 12
16

You can use sklearn_pandas.CategoricalImputer for the categorical columns. Details:

First, (from the book Hands-On Machine Learning with Scikit-Learn and TensorFlow) you can have subpipelines for numerical and string/categorical features, where each subpipeline's first transformer is a selector that takes a list of column names (and the full_pipeline.fit_transform() takes a pandas DataFrame):

class DataFrameSelector(BaseEstimator, TransformerMixin):
    def __init__(self, attribute_names):
        self.attribute_names = attribute_names
    def fit(self, X, y=None):
        return self
    def transform(self, X):
        return X[self.attribute_names].values

You can then combine these sub pipelines with sklearn.pipeline.FeatureUnion, for example:

full_pipeline = FeatureUnion(transformer_list=[
    ("num_pipeline", num_pipeline),
    ("cat_pipeline", cat_pipeline)
])

Now, in the num_pipeline you can simply use sklearn.preprocessing.Imputer(), but in the cat_pipline, you can use CategoricalImputer() from the sklearn_pandas package.

note: sklearn-pandas package can be installed with pip install sklearn-pandas, but it is imported as import sklearn_pandas

🌐
DataCamp
campus.datacamp.com › courses › dealing-with-missing-data-in-python › advanced-imputation-techniques
Imputing categorical values | Python
To summarize, the 3 steps in imputing missing categorical values include converting the non-missing categorical values to ordinal or numerical values, imputing the ordinal DataFrame and lastly converting back to categorical values.
Discussions

r - Missing values imputation for categorical variables in Python - Stack Overflow
I'm seeking for a good imputation method for this case. I have a dataframe with categorical variables and missing data like the following one: import pandas as pd var1 = ['a','a','a','c','e',... More on stackoverflow.com
🌐 stackoverflow.com
October 8, 2019
python - Imputation of missing values and dealing with categorical values - Data Science Stack Exchange
I tried PCA, but it also doesn't ... method be a good approach to deal with this? For missing values imputation I tried KNN and maximum likelihood but I am getting errors due to categorical variables.... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
May 23, 2017
Imputation of categorical variables in python/scikit - Stack Overflow
I have a csv file with 23 columns of categorical string variables i.e. Gender, Location, skillset, etc. Several of these columns have missing values. No column is missing more than 20% of its data... More on stackoverflow.com
🌐 stackoverflow.com
machine learning - knn imputation of categorical variables in python - Stack Overflow
I am trying to implement kNN from the fancyimpute module on a dataset. I was able to implement the code for continuous variables of the datasets using the code below: knn_impute2=KNN(k=3).complete( More on stackoverflow.com
🌐 stackoverflow.com
🌐
James LeDoux's Blog
jamesrledoux.com › code › imputation
Impute Missing Values - James LeDoux’s Blog
June 1, 2019 - # replace missing values with the column mean df_zero_imputed = df.fillna(0) df_zero_imputed · For categorical features, using mean, median, or zero-imputation doesn’t make much sense.
🌐
DataCamp
campus.datacamp.com › courses › dealing-with-missing-data-in-python › advanced-imputation-techniques
KNN imputation of categorical values | Python
Use the KNN() function from fancyimpute to impute the missing values in the ordinally encoded DataFrame users. Convert the ordinal values back to their respective categories using the ordinal encoder's .inverse_transform() method.
🌐
Trainindata
feature-engine.trainindata.com › en › 1.7.x › user_guide › imputation › CategoricalImputer.html
CategoricalImputer — 1.7.0
Feature-engine’s CategoricalImputer() can replace missing data in categorical variables with an arbitrary value, like the string ‘Missing’, or with the most frequent category. You can impute a subset of the categorical variables by passing their names to CategoricalImputer() in a list.
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to handle missing values of categorical variables?
How to Handle Missing Values of Categorical Variables?
April 4, 2025 - You can always impute them based on Mode in the case of categorical variables, just make sure you don’t have highly skewed class distributions.
Find elsewhere
🌐
Medium
medium.com › aiskunks › imputation-techniques-for-categorical-features-75c43a86af65
Crash Course in Data: Imputation techniques for Categorical features | by Akhilesh Dongre | AI Skunks | Medium
March 30, 2023 - We then create a copy of the dataset to use for imputation (df_imputed) and specify the columns with null values that we want to impute (cols_to_impute). We create an instance of the IterativeImputer class from sklearn and use it to impute the ...
🌐
Medium
medium.com › analytics-vidhya › ways-to-handle-categorical-column-missing-data-its-implementations-15dc4a56893
Ways To Handle Categorical Column Missing Data & Its Implementations | by Ganesh Dhasade | Analytics Vidhya | Medium
September 4, 2020 - ... Step 1: Find which category occurred most in each category using mode(). Step 2: Replace all NAN values in that column with that category. Step 3: Drop original columns and keep newly imputed columns.
🌐
Kaggle
kaggle.com › code › ilayaraja07 › 6-categorical-imputation
6. Categorical Imputation | Kaggle
August 13, 2022 - Explore and run machine learning code with Kaggle Notebooks | Using data from Data Cleaning - Feature Imputation
Top answer
1 of 2
3

You can use K nearest neighbors imputation.

Here's one example in R:

library(DMwR)

var1 = c('a','a','a','c','e',NA)
var2 = c('p1','p1','p1','p2','p3','p1')
var3 = c('o1','o1','o1','o2','o3','o1')

df = data.frame('v1'=var1,'v2'=var2,'v3'=var3)
df

knnOutput <- DMwR::knnImputation(df, k = 5) 
knnOutput

Output:

  v1 v2 v3
1  a p1 o1
2  a p1 o1
3  a p1 o1
4  c p2 o2
5  e p3 o3
6  a p1 o1

UPDATE:

KNN doesn't work well for large data sets. Two options for large data sets are Multinomial imputation and Naive Bayes imputation. Multinomial imputation is a little easier, because you don't need to convert the variables into dummy variables. The Naive Bayes implementation I have shown below is a little more work because it requires you to convert to dummy variables. Below, I show how to fit each of these in R:

# make data with 6M rows
var1 = rep(c('a','a','a','c','e',NA), 10**6)
var2 = rep(c('p1','p1','p1','p2','p3','p1'), 10**6)
var3 = rep(c('o1','o1','o1','o2','o3','o1'), 10**6)
df = data.frame('v1'=var1,'v2'=var2,'v3'=var3)

####################################################################
## Multinomial imputation
library(nnet)
# fit multinomial model on only complete rows
imputerModel = multinom(v1 ~ (v2+ v3)^2, data = df[!is.na(df$v1), ])

# predict missing data
predictions = predict(imputerModel, newdata = df[is.na(df$v1), ])

####################################################################
#### Naive Bayes
library(naivebayes)
library(fastDummies)
# convert to dummy variables
dummyVars <- fastDummies::dummy_cols(df, 
                                     select_columns = c("v2", "v3"), 
                                     ignore_na = TRUE)
head(dummyVars)

The dummy_cols function adds dummy variables to the existing data frame, so now we will use only columns 4:9 as our training data.

#     v1 v2 v3 v2_p1 v2_p2 v2_p3 v3_o1 v3_o2 v3_o3
# 1    a p1 o1     1     0     0     1     0     0
# 2    a p1 o1     1     0     0     1     0     0
# 3    a p1 o1     1     0     0     1     0     0
# 4    c p2 o2     0     1     0     0     1     0
# 5    e p3 o3     0     0     1     0     0     1
# 6 <NA> p1 o1     1     0     0     1     0     0
# create training set
X_train <- na.omit(dummyVars)[, 4:ncol(dummyVars)]
y_train <- na.omit(dummyVars)[, "v1"]

X_to_impute <- dummyVars[is.na(df$v1), 4:ncol(dummyVars)]


Naive_Bayes_Model=multinomial_naive_bayes(x = as.matrix(X_train), 
                                          y = y_train)

# predict missing data
Naive_Bayes_preds = predict(Naive_Bayes_Model, 
                                  newdata = as.matrix(X_to_impute))


# fill in predictions
df$multinom_preds[is.na(df$v1)] = as.character(predictions)
df$Naive_Bayes_preds[is.na(df$v1)] = as.character(Naive_Bayes_preds)
head(df, 15)


#         v1 v2 v3 multinom_preds Naive_Bayes_preds
#    1     a p1 o1           <NA>              <NA>
#    2     a p1 o1           <NA>              <NA>
#    3     a p1 o1           <NA>              <NA>
#    4     c p2 o2           <NA>              <NA>
#    5     e p3 o3           <NA>              <NA>
#    6 <NA> p1 o1              a                 a
#    7     a p1 o1           <NA>              <NA>
#    8     a p1 o1           <NA>              <NA>
#    9     a p1 o1           <NA>              <NA>
#    10    c p2 o2           <NA>              <NA>
#    11    e p3 o3           <NA>              <NA>
#    12 <NA> p1 o1              a                 a
#    13    a p1 o1           <NA>              <NA>
#    14    a p1 o1           <NA>              <NA>
#    15    a p1 o1           <NA>              <NA>
2 of 2
1

In case you have access to GPU's you can check out DataWig from AWS Labs to do deep learning-driven categorical imputation. You can experiment with batch sizes (depending on the available GPU memory) and hyperparameter optimization. You can specifically choose categorical encoders with embedding.

As a sidenote, there is also the algorithm MICE (Multivariate Imputation by Chained Equations). Miceforest is one example of a library that runs on CPU's by default. However, the backend uses LightGBM (Gradient Boosting Machine) for random forests classification. You can pass a couple of parameters to the .tune_parameters() function from miceforest when LightGBM was built for GPU's. For example, device="gpu",gpu_platform_id=0,gpu_device_id=0, etc. More info on how to optimize GPU-performance can be found here https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html.

Top answer
1 of 3
2

I think like @El Burro suggested, you I believe you should focus on feature transformation mainly. Use different techniques for different features. For straightforward features, such as occupation or gender for example, use one-hot encoding, while for others you can use some kind of hierarchical mapping-clustering (e.g. map values to groups defined by you, for example if those urls linked to products, make 30 different groups with similar types of products and map the urls to these groups. Then you can use again one-hot encoding for these mapped features).

For the textual features you mentioned, I'm pretty confident you could drop some of them. If you really want to exploit some kind of information generated from text though, for starters you could either do some tf-idf feature extraction, so as to generate features for each text-snippet or consider topic modeling (take a look at gensim for python if interested) and represent each text as a mixture of topics.

2 of 3
1

i have impression, the question was more about how to impute mixed (categorical and contineous) data, than about reduction of the number of features (although an implied problem), so i might give it another try:

Since you not want to use some baseline imputation like mean or median, you will have to fit different models for the imputation of different columns in your data.

There are some ML models that do support mixed input feature types. For example, XGB. XGB also comes with the benefit of being fast in training and offering a quite powerful, yet easily set up generalisation method to prevent overfitting. (via early-stopping the training when booster performance degrades on the validation holdout).

Also, XGB does not require feature normalisation and can handle correlated features as well as inbalanced Data well, without preprocessing/tranformation help. It also can handle missing data (NaN) in the predictors, appearing during training or inference (although it will not perform very well on predicting samples with NaN features, if their specific NaN pattern wasnt part of the training data).

So, basically, to impute column c of your data set X, you would train an XGBClassifier or XGBRegressor (depending on the type of c), to predict c, given X. Depending on the severety and the pattern of the missingnes of your data, you might want to pick a complete data portion (all features present) for validation - or, if not available, rely on xgboost implicit imputation.

Now you can simply train one Imputation model per column. (test data hold out would have to be imputed with a model trained on the training data portion only)

There are several ways of enhancing this process or implementing it more sophisticated: for example, start with the feature that has the fewest missing values, than, after training a model to predict those, use the complete/imputed feature vector as predictor for the next columns model and so on. This, i think goes under the Name Iterative Imputation in the literature.

Than there are more intricate, readily packaged frameworks, like MICE or MissForest, wich wrap and interprete the iterative imputation task in different ways and with different base learners. (At least missForest, i think, can handle mixed data.).

🌐
Trainindata
feature-engine.trainindata.com › en › latest › api_doc › imputation › CategoricalImputer.html
CategoricalImputer — 1.9.4
The CategoricalImputer() replaces missing data in categorical variables by an arbitrary value or by the most frequent category. The CategoricalImputer() imputes by default only categorical variables (type ‘object’ or ‘categorical’).
🌐
Kaggle
kaggle.com › c › house-prices-advanced-regression-techniques › discussion › 47915
Checking your browser - reCAPTCHA
January 20, 2018 - Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
Data Science Dojo
discuss.datasciencedojo.com › python
How to impute missing values of categorical data? - Python - Data Science Dojo Discussions
March 6, 2023 - My categorical dataset contains missing values. During data preprocessing, I want to remove missing values with the category that fits best. I’m a beginner and don’t know much about handling missing values and imputing categorical data. During the research, I found a method but I’m not ...
🌐
Packtpub
subscription.packtpub.com › book › data › 9781789806311 › 2 › ch02lvl1sec16 › implementing-mode-or-frequent-category-imputation
Imputing Missing Data | Python Feature Engineering Cookbook
To impute missing data with pandas in multiple categorical variables, in step 4 we created a for loop over the categorical variables A4 to A7, and for each variable, we calculated the most frequent value using the pandas mode() method in the train set. Then, we used this value to replace the missing values with pandas fillna() in the train and test sets.