Update:

  • This Mode function is for dataframes:
my_mode <- function (x, na.rm) {
  xtab <- table(x)
  xmode <- names(which(xtab == max(xtab)))
  if (length(xmode) > 1) xmode <- ">1 mode"
  return(xmode)
}

for (var in 1:ncol(cat_df)) {
  if (class(cat_df[,var])=="numeric") {
    cat_df[is.na(cat_df[,var]),var] <- mean(cat_df[,var], na.rm = TRUE)
  } else if (class(cat_df[,var]) %in% c("character", "factor")) {
    cat_df[is.na(cat_df[,var]),var] <- my_mode(cat_df[,var], na.rm = TRUE)
  }
}

This mode function is for vectors Try this and please let me know.

#define missing values in vector
values <- unique(cat_column)[!is.na(cat_column)]
# mode of cat_column
themode <- values[which.max(tabulate(match(cat_column, values)))] 
#assign missing vector
imputevector <- cat_column                                  
imputevector[is.na(imputevector)] <- themode
Answer from TarJae on Stack Overflow
🌐
ScienceDirect
sciencedirect.com › science › article › pii › S2352914823002289
A comparison of imputation methods for categorical data - ScienceDirect
October 17, 2023 - Imputation methods that can handle ... Imputation: This is one of the simplest and fastest method of dealing with missing data, where missing values are replaced with the mode of a categorical variable....
🌐
Statistics Globe
statisticsglobe.com › home › statistical methods for data analysis | research techniques & applications › mode imputation (how to impute categorical variables using r)
Mode Imputation (How to Impute Categorical Variables Using R)
March 15, 2022 - While category 2 is highly over-represented, all other categories are underrepresented. In other words: The distribution of our imputed data is highly biased! As you have seen, mode imputation is usually not a good idea. The method should only be used, if you have strong theoretical arguments (similar to mean imputation in case of continuous variables). You might say: OK, got it! But what should I do instead?! Recent research literature advises two imputation methods for categorical variables:
Discussions

r - Mode imputation for categorical variables in a dataframe - Stack Overflow
I have a data frame(cat_df) which has categorical variables only. I want to impute mode values to missing values in each variable. I tried the following code. But It's not working. Way -1 cat_df[is... More on stackoverflow.com
🌐 stackoverflow.com
python - Imputation of missing values for categories in pandas - Stack Overflow
Most of the time, you wouldn't want the same imputing strategy for all the columns. For example, you may want column mode for categorical variables and column mean or median for numeric columns. More on stackoverflow.com
🌐 stackoverflow.com
Imputation for categorical values (Specifically via KNN imputation)
I often treat missing as another categorical value. Missings are often missing for a reason - missingness contains information and maybe you don’t want to throw that away. More on reddit.com
🌐 r/datascience
9
8
June 18, 2022
Imputation whilst accounting for phylogenetic relatedness

You're going to have to assume a particular model of evolution to do this. There's no universal rule for defining how trait similarity should decay with phylogenetic distance--it depends on the organisms and traits in question.

Brownian Motion is a useful (universally used if not necessarily realistic) starting point. I know Rphylopars is supposed to be able to deal with missing species data (and I think return estimated species means and standard errors). Have you looked into using that to simulate your missing trait values?

More on reddit.com
🌐 r/rstats
10
2
April 26, 2021
🌐
Number Analytics
numberanalytics.com › home › blog › formal science › mastering mode imputation techniques for effective data analysis
Mastering Mode Imputation Techniques for Effective Data Analysis
March 13, 2025 - The algorithm for mode imputation is generally straightforward. However, when scaling up to large datasets or incorporating multiple variables, more nuanced methodologies may be necessary. Consider the following pseudocode representation: for each column in dataset: if data type is categorical: identify missing entries calculate mode_value = frequency_max(column) replace missing entries with mode_value else: continue with appropriate method (mean, median, etc.)
Top answer
1 of 2
1

Update:

  • This Mode function is for dataframes:
my_mode <- function (x, na.rm) {
  xtab <- table(x)
  xmode <- names(which(xtab == max(xtab)))
  if (length(xmode) > 1) xmode <- ">1 mode"
  return(xmode)
}

for (var in 1:ncol(cat_df)) {
  if (class(cat_df[,var])=="numeric") {
    cat_df[is.na(cat_df[,var]),var] <- mean(cat_df[,var], na.rm = TRUE)
  } else if (class(cat_df[,var]) %in% c("character", "factor")) {
    cat_df[is.na(cat_df[,var]),var] <- my_mode(cat_df[,var], na.rm = TRUE)
  }
}

This mode function is for vectors Try this and please let me know.

#define missing values in vector
values <- unique(cat_column)[!is.na(cat_column)]
# mode of cat_column
themode <- values[which.max(tabulate(match(cat_column, values)))] 
#assign missing vector
imputevector <- cat_column                                  
imputevector[is.na(imputevector)] <- themode
2 of 2
0

User Defined Function

Here is the mode function I use with an additional line to choose a single mode in the event there are actually multiple modes:

my_mode <- function(x) {
  ux <- unique(x)
  tab <- tabulate(match(x, ux))
  mode <- ux[tab == max(tab)]
  ifelse(length(mode) > 1, sample(mode, 1), mode)
}

# single mode
cat_col_1 <- c(1, 1, 2, NA)
cat_col_1
#> [1]  1  1  2 NA
cat_col_1[is.na(cat_col_1)] <- my_mode(cat_col_1)
cat_col_1
#> [1] 1 1 2 1

# random sample among multimodal
cat_col_2 <- c(1, 1, 2, 2, NA)
cat_col_2
#> [1]  1  1  2  2 NA
cat_col_2[is.na(cat_col_2)] <- my_mode(cat_col_2)
cat_col_2
#> [1] 1 1 2 2 2

DescTools::Mode()

But other folks have written mode functions. One possibility is in the DescTools package and is named Mode().

Because it returns multiple modes in the event there are more than one, you would need to decide what to do in that event.

Here is an example to randomly sample with replacement, the necessary number of modes to replace the missing values.

# single mode
cat_col_3 <- c(1, 1, 2, NA)
cat_col_3
#> [1]  1  1  2 NA
cat_col_3_modes <- DescTools::Mode(cat_col_3, na.rm = TRUE)
cat_col_3_nmiss <- sum(is.na(cat_col_3))
cat_col_3[is.na(cat_col_3)] <- sample(cat_col_3_modes, cat_col_3_nmiss, TRUE)
cat_col_3
#> [1] 1 1 2 1

# random sample among multimodal
cat_col_4 <- c(1, 1, 2, 2, NA, NA)
cat_col_4
#> [1]  1  1  2  2 NA NA
cat_col_4_modes <- DescTools::Mode(cat_col_4, na.rm = TRUE)
cat_col_4_nmiss <- sum(is.na(cat_col_4))
cat_col_4[is.na(cat_col_4)] <- sample(cat_col_4_modes, cat_col_4_nmiss, TRUE)
cat_col_4
#> [1] 1 1 2 2 2 1

Created on 2021-04-16 by the reprex package (v1.0.0)

🌐
APXML
apxml.com › courses › intro-data-cleaning-preprocessing › chapter-2-handling-missing-data › basic-imputation-mean-median-mode
Strategy 3: Basic Imputation (Mean/Median/Mode)
When to use it: Mode imputation is the standard choice for categorical columns (like 'Color', 'Country', 'Yes/No'). You cannot calculate a meaningful mean or median for non-numeric categories.
🌐
Analytics Vidhya
analyticsvidhya.com › home › how to handle missing values of categorical variables?
How to Handle Missing Values of Categorical Variables?
April 4, 2025 - For Example,1, Implement this method in a given dataset, we can delete the entire row which contains missing values(delete row-2). You can always impute them based on Mode in the case of categorical variables, just make sure you don’t have highly skewed class distributions.
🌐
Number Analytics
numberanalytics.com › blog › practical-mode-imputation-data-science-workflows
Practical Approaches to Mode Imputation in Data Science Workflows
Mode imputation is routinely applied to address these gaps without compromising the integrity of the analysis. For example, in epidemiological studies, the most common diagnosis is used to impute missing values, ensuring that the analysis reflects realistic patient distributions.
Find elsewhere
🌐
Number Analytics
numberanalytics.com › home › blog › data science › 10 key insights on mode imputation for data excellence
10 Key Insights on Mode Imputation for Data Excellence
March 18, 2025 - Adaptive Imputation: In scenarios where some features are numerical and others categorical, employ a hybrid strategy that applies mode imputation to categorical variables and mean/median imputation to numerical ones.
🌐
Trainindata
feature-engine.trainindata.com › en › 1.7.x › user_guide › imputation › CategoricalImputer.html
CategoricalImputer — 1.7.0
We can check that the variable has various modes like this: ... Replacing missing values in categorical features with a bespoke category is standard practice and perhaps the more natural thing to do. We’ll probably want to impute with the most frequent category when the percentage of missing values is small and the cardinality of the variable is low, not to introduce unnecessary noise.
🌐
ResearchGate
researchgate.net › publication › 374803197_A_comparison_of_imputation_methods_for_categorical_data
A comparison of imputation methods for categorical data
October 1, 2023 - To address this, zero imputation is applied to numerical features (MacNeil Vroomen et al.), and mode imputation is used for categorical features · (Memon et al., 2023), effectively filling the missing hourly values. The dataset is then resampled so that each hour appears only once per day, with numerical variables aggregated using the sum function and categorical variables using the mode function, resulting in a uniform and standardized temporal dataset (Andersen et al., 2021).
🌐
LinkedIn
linkedin.com › all › data cleaning
How do you choose the appropriate imputation method for missing values in categorical data?
October 22, 2023 - The type of categorical data affects the choice of imputation method because some methods preserve the order or ranking of the data, while others do not. For example, mean or median imputation is not suitable for nominal data, because it assigns a numerical value to a non-numerical category. Similarly, mode imputation is not suitable for ordinal data, because it ignores the order or ranking of the categories.
🌐
Statology
statology.org › home › how to choose between mean, median, or mode for imputation
How to Choose Between Mean, Median, or Mode for Imputation
October 6, 2025 - The mean is suitable when the data ... for many real-world datasets · The mode is most appropriate for categorical variables or when the most frequent value best represents the missing data...
🌐
SSRN
papers.ssrn.com › sol3 › papers.cfm
A Comparison of Imputation Methods for Categorical Data by Shaheen Memon, Robert Wamala, Ignace H. Kabano :: SSRN
July 26, 2023 - These databases contain mainly categorical variables. The questionable aspect is the best imputation method for categorical data. Materials and methods: We utilized data extracted from paper-based maternal health records from Kawempe National Referral Hospital, Uganda. We compared the following imputation methods for categorical data in an empirical analysis: Mode, K-Nearest Neighbors (KNN), Random Forest (RF), Sequential Hot-Deck (SHD), and Multiple Imputation by Chained Equations (MICE).
🌐
Number Analytics
numberanalytics.com › home › blog › data science › 5 proven mode imputation methods to boost data quality
5 Proven Mode Imputation Methods to Boost Data Quality
March 18, 2025 - L L associated with imputation. If we define the loss function as the difference between the true values and the imputed values, under the assumption of missing completely at random (MCAR), the simple mode imputation minimizes the expected loss for categorical variables.
🌐
Medium
medium.com › @rahultheogre › grouped-mode-imputation-for-categorical-data-4676d3095423
Grouped Mode Imputation for Categorical Data | by Rahul Sharma | Medium
September 18, 2025 - Grouped mode imputation is a context-aware statistical method for filling missing categorical values. Instead of using the overall most frequent value (global mode), it leverages relationships between categorical variables to make more intelligent ...
🌐
Medium
medium.com › aiskunks › imputation-techniques-for-categorical-features-75c43a86af65
Crash Course in Data: Imputation techniques for Categorical features | by Akhilesh Dongre | AI Skunks | Medium
March 30, 2023 - IterativeImputer: A strategy for imputing missing values by modelling each feature with missing values as a function of other features in a round-robin fashion.
🌐
SAS Support Communities
communities.sas.com › t5 › SAS-Programming › imputing-categorical-variables › td-p › 747726
Solved: imputing categorical variables - SAS Support Communities
June 14, 2021 - To determine the mode of a class variable, you can do this in PROC FREQ and then once you have found the mode, replace missings in your data set with the calculated mode. ... If I was right PROC STDIZE can't handle category variable.
🌐
O'Reilly
oreilly.com › library › view › python-feature-engineering › 9781789806311 › 25efd05f-b90b-42de-821d-1062a36a3268.xhtml
Implementing mode or frequent category imputation - Python Feature Engineering Cookbook [Book]
January 22, 2020 - Mode imputation consists of replacing missing values with the mode. We normally use this procedure in categorical variables, hence the frequent category imputation name. Frequent categories are estimated using the train set and then used to ...
Author: Soledad Galli
Published: 2020
Pages: 372
🌐
Medium
medium.com › @adnan.mazraeh1993 › advanced-missing-data-imputation-from-hero-to-zero-5-4978286448de
Advanced Missing Data Imputation: From Zero to Hero 5 | by Adnan Mazraeh | Medium
November 5, 2025 - Mode Imputation: One simple approach is to impute missing values with the mode (the most frequent category), but this can sometimes introduce bias if the missingness is not random.