To use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.
import pandas as pd
import numpy as np
from sklearn.base import TransformerMixin
class DataFrameImputer(TransformerMixin):
def __init__(self):
"""Impute missing values.
Columns of dtype object are imputed with the most frequent value
in column.
Columns of other types are imputed with mean of column.
"""
def fit(self, X, y=None):
self.fill = pd.Series([X[c].value_counts().index[0]
if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
index=X.columns)
return self
def transform(self, X, y=None):
return X.fillna(self.fill)
data = [
['a', 1, 2],
['b', 1, 1],
['b', 2, 2],
[np.nan, np.nan, np.nan]
]
X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)
print('before...')
print(X)
print('after...')
print(xt)
which prints,
before...
0 1 2
0 a 1 2
1 b 1 1
2 b 2 2
3 NaN NaN NaN
after...
0 1 2
0 a 1.000000 2.000000
1 b 1.000000 1.000000
2 b 2.000000 2.000000
3 b 1.333333 1.666667
Answer from sveitser on Stack OverflowTo use mean values for numeric columns and the most frequent value for non-numeric columns you could do something like this. You could further distinguish between integers and floats. I guess it might make sense to use the median for integer columns instead.
import pandas as pd
import numpy as np
from sklearn.base import TransformerMixin
class DataFrameImputer(TransformerMixin):
def __init__(self):
"""Impute missing values.
Columns of dtype object are imputed with the most frequent value
in column.
Columns of other types are imputed with mean of column.
"""
def fit(self, X, y=None):
self.fill = pd.Series([X[c].value_counts().index[0]
if X[c].dtype == np.dtype('O') else X[c].mean() for c in X],
index=X.columns)
return self
def transform(self, X, y=None):
return X.fillna(self.fill)
data = [
['a', 1, 2],
['b', 1, 1],
['b', 2, 2],
[np.nan, np.nan, np.nan]
]
X = pd.DataFrame(data)
xt = DataFrameImputer().fit_transform(X)
print('before...')
print(X)
print('after...')
print(xt)
which prints,
before...
0 1 2
0 a 1 2
1 b 1 1
2 b 2 2
3 NaN NaN NaN
after...
0 1 2
0 a 1.000000 2.000000
1 b 1.000000 1.000000
2 b 2.000000 2.000000
3 b 1.333333 1.666667
You can use sklearn_pandas.CategoricalImputer for the categorical columns. Details:
First, (from the book Hands-On Machine Learning with Scikit-Learn and TensorFlow) you can have subpipelines for numerical and string/categorical features, where each subpipeline's first transformer is a selector that takes a list of column names (and the full_pipeline.fit_transform() takes a pandas DataFrame):
class DataFrameSelector(BaseEstimator, TransformerMixin):
def __init__(self, attribute_names):
self.attribute_names = attribute_names
def fit(self, X, y=None):
return self
def transform(self, X):
return X[self.attribute_names].values
You can then combine these sub pipelines with sklearn.pipeline.FeatureUnion, for example:
full_pipeline = FeatureUnion(transformer_list=[
("num_pipeline", num_pipeline),
("cat_pipeline", cat_pipeline)
])
Now, in the num_pipeline you can simply use sklearn.preprocessing.Imputer(), but in the cat_pipline, you can use CategoricalImputer() from the sklearn_pandas package.
note: sklearn-pandas package can be installed with pip install sklearn-pandas, but it is imported as import sklearn_pandas
r - Missing values imputation for categorical variables in Python - Stack Overflow
python - Imputation of missing values and dealing with categorical values - Data Science Stack Exchange
Imputation of categorical variables in python/scikit - Stack Overflow
machine learning - knn imputation of categorical variables in python - Stack Overflow
You can use K nearest neighbors imputation.
Here's one example in R:
library(DMwR)
var1 = c('a','a','a','c','e',NA)
var2 = c('p1','p1','p1','p2','p3','p1')
var3 = c('o1','o1','o1','o2','o3','o1')
df = data.frame('v1'=var1,'v2'=var2,'v3'=var3)
df
knnOutput <- DMwR::knnImputation(df, k = 5)
knnOutput
Output:
v1 v2 v3
1 a p1 o1
2 a p1 o1
3 a p1 o1
4 c p2 o2
5 e p3 o3
6 a p1 o1
UPDATE:
KNN doesn't work well for large data sets. Two options for large data sets are Multinomial imputation and Naive Bayes imputation. Multinomial imputation is a little easier, because you don't need to convert the variables into dummy variables. The Naive Bayes implementation I have shown below is a little more work because it requires you to convert to dummy variables. Below, I show how to fit each of these in R:
# make data with 6M rows
var1 = rep(c('a','a','a','c','e',NA), 10**6)
var2 = rep(c('p1','p1','p1','p2','p3','p1'), 10**6)
var3 = rep(c('o1','o1','o1','o2','o3','o1'), 10**6)
df = data.frame('v1'=var1,'v2'=var2,'v3'=var3)
####################################################################
## Multinomial imputation
library(nnet)
# fit multinomial model on only complete rows
imputerModel = multinom(v1 ~ (v2+ v3)^2, data = df[!is.na(df$v1), ])
# predict missing data
predictions = predict(imputerModel, newdata = df[is.na(df$v1), ])
####################################################################
#### Naive Bayes
library(naivebayes)
library(fastDummies)
# convert to dummy variables
dummyVars <- fastDummies::dummy_cols(df,
select_columns = c("v2", "v3"),
ignore_na = TRUE)
head(dummyVars)
The dummy_cols function adds dummy variables to the existing data frame, so now we will use only columns 4:9 as our training data.
# v1 v2 v3 v2_p1 v2_p2 v2_p3 v3_o1 v3_o2 v3_o3
# 1 a p1 o1 1 0 0 1 0 0
# 2 a p1 o1 1 0 0 1 0 0
# 3 a p1 o1 1 0 0 1 0 0
# 4 c p2 o2 0 1 0 0 1 0
# 5 e p3 o3 0 0 1 0 0 1
# 6 <NA> p1 o1 1 0 0 1 0 0
# create training set
X_train <- na.omit(dummyVars)[, 4:ncol(dummyVars)]
y_train <- na.omit(dummyVars)[, "v1"]
X_to_impute <- dummyVars[is.na(df$v1), 4:ncol(dummyVars)]
Naive_Bayes_Model=multinomial_naive_bayes(x = as.matrix(X_train),
y = y_train)
# predict missing data
Naive_Bayes_preds = predict(Naive_Bayes_Model,
newdata = as.matrix(X_to_impute))
# fill in predictions
df$multinom_preds[is.na(df$v1)] = as.character(predictions)
df$Naive_Bayes_preds[is.na(df$v1)] = as.character(Naive_Bayes_preds)
head(df, 15)
# v1 v2 v3 multinom_preds Naive_Bayes_preds
# 1 a p1 o1 <NA> <NA>
# 2 a p1 o1 <NA> <NA>
# 3 a p1 o1 <NA> <NA>
# 4 c p2 o2 <NA> <NA>
# 5 e p3 o3 <NA> <NA>
# 6 <NA> p1 o1 a a
# 7 a p1 o1 <NA> <NA>
# 8 a p1 o1 <NA> <NA>
# 9 a p1 o1 <NA> <NA>
# 10 c p2 o2 <NA> <NA>
# 11 e p3 o3 <NA> <NA>
# 12 <NA> p1 o1 a a
# 13 a p1 o1 <NA> <NA>
# 14 a p1 o1 <NA> <NA>
# 15 a p1 o1 <NA> <NA>
In case you have access to GPU's you can check out DataWig from AWS Labs to do deep learning-driven categorical imputation. You can experiment with batch sizes (depending on the available GPU memory) and hyperparameter optimization. You can specifically choose categorical encoders with embedding.
As a sidenote, there is also the algorithm MICE (Multivariate Imputation by Chained Equations). Miceforest is one example of a library that runs on CPU's by default. However, the backend uses LightGBM (Gradient Boosting Machine) for random forests classification. You can pass a couple of parameters to the .tune_parameters() function from miceforest when LightGBM was built for GPU's. For example, device="gpu",gpu_platform_id=0,gpu_device_id=0, etc. More info on how to optimize GPU-performance can be found here https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html.
I think like @El Burro suggested, you I believe you should focus on feature transformation mainly. Use different techniques for different features. For straightforward features, such as occupation or gender for example, use one-hot encoding, while for others you can use some kind of hierarchical mapping-clustering (e.g. map values to groups defined by you, for example if those urls linked to products, make 30 different groups with similar types of products and map the urls to these groups. Then you can use again one-hot encoding for these mapped features).
For the textual features you mentioned, I'm pretty confident you could drop some of them. If you really want to exploit some kind of information generated from text though, for starters you could either do some tf-idf feature extraction, so as to generate features for each text-snippet or consider topic modeling (take a look at gensim for python if interested) and represent each text as a mixture of topics.
i have impression, the question was more about how to impute mixed (categorical and contineous) data, than about reduction of the number of features (although an implied problem), so i might give it another try:
Since you not want to use some baseline imputation like mean or median, you will have to fit different models for the imputation of different columns in your data.
There are some ML models that do support mixed input feature types. For example, XGB. XGB also comes with the benefit of being fast in training and offering a quite powerful, yet easily set up generalisation method to prevent overfitting. (via early-stopping the training when booster performance degrades on the validation holdout).
Also, XGB does not require feature normalisation and can handle correlated features as well as inbalanced Data well, without preprocessing/tranformation help. It also can handle missing data (NaN) in the predictors, appearing during training or inference (although it will not perform very well on predicting samples with NaN features, if their specific NaN pattern wasnt part of the training data).
So, basically, to impute column c of your data set X, you would train an XGBClassifier or XGBRegressor (depending on the type of c), to predict c, given X. Depending on the severety and the pattern of the missingnes of your data, you might want to pick a complete data portion (all features present) for validation - or, if not available, rely on xgboost implicit imputation.
Now you can simply train one Imputation model per column. (test data hold out would have to be imputed with a model trained on the training data portion only)
There are several ways of enhancing this process or implementing it more sophisticated: for example, start with the feature that has the fewest missing values, than, after training a model to predict those, use the complete/imputed feature vector as predictor for the next columns model and so on. This, i think goes under the Name Iterative Imputation in the literature.
Than there are more intricate, readily packaged frameworks, like MICE or MissForest, wich wrap and interprete the iterative imputation task in different ways and with different base learners. (At least missForest, i think, can handle mixed data.).
I was able to impute the categorical variables using the steps listed below. I will gladly welcome any omissions or program that can perform such tasks automatically
Step1: Subsets the object's data types(all) into another container
Step2: Change np.NaN into an object data type, say None. Now, the container is made up of only objects data types
Step3: Change the entire container into categorical datasets
Step4: Encode the data set(i am using .cat.codes)
Step5: Change back the value of encoded None into np.NaN
Step5: Use KNN (from fancyimpute) to impute the missing values
Step6: Re-map the encoded dataset to its initial names
Imputer works only on numbers. You can convert the 'sex' column to numbers 1 and 0 using the map function
df.sex=df.sex.map({'female':1,'male':0})
After this you can use Imputer to fill all missing values with 1 or 0 and use the map function again to convert 'sex' back to string values (if you need to).