There is no need to use the SimpleImputer.
DataFrame.fillna() can do the work as well
For the second column, use
column.fillna(column.mean(), inplace=True)For the third column, use
column.fillna(constant, inplace=True)
Of course, you will need to replace column with your DataFrame's column you want to change and constant with your desired constant.
Edit
Since the use of inplace is discouraged and will be deprecated, the syntax should be
column = column.fillna(column.mean())
Answer from liakoyras on Stack OverflowThere is no need to use the SimpleImputer.
DataFrame.fillna() can do the work as well
For the second column, use
column.fillna(column.mean(), inplace=True)For the third column, use
column.fillna(constant, inplace=True)
Of course, you will need to replace column with your DataFrame's column you want to change and constant with your desired constant.
Edit
Since the use of inplace is discouraged and will be deprecated, the syntax should be
column = column.fillna(column.mean())
Following Dan's advice, an example of using ColumnTransformer and SimpleImputer to backfill the columns is:
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
A = [[7,2,3],[4,np.nan,6],[10,5,np.nan]]
column_trans = ColumnTransformer(
[('imp_col1', SimpleImputer(strategy='mean'), [1]),
('imp_col2', SimpleImputer(strategy='constant', fill_value=29), [2])],
remainder='passthrough')
print(column_trans.fit_transform(A)[:, [2,0,1]])
# [[7 2.0 3]
# [4 3.5 6]
# [10 5.0 29]]
This approach helps with constructing pipelines which are more suitable for larger applications.
You can use numpy.ravel:
from sklearn.preprocessing import Imputer
imp = Imputer(missing_values=0, strategy="mean", axis=0)
df["C"] = imp.fit_transform(df[["C"]]).ravel()
If you have a dataframe with missing data in multiple columns, and you want to impute a specific column based on the others, you can impute everything and take that specific column that you want:
from sklearn.impute import KNNImputer
import pandas as pd
imputer = KNNImputer()
imputed_data = imputer.fit_transform(df) # impute all the missing data
df_temp = pd.DataFrame(imputed_data)
df_temp.columns = df.columns
df['COL_TO_IMPUTE'] = df_temp['COL_TO_IMPUTE'] # update only the desired column
Another method would be to transform all the missing data in the desired column to a unique character that is not contained in the other columns, say # if the data is strings (or max + 1 if the data is numeric), and then tell the imputer that your missing data is #:
from sklearn.impute import KNNImputer
import pandas as pd
cols_backup = df.columns
df['COL_TO_IMPUTE'].fillna('#', inplace=True) # replace all missing data in desired column with with '#'
imputer = KNNImputer(missing_values='#') # tell the imputer to consider only '#' as missing data
imputed_data = imputer.fit_transform(df) # impute all '#'
df = pd.DataFrame(data=imputed_data, columns=cols_backup)