As quickly sketched in the comment there are a couple of considerations to be done on your example:
method
.fit_transform()generally returns either a sparse matrix or a numpy array. Returning a sparse matrix serves the purpose of saving memory; think to the example where you one-hot-encode a categorical attribute with many categories. You'll end up having a matrix with many columns and a single non-zero entry per row; with a sparse matrix you can store the location of the non-zero element only. In these situation you can call.toarray()on the output of.fit_transform()to get a numpy array back to be passed to thepd.DataFrameconstructor.Actually, on a five-rows dataset similar to the one you provided
df = pd.DataFrame({ 'department': ['operations', 'operations', 'support', 'logistics', 'sales'], 'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203], 'projects': [3, 3, 3, 4, 3], 'salary': ['low', 'medium', 'medium', 'low', 'high'], 'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203], 'bonus': [0, 0, 0, 0, 1], 'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083], 'left': [0, 0, 1, 0, 0] }) ord_features = ["salary"] ordinal_transformer = OrdinalEncoder() cat_features = ["department"] categorical_transformer = OneHotEncoder(handle_unknown="ignore") ct = ColumnTransformer(transformers=[ ("ord", ordinal_transformer, ord_features), ("cat", categorical_transformer, cat_features), ])I can't reproduce your issue (namely, I directly obtain a numpy array), but basically
pd.DataFrame(ct.fit_transform(df).toarray())should work for your case. This is the output you would get:
As you can see, with respect to your expected output, this only contains the transformed (ordinally encoded) salary column as first column and the transformed (one-hot-encoded) department column from the second to the last column. That's because, as you can see within the docs, parameter
remainderis set to'drop'by default, which implies that all columns which are not subject to transformation are dropped. To avoid this, you should set it to'passthrough'; this will help you to transform the columns you need and keep the other untouched.ct = ColumnTransformer(transformers=[ ("ord", ordinal_transformer, ord_features), ("cat", categorical_transformer, cat_features )], remainder='passthrough' )This would be the output of your
pd.DataFrame(ct.fit_transform(df).toarray())in such a case:
Again, as you can see also column order is not the one you would expect after the transformation. Long story short, that's because in a
ColumnTransformer
The order of the columns in the transformed feature matrix follows the order of how the columns are specified in the transformers list. Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers.
I would aggest reading Preserve column order after applying sklearn.compose.ColumnTransformer at this proposal.
- Eventually, for what concerns column names you should probably apply a custom solution passing what you want directly to the
columnsparameter to be passed to thepd.DataFrameconstructor. Indeed,OrdinalEncoder(differently fromOneHotEncoder) does not provide a.get_feature_names_out()method that makes it generally easy to passcolumns=ct.get_feature_names_out()to thepd.DataFrameconstructor. See ColumnTransformer & Pipeline with OHE - Is the OHE encoded field retained or removed after ct is performed? for an example of its usage.
Update 10/2022 - sklearn version 1.2.dev0
With sklearn version 1.2.0 it will be possible to solve the problem of returning a DataFrame when transforming a ColumnTransformer instance much more easily. Such version has not been released yet, but you can test the following in dev (version 1.2.dev0), by installing the nightly builds as such:
pip install --pre --extra-index https://pypi.anaconda.org/scipy-wheels-nightly/simple scikit-learn -U
The ColumnTransformer (and other transformers as well) now exposes a .set_output() method which gives the possibility to configure a transformer to output pandas DataFrames, by passing parameter transform='pandas' to it.
Therefore, the example becomes:
import pandas as pd
from sklearn.preprocessing import LabelEncoder, OneHotEncoder, OrdinalEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.ensemble import RandomForestClassifier
df = pd.DataFrame({
'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
'projects': [3, 3, 3, 4, 3],
'salary': ['low', 'medium', 'medium', 'low', 'high'],
'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
'bonus': [0, 0, 0, 0, 1],
'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
'left': [0, 0, 1, 0, 0]
})
ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()
cat_features = ["department"]
categorical_transformer = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
ct = ColumnTransformer(transformers=[
("ord", ordinal_transformer, ord_features),
("cat", categorical_transformer, cat_features )],
remainder='passthrough'
)
ct.set_output('pandas')
df_pandas = ct.fit_transform(df)
df_pandas

The output also becomes much easier to read as it has proper column names (indeed, at each step, the transformers of which ColumnTransformer is made of do have the attribute feature_names_in_; so you don't lose column names anymore while transforming the input).
Last note. Observe that the example now requires parameter sparse_output=False to be passed to the OneHotEncoder instance in order to work.
As quickly sketched in the comment there are a couple of considerations to be done on your example:
method
.fit_transform()generally returns either a sparse matrix or a numpy array. Returning a sparse matrix serves the purpose of saving memory; think to the example where you one-hot-encode a categorical attribute with many categories. You'll end up having a matrix with many columns and a single non-zero entry per row; with a sparse matrix you can store the location of the non-zero element only. In these situation you can call.toarray()on the output of.fit_transform()to get a numpy array back to be passed to thepd.DataFrameconstructor.Actually, on a five-rows dataset similar to the one you provided
df = pd.DataFrame({ 'department': ['operations', 'operations', 'support', 'logistics', 'sales'], 'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203], 'projects': [3, 3, 3, 4, 3], 'salary': ['low', 'medium', 'medium', 'low', 'high'], 'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203], 'bonus': [0, 0, 0, 0, 1], 'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083], 'left': [0, 0, 1, 0, 0] }) ord_features = ["salary"] ordinal_transformer = OrdinalEncoder() cat_features = ["department"] categorical_transformer = OneHotEncoder(handle_unknown="ignore") ct = ColumnTransformer(transformers=[ ("ord", ordinal_transformer, ord_features), ("cat", categorical_transformer, cat_features), ])I can't reproduce your issue (namely, I directly obtain a numpy array), but basically
pd.DataFrame(ct.fit_transform(df).toarray())should work for your case. This is the output you would get:
As you can see, with respect to your expected output, this only contains the transformed (ordinally encoded) salary column as first column and the transformed (one-hot-encoded) department column from the second to the last column. That's because, as you can see within the docs, parameter
remainderis set to'drop'by default, which implies that all columns which are not subject to transformation are dropped. To avoid this, you should set it to'passthrough'; this will help you to transform the columns you need and keep the other untouched.ct = ColumnTransformer(transformers=[ ("ord", ordinal_transformer, ord_features), ("cat", categorical_transformer, cat_features )], remainder='passthrough' )This would be the output of your
pd.DataFrame(ct.fit_transform(df).toarray())in such a case:
Again, as you can see also column order is not the one you would expect after the transformation. Long story short, that's because in a
ColumnTransformer
The order of the columns in the transformed feature matrix follows the order of how the columns are specified in the transformers list. Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers.
I would aggest reading Preserve column order after applying sklearn.compose.ColumnTransformer at this proposal.
- Eventually, for what concerns column names you should probably apply a custom solution passing what you want directly to the
columnsparameter to be passed to thepd.DataFrameconstructor. Indeed,OrdinalEncoder(differently fromOneHotEncoder) does not provide a.get_feature_names_out()method that makes it generally easy to passcolumns=ct.get_feature_names_out()to thepd.DataFrameconstructor. See ColumnTransformer & Pipeline with OHE - Is the OHE encoded field retained or removed after ct is performed? for an example of its usage.
Update 10/2022 - sklearn version 1.2.dev0
With sklearn version 1.2.0 it will be possible to solve the problem of returning a DataFrame when transforming a ColumnTransformer instance much more easily. Such version has not been released yet, but you can test the following in dev (version 1.2.dev0), by installing the nightly builds as such:
pip install --pre --extra-index https://pypi.anaconda.org/scipy-wheels-nightly/simple scikit-learn -U
The ColumnTransformer (and other transformers as well) now exposes a .set_output() method which gives the possibility to configure a transformer to output pandas DataFrames, by passing parameter transform='pandas' to it.
Therefore, the example becomes:
import pandas as pd
from sklearn.preprocessing import LabelEncoder, OneHotEncoder, OrdinalEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.ensemble import RandomForestClassifier
df = pd.DataFrame({
'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
'projects': [3, 3, 3, 4, 3],
'salary': ['low', 'medium', 'medium', 'low', 'high'],
'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
'bonus': [0, 0, 0, 0, 1],
'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
'left': [0, 0, 1, 0, 0]
})
ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()
cat_features = ["department"]
categorical_transformer = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
ct = ColumnTransformer(transformers=[
("ord", ordinal_transformer, ord_features),
("cat", categorical_transformer, cat_features )],
remainder='passthrough'
)
ct.set_output('pandas')
df_pandas = ct.fit_transform(df)
df_pandas

The output also becomes much easier to read as it has proper column names (indeed, at each step, the transformers of which ColumnTransformer is made of do have the attribute feature_names_in_; so you don't lose column names anymore while transforming the input).
Last note. Observe that the example now requires parameter sparse_output=False to be passed to the OneHotEncoder instance in order to work.
This answer skips the workaround and directly provides a solution for scikit-learn version 1.2+
From sklearn version 1.2 on, transformers can return a pandas DataFrame directly without further handling. It is done with set_output, which can be configured per estimator by calling the set_output method or globally by setting set_config(transform_output="pandas"). See Release Highlights for scikit-learn 1.2 - Pandas output with set_output API
In your case the solution would be:
ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()
cat_features = ["department"]
categorical_transformer = OneHotEncoder(handle_unknown="ignore")
ct = ColumnTransformer(
transformers=[
("ord", ordinal_transformer, ord_features),
("cat", categorical_transformer, cat_features ),
]
)
# Add the following line to your code
ct.set_output(transform="pandas")
df_new = ct.fit_transform(df)
df_new
FunctionTransformer gets numpy instead of pandas after ColumnTransformer
python - Appending the ColumnTransformer() result to the original data within a pipeline? - Stack Overflow
Ch2: returning a dataframe after the ColumnTransformer
python - Is it possible for ColumnTransformer or Pipeline to return a dataframe? - Stack Overflow
Since sklearn Version 1.2, set_output can be configured per estimator by calling the set_output method or globally by setting set_config(transform_output="pandas")
See Release Highlights for scikit-learn 1.2 - Pandas output with set_output API
Example for set_output():
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler().set_output(transform="pandas")
Example for set_config():
from sklearn import set_config
set_config(transform_output="pandas")
Might be late but for anyone with the same question the answer (as almost everything with Scikit-learn) is the usage of Pipelines
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import FunctionTransformer
from sklearn.pipeline import Pipeline
import pandas as pd
df = pd.DataFrame(dict(
x=[1, 2, np.nan],
y=[2, np.nan, 0]
))
imputer = Pipeline([("imputer", SimpleImputer()),
("pandarizer",FunctionTransformer(lambda x: pd.DataFrame(x, columns = ["x", "y"])))])
imputer.fit_transform(df)
One way to do it would be using a dummy transformer that just returns the transformed column with its original value:
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import PowerTransformer
np.random.seed(1714)
class NoTransformer(BaseEstimator, TransformerMixin):
def fit(self, X, y=None):
return self
def transform(self, X):
assert isinstance(X, pd.DataFrame)
return X
I'm adding an id column to the dataset so I can show the use of the remainder parameter in ColumnTransformer(), which I find very useful.
df = pd.DataFrame(np.hstack((np.arange(10).reshape((10, 1)),
np.random.randint(1,100,size=(10, 3)))),
columns=["id"] + list('rfm'))
Using remainder with the value passthrough (by default the value is drop) one can retain the columns that are not transformed; from the docs.
And using the NoTransformer() dummy class we can transform the columns 'r', 'f', 'm' to have the same value.
column_trans = ColumnTransformer(
[('r_original', NoTransformer(), ['r']),
('f_original', NoTransformer(), ['f']),
('m_original', NoTransformer(), ['m']),
('r_std', StandardScaler(), ['r']),
('f_std', StandardScaler(), ['f']),
('m_std', StandardScaler(), ['m']),
('r_boxcox', PowerTransformer(method='box-cox'), ['r']),
('f_boxcox', PowerTransformer(method='box-cox'), ['f']),
('m_boxcox', PowerTransformer(method='box-cox'), ['m']),
], remainder="passthrough")
A tip if you want to transform many more columns: the fitted ColumnTransformer() class (column_trans in your case) has a transformers_ method that lets you access the names ['r_std', 'f_std', 'm_std', 'r_boxcox', 'f_boxcox', 'm_boxcox'] programmatically:
column_trans.transformers_
#[('r_original', NoTransformer(), ['r']),
# ('f_original', NoTransformer(), ['f']),
# ('m_original', NoTransformer(), ['m']),
# ('r_std', StandardScaler(copy=True, with_mean=True, with_std=True), ['r']),
# ('f_std', StandardScaler(copy=True, with_mean=True, with_std=True), ['f']),
# ('m_std', StandardScaler(copy=True, with_mean=True, with_std=True), ['m']),
# ('r_boxcox',
# PowerTransformer(copy=True, method='box-cox', standardize=True),
# ['r']),
# ('f_boxcox',
# PowerTransformer(copy=True, method='box-cox', standardize=True),
# ['f']),
# ('m_boxcox',
# PowerTransformer(copy=True, method='box-cox', standardize=True),
# ['m']),
# ('remainder', 'passthrough', [0])]
Finally, I think your code could be simplified like this:
column_trans_2 = ColumnTransformer(
([
('original', NoTransformer(), ['r', 'f', 'm']),
('std', StandardScaler(), ['r', 'f', 'm']),
('boxcox', PowerTransformer(method='box-cox'), ['r', 'f', 'm']),
]), remainder="passthrough")
transformed_2 = column_trans_2.fit_transform(df)
column_trans_2.transformers_
#[('std',
# StandardScaler(copy=True, with_mean=True, with_std=True),
# ['r', 'f', 'm']),
# ('boxcox',
# PowerTransformer(copy=True, method='box-cox', standardize=True),
# ['r', 'f', 'm'])]
And assign the column names programmatically through transformers_:
new_col_names = []
for i in range(len(column_trans_2.transformers)):
new_col_names += [column_trans_2.transformers[i][0] + '_' + s for s in column_trans_2.transformers[i][2]]
# The non-transformed columns ('id' in this case) will be appended on the right of
# the array and do not show up in the 'transformers_' method.
# Add the id columns to the col_names manually
new_col_names += ['id']
# ['original_r', 'original_f', 'original_m', 'std_r', 'std_f', 'std_m', 'boxcox_r',
# 'boxcox_f', 'boxcox_m', 'id']
pd.DataFrame(transformed_2, columns=new_col_names)
Yes, there is a way to do this which luckily is included in SKLearn. In the original documentation of ColumnTransformer you can find a confusing but useful line, which is the following:
transformer{‘drop’, ‘passthrough’} or estimator
Estimator must support fit and transform. Special-cased strings ‘drop’ and ‘passthrough’ are accepted as well, to indicate to drop the columns or to pass them through untransformed, respectively.
This means that if you want to keep a column during ColumnTransformer or drop a column during ColumnTransformer, you can simply indicate it using one of the two special-cased strings, just like this:
column_trans = ColumnTransformer(
[('r_std', StandardScaler(), ['r']),
('f_std', StandardScaler(), ['f']),
('m_std', StandardScaler(), ['m']),
('r_boxcox', PowerTransformer(method='box-cox'), ['r']),
('f_boxcox', PowerTransformer(method='box-cox'), ['f']),
('m_boxcox', PowerTransformer(method='box-cox'), ['m']),
('col_keep', 'passthrough', ['r','f','m'])
])
If you then use the ColumnTransformer, those 3 columns will be kept and not dropped. Alternatively, if you use 'drop' instead of 'passthrough', you can selectively drop certain columns. This in combination with remainder='passthrough' would allow you to drop some columns and keep all of the others. I hope you find this useful!