As quickly sketched in the comment there are a couple of considerations to be done on your example:

  • method .fit_transform() generally returns either a sparse matrix or a numpy array. Returning a sparse matrix serves the purpose of saving memory; think to the example where you one-hot-encode a categorical attribute with many categories. You'll end up having a matrix with many columns and a single non-zero entry per row; with a sparse matrix you can store the location of the non-zero element only. In these situation you can call .toarray() on the output of .fit_transform() to get a numpy array back to be passed to the pd.DataFrame constructor.

    Actually, on a five-rows dataset similar to the one you provided

    df = pd.DataFrame({
        'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
        'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
        'projects': [3, 3, 3, 4, 3],
        'salary': ['low', 'medium', 'medium', 'low', 'high'],
        'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
        'bonus': [0, 0, 0, 0, 1],
        'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
        'left': [0, 0, 1, 0, 0]
    })
    
    ord_features = ["salary"]
    ordinal_transformer = OrdinalEncoder()
    
    cat_features = ["department"]
    categorical_transformer = OneHotEncoder(handle_unknown="ignore")
    
    ct = ColumnTransformer(transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features),
    ])
    

    I can't reproduce your issue (namely, I directly obtain a numpy array), but basically pd.DataFrame(ct.fit_transform(df).toarray()) should work for your case. This is the output you would get:

  • As you can see, with respect to your expected output, this only contains the transformed (ordinally encoded) salary column as first column and the transformed (one-hot-encoded) department column from the second to the last column. That's because, as you can see within the docs, parameter remainder is set to 'drop' by default, which implies that all columns which are not subject to transformation are dropped. To avoid this, you should set it to 'passthrough'; this will help you to transform the columns you need and keep the other untouched.

    ct = ColumnTransformer(transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features )],
        remainder='passthrough'
    )
    

    This would be the output of your pd.DataFrame(ct.fit_transform(df).toarray()) in such a case:

  • Again, as you can see also column order is not the one you would expect after the transformation. Long story short, that's because in a ColumnTransformer

The order of the columns in the transformed feature matrix follows the order of how the columns are specified in the transformers list. Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers.

I would aggest reading Preserve column order after applying sklearn.compose.ColumnTransformer at this proposal.

  • Eventually, for what concerns column names you should probably apply a custom solution passing what you want directly to the columns parameter to be passed to the pd.DataFrame constructor. Indeed, OrdinalEncoder (differently from OneHotEncoder) does not provide a .get_feature_names_out() method that makes it generally easy to pass columns=ct.get_feature_names_out() to the pd.DataFrame constructor. See ColumnTransformer & Pipeline with OHE - Is the OHE encoded field retained or removed after ct is performed? for an example of its usage.

Update 10/2022 - sklearn version 1.2.dev0

With sklearn version 1.2.0 it will be possible to solve the problem of returning a DataFrame when transforming a ColumnTransformer instance much more easily. Such version has not been released yet, but you can test the following in dev (version 1.2.dev0), by installing the nightly builds as such:

pip install --pre --extra-index https://pypi.anaconda.org/scipy-wheels-nightly/simple scikit-learn -U

The ColumnTransformer (and other transformers as well) now exposes a .set_output() method which gives the possibility to configure a transformer to output pandas DataFrames, by passing parameter transform='pandas' to it.

Therefore, the example becomes:

import pandas as pd
from sklearn.preprocessing import LabelEncoder, OneHotEncoder, OrdinalEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.ensemble import RandomForestClassifier

df = pd.DataFrame({
    'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
    'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
    'projects': [3, 3, 3, 4, 3],
    'salary': ['low', 'medium', 'medium', 'low', 'high'],
    'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
    'bonus': [0, 0, 0, 0, 1],
    'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
    'left': [0, 0, 1, 0, 0]
})

ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()

cat_features = ["department"]
categorical_transformer = OneHotEncoder(sparse_output=False, handle_unknown="ignore")

ct = ColumnTransformer(transformers=[
    ("ord", ordinal_transformer, ord_features),
    ("cat", categorical_transformer, cat_features )],
    remainder='passthrough'
)

ct.set_output('pandas')
df_pandas = ct.fit_transform(df)
df_pandas

The output also becomes much easier to read as it has proper column names (indeed, at each step, the transformers of which ColumnTransformer is made of do have the attribute feature_names_in_; so you don't lose column names anymore while transforming the input).

Last note. Observe that the example now requires parameter sparse_output=False to be passed to the OneHotEncoder instance in order to work.

Answer from amiola on Stack Overflow
Top answer
1 of 4
24

As quickly sketched in the comment there are a couple of considerations to be done on your example:

  • method .fit_transform() generally returns either a sparse matrix or a numpy array. Returning a sparse matrix serves the purpose of saving memory; think to the example where you one-hot-encode a categorical attribute with many categories. You'll end up having a matrix with many columns and a single non-zero entry per row; with a sparse matrix you can store the location of the non-zero element only. In these situation you can call .toarray() on the output of .fit_transform() to get a numpy array back to be passed to the pd.DataFrame constructor.

    Actually, on a five-rows dataset similar to the one you provided

    df = pd.DataFrame({
        'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
        'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
        'projects': [3, 3, 3, 4, 3],
        'salary': ['low', 'medium', 'medium', 'low', 'high'],
        'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
        'bonus': [0, 0, 0, 0, 1],
        'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
        'left': [0, 0, 1, 0, 0]
    })
    
    ord_features = ["salary"]
    ordinal_transformer = OrdinalEncoder()
    
    cat_features = ["department"]
    categorical_transformer = OneHotEncoder(handle_unknown="ignore")
    
    ct = ColumnTransformer(transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features),
    ])
    

    I can't reproduce your issue (namely, I directly obtain a numpy array), but basically pd.DataFrame(ct.fit_transform(df).toarray()) should work for your case. This is the output you would get:

  • As you can see, with respect to your expected output, this only contains the transformed (ordinally encoded) salary column as first column and the transformed (one-hot-encoded) department column from the second to the last column. That's because, as you can see within the docs, parameter remainder is set to 'drop' by default, which implies that all columns which are not subject to transformation are dropped. To avoid this, you should set it to 'passthrough'; this will help you to transform the columns you need and keep the other untouched.

    ct = ColumnTransformer(transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features )],
        remainder='passthrough'
    )
    

    This would be the output of your pd.DataFrame(ct.fit_transform(df).toarray()) in such a case:

  • Again, as you can see also column order is not the one you would expect after the transformation. Long story short, that's because in a ColumnTransformer

The order of the columns in the transformed feature matrix follows the order of how the columns are specified in the transformers list. Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers.

I would aggest reading Preserve column order after applying sklearn.compose.ColumnTransformer at this proposal.

  • Eventually, for what concerns column names you should probably apply a custom solution passing what you want directly to the columns parameter to be passed to the pd.DataFrame constructor. Indeed, OrdinalEncoder (differently from OneHotEncoder) does not provide a .get_feature_names_out() method that makes it generally easy to pass columns=ct.get_feature_names_out() to the pd.DataFrame constructor. See ColumnTransformer & Pipeline with OHE - Is the OHE encoded field retained or removed after ct is performed? for an example of its usage.

Update 10/2022 - sklearn version 1.2.dev0

With sklearn version 1.2.0 it will be possible to solve the problem of returning a DataFrame when transforming a ColumnTransformer instance much more easily. Such version has not been released yet, but you can test the following in dev (version 1.2.dev0), by installing the nightly builds as such:

pip install --pre --extra-index https://pypi.anaconda.org/scipy-wheels-nightly/simple scikit-learn -U

The ColumnTransformer (and other transformers as well) now exposes a .set_output() method which gives the possibility to configure a transformer to output pandas DataFrames, by passing parameter transform='pandas' to it.

Therefore, the example becomes:

import pandas as pd
from sklearn.preprocessing import LabelEncoder, OneHotEncoder, OrdinalEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.ensemble import RandomForestClassifier

df = pd.DataFrame({
    'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
    'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
    'projects': [3, 3, 3, 4, 3],
    'salary': ['low', 'medium', 'medium', 'low', 'high'],
    'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
    'bonus': [0, 0, 0, 0, 1],
    'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
    'left': [0, 0, 1, 0, 0]
})

ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()

cat_features = ["department"]
categorical_transformer = OneHotEncoder(sparse_output=False, handle_unknown="ignore")

ct = ColumnTransformer(transformers=[
    ("ord", ordinal_transformer, ord_features),
    ("cat", categorical_transformer, cat_features )],
    remainder='passthrough'
)

ct.set_output('pandas')
df_pandas = ct.fit_transform(df)
df_pandas

The output also becomes much easier to read as it has proper column names (indeed, at each step, the transformers of which ColumnTransformer is made of do have the attribute feature_names_in_; so you don't lose column names anymore while transforming the input).

Last note. Observe that the example now requires parameter sparse_output=False to be passed to the OneHotEncoder instance in order to work.

2 of 4
14

This answer skips the workaround and directly provides a solution for scikit-learn version 1.2+

From sklearn version 1.2 on, transformers can return a pandas DataFrame directly without further handling. It is done with set_output, which can be configured per estimator by calling the set_output method or globally by setting set_config(transform_output="pandas"). See Release Highlights for scikit-learn 1.2 - Pandas output with set_output API

In your case the solution would be:

ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()


cat_features = ["department"]
categorical_transformer = OneHotEncoder(handle_unknown="ignore")

ct = ColumnTransformer(
    transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features ),
           ]
)

# Add the following line to your code
ct.set_output(transform="pandas")

df_new = ct.fit_transform(df)
df_new
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.1 documentation
Integers are interpreted as positional columns, while strings can reference DataFrame columns by name. A scalar string or int should be used where transformer expects X to be a 1d array-like (vector), otherwise a 2d array will be passed to the transformer. A callable is passed the input data X and can return any of the above.
Discussions

FunctionTransformer gets numpy instead of pandas after ColumnTransformer
If you run it this way, the DataFrame is printed. If you uncomment ColumnTransformer, then a numpy array is printed. I do not suspect it to be a bug. Seems I just cannot guess the right return format for ColumnTransformer. More on github.com
🌐 github.com
2
1
python - Appending the ColumnTransformer() result to the original data within a pipeline? - Stack Overflow
This is my input data: This is the desired output with transformations applied to the columns r, f, and m and the result is appended to the original data Here's the code: import pandas as pd import More on stackoverflow.com
🌐 stackoverflow.com
Ch2: returning a dataframe after the ColumnTransformer
Dear Ageron, Thank you for this book, it teaches me a lot! I have a problem after successfully running the ColumnTransformer in Chapter two. It returns an array, while I want to turn it to a data f... More on github.com
🌐 github.com
1
October 31, 2019
python - Is it possible for ColumnTransformer or Pipeline to return a dataframe? - Stack Overflow
If you want to convert the output to a DataFrame, it might be convenient to maintain the column labels somewhere. Alexander L. Hayes – Alexander L. Hayes · 2021-05-03 14:18:09 +00:00 Commented May 3, 2021 at 14:18 · Yes, I also believe that the kNN imputation is the bottleneck. Regarding to my questions: I think what I am really looking for is how to convert the output of ColumnTransformer ... More on stackoverflow.com
🌐 stackoverflow.com
Top answer
1 of 3
9

One way to do it would be using a dummy transformer that just returns the transformed column with its original value:

import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import PowerTransformer    

np.random.seed(1714)

class NoTransformer(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        return self

    def transform(self, X):
        assert isinstance(X, pd.DataFrame)
        return X

I'm adding an id column to the dataset so I can show the use of the remainder parameter in ColumnTransformer(), which I find very useful.

df = pd.DataFrame(np.hstack((np.arange(10).reshape((10, 1)),
                             np.random.randint(1,100,size=(10, 3)))),
                  columns=["id"] + list('rfm'))

Using remainder with the value passthrough (by default the value is drop) one can retain the columns that are not transformed; from the docs.

And using the NoTransformer() dummy class we can transform the columns 'r', 'f', 'm' to have the same value.

column_trans = ColumnTransformer(
    [('r_original', NoTransformer(), ['r']),
     ('f_original', NoTransformer(), ['f']),
     ('m_original', NoTransformer(), ['m']),
     ('r_std', StandardScaler(), ['r']),
     ('f_std', StandardScaler(), ['f']),
     ('m_std', StandardScaler(), ['m']),
     ('r_boxcox', PowerTransformer(method='box-cox'), ['r']),
     ('f_boxcox', PowerTransformer(method='box-cox'), ['f']),
     ('m_boxcox', PowerTransformer(method='box-cox'), ['m']),
    ], remainder="passthrough")

A tip if you want to transform many more columns: the fitted ColumnTransformer() class (column_trans in your case) has a transformers_ method that lets you access the names ['r_std', 'f_std', 'm_std', 'r_boxcox', 'f_boxcox', 'm_boxcox'] programmatically:

column_trans.transformers_

#[('r_original', NoTransformer(), ['r']),
# ('f_original', NoTransformer(), ['f']),
# ('m_original', NoTransformer(), ['m']),
# ('r_std', StandardScaler(copy=True, with_mean=True, with_std=True), ['r']),
# ('f_std', StandardScaler(copy=True, with_mean=True, with_std=True), ['f']),
# ('m_std', StandardScaler(copy=True, with_mean=True, with_std=True), ['m']),
# ('r_boxcox',
#  PowerTransformer(copy=True, method='box-cox', standardize=True),
#  ['r']),
# ('f_boxcox',
#  PowerTransformer(copy=True, method='box-cox', standardize=True),
#  ['f']),
# ('m_boxcox',
#  PowerTransformer(copy=True, method='box-cox', standardize=True),
#  ['m']),
# ('remainder', 'passthrough', [0])]


Finally, I think your code could be simplified like this:

column_trans_2 = ColumnTransformer(
    ([
     ('original', NoTransformer(), ['r', 'f', 'm']),
     ('std', StandardScaler(), ['r', 'f', 'm']),
     ('boxcox', PowerTransformer(method='box-cox'), ['r', 'f', 'm']),
    ]), remainder="passthrough")

transformed_2 = column_trans_2.fit_transform(df)
column_trans_2.transformers_

#[('std',
#  StandardScaler(copy=True, with_mean=True, with_std=True),
#  ['r', 'f', 'm']),
# ('boxcox',
#  PowerTransformer(copy=True, method='box-cox', standardize=True),
#  ['r', 'f', 'm'])]

And assign the column names programmatically through transformers_:

new_col_names = []
for i in range(len(column_trans_2.transformers)):
    new_col_names += [column_trans_2.transformers[i][0] + '_' + s for s in column_trans_2.transformers[i][2]]
# The non-transformed columns ('id' in this case) will be appended on the right of
# the array and do not show up in the 'transformers_' method.
# Add the id columns to the col_names manually
new_col_names += ['id']

# ['original_r', 'original_f', 'original_m', 'std_r', 'std_f', 'std_m', 'boxcox_r',
#  'boxcox_f', 'boxcox_m', 'id']


pd.DataFrame(transformed_2, columns=new_col_names)
2 of 3
4

Yes, there is a way to do this which luckily is included in SKLearn. In the original documentation of ColumnTransformer you can find a confusing but useful line, which is the following:

transformer{‘drop’, ‘passthrough’} or estimator

Estimator must support fit and transform. Special-cased strings ‘drop’ and ‘passthrough’ are accepted as well, to indicate to drop the columns or to pass them through untransformed, respectively.

This means that if you want to keep a column during ColumnTransformer or drop a column during ColumnTransformer, you can simply indicate it using one of the two special-cased strings, just like this:

column_trans = ColumnTransformer(
[('r_std', StandardScaler(), ['r']),
 ('f_std', StandardScaler(), ['f']),
 ('m_std', StandardScaler(), ['m']),
 ('r_boxcox', PowerTransformer(method='box-cox'), ['r']),
 ('f_boxcox', PowerTransformer(method='box-cox'), ['f']),
 ('m_boxcox', PowerTransformer(method='box-cox'), ['m']),
 ('col_keep', 'passthrough', ['r','f','m'])
])

If you then use the ColumnTransformer, those 3 columns will be kept and not dropped. Alternatively, if you use 'drop' instead of 'passthrough', you can selectively drop certain columns. This in combination with remainder='passthrough' would allow you to drop some columns and keep all of the others. I hope you find this useful!

🌐
Datasciencebyexample
datasciencebyexample.com › 2023 › 02 › 14 › sklearn-columntransformer-one-columns-to-many-columns
Transforming One or More Columns of a Pandas DataFrame using ColumnTransformer | DataScienceTribe
In this example, the CustomTransformer class takes two input columns (‘A’ and ‘B’) and transforms them into two output columns (‘A_squared’ and ‘B_sqrt’) in a pandas DataFrame. The ColumnTransformer applies this transformer to columns ‘A’ and ‘B’ of the input data, and preserves column ‘C’. The “passthrough” option has been used to preserve the remaining column ‘C’ in its original form.
Find elsewhere
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.5.2 documentation
Integers are interpreted as positional columns, while strings can reference DataFrame columns by name. A scalar string or int should be used where transformer expects X to be a 1d array-like (vector), otherwise a 2d array will be passed to the transformer. A callable is passed the input data X and can return any of the above.
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - Perhaps the Pipeline is needed to execute in a sequence while ColumnTransformer doesn’t do that ? Another issue I observed was while doing transformations of train and valid datasets. The resultant train dataset returned a scipy.sparse.csr.csr_matrix while the valid data just returned an ndarray.
🌐
GitHub
github.com › ageron › handson-ml › issues › 507
Ch2: returning a dataframe after the ColumnTransformer · Issue #507 · ageron/handson-ml
October 31, 2019 - Thank you for this book, it teaches me a lot! I have a problem after successfully running the ColumnTransformer in Chapter two. It returns an array, while I want to turn it to a data frame for better accessibility. Yet, it has 17 columns after the transformation, I aware that the extra columns should be the one-hot-encoding columns, but it does not appear in the .transformers_ attribute..
Author: ageron
🌐
Linux find Examples
queirozf.com › entries › scikit-learn-pipelines-custom-pipelines-and-pandas-integration
Scikit-learn Pipelines: Custom Transformers and Pandas integration
August 15, 2020 - import pandas as pd from sklearn.pipeline import Pipeline class SelectColumnsTransformer(): def __init__(self, columns=None): self.columns = columns def transform(self, X, **transform_params): cpy_df = X[self.columns].copy() return cpy_df def fit(self, X, y=None, **fit_params): return self df = pd.DataFrame({ 'name':['alice','bob','charlie','david','edward'], 'age':[24,32,np.nan,38,20] }) # create a pipeline with a single transformer pipe = Pipeline([ ('selector', SelectColumnsTransformer(["name"])) ]) pipe.fit_transform(df) ... import pandas as pd from sklearn.compose import ColumnTransformer
🌐
Bartosz Mikulski
mikulskibartosz.name › preprocessing-the-input-pandas-dataframe-using-columntransformer-in-scikit-learn
Preprocessing the input Pandas DataFrame using ColumnTransformer in Scikit-learn – Bartosz Mikulski | Principal AI Engineer & MLOps Architect
March 25, 2019 - from sklearn.compose import ColumnTransformer preprocessor = ColumnTransformer( remainder = 'passthrough', transformers=[ ('numeric', numeric_transformer, numeric_features), ('categorical', categorical_transformer, categorical_features), ('remove', 'drop', to_be_removed) ])
🌐
Kaggle
kaggle.com › code › ksvmuralidhar › columntransformer-pipeline-simplified
ColumnTransformer & Pipeline Simplified
October 7, 2020 - INTRODUCTION TO COLUMNTRANSFORMER · This Notebook has been released under the Apache 2.0 open source license. Input1 file · arrow_right_alt · Output0 files · arrow_right_alt · Logs18.5 second run - successful · arrow_right_alt · Comments8 comments ·
🌐
Stack Overflow
stackoverflow.com › questions › 67367641 › is-it-possible-for-columntransformer-or-pipeline-to-return-a-dataframe
python - Is it possible for ColumnTransformer or Pipeline to return a dataframe? - Stack Overflow
## numerical features numerical_transformer = make_pipeline( KNNImputer(n_neighbors=3), StandardScaler() ) ## Categorical features from sklearn.impute import SimpleImputer categorical_transformer = make_pipeline( SimpleImputer(strategy='most_frequent'), OneHotEncoder(handle_unknown='ignore') ) ## Combine both steps from sklearn.compose import ColumnTransformer preprocessor = ColumnTransformer( transformers=[ ('numerical', numerical_transformer, numerical_cols), ('categorical', categorical_transformer, categorical_cols) ])
🌐
Dask
ml.dask.org › modules › generated › dask_ml.compose.ColumnTransformer.html
dask_ml.compose.ColumnTransformer — dask-ml 2025.1.1 documentation
class dask_ml.compose.ColumnTransformer(transformers, remainder='drop', sparse_threshold=0.3, n_jobs=1, transformer_weights=None, preserve_dataframe=True)¶
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › miscellaneous › plot_set_output.html
Introducing the set_output API — scikit-learn 1.9.0 documentation
This example will demonstrate the set_output API to configure transformers to output pandas DataFrames. set_output can be configured per estimator by calling the set_output method or globally by setting set_config(transform_output="pandas").
🌐
Fedorkobak
fedorkobak.github.io › python › sklearn › data_transform › pack_object › columns_transformer.html
Columns transformer — Python
np.random.seed(10) sample_size = 10 df = pd.DataFrame({ "col1" : np.random.uniform(5, 10, sample_size), "col2" : np.random.normal(5, 10, sample_size) }) df · The following cell shows the column transformer that passes both columns unchanged by both options. It also adds a standard scaler to demonstrate that the other output remains unaffected. col_transform = ColumnTransformer( transformers = [ ("dummy", FunctionTransformer(lambda x: x), ["col1", "col2"]), ("passthrough", "passthrough", ["col1", "col2"]), ("standart_scaler", StandardScaler(), ["col1", "col2"]), ] ) col_transform.set_output(transform="pandas") col_transform.fit_transform(df)
🌐
Sklearn
sklearn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.8.0 documentation - sklearn
Integers are interpreted as positional columns, while strings can reference DataFrame columns by name. A scalar string or int should be used where transformer expects X to be a 1d array-like (vector), otherwise a 2d array will be passed to the transformer. A callable is passed the input data X and can return any of the above.
🌐
The Neural Base
theneuralbase.com › home › pandas for ml › intermediate course › columntransformer with pandas
ColumnTransformer with pandas | Pandas For Ml Intermediate Course | The Neural Base
Developers often pass a pandas DataFrame directly to ColumnTransformer.fit_transform() and expect a DataFrame back: you get a numpy array instead.