🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.1 documentation
Valid parameter keys can be listed ... of ColumnTransformer. ... Estimator parameters. ... This estimator. ... Transform X separately by each transformer, concatenate results. ... The data to be transformed by subset. ... Parameters to be passed to the underlying transformers’ transform method. You can only pass this if metadata routing is enabled, which you can enable using sklearn.set_confi...
🌐
GeeksforGeeks
geeksforgeeks.org › machine learning › using-columntransformer-in-scikit-learn-for-data-preprocessing
Using ColumnTransformer in Scikit-Learn for Data Preprocessing - GeeksforGeeks
July 23, 2025 - ColumnTransformer :The ColumnTransformer is a powerful tool in sklearn for applying different preprocessing transformations to specific columns within a dataset.
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
from sklearn.compose import make_column_selector as selector preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, selector(dtype_exclude="category")), ("cat", categorical_transformer, selector(dtype_include="category")), ] ) clf = Pipeline( steps=[("preprocessor", preprocessor), ("classifier", LogisticRegression())] ) clf.fit(X_train, y_train) print("model score: %.3f" % clf.score(X_test, y_test)) clf
🌐
Kaggle
kaggle.com › code › ksvmuralidhar › columntransformer-pipeline-simplified
ColumnTransformer & Pipeline Simplified
October 7, 2020 - INTRODUCTION TO COLUMNTRANSFORMER · This Notebook has been released under the Apache 2.0 open source license. Input1 file · arrow_right_alt · Output0 files · arrow_right_alt · Logs18.5 second run - successful · arrow_right_alt · Comments8 comments ·
🌐
Medium
medium.com › @vinodkumargr › 11-column-transformer-in-ml-sklearn-column-transformer-in-machine-learning-48479f8cb48f
11] Column Transformer in ML: Sklearn Column Transformer in Machine Learning | by Vinod Kumar G R | Medium
January 31, 2024 - # Combine Transformers using ColumnTransformer preprocessor = ColumnTransformer( transformers=[ ('SI', SimpleImputer(), imputer_features), ('SS', StandardScaler(), scaling_features), ('OHE', OneHotEncoder(sparse=False), categorical_features], remainder='passthrough' ) # In the transformer, # SI is a given name for that transformation # SimpleImputer is a sklearn library for transformation # imputer_features are the columns/features
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-6-building-pipelines › columntransformer-pipelines
ColumnTransformer for Complex Scikit-learn Pipelines
This is where Scikit-learn's ColumnTransformer comes into play. It allows you to apply different transformers to different columns of your input data in parallel.
🌐
scikit-learn
scikit-learn.org › 0.23 › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 0.23.2 documentation
This example illustrates how to apply different preprocessing and feature extraction pipelines to different subsets of features, using sklearn.compose.ColumnTransformer.
🌐
ONNX
onnx.ai › sklearn-onnx › auto_examples › plot_complex_pipeline.html
Convert a pipeline with ColumnTransformer - sklearn-onnx 1.20.0 documentation
scikit-learn recently shipped ColumnTransformer which lets the user define complex pipeline where each column may be preprocessed with a different transformer. sklearn-onnx still works in this case as shown in Section Convert complex pipelines.
🌐
Medium
medium.com › @paghadalsneh › column-transformer-in-sklearn-434325c365f2
Column Transformer in Sklearn. In the world of data science… | by Sneh Paghdal | Medium
January 25, 2025 - In this blog post, we’ll explore how to use ColumnTransformer to preprocess a dataset effectively. import numpy as np import pandas as pd from sklearn.impute import SimpleImputer from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder from sklearn.model_selection import train_test_split from sklearn.compose import ColumnTransformer
Find elsewhere
🌐
DEV Community
dev.to › vikas_gulia › columntransformer-and-pipelines-in-scikit-learn-clean-scalable-and-powerful-preprocessing-3mff
ColumnTransformer and Pipelines in Scikit-Learn: Clean, Scalable, and Powerful Preprocessing - DEV Community
June 26, 2025 - ColumnTransformer is a powerful utility in Scikit-learn that allows you to apply different transformations to different columns in a clean and efficient way.
Top answer
1 of 4
24

As quickly sketched in the comment there are a couple of considerations to be done on your example:

  • method .fit_transform() generally returns either a sparse matrix or a numpy array. Returning a sparse matrix serves the purpose of saving memory; think to the example where you one-hot-encode a categorical attribute with many categories. You'll end up having a matrix with many columns and a single non-zero entry per row; with a sparse matrix you can store the location of the non-zero element only. In these situation you can call .toarray() on the output of .fit_transform() to get a numpy array back to be passed to the pd.DataFrame constructor.

    Actually, on a five-rows dataset similar to the one you provided

    df = pd.DataFrame({
        'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
        'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
        'projects': [3, 3, 3, 4, 3],
        'salary': ['low', 'medium', 'medium', 'low', 'high'],
        'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
        'bonus': [0, 0, 0, 0, 1],
        'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
        'left': [0, 0, 1, 0, 0]
    })
    
    ord_features = ["salary"]
    ordinal_transformer = OrdinalEncoder()
    
    cat_features = ["department"]
    categorical_transformer = OneHotEncoder(handle_unknown="ignore")
    
    ct = ColumnTransformer(transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features),
    ])
    

    I can't reproduce your issue (namely, I directly obtain a numpy array), but basically pd.DataFrame(ct.fit_transform(df).toarray()) should work for your case. This is the output you would get:

  • As you can see, with respect to your expected output, this only contains the transformed (ordinally encoded) salary column as first column and the transformed (one-hot-encoded) department column from the second to the last column. That's because, as you can see within the docs, parameter remainder is set to 'drop' by default, which implies that all columns which are not subject to transformation are dropped. To avoid this, you should set it to 'passthrough'; this will help you to transform the columns you need and keep the other untouched.

    ct = ColumnTransformer(transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features )],
        remainder='passthrough'
    )
    

    This would be the output of your pd.DataFrame(ct.fit_transform(df).toarray()) in such a case:

  • Again, as you can see also column order is not the one you would expect after the transformation. Long story short, that's because in a ColumnTransformer

The order of the columns in the transformed feature matrix follows the order of how the columns are specified in the transformers list. Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers.

I would aggest reading Preserve column order after applying sklearn.compose.ColumnTransformer at this proposal.

  • Eventually, for what concerns column names you should probably apply a custom solution passing what you want directly to the columns parameter to be passed to the pd.DataFrame constructor. Indeed, OrdinalEncoder (differently from OneHotEncoder) does not provide a .get_feature_names_out() method that makes it generally easy to pass columns=ct.get_feature_names_out() to the pd.DataFrame constructor. See ColumnTransformer & Pipeline with OHE - Is the OHE encoded field retained or removed after ct is performed? for an example of its usage.

Update 10/2022 - sklearn version 1.2.dev0

With sklearn version 1.2.0 it will be possible to solve the problem of returning a DataFrame when transforming a ColumnTransformer instance much more easily. Such version has not been released yet, but you can test the following in dev (version 1.2.dev0), by installing the nightly builds as such:

pip install --pre --extra-index https://pypi.anaconda.org/scipy-wheels-nightly/simple scikit-learn -U

The ColumnTransformer (and other transformers as well) now exposes a .set_output() method which gives the possibility to configure a transformer to output pandas DataFrames, by passing parameter transform='pandas' to it.

Therefore, the example becomes:

import pandas as pd
from sklearn.preprocessing import LabelEncoder, OneHotEncoder, OrdinalEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.ensemble import RandomForestClassifier

df = pd.DataFrame({
    'department': ['operations', 'operations', 'support', 'logistics', 'sales'],
    'review': [0.577569, 0.751900, 0.722548, 0.675158, 0.676203],
    'projects': [3, 3, 3, 4, 3],
    'salary': ['low', 'medium', 'medium', 'low', 'high'],
    'satisfaction': [0.626759, 0.751900, 0.722548, 0.675158, 0.676203],
    'bonus': [0, 0, 0, 0, 1],
    'avg_hrs_month': [180.866070, 182.708149, 184.416084, 188.707545, 179.821083],
    'left': [0, 0, 1, 0, 0]
})

ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()

cat_features = ["department"]
categorical_transformer = OneHotEncoder(sparse_output=False, handle_unknown="ignore")

ct = ColumnTransformer(transformers=[
    ("ord", ordinal_transformer, ord_features),
    ("cat", categorical_transformer, cat_features )],
    remainder='passthrough'
)

ct.set_output('pandas')
df_pandas = ct.fit_transform(df)
df_pandas

The output also becomes much easier to read as it has proper column names (indeed, at each step, the transformers of which ColumnTransformer is made of do have the attribute feature_names_in_; so you don't lose column names anymore while transforming the input).

Last note. Observe that the example now requires parameter sparse_output=False to be passed to the OneHotEncoder instance in order to work.

2 of 4
14

This answer skips the workaround and directly provides a solution for scikit-learn version 1.2+

From sklearn version 1.2 on, transformers can return a pandas DataFrame directly without further handling. It is done with set_output, which can be configured per estimator by calling the set_output method or globally by setting set_config(transform_output="pandas"). See Release Highlights for scikit-learn 1.2 - Pandas output with set_output API

In your case the solution would be:

ord_features = ["salary"]
ordinal_transformer = OrdinalEncoder()


cat_features = ["department"]
categorical_transformer = OneHotEncoder(handle_unknown="ignore")

ct = ColumnTransformer(
    transformers=[
        ("ord", ordinal_transformer, ord_features),
        ("cat", categorical_transformer, cat_features ),
           ]
)

# Add the following line to your code
ct.set_output(transform="pandas")

df_new = ct.fit_transform(df)
df_new
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › main › sklearn › compose › _column_transformer.py
scikit-learn/sklearn/compose/_column_transformer.py at main · scikit-learn/scikit-learn
ColumnTransformer : Class that allows combining the · outputs of multiple transformer objects used on column subsets · of the data into a single feature space. · Examples · -------- >>> from sklearn.preprocessing ...
Author: scikit-learn
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - Thankfully, the scikit-learn Python machine learning library provides the ColumnTransformer that allows you to selectively apply data transforms to different columns in your dataset.
🌐
YouTube
youtube.com › watch
Column Transformer in Machine Learning | How to use Sklearn ColumnTransformer | Data Preprocessing - YouTube
Column Transformer is a sciket-learn class used to create and apply separate transformers for numerical and categorical data. It is a class in the scikit-lea...
Published: May 29, 2023
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.ensemble.RandomForestClassifier.html
RandomForestClassifier — scikit-learn 1.9.1 documentation
>>> from sklearn.ensemble import RandomForestClassifier >>> from sklearn.datasets import make_classification >>> X, y = make_classification(n_samples=1000, n_features=4, ... n_informative=2, n_redundant=0, ...
🌐
Medium
medium.com › @joh.haupt › categorical-variables-and-columntransformer-in-scikit-learn-c27f9d6a7a13
Categorical Variables and ColumnTransformer in scikit-learn | by Johannes Haupt | Medium
April 3, 2019 - The ColumnTransformer looks like a sklearn pipepline with an additional argument to select the columns for each transformation.
🌐
Medium
medium.com › mlearning-ai › neat-data-preprocessing-with-pipeline-and-columntransformer-2a0468865b6b
Neat data preprocessing with Pipeline and ColumnTransformer | by Yannawut Kimnaruk | MLearning.ai | Medium
May 25, 2022 - from sklearn.compose import ColumnTransformercol_trans = ColumnTransformer(transformers=[ ('num_pipeline',num_pipeline,num_cols), ('cat_pipeline',cat_pipeline,cat_cols) ], remainder='drop', n_jobs=-1)
🌐
Udemy
udemy.com › development
Machine Learning A-Z [2026]: ML, DL, AI with AWS, Python & R
June 13, 2026 - Learn how to effectively handle categorical data in Python using one-hot encoding with ColumnTransformer. This comprehensive guide covers essential data preprocessing techniques for machine learning models.
Rating: 4.5 ​ - ​ 206K votes
🌐
Hg95
hg95.github.io › sklearn-notes › compose › ColumnTransformer.html
ColumnTransformer · python 学习记录
ColumnTransformer · Update time: 2020-07-15 · results matching "" · No results matching ""