🌐
Analytics Vidhya
analyticsvidhya.com › home › understanding column transformer and machine learning pipelines
Understanding Column Transformer and Machine Learning Pipelines
October 16, 2024 - And one extra transformer is our Model so the total transformation steps are 5. When working with Pipelines While creating a Column transformer it’s suggested to pass the index of columns rather than its name because after transformation it’s converted into Numpy Array and the array does not have any column names.
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
Use ColumnTransformer by selecting column by names · We will train our classifier with the following features: ... We create the preprocessing pipelines for both numeric and categorical data. Note that pclass could either be treated as a categorical or numeric feature. numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
🌐
ONNX
onnx.ai › sklearn-onnx › auto_examples › plot_complex_pipeline.html
Convert a pipeline with ColumnTransformer - sklearn-onnx 1.20.0 documentation
scikit-learn recently shipped ColumnTransformer which lets the user define complex pipeline where each column may be preprocessed with a different transformer.
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-6-building-pipelines › columntransformer-pipelines
ColumnTransformer for Complex Scikit-learn Pipelines
It allows you to apply different transformers to different columns of your input data in parallel. The results from applying each transformer are then concatenated horizontally to form the final transformed dataset, which can then be passed to the next step in a larger pipeline, typically an ...
🌐
Towards Data Science
towardsdatascience.com › home › latest › improve your data preprocessing with columntransformer and pipelines
Improve Your Data Preprocessing with ColumnTransformer and Pipelines | Towards Data Science
March 5, 2025 - Sklearn’s Pipelines with ColumnTransformer is an easy way to apply transformation rules in a standard manner, creating a more organized and clean code.
🌐
Kaggle
kaggle.com › code › ksvmuralidhar › columntransformer-pipeline-simplified
ColumnTransformer & Pipeline Simplified
October 7, 2020 - INTRODUCTION TO COLUMNTRANSFORMER · This Notebook has been released under the Apache 2.0 open source license. Input1 file · arrow_right_alt · Output0 files · arrow_right_alt · Logs18.5 second run - successful · arrow_right_alt · Comments8 comments ·
🌐
Medium
medium.com › @zuohaibashraf › column-transformer-and-pipelines-398d66e28f5c
Column Transformer and Pipelines. Column Transformer and Pipelines in… | by Zuhaib Ashraf | Medium
July 14, 2023 - Rewriting all the preprocessing steps each time can be time-consuming. To save time and effort, pipelines are employed. Machine learning pipelines enable us to execute all the preprocessing steps sequentially, and with the help of a Column Transformer, this can be achieved with just a single line of code.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.1 documentation
List of (name, transformer, columns) tuples specifying the transformer objects to be applied to subsets of the data. ... Like in Pipeline and FeatureUnion, this allows the transformer and its parameters to be set using set_params and searched ...
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - Once the transformer is defined, it can be used to transform a dataset. ... A ColumnTransformer can also be used in a Pipeline to selectively prepare the columns of your dataset before fitting a model on the transformed data.
Find elsewhere
🌐
Amir Masoud Sefidian
sefidian.com › 2022 › 08 › 30 › a-tutorial-on-scikit-learn-pipeline-columntransformer-and-featureunion
A tutorial on Scikit-Learn Pipeline, ColumnTransformer, and FeatureUnion
March 27, 2023 - Wouldn’t it be nice to also transform the numerical column too? In particular, let’s impute missing values with median size and scale it between 0 and 1: #Define classification pipeline cat_pipe = Pipeline([('imputer', SimpleImputer(strategy='constant', fill_value='missing')), ('encoder', OneHotEncoder(handle_unknown='ignore', sparse=False))]) #Define value pipeline num_pipe = Pipeline([('imputer', SimpleImputer(strategy='median')), ('scaler', MinMaxScaler())]) #Make columntransformer fit training data preprocessor = ColumnTransformer(transformers=[('cat', cat_pipe, categorical), ('num', n
🌐
Medium
medium.com › @abhaysingh71711 › building-smarter-ml-pipelines-with-column-transformers-895904e97254
Building Smarter ML Pipelines with Column Transformers | by Abhay singh | Medium
September 24, 2024 - This can be cumbersome and error-prone. ColumnTransformer simplifies this by handling all transformations in a single step. Pipeline Integration: ColumnTransformer integrates seamlessly with scikit-learn’s pipeline, enabling you to create more readable and maintainable machine learning workflows.
Top answer
1 of 1
1

There is mismatch between the columns that the function transformer returns, and its feature_names_out= name. They are clashing, both in terms of how many features are returned and their name. I firstly made sure that each function transformer only returns the one column (or more) that it's supposed to, and also changed the names to make them consistent. This way, the number of features and names defined in feature_names_out= matches the actual returned data for each transformer.

A second issue was that new features are created in the ColumnTransformer, and an attempt is made to access those variables by other parts of the same ColumnTransformer. This won't work as the column transformer runs things in parallel. What you'd need to do is chain the column transformers in a sequence, such that the next column transformer can access the new variables created from the earlier one. I have made this change to the code.

I made some other minor changes, including forcing transformers to return pandas dataframes rather than numpy arrays.

The new preprocessor:

Its output column names:

Index(['Sex_female', 'Sex_male', 'Embarked_C', 'Embarked_Q', 'Embarked_S',
       'traveling_category_A', 'traveling_category_B', 'traveling_category_C',
       'age_interval_(0, 10]', 'age_interval_(10, 20]',
       'age_interval_(20, 30]', 'age_interval_(30, 40]',
       'age_interval_(40, 50]', 'age_interval_(50, 60]',
       'age_interval_(60, 70]', 'age_interval_(70, 100]', 'Pclass', 'Fare',
       'PassengerId', 'Survived', 'Name', 'Ticket', 'Cabin'],
      dtype='object')

Code:

import pandas as pd
titanic_data = pd.read_csv('../titanic.csv')

from sklearn.pipeline import make_pipeline, FunctionTransformer
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OrdinalEncoder, OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer

from sklearn import set_config
set_config(transform_output='pandas')

#
# Sum relatives
#
def sum_name(function_transformer, feature_names_in):
    return ["total_relatives"]  # feature names out

def sum_relatives(X):
    X_copy = X.copy()
    X_copy['total_relatives'] = X_copy['SibSp'] + X_copy['Parch']
    return X_copy[['total_relatives']]

total_relatives_pipeline = make_pipeline(
    FunctionTransformer(sum_relatives, feature_names_out=sum_name)
)

#
#Categorize travel
#
def cat_travel_name(function_transformer, feature_names_in):
    return ["traveling_category"]  # feature names out

def categorize_travel(X):
    X_copy = X.copy()

    conditions = [
        (X_copy['total_relatives'] == 0),
        (X_copy['total_relatives'] >= 1) & (X_copy['total_relatives'] <= 3),
        (X_copy['total_relatives'] >= 4)
    ]
    categories = ['A', 'B', 'C']

    X_copy['traveling_category'] = np.select(conditions, categories, default='Unknown')
    return X_copy[['traveling_category']]

travel_category_pipeline = make_pipeline(
    FunctionTransformer(categorize_travel, feature_names_out=cat_travel_name)
)

#
# Ordinal encoder
#
class_order = [[1, 2, 3]]

ord_pipeline = make_pipeline(
    OrdinalEncoder(categories=class_order)    
    )

cat_pipeline = make_pipeline(
    SimpleImputer(strategy="most_frequent"),
    OneHotEncoder(handle_unknown="ignore", sparse_output=False)
    )

#Numerical for fare
fare_pipeline = make_pipeline(
        StandardScaler()
    )

#
# Age transformer
#
def interval_name(function_transformer, feature_names_in):
    return ["age_interval"]  # feature names out

def age_transformer(X):
    X_copy = X.copy()
    median_age_by_class = X_copy.groupby('Pclass')['Age'].median().reset_index()
    median_age_by_class.columns = ['Pclass', 'median_age']
    for index, row in median_age_by_class.iterrows():
        class_value = row['Pclass']
        median_age = row['median_age']
        X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = \
            X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age)
    bins = [0, 10, 20, 30, 40, 50, 60, 70, 100]
    X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins)
    return X_copy[['age_interval']]

def age_processor():
    return make_pipeline(
        FunctionTransformer(age_transformer, feature_names_out=interval_name),
        )

#
# Column transformers
#
preprocessing_initial = ColumnTransformer([
        ("ord", ord_pipeline, ['Pclass']),
        ("age_processing", age_processor(), ['Pclass', 'Age']),
        ("num", fare_pipeline, ['Fare']),
        ("total_relatives", total_relatives_pipeline, ['SibSp', 'Parch'])],
        remainder='passthrough',
        verbose_feature_names_out=False
)

preprocessing_travel_category = ColumnTransformer(
    [("travel_category", travel_category_pipeline, ['total_relatives'])],
    remainder='passthrough',
    verbose_feature_names_out=False
)

preprocessing_cat = ColumnTransformer(
    [("cat", cat_pipeline, ['Sex', 'Embarked', 'traveling_category', 'age_interval'])],
    remainder='passthrough',
    verbose_feature_names_out=False
)

#Final transformer
preprocessor = make_pipeline(
    preprocessing_initial,
    preprocessing_travel_category,
    preprocessing_cat
)

#Run
preprocessor.fit(titanic_data)
preprocessor.fit_transform(titanic_data).columns
🌐
LinkedIn
linkedin.com › pulse › column-transformer-pipelines-machine-learning-zuhaib-ashraf
Column Transformer and Pipelines in Machine Learning
July 14, 2023 - Rewriting all the preprocessing steps each time can be time-consuming. To save time and effort, pipelines are employed. Machine learning pipelines enable us to execute all the preprocessing steps sequentially, and with the help of a Column Transformer, this can be achieved with just a single line of code.
🌐
Medium
medium.com › @pujalabhanuprakash › understanding-the-difference-between-column-transformation-and-pipeline-in-scikit-learn-4b7fb252b52e
Understanding the Difference between Column Transformation and Pipeline in Scikit-Learn” | by BHANUPRAKASH PUJALA | Medium
April 19, 2023 - The column transformer is useful for applying different transformations to specific columns, while the pipeline simplifies the machine learning workflow by combining data pre-processing and model training into a single object.
🌐
Medium
yannawut.medium.com › neat-data-preprocessing-with-pipeline-and-columntransformer-2a0468865b6b
Neat data preprocessing with Pipeline and ColumnTransformer | by Yannawut Kimnaruk | Medium
May 25, 2022 - Pass numerical columns through the numerical pipeline and pass categorical columns through the categorical pipeline created in step 3. remainder=’drop’ is specified to ignore other columns in a dataframe. n_job = -1 means using all processors to run in parallel. from sklearn.compose import ColumnTransformercol_trans = ColumnTransformer(transformers=[ ('num_pipeline',num_pipeline,num_cols), ('cat_pipeline',cat_pipeline,cat_cols) ], remainder='drop', n_jobs=-1)
🌐
GitHub
github.com › scikit-learn › scikit-learn › discussions › 24261
How can I perform multiple transformations of columns with some columns being same across the transformations · scikit-learn/scikit-learn · Discussion #24261
You could also use a pipeline with 2 ColumnTransformers, one after the other if needed. In this last case, tracking the column name-indices is not easy due to the reordering but this is feasible if really you want to. preprocessor_1 = make_column_transformer( ("target_encoder", TargetEncoder(), ["c1", "c2", "c3"]), ("imputer", SimpleImputer(), ["c4"]), ("others", "passthrough", [f"c{i}" for i in range(5, 9)], ) model = make_pipeline( preprocessor_1, StandardScaler(), Predictor() ) So if you would like to add `ColumnTransformer` instead of only a `StandardScaler`, this is where you would need to provide the indices of the reordered columns because the columns names have been dropped then.
Author: scikit-learn
🌐
DEV Community
dev.to › vikas_gulia › columntransformer-and-pipelines-in-scikit-learn-clean-scalable-and-powerful-preprocessing-3mff
ColumnTransformer and Pipelines in Scikit-Learn: Clean, Scalable, and Powerful Preprocessing - DEV Community
June 26, 2025 - Use remainder='passthrough' in ColumnTransformer if you want to keep unprocessed columns. Set handle_unknown='ignore' in OneHotEncoder to avoid crashing on unseen categories during inference. Always split data before fitting your pipeline to avoid data leakage.
🌐
scikit-learn
scikit-learn.org › 1.5 › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.5.2 documentation
Use ColumnTransformer by selecting column by names · We will train our classifier with the following features: ... We create the preprocessing pipelines for both numeric and categorical data. Note that pclass could either be treated as a categorical or numeric feature. numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore")), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )