I spot my error, I was switching the instantiation of the classes: the custom transformers have to be instantiated inside the ColumnTransformer, while the ColumnTransformer does not have to be instantiated inside the pipeline.

The correct code is the following:

transformation_pipeline = ColumnTransformer([
    ('adoption', TransformAdoptionFeatures(), features_adoption),
    ('census', TransformCensusFeaturesRegr(), features_census),
    ('climate', TransformClimateFeatures(), features_climate),
    ('soil', TransformSoilFeatures(), features_soil),
    ('economic', TransformEconomicFeatures(), features_economic)
],
    remainder='drop')

full_pipeline_stand = Pipeline([
    ('transformation', transformation_pipeline),
    ('scaling', StandardScaler())
])
Answer from giacrava on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
🌐
Analytics Vidhya
analyticsvidhya.com › home › understanding column transformer and machine learning pipelines
Understanding Column Transformer and Machine Learning Pipelines
October 16, 2024 - When working with Pipelines While creating a Column transformer it’s suggested to pass the index of columns rather than its name because after transformation it’s converted into Numpy Array and the array does not have any column names. #1st Imputation Transformer trf1 = ColumnTransformer([ ('impute_age',SimpleImputer(),[2]), ('impute_embarked',SimpleImputer(strategy='most_frequent'),[6]) ],remainder='passthrough') #2nd One Hot Encoding trf2 = ColumnTransformer([ ('ohe_sex_embarked', OneHotEncoder(sparse=False, handle_unknown='ignore'),[1,6]) ], remainder='passthrough') #3rd Scaling trf3 = ColumnTransformer([ ('scale', MinMaxScaler(), slice(0,10)) ]) #4th Feature selection trf4 = SelectKBest(score_func=chi2,k=8) #5th Model trf5 = DecisionTreeClassifier()
🌐
ONNX
onnx.ai › sklearn-onnx › auto_examples › plot_complex_pipeline.html
Convert a pipeline with ColumnTransformer - sklearn-onnx 1.20.0 documentation
scikit-learn recently shipped ColumnTransformer which lets the user define complex pipeline where each column may be preprocessed with a different transformer.
🌐
Kaggle
kaggle.com › code › ksvmuralidhar › columntransformer-pipeline-simplified
ColumnTransformer & Pipeline Simplified
October 7, 2020 - INTRODUCTION TO COLUMNTRANSFORMER · This Notebook has been released under the Apache 2.0 open source license. Input1 file · arrow_right_alt · Output0 files · arrow_right_alt · Logs18.5 second run - successful · arrow_right_alt · Comments8 comments ·
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.1 documentation
Returns the parameters given in the constructor as well as the estimators contained within the transformers of the ColumnTransformer.
🌐
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-6-building-pipelines › columntransformer-pipelines
ColumnTransformer for Complex Scikit-learn Pipelines
This is where Scikit-learn's ColumnTransformer comes into play. It allows you to apply different transformers to different columns of your input data in parallel. The results from applying each transformer are then concatenated horizontally to form the final transformed dataset, which can then ...
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - We can then use these lists in the ColumnTransformer to one hot encode the categorical variables, which should just be the first column. We can also use the list of numerical columns to normalize the remaining data. Next, we can define our SVR model and define a Pipeline that first uses the ColumnTransformer, then fits the model on the prepared dataset.
Find elsewhere
🌐
Towards Data Science
towardsdatascience.com › home › latest › improve your data preprocessing with columntransformer and pipelines
Improve Your Data Preprocessing with ColumnTransformer and Pipelines | Towards Data Science
March 5, 2025 - You can also put a ColumnTransformer inside a Pipeline, because it is also a simple transformer object, and this loop can go on as much as you need.
🌐
Medium
medium.com › @zuohaibashraf › column-transformer-and-pipelines-398d66e28f5c
Column Transformer and Pipelines. Column Transformer and Pipelines in… | by Zuhaib Ashraf | Medium
July 14, 2023 - Machine learning pipelines enable us to execute all the preprocessing steps sequentially, and with the help of a Column Transformer, this can be achieved with just a single line of code.
🌐
freeCodeCamp
freecodecamp.org › news › machine-learning-pipeline
How to Improve Machine Learning Code Quality with Scikit-learn Pipeline and ColumnTransformer
September 8, 2022 - You use the pipeline for multiple transformations of the same columns. On the other hand, you use the ColumnTransformer** to transform each column set separately before combining them later.
🌐
Medium
yannawut.medium.com › neat-data-preprocessing-with-pipeline-and-columntransformer-2a0468865b6b
Neat data preprocessing with Pipeline and ColumnTransformer | by Yannawut Kimnaruk | Medium
May 25, 2022 - Pipeline: Use for multiple transformations of the same columns. ColumnTransformer: Use to transform each column set differently.
🌐
Medium
medium.com › @abhaysingh71711 › building-smarter-ml-pipelines-with-column-transformers-895904e97254
Building Smarter ML Pipelines with Column Transformers | by Abhay singh | Medium
September 24, 2024 - When working with Pipelines While creating a Column transformer it’s suggested to pass the index of columns rather than its name because after transformation it’s converted into Numpy Array and the array does not have any column names. #1st Imputation Transformer trf1 = ColumnTransformer([ ('impute_age',SimpleImputer(),[2]), ('impute_embarked',SimpleImputer(strategy='most_frequent'),[6]) ],remainder='passthrough') #2nd One Hot Encoding trf2 = ColumnTransformer([ ('ohe_sex_embarked', OneHotEncoder(sparse=False, handle_unknown='ignore'),[1,6]) ], remainder='passthrough') #3rd Scaling trf3 = ColumnTransformer([ ('scale', MinMaxScaler(), slice(0,10)) ]) #4th Feature selection trf4 = SelectKBest(score_func=chi2,k=8) #5th Model trf5 = DecisionTreeClassifier()
🌐
Amir Masoud Sefidian
sefidian.com › 2022 › 08 › 30 › a-tutorial-on-scikit-learn-pipeline-columntransformer-and-featureunion
A tutorial on Scikit-Learn Pipeline, ColumnTransformer, and FeatureUnion
March 27, 2023 - Once trained, this Pipeline object can be used for smoother deployment. In the previous example, we imputed and encoded all columns the same way. However, we often need to apply different sets of transformers to different groups of columns. For instance, we would want to apply OneHotEncoder to only categorical columns but not to numerical columns. This is where ColumnTransformer comes in.
Top answer
1 of 1
1

Okay this seems to kinda work, but I still would like to have better column names:


from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import FunctionTransformer

def take_cube(data, col):
    data["speed_cube"] = np.power(data[col], 3)
    return data[['speed_cube']]

def sin_angle(data, speed, acc):
    data["angle"] = data[speed] * np.sin(data[acc])
    return data[['angle']]

preprocessor = ColumnTransformer(
    transformers=[
        ("step1", FunctionTransformer(take_cube, kw_args={'col': 'speed'}, validate=False), ['speed']),
        ("step2", FunctionTransformer(sin_angle, kw_args={'speed': 'speed', 'acc': 'acceleration'}, validate=False), ['speed', 'acceleration']),
    ],
    remainder="passthrough"
).set_output(transform="pandas")
new_data = preprocessor.fit_transform(data)

print(new_data)

Output:


                   step1__speed_cube  step2__angle  remainder__heart-rate  \
                                                                              
2020-08-18 14:43:19          80.901828      0.380109                  102.0   
2020-08-18 14:43:20          81.520685      0.364660                  103.0   
2020-08-18 14:43:21          85.707790      0.103161                  105.0   
2020-08-18 14:43:22          87.824421      0.007112                  106.0   
2020-08-18 14:43:23          87.587538      0.506943                  106.0   
...                                ...           ...                    ...   
2020-09-13 14:55:57           1.170905      0.024661                  130.0   
2020-09-13 14:55:58           0.569723      0.021386                  130.0   
2020-09-13 14:55:59           0.233745     -0.103366                  129.0   
2020-09-13 14:56:00           0.000000     -0.000000                  130.0   
2020-09-13 14:56:01           0.000000     -0.000000                  130.0   

                     remainder__cadence  remainder__slope  remainder__angle  
                                                                             
2020-08-18 14:43:19                64.0         -0.033870         -0.033857  
2020-08-18 14:43:20                64.0         -0.033571         -0.033559  
2020-08-18 14:43:21                66.0         -0.033223         -0.033210  
2020-08-18 14:43:22                66.0         -0.032908         -0.032896  
2020-08-18 14:43:23                67.0          0.000000          0.000000  
...                                 ...               ...               ...  
2020-09-13 14:55:57                 0.0          0.000000          0.000000  
2020-09-13 14:55:58                 0.0          0.000000          0.000000  
2020-09-13 14:55:59                 0.0          0.000000          0.000000  
2020-09-13 14:56:00                 0.0          0.000000          0.000000  
2020-09-13 14:56:01                 0.0          0.000000          0.000000  

[38254 rows x 6 columns]```

Notice that I had to call `step1` and `step2` on the ColumnTransformer otherwise the naming of the column would look something like `cube__speed_cube` for the first step...
Top answer
1 of 1
1

There is mismatch between the columns that the function transformer returns, and its feature_names_out= name. They are clashing, both in terms of how many features are returned and their name. I firstly made sure that each function transformer only returns the one column (or more) that it's supposed to, and also changed the names to make them consistent. This way, the number of features and names defined in feature_names_out= matches the actual returned data for each transformer.

A second issue was that new features are created in the ColumnTransformer, and an attempt is made to access those variables by other parts of the same ColumnTransformer. This won't work as the column transformer runs things in parallel. What you'd need to do is chain the column transformers in a sequence, such that the next column transformer can access the new variables created from the earlier one. I have made this change to the code.

I made some other minor changes, including forcing transformers to return pandas dataframes rather than numpy arrays.

The new preprocessor:

Its output column names:

Index(['Sex_female', 'Sex_male', 'Embarked_C', 'Embarked_Q', 'Embarked_S',
       'traveling_category_A', 'traveling_category_B', 'traveling_category_C',
       'age_interval_(0, 10]', 'age_interval_(10, 20]',
       'age_interval_(20, 30]', 'age_interval_(30, 40]',
       'age_interval_(40, 50]', 'age_interval_(50, 60]',
       'age_interval_(60, 70]', 'age_interval_(70, 100]', 'Pclass', 'Fare',
       'PassengerId', 'Survived', 'Name', 'Ticket', 'Cabin'],
      dtype='object')

Code:

import pandas as pd
titanic_data = pd.read_csv('../titanic.csv')

from sklearn.pipeline import make_pipeline, FunctionTransformer
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OrdinalEncoder, OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer

from sklearn import set_config
set_config(transform_output='pandas')

#
# Sum relatives
#
def sum_name(function_transformer, feature_names_in):
    return ["total_relatives"]  # feature names out

def sum_relatives(X):
    X_copy = X.copy()
    X_copy['total_relatives'] = X_copy['SibSp'] + X_copy['Parch']
    return X_copy[['total_relatives']]

total_relatives_pipeline = make_pipeline(
    FunctionTransformer(sum_relatives, feature_names_out=sum_name)
)

#
#Categorize travel
#
def cat_travel_name(function_transformer, feature_names_in):
    return ["traveling_category"]  # feature names out

def categorize_travel(X):
    X_copy = X.copy()

    conditions = [
        (X_copy['total_relatives'] == 0),
        (X_copy['total_relatives'] >= 1) & (X_copy['total_relatives'] <= 3),
        (X_copy['total_relatives'] >= 4)
    ]
    categories = ['A', 'B', 'C']

    X_copy['traveling_category'] = np.select(conditions, categories, default='Unknown')
    return X_copy[['traveling_category']]

travel_category_pipeline = make_pipeline(
    FunctionTransformer(categorize_travel, feature_names_out=cat_travel_name)
)

#
# Ordinal encoder
#
class_order = [[1, 2, 3]]

ord_pipeline = make_pipeline(
    OrdinalEncoder(categories=class_order)    
    )

cat_pipeline = make_pipeline(
    SimpleImputer(strategy="most_frequent"),
    OneHotEncoder(handle_unknown="ignore", sparse_output=False)
    )

#Numerical for fare
fare_pipeline = make_pipeline(
        StandardScaler()
    )

#
# Age transformer
#
def interval_name(function_transformer, feature_names_in):
    return ["age_interval"]  # feature names out

def age_transformer(X):
    X_copy = X.copy()
    median_age_by_class = X_copy.groupby('Pclass')['Age'].median().reset_index()
    median_age_by_class.columns = ['Pclass', 'median_age']
    for index, row in median_age_by_class.iterrows():
        class_value = row['Pclass']
        median_age = row['median_age']
        X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = \
            X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age)
    bins = [0, 10, 20, 30, 40, 50, 60, 70, 100]
    X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins)
    return X_copy[['age_interval']]

def age_processor():
    return make_pipeline(
        FunctionTransformer(age_transformer, feature_names_out=interval_name),
        )

#
# Column transformers
#
preprocessing_initial = ColumnTransformer([
        ("ord", ord_pipeline, ['Pclass']),
        ("age_processing", age_processor(), ['Pclass', 'Age']),
        ("num", fare_pipeline, ['Fare']),
        ("total_relatives", total_relatives_pipeline, ['SibSp', 'Parch'])],
        remainder='passthrough',
        verbose_feature_names_out=False
)

preprocessing_travel_category = ColumnTransformer(
    [("travel_category", travel_category_pipeline, ['total_relatives'])],
    remainder='passthrough',
    verbose_feature_names_out=False
)

preprocessing_cat = ColumnTransformer(
    [("cat", cat_pipeline, ['Sex', 'Embarked', 'traveling_category', 'age_interval'])],
    remainder='passthrough',
    verbose_feature_names_out=False
)

#Final transformer
preprocessor = make_pipeline(
    preprocessing_initial,
    preprocessing_travel_category,
    preprocessing_cat
)

#Run
preprocessor.fit(titanic_data)
preprocessor.fit_transform(titanic_data).columns
🌐
Jorisvandenbossche
jorisvandenbossche.github.io › blog › 2018 › 05 › 28 › scikit-learn-columntransformer
Introducing the ColumnTransformer: applying different transformations to different features in a scikit-learn pipeline | Joris Van den Bossche
May 28, 2018 - Here we will create again a ColumnTransformer, but now using more advanced features: we use a mask to select the column subsets based on the dtypes, and we use another pipeline to combine imputation and scaling for the numerical features:
🌐
Medium
medium.com › @pujalabhanuprakash › understanding-the-difference-between-column-transformation-and-pipeline-in-scikit-learn-4b7fb252b52e
Understanding the Difference between Column Transformation and Pipeline in Scikit-Learn” | by BHANUPRAKASH PUJALA | Medium
April 19, 2023 - The column transformer is useful for applying different transformations to specific columns, while the pipeline simplifies the machine learning workflow by combining data pre-processing and model training into a single object.
🌐
Dataschool
mlbook.dataschool.io › ch04.html
4 Improving your workflow with ColumnTransformer and Pipeline – Master Machine Learning with scikit-learn
ColumnTransformer will make it easy to apply different preprocessing steps to different columns. Pipeline will make it easy to apply the same workflow to training data and new data.