I spot my error, I was switching the instantiation of the classes: the custom transformers have to be instantiated inside the ColumnTransformer, while the ColumnTransformer does not have to be instantiated inside the pipeline.
The correct code is the following:
transformation_pipeline = ColumnTransformer([
('adoption', TransformAdoptionFeatures(), features_adoption),
('census', TransformCensusFeaturesRegr(), features_census),
('climate', TransformClimateFeatures(), features_climate),
('soil', TransformSoilFeatures(), features_soil),
('economic', TransformEconomicFeatures(), features_economic)
],
remainder='drop')
full_pipeline_stand = Pipeline([
('transformation', transformation_pipeline),
('scaling', StandardScaler())
])
Answer from giacrava on Stack Overflowscikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
Finally, the preprocessing pipeline is integrated in a full prediction pipeline using Pipeline, together with a simple classification model. # Authors: The scikit-learn developers # SPDX-License-Identifier: BSD-3-Clause · import numpy as np from sklearn.compose import ColumnTransformer from sklearn.datasets import fetch_openml from sklearn.feature_selection import SelectPercentile, chi2 from sklearn.impute import SimpleImputer from sklearn.linear_model import LogisticRegression from sklearn.model_selection import RandomizedSearchCV, train_test_split from sklearn.pipeline import Pipeline from sklearn.preprocessing import OneHotEncoder, StandardScaler np.random.seed(0)
03:12
Passthrough some columns and drop others in a ColumnTransformer ...
02:19
Get the feature names output by a ColumnTransformer - YouTube
03:24
Use ColumnTransformer to apply different preprocessing to different ...
04:07
Seven ways to select columns using ColumnTransformer - YouTube
27:59
How do I encode categorical features using scikit-learn? - YouTube
14:40
Scikit learn Pipelines, Column Transformer and Functional Transformer ...
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.1 documentation
Valid parameter keys can be listed with get_params(). Note that you can directly set the parameters of the estimators contained in transformers of ColumnTransformer. ... Estimator parameters. ... This estimator. ... Transform X separately by each transformer, concatenate results. ... The data to be transformed by subset. ... Parameters to be passed to the underlying transformers’ transform method. You can only pass this if metadata routing is enabled, which you can enable using sklearn.set_config(enable_metadata_routing=True).
ONNX
onnx.ai › sklearn-onnx › auto_examples › plot_complex_pipeline.html
Convert a pipeline with ColumnTransformer - sklearn-onnx 1.20.0 documentation
scikit-learn recently shipped ColumnTransformer which lets the user define complex pipeline where each column may be preprocessed with a different transformer. sklearn-onnx still works in this case as shown in Section Convert complex pipelines.
Amir Masoud Sefidian
sefidian.com › 2022 › 08 › 30 › a-tutorial-on-scikit-learn-pipeline-columntransformer-and-featureunion
A tutorial on Scikit-Learn Pipeline, ColumnTransformer, and FeatureUnion
March 27, 2023 - Once trained, this Pipeline object can be used for smoother deployment. In the previous example, we imputed and encoded all columns the same way. However, we often need to apply different sets of transformers to different groups of columns. For instance, we would want to apply OneHotEncoder to only categorical columns but not to numerical columns. This is where ColumnTransformer comes in.
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-6-building-pipelines › columntransformer-pipelines
ColumnTransformer for Complex Scikit-learn Pipelines
This is where Scikit-learn's ColumnTransformer comes into play. It allows you to apply different transformers to different columns of your input data in parallel. The results from applying each transformer are then concatenated horizontally to form the final transformed dataset, which can then ...
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer.html
Column Transformer with Heterogeneous Data Sources — scikit-learn 1.9.0 documentation
pipeline = Pipeline( [ # Extract subject & body ("subjectbody", subject_body_transformer), # Use ColumnTransformer to combine the subject and body features ( "union", ColumnTransformer( [ # bag-of-words for subject (col 0) ("subject", TfidfVectorizer(min_df=50), 0), # bag-of-words with decomposition for body (col 1) ( "body_bow", Pipeline( [ ("tfidf", TfidfVectorizer()), ("best", PCA(n_components=50, svd_solver="arpack")), ] ), 1, ), # Pipeline for pulling text stats from post's body ( "body_stats", Pipeline( [ ( "stats", text_stats_transformer, ), # returns a list of dicts ( "vect", DictVectorizer(), ), # list of dicts -> feature matrix ] ), 1, ), ], # weight above ColumnTransformer features transformer_weights={ "subject": 0.8, "body_bow": 0.5, "body_stats": 1.0, }, ), ), # Use an SVC classifier on the combined features ("svc", LinearSVC(dual=False)), ], verbose=True, )
DEV Community
dev.to › vikas_gulia › columntransformer-and-pipelines-in-scikit-learn-clean-scalable-and-powerful-preprocessing-3mff
ColumnTransformer and Pipelines in Scikit-Learn: Clean, Scalable, and Powerful Preprocessing - DEV Community
June 26, 2025 - Once you’ve defined your ColumnTransformer, you can chain it with a machine learning model using Pipeline. from sklearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression model = Pipeline([ ('preprocessor', preprocessor), ('classifier', LogisticRegression()) ]) Now your entire workflow—preprocessing + model training—is wrapped in one object.
scikit-learn
scikit-learn.org › 1.5 › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.5.2 documentation
from sklearn.compose import make_column_selector as selector preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, selector(dtype_exclude="category")), ("cat", categorical_transformer, selector(dtype_include="category")), ] ) clf = Pipeline( steps=[("preprocessor", preprocessor), ("classifier", LogisticRegression())] ) clf.fit(X_train, y_train) print("model score: %.3f" % clf.score(X_test, y_test)) clf
Jorisvandenbossche
jorisvandenbossche.github.io › blog › 2018 › 05 › 28 › scikit-learn-columntransformer
Introducing the ColumnTransformer: applying different transformations to different features in a scikit-learn pipeline | Joris Van den Bossche
May 28, 2018 - Now let's show a full example where we integrate the ColumnTransformer in a prediction pipeline. ... from sklearn.model_selection import train_test_split from sklearn.pipeline import make_pipeline from sklearn.linear_model import LogisticRegression from sklearn.impute import SimpleImputer
scikit-learn
scikit-learn.org › stable › modules › compose.html
8.1. Pipelines and composite estimators — scikit-learn 1.9.1 documentation
The ColumnTransformer helps performing different transformations for different columns of the data, within a Pipeline that is safe from data leakage and that can be parametrized.
Medium
medium.com › @vinodkumargr › 11-column-transformer-in-ml-sklearn-column-transformer-in-machine-learning-48479f8cb48f
11] Column Transformer in ML: Sklearn Column Transformer in Machine Learning | by Vinod Kumar G R | Medium
January 31, 2024 - # Combine Transformers using ColumnTransformer preprocessor = ColumnTransformer( transformers=[ ('SI', SimpleImputer(), imputer_features), ('SS', StandardScaler(), scaling_features), ('OHE', OneHotEncoder(sparse=False), categorical_features], remainder='passthrough' ) # In the transformer, # SI is a given name for that transformation # SimpleImputer is a sklearn library for transformation # imputer_features are the columns/features
scikit-learn
scikit-learn.org › 0.21 › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 0.21.3 documentation
This example illustrates how to apply different preprocessing and feature extraction pipelines to different subsets of features, using sklearn.compose.ColumnTransformer.

