I spot my error, I was switching the instantiation of the classes: the custom transformers have to be instantiated inside the ColumnTransformer, while the ColumnTransformer does not have to be instantiated inside the pipeline.
The correct code is the following:
transformation_pipeline = ColumnTransformer([
('adoption', TransformAdoptionFeatures(), features_adoption),
('census', TransformCensusFeaturesRegr(), features_census),
('climate', TransformClimateFeatures(), features_climate),
('soil', TransformSoilFeatures(), features_soil),
('economic', TransformEconomicFeatures(), features_economic)
],
remainder='drop')
full_pipeline_stand = Pipeline([
('transformation', transformation_pipeline),
('scaling', StandardScaler())
])
Answer from giacrava on Stack Overflowscikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
Analytics Vidhya
analyticsvidhya.com › home › understanding column transformer and machine learning pipelines
Understanding Column Transformer and Machine Learning Pipelines
October 16, 2024 - When working with Pipelines While creating a Column transformer it’s suggested to pass the index of columns rather than its name because after transformation it’s converted into Numpy Array and the array does not have any column names. #1st Imputation Transformer trf1 = ColumnTransformer([ ('impute_age',SimpleImputer(),[2]), ('impute_embarked',SimpleImputer(strategy='most_frequent'),[6]) ],remainder='passthrough') #2nd One Hot Encoding trf2 = ColumnTransformer([ ('ohe_sex_embarked', OneHotEncoder(sparse=False, handle_unknown='ignore'),[1,6]) ], remainder='passthrough') #3rd Scaling trf3 = ColumnTransformer([ ('scale', MinMaxScaler(), slice(0,10)) ]) #4th Feature selection trf4 = SelectKBest(score_func=chi2,k=8) #5th Model trf5 = DecisionTreeClassifier()
12:19
Learn how practically use Pipeline & column transformers in Machine ...
02:19
Get the feature names output by a ColumnTransformer - YouTube
- YouTube
03:24
Use ColumnTransformer to apply different preprocessing to different ...
27:20
Using Column Transformer and Pipeline to handle data with missing ...
27:59
How do I encode categorical features using scikit-learn? - YouTube
ONNX
onnx.ai › sklearn-onnx › auto_examples › plot_complex_pipeline.html
Convert a pipeline with ColumnTransformer - sklearn-onnx 1.20.0 documentation
scikit-learn recently shipped ColumnTransformer which lets the user define complex pipeline where each column may be preprocessed with a different transformer.
APXML
apxml.com › courses › getting-started-with-scikit-learn › chapter-6-building-pipelines › columntransformer-pipelines
ColumnTransformer for Complex Scikit-learn Pipelines
This is where Scikit-learn's ColumnTransformer comes into play. It allows you to apply different transformers to different columns of your input data in parallel. The results from applying each transformer are then concatenated horizontally to form the final transformed dataset, which can then ...
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - We can then use these lists in the ColumnTransformer to one hot encode the categorical variables, which should just be the first column. We can also use the list of numerical columns to normalize the remaining data. Next, we can define our SVR model and define a Pipeline that first uses the ColumnTransformer, then fits the model on the prepared dataset.
Medium
medium.com › @abhaysingh71711 › building-smarter-ml-pipelines-with-column-transformers-895904e97254
Building Smarter ML Pipelines with Column Transformers | by Abhay singh | Medium
September 24, 2024 - When working with Pipelines While creating a Column transformer it’s suggested to pass the index of columns rather than its name because after transformation it’s converted into Numpy Array and the array does not have any column names. #1st Imputation Transformer trf1 = ColumnTransformer([ ('impute_age',SimpleImputer(),[2]), ('impute_embarked',SimpleImputer(strategy='most_frequent'),[6]) ],remainder='passthrough') #2nd One Hot Encoding trf2 = ColumnTransformer([ ('ohe_sex_embarked', OneHotEncoder(sparse=False, handle_unknown='ignore'),[1,6]) ], remainder='passthrough') #3rd Scaling trf3 = ColumnTransformer([ ('scale', MinMaxScaler(), slice(0,10)) ]) #4th Feature selection trf4 = SelectKBest(score_func=chi2,k=8) #5th Model trf5 = DecisionTreeClassifier()
Amir Masoud Sefidian
sefidian.com › 2022 › 08 › 30 › a-tutorial-on-scikit-learn-pipeline-columntransformer-and-featureunion
A tutorial on Scikit-Learn Pipeline, ColumnTransformer, and FeatureUnion
March 27, 2023 - Once trained, this Pipeline object can be used for smoother deployment. In the previous example, we imputed and encoded all columns the same way. However, we often need to apply different sets of transformers to different groups of columns. For instance, we would want to apply OneHotEncoder to only categorical columns but not to numerical columns. This is where ColumnTransformer comes in.
Jorisvandenbossche
jorisvandenbossche.github.io › blog › 2018 › 05 › 28 › scikit-learn-columntransformer
Introducing the ColumnTransformer: applying different transformations to different features in a scikit-learn pipeline | Joris Van den Bossche
May 28, 2018 - Here we will create again a ColumnTransformer, but now using more advanced features: we use a mask to select the column subsets based on the dtypes, and we use another pipeline to combine imputation and scaling for the numerical features:
