What if I have a column called "numpy_array" how can I get that one only passed through?

from sklearn.compose import ColumnTransformer

ct = ColumnTransformer(
    transformers=[
        ('np_array_transform', 'passthrough', ['numpy_array']),
    ],
    remainder='drop',
)
Answer from Sanjar Adilov on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.1 documentation
Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers. ... >>> import numpy as np >>> from sklearn.compose import ColumnTransformer >>> from sklearn.preprocessing import Normalizer >>> ct = ColumnTransformer( ...
🌐
MachineLearningMastery
machinelearningmastery.com › home › blog › how to use the columntransformer for data preparation
How to Use the ColumnTransformer for Data Preparation - MachineLearningMastery.com
December 31, 2020 - For example, the ColumnTransformer below applies a OneHotEncoder to columns 0 and 1. The example below applies a SimpleImputer with median imputing for numerical columns 0 and 1, and SimpleImputer with most frequent imputing to categorical columns ...
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.5.2 documentation
Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers. ... >>> import numpy as np >>> from sklearn.compose import ColumnTransformer >>> from sklearn.preprocessing import Normalizer >>> ct = ColumnTransformer( ...
🌐
Medium
medium.com › @noorfatimaafzalbutt › the-magic-of-column-transformer-in-machine-learning-5783c412e44c
The Magic of Column Transformer in Machine Learning | by Noor Fatima | Medium
June 27, 2024 - The transformer object (SimpleImputer, OrdinalEncoder, OneHotEncoder) for each respective transformation. The list of column names (['fever'], ['cough'], ['gender', 'city']) to which the transformer should be applied.
🌐
Analytics Vidhya
analyticsvidhya.com › home › understanding column transformer and machine learning pipelines
Understanding Column Transformer and Machine Learning Pipelines
October 16, 2024 - import pandas as pd from sklearn.impute import SimpleImputer from sklearn.preprocessing import OneHotEncoder from sklearn.preprocessing import OrdinalEncoder from sklearn.model_selection import train_test_split df = pd.read_csv('covid_toy.csv') X_train,X_test,y_train,y_test = train_test_split(df.drop(columns=['has_covid']),df['has_covid'],test_size=0.2) #create Transformer from sklearn.compose import ColumnTransformer transformer = ColumnTransformer(transformers=[ ('tnf1',SimpleImputer(),['fever']), ('tnf2',OrdinalEncoder(categories=[['Mild','Strong']]),['cough']), ('tnf3',OneHotEncoder(sparse=False,drop='first'),['gender','city']) ],remainder='passthrough') x_train_transform = transformer.fit_transform(X_train) print(x_train_transform[:5])
🌐
YouTube
youtube.com › watch
Passthrough some columns and drop others in a ColumnTransformer - YouTube
In a ColumnTransformer, you can use the strings 'passthrough' and 'drop' in place of a transformer. Useful if you need to passthrough some columns and drop o...
Published: September 30, 2021
🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.compose.ColumnTransformer.html
ColumnTransformer — scikit-learn 1.9.dev0 documentation
Columns of the original feature matrix that are not specified are dropped from the resulting transformed feature matrix, unless specified in the passthrough keyword. Those columns specified with passthrough are added at the right to the output of the transformers. ... >>> import numpy as np >>> from sklearn.compose import ColumnTransformer >>> from sklearn.preprocessing import Normalizer >>> ct = ColumnTransformer( ...
Find elsewhere
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
Top answer
1 of 9
32

It is a bit strange to encode continuous data as Salary. It makes no sense unless you have binned your salary to certain ranges/categories. If I were you I would do:

import pandas as pd
import numpy as np

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder



numeric_features = ['Salary']
numeric_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler())])

categorical_features = ['Age','Country']
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='constant', fill_value='missing')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))])

preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, numeric_features),
        ('cat', categorical_transformer, categorical_features)])

from here you can pipe it with a classifier e.g.

clf = Pipeline(steps=[('preprocessor', preprocessor),
                  ('classifier', LogisticRegression(solver='lbfgs'))])  
                  

Use it as so:

clf.fit(X_train,y_train)

this will apply the preprocessor and then pass transformed data to the predictor.

Updates:

If we want to select data types on the fly, we can modify our preprocessor to use column selector by data dtypes:

from sklearn.compose import make_column_selector as selector

preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, selector(dtype_include="numeric")),
        ('cat', categorical_transformer, selector(dtype_include="category"))])

Using GridSearch

param_grid = {
    'preprocessor__num__imputer__strategy': ['mean', 'median'],
    'classifier__C': [0.1, 1.0, 10, 100],
    'classifier__solver': ['lbfgs', 'sag'],
}

grid_search = GridSearchCV(clf, param_grid, cv=10)
grid_search.fit(X_train,y_train)

Getting names of features


preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, selector(dtype_include="numeric")),
        ('cat', categorical_transformer, selector(dtype_include="category"))],
    verbose_feature_names_out=False, # added this line
)

# now we can access feature names with

clf[:-1]. get_feature_names_out() # step before estimator

2 of 9
13

I think the poster is not trying to transform the Age and Salary. From the documentation (https://scikit-learn.org/stable/modules/generated/sklearn.compose.make_column_transformer.html), you ColumnTransformer (and make_column_transformer) only columns specified in the transformer (i.e., [0] in your example). You should set remainder="passthrough" to get the rest of the columns. In other words:

preprocessor = make_column_transformer( (OneHotEncoder(),[0]),remainder="passthrough")
x = preprocessor.fit_transform(x)
🌐
YouTube
youtube.com › watch
Use ColumnTransformer to apply different preprocessing to different columns - YouTube
Use ColumnTransformer to apply different preprocessing to different columns:- select from DataFrame columns by name- passthrough or drop unspecified columnsR...
Published: October 13, 2020
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.compose.make_column_transformer.html
make_column_transformer — scikit-learn 1.9.1 documentation
This is a shorthand for the ColumnTransformer constructor; it does not require, and does not permit, naming the transformers. Instead, they will be given names automatically based on their types. It also does not allow weighting with transformer_weights. Read more in the User Guide. ... Tuples of the form (transformer, columns) specifying the transformer objects to be applied to subsets of the data. transformer{‘drop’, ‘passthrough...
🌐
Codefinity
codefinity.com › courses › v2 › a65bbc96-309e-4df9-a790-a1eb8c815a1c › 5d22c5e7-b4d7-42a1-9273-04eb67cc094a › a6fecd34-e163-40b8-b681-45e1fdef30c8
Learn ColumnTransformer | Pipelines
For example, consider the exams.csv file. It contains several nominal columns ('gender', 'race/ethnicity', 'lunch', 'test preparation course') and one ordinal column, 'parental level of education'.
🌐
Jorisvandenbossche
jorisvandenbossche.github.io › blog › 2018 › 05 › 28 › scikit-learn-columntransformer
Introducing the ColumnTransformer: applying different transformations to different features in a scikit-learn pipeline | Joris Van den Bossche
May 28, 2018 - Further, the ColumnTransformer allows you to specify whether to drop or pass through other columns that were not specified. See the development docs for more details. This is new functionality in scikit-learn, so you are very welcome to try out the development version, experiment with it in your use cases, and provide feedback! I am sure there are ways to further improve this functionality (the PR) The rest of the post shows a more complete example of using the ColumnTransformer in a scikit-learn pipeline.
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › main › sklearn › compose › _column_transformer.py
scikit-learn/sklearn/compose/_column_transformer.py at main · scikit-learn/scikit-learn
Those columns specified with `passthrough` are added at the right to the output of the transformers. · Examples · -------- >>> import numpy as np · >>> from sklearn.compose import ColumnTransformer · >>> from sklearn.preprocessing import Normalizer ·
Author: scikit-learn
🌐
Snowflake Documentation
docs.snowflake.com › en › developer-guide › snowpark-ml › reference › 1.3.1 › api › modeling › snowflake.ml.modeling.compose.ColumnTransformer
snowflake.ml.modeling.compose.ColumnTransformer | Snowflake Documentation
set_input_cols(input_cols: Optional[Union[str, Iterable[str]]]) → ColumnTransformer¶ · Input columns setter. ... Label column setter. ... Output columns setter. ... Set the parameters of this transformer. The method works on simple transformers as well as on nested objects. The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object. ... SnowflakeMLException: Invalid parameter keys. set_passthrough_cols(passthrough_cols: Optional[Union[str, Iterable[str]]]) → Base¶
🌐
Kaggle
kaggle.com › code › ksvmuralidhar › columntransformer-pipeline-simplified
ColumnTransformer & Pipeline Simplified
October 7, 2020 - INTRODUCTION TO COLUMNTRANSFORMER · This Notebook has been released under the Apache 2.0 open source license. Input1 file · arrow_right_alt · Output0 files · arrow_right_alt · Logs18.5 second run - successful · arrow_right_alt · Comments8 comments ·
🌐
GitHub
github.com › scikit-learn › scikit-learn › issues › 25422
``ColumnTransformer`` not honoring passthrough on `transform` with pandas when passthroughs not present in `fit` data · Issue #25422 · scikit-learn/scikit-learn
January 17, 2023 - Ideally when remainder='passthrough', I should be able to pass in any dataframe to transform provided it has the required columns to transform and nothing should get dropped regardless of whether or not the passthroughs were present in the fit data. If this is intentional to avoid unexpected columns during the transform, then I suggest making this clear in the docstring. import pandas as pd import sklearn from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder # Make a mixed-type dataset including passthrough columns df = pd.DataFrame( { "a
Author: scikit-learn
🌐
Towards Data Science
towardsdatascience.com › home › latest › improve your data preprocessing with columntransformer and pipelines
Improve Your Data Preprocessing with ColumnTransformer and Pipelines | Towards Data Science
March 5, 2025 - The replacement operation only occurred in the column specified, while the remained stayed untouched (as specified by remainder="passthrough" ). The pandas DataFrame was also replaced by a Numpy Array, as this is the default behavior of Sklearn’s transformers. Let’s see a more complex example.
🌐
Sktime
sktime.org › en › stable › api_reference › auto_generated › sktime.transformations.panel.compose.ColumnTransformer.html
ColumnTransformer — sktime documentation
By default, only the specified ... columns are dropped. (default of "drop"). By specifying remainder="passthrough", all remaining columns that were not specified in transformations will be automatically passed through....