The preferred practice is not to use StandardScaler on one-hot-encoded columns. The first example below demonstrates the application of OHE on the categorical variables and StandardScaler on the numeric columns. The second example, shows the sequential application of OHE on selected columns and StandardScaler on all columns, but this is not recommended.

Example_1:

import numpy as np
import pandas as pd

from sklearn.preprocessing import OneHotEncoder
from sklearn.preprocessing import StandardScaler
from sklearn.compose import ColumnTransformer
from sklearn.compose import make_column_transformer
from sklearn.pipeline import Pipeline

df = pd.DataFrame({'Cat_Var': np.random.choice(['a', 'b'], size=5),
                   'Num_Var': np.arange(5)})

cat_cols = ['Cat_Var']
num_cols = ['Num_Var']

col_transformer = make_column_transformer(
        (OneHotEncoder(), cat_cols),
        remainder=StandardScaler())

X = col_transformer.fit_transform(df)

Output:

df
Out[57]: 
  Cat_Var  Num_Var
0       b        0
1       a        1
2       b        2
3       a        3
4       a        4

X
Out[58]: 
array([[ 0.        ,  1.        , -1.41421356],
       [ 1.        ,  0.        , -0.70710678],
       [ 0.        ,  1.        ,  0.        ],
       [ 1.        ,  0.        ,  0.70710678],
       [ 1.        ,  0.        ,  1.41421356]])

Example 2:

col_transformer_2 = ColumnTransformer(
        [('cat_transform', OneHotEncoder(), cat_cols)],
        remainder='passthrough'
        )

pipe = Pipeline(
        [
         ('col_tranform', col_transformer_2),
         ('standard_scaler', StandardScaler())
         ])

X_2 = pipe.fit_transform(df)

Output:

X_2
Out[62]: 
array([[-1.22474487,  1.22474487, -1.41421356],
       [ 0.81649658, -0.81649658, -0.70710678],
       [-1.22474487,  1.22474487,  0.        ],
       [ 0.81649658, -0.81649658,  0.70710678],
       [ 0.81649658, -0.81649658,  1.41421356]])
Answer from KRKirov on Stack Overflow
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.9.1 documentation
numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
🌐
GeeksforGeeks
geeksforgeeks.org › prediction-using-columntransformer-onehotencoder-and-pipeline
Prediction using ColumnTransformer, OneHotEncoder and Pipeline | GeeksforGeeks
July 17, 2020 - In this tutorial, we'll predict insurance premium costs for each customer having various features, using ColumnTransformer, OneHotEncoder and Pipeline.
Discussions

python - How to use sklearn Column Transformer? - Stack Overflow
I'm trying to convert categorical ... and was able to convert the categorical value. But i'm getting warning like OneHotEncoder 'categorical_features' keyword is deprecated "use the ColumnTransformer instead." So how i can use ColumnTransformer to achieve same result ... More on stackoverflow.com
🌐 stackoverflow.com
July 2, 2019
python - sklearn OneHotEncoder with ColumnTransformer resulting in sparse Matrix in place of creating dummies - Stack Overflow
I am trying to convert categorical value to integer using OneHotEncoder and ColumnTransformer. My understanding is it should create dummies for category columns like pd.get_dummies. My file is havi... More on stackoverflow.com
🌐 stackoverflow.com
Scikit-learn ColumnTransformer + OneHotEncoder - Stack Overflow
scikit-learn: ColumnTransformer and OneHotEncoder – how to err out for all new categorical levels across all fields? More on stackoverflow.com
🌐 stackoverflow.com
python - ML ColumnTransformer OneHotEncoder - Stack Overflow
When converting categorical data in first column of my dataframe I am getting strange behavior of ColumnTransformer with OneHotEncoder. the behavior occurs when I add one row to my csv file. the in... More on stackoverflow.com
🌐 stackoverflow.com
Top answer
1 of 3
5

The preferred practice is not to use StandardScaler on one-hot-encoded columns. The first example below demonstrates the application of OHE on the categorical variables and StandardScaler on the numeric columns. The second example, shows the sequential application of OHE on selected columns and StandardScaler on all columns, but this is not recommended.

Example_1:

import numpy as np
import pandas as pd

from sklearn.preprocessing import OneHotEncoder
from sklearn.preprocessing import StandardScaler
from sklearn.compose import ColumnTransformer
from sklearn.compose import make_column_transformer
from sklearn.pipeline import Pipeline

df = pd.DataFrame({'Cat_Var': np.random.choice(['a', 'b'], size=5),
                   'Num_Var': np.arange(5)})

cat_cols = ['Cat_Var']
num_cols = ['Num_Var']

col_transformer = make_column_transformer(
        (OneHotEncoder(), cat_cols),
        remainder=StandardScaler())

X = col_transformer.fit_transform(df)

Output:

df
Out[57]: 
  Cat_Var  Num_Var
0       b        0
1       a        1
2       b        2
3       a        3
4       a        4

X
Out[58]: 
array([[ 0.        ,  1.        , -1.41421356],
       [ 1.        ,  0.        , -0.70710678],
       [ 0.        ,  1.        ,  0.        ],
       [ 1.        ,  0.        ,  0.70710678],
       [ 1.        ,  0.        ,  1.41421356]])

Example 2:

col_transformer_2 = ColumnTransformer(
        [('cat_transform', OneHotEncoder(), cat_cols)],
        remainder='passthrough'
        )

pipe = Pipeline(
        [
         ('col_tranform', col_transformer_2),
         ('standard_scaler', StandardScaler())
         ])

X_2 = pipe.fit_transform(df)

Output:

X_2
Out[62]: 
array([[-1.22474487,  1.22474487, -1.41421356],
       [ 0.81649658, -0.81649658, -0.70710678],
       [-1.22474487,  1.22474487,  0.        ],
       [ 0.81649658, -0.81649658,  0.70710678],
       [ 0.81649658, -0.81649658,  1.41421356]])
2 of 3
2

Making some additions to KRKirov's answer as it might be useful as well.

As you are using make_column_transformer, it do preprocessing your features in the given order. So the remainder (all features that are untouched during processing) should come at the end.

The problem with your code is that you pass the remainder parameter at the middle, so all features that remained from process goes to there. And you can't do processing after that. So first, do all the special processing first then do processing with other features with remainder parameter.

Here I will explain it with the codes.

(1) importing

import pandas as pd
from sklearn.compose import make_column_transformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
onhe = OneHotEncoder()
scaler = StandardScaler()

(2) I will create basic DataFrame

df = pd.DataFrame({'sex':['m', 'f','f','m'],
                       'age':[45,25,10,31], 
                       'married':['y','y','n','y'],
                       'salary':[1000,300,370,500],
                       'child':[5, 1,0,3]})
print(df)

(3) Let's say we want to do encoding for sex and married columns, standard scaling for age and salary columns and leave the child column as it is.

transforming = make_column_transformer((onhe,['sex','married']),
 (scaler,['age', 'salary']),
    remainder = 'passthrough')
     
processed_df = transforming.fit_transform(df)
print(processed_df)

Note that remainder is being assigned at the end of the process. What is more, if you want to do scaler in all remaining features ('age','salary', 'child'), then you can use:

transforming_1 = make_column_transformer((onhe, ['sex', 'married']), remainder = scaler)
processed_df_1 = transforming_1.fit_transform(df)
print(processed_df_1)

It will encode two given columns then do StandardScaling for all remaining columns.

And when it comes to your situation, your code (from which you got an error), should look like this:

trans_cols= make_column_transformer((OneHotEncoder(),['job', 'marital', 'education', 'default','housing','loan','contact','month','poutcome']),(StandardScaler(),['age', 'job', 'marital', 'education', 'default', 'balance','housing','loan', 'contact', 'month', 'duration','campaign', 'pdays', 'previous','poutcome']),remainder='passthrough')
🌐
Towards Data Science
towardsdatascience.com › home › latest › columntransformer in scikit for labelencoding and onehotencoding in machine learning
ColumnTransformer in SciKit for LabelEncoding and OneHotEncoding in Machine Learning | Towards Data Science
January 23, 2025 - The developers of the library might ... library called the ColumnTransformer, which will basically combine LabelEncoding and OneHotEncoding into just one line of code....
🌐
datagy
datagy.io › home › python posts › one-hot encoding in scikit-learn with onehotencoder
One-Hot Encoding in Scikit-Learn with OneHotEncoder • datagy
April 14, 2024 - The function generates ColumnTransformer objects for you and handles the transformations. This allows us to simply pass in a list of transformations we want to do and the columns to which we want to apply them.
🌐
Kaggle
kaggle.com › code › gurdit559 › column-transformer-one-hot-encoding
Column Transformer-One hot encoding | Kaggle
September 12, 2019 - Explore and run AI code with Kaggle Notebooks | Using data from [Private Datasource]
Find elsewhere
Top answer
1 of 9
32

It is a bit strange to encode continuous data as Salary. It makes no sense unless you have binned your salary to certain ranges/categories. If I were you I would do:

import pandas as pd
import numpy as np

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder



numeric_features = ['Salary']
numeric_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler())])

categorical_features = ['Age','Country']
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='constant', fill_value='missing')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))])

preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, numeric_features),
        ('cat', categorical_transformer, categorical_features)])

from here you can pipe it with a classifier e.g.

clf = Pipeline(steps=[('preprocessor', preprocessor),
                  ('classifier', LogisticRegression(solver='lbfgs'))])  
                  

Use it as so:

clf.fit(X_train,y_train)

this will apply the preprocessor and then pass transformed data to the predictor.

Updates:

If we want to select data types on the fly, we can modify our preprocessor to use column selector by data dtypes:

from sklearn.compose import make_column_selector as selector

preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, selector(dtype_include="numeric")),
        ('cat', categorical_transformer, selector(dtype_include="category"))])

Using GridSearch

param_grid = {
    'preprocessor__num__imputer__strategy': ['mean', 'median'],
    'classifier__C': [0.1, 1.0, 10, 100],
    'classifier__solver': ['lbfgs', 'sag'],
}

grid_search = GridSearchCV(clf, param_grid, cv=10)
grid_search.fit(X_train,y_train)

Getting names of features


preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, selector(dtype_include="numeric")),
        ('cat', categorical_transformer, selector(dtype_include="category"))],
    verbose_feature_names_out=False, # added this line
)

# now we can access feature names with

clf[:-1]. get_feature_names_out() # step before estimator

2 of 9
13

I think the poster is not trying to transform the Age and Salary. From the documentation (https://scikit-learn.org/stable/modules/generated/sklearn.compose.make_column_transformer.html), you ColumnTransformer (and make_column_transformer) only columns specified in the transformer (i.e., [0] in your example). You should set remainder="passthrough" to get the rest of the columns. In other words:

preprocessor = make_column_transformer( (OneHotEncoder(),[0]),remainder="passthrough")
x = preprocessor.fit_transform(x)
🌐
Kaggle
kaggle.com › code › pritampatil004 › label-encoder-one-hot-encoder-columntransformer
Label Encoder, One Hot Encoder, ColumnTransformer
Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
Stack Overflow
stackoverflow.com › questions › 55217933 › scikit-learn-columntransformer-onehotencoder
Scikit-learn ColumnTransformer + OneHotEncoder - Stack Overflow
from sklearn.preprocessing import OrdinalEncoder from sklearn.preprocessing import OneHotEncoder from sklearn.compose import ColumnTransformer import numpy as np X = [['male', 0, 3], ['male', 1, 0], ['female', 2, 1], ['female', 0, 2]] # 字符串编码为整形 sex_enc = OrdinalEncoder(dtype = np.int) # 独热编码 one_hot_enc = OneHotEncoder(sparse=False, handle_unknown='ignore', dtype=np.int) # 对第0列的字符串做整形转换, 然后对所有列做one-hot col_transformer = ColumnTransformer(transformers = [('sex_enc', sex_enc, [0]), ('one_hot_enc', one_hot_enc, [0])]) # 训练编码 col_transformer.fit(X) X_trans = col_transformer.transform(X) print(X_trans)
🌐
ONNX
onnx.ai › sklearn-onnx › auto_examples › plot_complex_pipeline.html
Convert a pipeline with ColumnTransformer - sklearn-onnx 1.20.0 documentation
Pipeline(steps=[('preprocessor', ColumnTransformer(transformers=[('num', Pipeline(steps=[('imputer', SimpleImputer(strategy='median')), ('scaler', StandardScaler())]), ['age', 'fare']), ('cat', Pipeline(steps=[('onehot', OneHotEncoder(handle_unknown='ignore'))]), ['embarked', 'sex', 'pclass'])])), ('classifier', LogisticRegression())])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
🌐
Stack Overflow
stackoverflow.com › questions › 78274904 › ml-columntransformer-onehotencoder
python - ML ColumnTransformer OneHotEncoder - Stack Overflow
The latter is the display for a scipy sparse array. OneHotEncoder produces sparse arrays as output, and ColumnTransformer uses a sparse array when the overall density of values of its output is below the parameter sparse_threshold.
🌐
GitHub
github.com › scikit-learn › scikit-learn › discussions › 27830
Sklearn OneHotEncoder not encoding correctly with Column Transformer · scikit-learn/scikit-learn · Discussion #27830
November 22, 2023 - train_df = pd.read_csv('train.csv') x_df = pd.DataFrame(np.array(train_df['MSZoning']).reshape(-1, 1)) cols_mapper = dict() cols_mapper[0] = 'MSZoning' x_df = x_df.rename(columns=cols_mapper) ct = make_column_transformer( (categorical_imputer, ['MSZoning']), (OneHotEncoder(handle_unknown='ignore'), ['MSZoning']), remainder='passthrough') enc_x = ct.fit_transform(x_df) enc_x
Author: scikit-learn
🌐
Analytics Vidhya
analyticsvidhya.com › home › understanding column transformer and machine learning pipelines
Understanding Column Transformer and Machine Learning Pipelines
October 16, 2024 - #1st Imputation Transformer trf1 = ColumnTransformer([ ('impute_age',SimpleImputer(),[2]), ('impute_embarked',SimpleImputer(strategy='most_frequent'),[6]) ],remainder='passthrough') #2nd One Hot Encoding trf2 = ColumnTransformer([ ('ohe_sex_embarked', OneHotEncoder(sparse=False, handle_unknown='ignore'),[1,6]) ], remainder='passthrough') #3rd Scaling trf3 = ColumnTransformer([ ('scale', MinMaxScaler(), slice(0,10)) ]) #4th Feature selection trf4 = SelectKBest(score_func=chi2,k=8) #5th Model trf5 = DecisionTreeClassifier()
🌐
Bait509-ubc
bait509-ubc.github.io › BAIT509 › lectures › lecture5.html
5. Preprocessing Categorical Features and Column Transformer — BAIT 509<br>Business Applications of Machine Learning
Explain handle_unknown="ignore" hyperparameter of scikit-learn’s OneHotEncoder. Use the scikit-learn ColumnTransformer function to implement preprocessing functions such as MinMaxScaler and OneHotEncoder to numeric and categorical features simultaneously.
🌐
Stack Overflow
stackoverflow.com › questions › 60762985 › sklearn-onehotencoding-with-columntransformer
python - sklearn OneHotEncoding with ColumnTransformer - Stack Overflow
March 20, 2020 - l have categorical data on column 8 of my dataset. l wish to encode this data and l am using ColumnTransformer. The first time l tried to use this method l used the code: Copyfrom sklearn.preprocessing import OneHotEncoder from sklearn.preprocessing import LabelEncoder from sklearn.compose import ColumnTransformer #encoding categorical data for dept column(independent variables) ct = ColumnTransformer(transformers=[('one_hot_encoder',OneHotEncoder(categories='auto'), [0])], remainder='passthrough') X = np.array(ct.fit_transform(X), dtype=np.float)
🌐
scikit-learn
scikit-learn.org › 1.5 › auto_examples › compose › plot_column_transformer_mixed_types.html
Column Transformer with Mixed Types — scikit-learn 1.5.2 documentation
numeric_features = ["age", "fare"] numeric_transformer = Pipeline( steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())] ) categorical_features = ["embarked", "sex", "pclass"] categorical_transformer = Pipeline( steps=[ ("encoder", OneHotEncoder(handle_unknown="ignore")), ("selector", SelectPercentile(chi2, percentile=50)), ] ) preprocessor = ColumnTransformer( transformers=[ ("num", numeric_transformer, numeric_features), ("cat", categorical_transformer, categorical_features), ] )
🌐
Jorisvandenbossche
jorisvandenbossche.github.io › blog › 2018 › 05 › 28 › scikit-learn-columntransformer
Introducing the ColumnTransformer: applying different transformations to different features in a scikit-learn pipeline | Joris Van den Bossche
May 28, 2018 - Here we will create again a ColumnTransformer, but now using more advanced features: we use a mask to select the column subsets based on the dtypes, and we use another pipeline to combine imputation and scaling for the numerical features: ... preprocess = make_column_transformer( (numerical_features, make_pipeline(SimpleImputer(), StandardScaler())), (categorical_features, OneHotEncoder()))