To center the data (make it have zero mean and unit standard error), you subtract the mean and then divide the result by the standard deviation:

You do that on the training set of the data. But then you have to apply the same transformation to your test set (e.g. in cross-validation), or to newly obtained examples before forecasting. But you have to use the exact same two parameters and (values) that you used for centering the training set.

Hence, every scikit-learn's transform's fit() just calculates the parameters (e.g. and in case of StandardScaler) and saves them as an internal object's state. Afterwards, you can call its transform() method to apply the transformation to any particular set of examples.

fit_transform() joins these two steps and is used for the initial fitting of parameters on the training set , while also returning the transformed . Internally, the transformer object just calls first fit() and then transform() on the same data.

Answer from K3---rnc on Stack Exchange
Top answer
1 of 10
251

To center the data (make it have zero mean and unit standard error), you subtract the mean and then divide the result by the standard deviation:

You do that on the training set of the data. But then you have to apply the same transformation to your test set (e.g. in cross-validation), or to newly obtained examples before forecasting. But you have to use the exact same two parameters and (values) that you used for centering the training set.

Hence, every scikit-learn's transform's fit() just calculates the parameters (e.g. and in case of StandardScaler) and saves them as an internal object's state. Afterwards, you can call its transform() method to apply the transformation to any particular set of examples.

fit_transform() joins these two steps and is used for the initial fitting of parameters on the training set , while also returning the transformed . Internally, the transformer object just calls first fit() and then transform() on the same data.

2 of 10
36

The following explanation is based on fit_transform of Imputer class, but the idea is the same for fit_transform of other scikit_learn classes like MinMaxScaler.


transform replaces the missing values with a number. By default this number is the means of columns of some data that you choose. Consider the following example:

imp = Imputer()
# calculating the means
imp.fit([
         [1,      3], 
         [np.nan, 2], 
         [8,      5.5]
        ])

Now the imputer have learned to use a mean for the first column and mean for the second column when it gets applied to a two-column data:

X = [[np.nan, 11], 
     [4,      np.nan], 
     [8,      2],
     [np.nan, 1]]
print(imp.transform(X))

we get

[[4.5, 11], 
 [4, 3.5],
 [8, 2],
 [4.5, 1]]

So by fit the imputer calculates the means of columns from some data, and by transform it applies those means to some data (which is just replacing missing values with the means). If both these data are the same (i.e. the data for calculating the means and the data that means are applied to) you can use fit_transform which is basically a fit followed by a transform.

Now your questions:

Why we might need to transform data?

"For various reasons, many real world datasets contain missing values, often encoded as blanks, NaNs or other placeholders. Such datasets however are incompatible with scikit-learn estimators which assume that all values in an array are numerical" (source)

What does it mean fitting model on training data and transforming to test data?

The fit of an imputer has nothing to do with fit used in model fitting. So using imputer's fit on training data just calculates means of each column of training data. Using transform on test data then replaces missing values of test data with means that were calculated from training data.

Discussions

python - sklearn.impute SimpleImputer: why does transform() need fit_transform() first? - Stack Overflow
NotFittedError: This SimpleImputer instance is not fitted yet. Call 'fit' with appropriate arguments before using this method. ... If you call the fit_transform only, there's no need to call transform at all. fit_transform first fits the data and then transforms it. More on stackoverflow.com
🌐 stackoverflow.com
machine learning - What is the difference between "fit_transform" and "transform" methods when using "SimpleImputer"? - Data Science Stack Exchange
I have following code, I am not able to understand the difference between use of fit_transform() and transform() method in this case. In the following code, when I am using fit_transform(X_valid) More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
machine learning - What is difference between fit, transform and fit_transform in python when using sklearn? - Stack Overflow
You need to think of X as a matrix that you have to transform in order to have no more NaN (missing values). Refer to the documentation for more information. ... Save this answer. ... Show activity on this post. Your comments tell you the difference. It is saying that if you don't use imputer.fit, ... More on stackoverflow.com
🌐 stackoverflow.com
machine learning - Should I use transformer.fit_transform(X_test, y_test) or not? - Cross Validated
My question more specifically is when I want to use transformers in the data processing step. Let's say I have a dataset already splitted into train/test and there are nan values ​​in both and in the same feature. I chose to use the SimpleImputer(strategy='mean') class to fill these fields. In all the websites I read and videos I watched it was stated that I should use fit... More on stats.stackexchange.com
🌐 stats.stackexchange.com
🌐
Kaggle
kaggle.com › questions-and-answers › 58368
fit, transform and fit_transform | Kaggle
First simple imputer captured the mean by fit() function in line 2 Then it is used to transform the X, it applied whatever it learned from line 2 If we use fit_transform on X, it will capture the mean of X and will fill the NaN with mean of X
🌐
Kaggle
kaggle.com › getting-started › 159827
Using .fit_transform() vs .transform() | Kaggle
What is the Difference between .fit_transform() and .transform()
🌐
CloudxLab
cloudxlab.com › assessment › displayslide › 7554 › simpleimputer
SimpleImputer | Automated hands-on| CloudxLab
Otherwise, it is advised to use the fit_transform() method. The transform method returns a numpy array. So, we convert it back to a pandas DataFrame. ... Import SimpleImputer from sklearn.impute.
🌐
Stack Exchange
datascience.stackexchange.com › questions › 66473 › what-is-the-difference-between-fit-transform-and-transform-methods-when-usin
machine learning - What is the difference between "fit_transform" and "transform" methods when using "SimpleImputer"? - Data Science Stack Exchange
In the following code, when I am using fit_transform(X_valid) instead of transform(X_valid) it gives me different outputs. from sklearn.impute import SimpleImputer # Imputation my_imputer = SimpleImputer() imputed_X_train = pd.DataFrame(my_imputer.fit_transform(X_train)) imputed_X_valid = pd.DataFrame(my_imputer.transform(X_valid)) machine-learning ·
Find elsewhere
🌐
Kaggle
kaggle.com › c › home-data-for-ml-course › discussion › 133076
Housing Prices Competition for Kaggle Learn Users
March 1, 2020 - Apply what you learned in the Machine Learning course on Kaggle Learn alongside others in the course.
🌐
Towards Data Science
towardsdatascience.com › home › latest › fit vs. transform in scikit libraries for machine learning
Fit vs. Transform in SciKit libraries for Machine Learning | Towards Data Science
January 9, 2025 - This step of calculating that value is called the fit() method. Next, the transform() method will just replace the NaNs in the column with the newly calculated value, and return the new dataset.
🌐
Kaggle
kaggle.com › learn-forum › 159827
Checking your browser - reCAPTCHA
June 22, 2020 - Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
Towards Data Science
towardsdatascience.com › home › data science › what and why behind fit_transform() vs transform() in scikit-learn !
What and why behind fit_transform() vs transform() in scikit-learn ! | Towards Data Science
August 25, 2020 - The fit method is calculating the mean and variance of each of the features present in our data. The transform method is transforming all the features using the respective mean and variance.
🌐
tutorialpedia
tutorialpedia.org › blog › sklearn-impute-simpleimputer-why-does-transform-need-fit-transform-first
Why Does sklearn.impute SimpleImputer's transform() Require fit_transform() First? Explaining the NotFittedError
Scikit-learn’s SimpleImputer is a go-to tool for this task, allowing you to fill missing values with statistics like mean, median, or mode. But if you’ve ever tried to call transform() directly on a SimpleImputer instance without first running fit() or fit_transform(), you’ve likely encountered a frustrating error: NotFittedError: This SimpleImputer instance is not fitted yet.
🌐
Physics Forums
physicsforums.com › other sciences › programming and computer science
Fit_transform() vs. transform() • Physics Forums
December 2, 2017 - ... Some participants explain that fit_transform() combines the actions of fitting and transforming the data, while transform() is used to apply the transformation to new data based on the previously fitted model.
🌐
LinkedIn
linkedin.com › pulse › difference-between-fit-transform-fittransform-predict-aman-agarwal
Difference between fit(), transform(), fit_transform() and predict() methods in scikit-learn
October 7, 2018 - fit() - It is used for calculating the initial filling of parameters on the training data (like mean of the column values) and saves them as an internal objects state · transform() - Use the above calculated values and return modified training data
🌐
YouTube
youtube.com › watch
fit vs transform vs fit_transform | fit vs fit_transform | fit and fit_transofrm in sklearn - YouTube
fit vs transform vs fit_transform | fit vs fit_transform | fit and fit_transofrm in sklearn#machinelearning #datascience #unfolddatascience Welcome! I'm Aman...
Published: December 8, 2022
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.impute.SimpleImputer.html
SimpleImputer — scikit-learn 1.9.1 documentation
Columns which only contained missing values at fit are discarded upon transform if strategy is not "constant". In a prediction context, simple imputation usually performs poorly when associated with a weak learner. However, with a powerful learner, it can lead to as good or better performance than complex imputation such as IterativeImputer or KNNImputer. ... >>> import numpy as np >>> from sklearn.impute import SimpleImputer >>> imp_mean = SimpleImputer(missing_values=np.nan, strategy='mean') >>> imp_mean.fit([[7, 2, 3], [4, np.nan, 6], [10, 5, 9]]) SimpleImputer() >>> X = [[np.nan, 2, 3], [4, np.nan, 6], [10, np.nan, 9]] >>> print(imp_mean.transform(X)) [[ 7.
🌐
Kaggle
kaggle.com › questions-and-answers › 154449
fit_transform vs tranform
Checking your browser before accessing www.kaggle.com · Click here if you are not automatically redirected after 5 seconds
🌐
scikit-learn
scikit-learn.org › 0.20 › modules › generated › sklearn.impute.SimpleImputer.html
sklearn.impute.SimpleImputer — scikit-learn 0.20.4 documentation
class sklearn.impute.SimpleImputer(missing_values=nan, strategy='mean', fill_value=None, verbose=0, copy=True)[source]¶ · Imputation transformer for completing missing values. Read more in the User Guide. ... Columns which only contained missing values at fit are discarded upon transform if strategy is not “constant”.