To center the data (make it have zero mean and unit standard error), you subtract the mean and then divide the result by the standard deviation:
You do that on the training set of the data. But then you have to apply the same transformation to your test set (e.g. in cross-validation), or to newly obtained examples before forecasting. But you have to use the exact same two parameters and
(values) that you used for centering the training set.
Hence, every scikit-learn's transform's fit() just calculates the parameters (e.g. and
in case of StandardScaler) and saves them as an internal object's state. Afterwards, you can call its
transform() method to apply the transformation to any particular set of examples.
fit_transform() joins these two steps and is used for the initial fitting of parameters on the training set , while also returning the transformed
. Internally, the transformer object just calls first
fit() and then transform() on the same data.
To center the data (make it have zero mean and unit standard error), you subtract the mean and then divide the result by the standard deviation:
You do that on the training set of the data. But then you have to apply the same transformation to your test set (e.g. in cross-validation), or to newly obtained examples before forecasting. But you have to use the exact same two parameters and
(values) that you used for centering the training set.
Hence, every scikit-learn's transform's fit() just calculates the parameters (e.g. and
in case of StandardScaler) and saves them as an internal object's state. Afterwards, you can call its
transform() method to apply the transformation to any particular set of examples.
fit_transform() joins these two steps and is used for the initial fitting of parameters on the training set , while also returning the transformed
. Internally, the transformer object just calls first
fit() and then transform() on the same data.
The following explanation is based on fit_transform of Imputer class, but the idea is the same for fit_transform of other scikit_learn classes like MinMaxScaler.
transform replaces the missing values with a number. By default this number is the means of columns of some data that you choose.
Consider the following example:
imp = Imputer()
# calculating the means
imp.fit([
[1, 3],
[np.nan, 2],
[8, 5.5]
])
Now the imputer have learned to use a mean for the first column and mean
for the second column when it gets applied to a two-column data:
X = [[np.nan, 11],
[4, np.nan],
[8, 2],
[np.nan, 1]]
print(imp.transform(X))
we get
[[4.5, 11],
[4, 3.5],
[8, 2],
[4.5, 1]]
So by fit the imputer calculates the means of columns from some data, and by transform it applies those means to some data (which is just replacing missing values with the means). If both these data are the same (i.e. the data for calculating the means and the data that means are applied to) you can use fit_transform which is basically a fit followed by a transform.
Now your questions:
Why we might need to transform data?
"For various reasons, many real world datasets contain missing values, often encoded as blanks, NaNs or other placeholders. Such datasets however are incompatible with scikit-learn estimators which assume that all values in an array are numerical" (source)
What does it mean fitting model on training data and transforming to test data?
The fit of an imputer has nothing to do with fit used in model fitting.
So using imputer's fit on training data just calculates means of each column of training data. Using transform on test data then replaces missing values of test data with means that were calculated from training data.
In scikit-learn estimator api,
fit() : used for generating learning model parameters from training data
transform() :
parameters generated from fit() method,applied upon model to generate transformed data set.
fit_transform() :
combination of fit() and transform() api on same data set

Checkout Chapter-4 from this book & answer from stackexchange for more clarity
These methods are used to center/feature scale of a given data. It basically helps to normalize the data within a particular range
For this, we use Z-score method.

We do this on the training set of data.
1.Fit(): Method calculates the parameters μ and σ and saves them as internal objects.
2.Transform(): Method using these calculated parameters apply the transformation to a particular dataset.
3.Fit_transform(): joins the fit() and transform() method for transformation of dataset.
Code snippet for Feature Scaling/Standardisation(after train_test_split).
from sklearn.preprocessing import StandardScaler
sc = StandardScaler()
sc.fit_transform(X_train)
sc.transform(X_test)
We apply the same(training set same two parameters μ and σ (values)) parameter transformation on our testing set.
python - sklearn.impute SimpleImputer: why does transform() need fit_transform() first? - Stack Overflow
machine learning - What is the difference between "fit_transform" and "transform" methods when using "SimpleImputer"? - Data Science Stack Exchange
machine learning - What is difference between fit, transform and fit_transform in python when using sklearn? - Stack Overflow
machine learning - Should I use transformer.fit_transform(X_test, y_test) or not? - Cross Validated
The confusing part is fit and transform.
#here fit method will calculate the required parameters (In this case mean)
#and store it in the impute object
imputer = imputer.fit(X[:, 1:3])
X[:, 1:3]=imputer.transform(X[:, 1:3])
#imputer.transform will actually do the work of replacement of nan with mean.
#This can be done in one step using fit_transform
Imputer is used to replace missing values. The fit method calculates the parameters while the fit_transform method changes the data to replace those NaN with the mean and outputs a new matrix X.
# Imports library
from sklearn.preprocessing import Imputer
# Create a new instance of the Imputer object
# Missing values are replaced with NaN
# Missing values are replaced by the mean later on
# The axis determines whether you want to move column or row wise
imputer = Imputer(missing_values='NaN', strategy='mean',axis=0)
# Fit the imputer to X
imputer = imputer.fit(X[:, 1:3])
# Replace in the original matrix X
# with the new values after the transformation of X
X[:, 1:3]=imputer.transform(X[:, 1:3])
I commented out the code for you, I hope this will make a bit more sense. You need to think of X as a matrix that you have to transform in order to have no more NaN (missing values).
Refer to the documentation for more information.
The whole point of training a model is to use information about the training data to make inferences about some larger population. It's not data leakage to use information about training data to make predictions -- that's how models work.
If you use information about the test data to make predictions about the test data, then you're letting your model cheat on its test.
No.
I'm assuming that you understand what transform and fit methods (or concepts) are individually.
You can use fit_transform() on training dataset as you are supposed to make your model/scaler to learn from the characteristics of your training dataset.
But you can only use transform() with your test dataset. If you use fit_transform() with test dataset, you would basically fit your model/scaler to learn from the test data too, which causes data leakage.
Remember, test dataset is supposed to be an unseen dataset!