Suppose you predict loan default from annual income (values around 500,000) and number of dependants (values 0–8). To a distance-based or gradient-based algorithm, income looks 100,000 times more important — purely because of its units. Feature scaling puts features on comparable scales. It takes one line of code, yet forgetting it is among the most common reasons a model underperforms.
Which algorithms need scaling?#
| Needs scaling | Why |
|---|---|
| k-NN, k-means, SVM (RBF), DBSCAN | Distances are dominated by large-scale features |
| Linear/logistic regression with regularisation | Penalties treat all coefficients equally |
| PCA | Variance is dominated by large-scale features |
| Neural networks, any gradient descent | Poor conditioning slows or destabilises optimisation |
| Does not need scaling | Why |
|---|---|
| Decision trees, random forests, gradient boosting | Splits depend only on the ordering of values |
| Naive Bayes (per-feature distributions) | Each feature is modelled separately |
Why gradient descent cares#
If features have very different scales, the loss surface becomes a long narrow valley: the Hessian is ill-conditioned, with condition number $\kappa$ roughly proportional to the ratio of feature variances. A learning rate small enough for the steep direction crawls along the flat one. Standardising makes the valley rounder, allowing larger steps and faster convergence.
Standardisation (z-score)#
Each feature gets mean 0 and standard deviation 1. It is the default for most models, preserves the shape of the distribution and does not bound values. It is sensitive to outliers because they inflate $\sigma$.
Min–max scaling#
Maps values to $[0, 1]$. Useful when an algorithm expects bounded inputs (image pixels are often scaled this way). Very sensitive to outliers: one extreme value squashes everything else into a tiny interval.
Robust scaling#
Uses the median and interquartile range, which ignore extreme values. Choose it when features contain outliers you do not want to remove.
Other transforms#
- Max-abs scaling divides by the maximum absolute value — keeps zeros as zeros, suitable for sparse data such as TF-IDF matrices.
- Power transforms (Box–Cox for positive data, Yeo–Johnson for any data) reduce skewness towards a Gaussian shape.
- Quantile transform maps each feature to a uniform or normal distribution by ranks — very robust, but distorts distances and linear relationships.
- Unit-norm (row) normalisation scales each sample to length 1 — used for text vectors and embeddings when only direction matters (cosine similarity).
import numpy as np
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler, PowerTransformer
rng = np.random.default_rng(0)
income = np.r_[rng.lognormal(10, 0.5, 995), [5e6, 7e6, 8e6, 9e6, 1e7]].reshape(-1, 1) # with outliers
for name, sc in [("standard", StandardScaler()), ("min-max", MinMaxScaler()),
("robust", RobustScaler()), ("yeo-johnson", PowerTransformer())]:
z = sc.fit_transform(income).ravel()
print(f"{name:<12} median={np.median(z):8.3f} IQR={np.subtract(*np.percentile(z, [75, 25])):8.3f} max={z.max():9.2f}")Notice how min–max scaling compresses the typical incomes into a tiny range because of five extreme values, while robust scaling and the power transform keep them spread out.
Scaling without leakage#
from sklearn.pipeline import make_pipeline
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
print("SVM without scaling:", cross_val_score(SVC(), X, y, cv=5).mean().round(3))
print("SVM with scaling: ", cross_val_score(make_pipeline(StandardScaler(), SVC()), X, y, cv=5).mean().round(3))Scaling targets#
For regression with neural networks, scaling the target (e.g. standardising, or log-transforming skewed targets) often stabilises training. Remember to invert the transformation for predictions — scikit-learn's TransformedTargetRegressor handles this.
Normalisation inside neural networks#
Deep networks extend the idea internally with batch normalisation, layer normalisation and related techniques that keep activations well scaled layer by layer — covered in the deep learning track.