📈 Machine Learning · Lecture 33 of 47

Feature Scaling and Normalisation: Standardisation, Min–Max and Robust Scaling

Many algorithms silently assume features share a scale. We explain which models need scaling and why, compare standardisation, min–max, robust and quantile scaling, and show how to apply them without leakage.

Suppose you predict loan default from annual income (values around 500,000) and number of dependants (values 0–8). To a distance-based or gradient-based algorithm, income looks 100,000 times more important — purely because of its units. Feature scaling puts features on comparable scales. It takes one line of code, yet forgetting it is among the most common reasons a model underperforms.

Which algorithms need scaling?#

Needs scalingWhy
k-NN, k-means, SVM (RBF), DBSCANDistances are dominated by large-scale features
Linear/logistic regression with regularisationPenalties treat all coefficients equally
PCAVariance is dominated by large-scale features
Neural networks, any gradient descentPoor conditioning slows or destabilises optimisation
Does not need scalingWhy
Decision trees, random forests, gradient boostingSplits depend only on the ordering of values
Naive Bayes (per-feature distributions)Each feature is modelled separately

Why gradient descent cares#

If features have very different scales, the loss surface becomes a long narrow valley: the Hessian is ill-conditioned, with condition number $\kappa$ roughly proportional to the ratio of feature variances. A learning rate small enough for the steep direction crawls along the flat one. Standardising makes the valley rounder, allowing larger steps and faster convergence.

Standardisation (z-score)#

$$ x' = \frac{x - \mu}{\sigma} $$

Each feature gets mean 0 and standard deviation 1. It is the default for most models, preserves the shape of the distribution and does not bound values. It is sensitive to outliers because they inflate $\sigma$.

Min–max scaling#

$$ x' = \frac{x - x_{\min}}{x_{\max} - x_{\min}} $$

Maps values to $[0, 1]$. Useful when an algorithm expects bounded inputs (image pixels are often scaled this way). Very sensitive to outliers: one extreme value squashes everything else into a tiny interval.

Robust scaling#

$$ x' = \frac{x - \text{median}}{\text{IQR}} $$

Uses the median and interquartile range, which ignore extreme values. Choose it when features contain outliers you do not want to remove.

Other transforms#

  • Max-abs scaling divides by the maximum absolute value — keeps zeros as zeros, suitable for sparse data such as TF-IDF matrices.
  • Power transforms (Box–Cox for positive data, Yeo–Johnson for any data) reduce skewness towards a Gaussian shape.
  • Quantile transform maps each feature to a uniform or normal distribution by ranks — very robust, but distorts distances and linear relationships.
  • Unit-norm (row) normalisation scales each sample to length 1 — used for text vectors and embeddings when only direction matters (cosine similarity).
python
import numpy as np
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler, PowerTransformer

rng = np.random.default_rng(0)
income = np.r_[rng.lognormal(10, 0.5, 995), [5e6, 7e6, 8e6, 9e6, 1e7]].reshape(-1, 1)  # with outliers

for name, sc in [("standard", StandardScaler()), ("min-max", MinMaxScaler()),
                 ("robust", RobustScaler()), ("yeo-johnson", PowerTransformer())]:
    z = sc.fit_transform(income).ravel()
    print(f"{name:<12} median={np.median(z):8.3f}  IQR={np.subtract(*np.percentile(z, [75, 25])):8.3f}  max={z.max():9.2f}")

Notice how min–max scaling compresses the typical incomes into a tiny range because of five extreme values, while robust scaling and the power transform keep them spread out.

Scaling without leakage#

python
from sklearn.pipeline import make_pipeline
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_breast_cancer

X, y = load_breast_cancer(return_X_y=True)
print("SVM without scaling:", cross_val_score(SVC(), X, y, cv=5).mean().round(3))
print("SVM with scaling:   ", cross_val_score(make_pipeline(StandardScaler(), SVC()), X, y, cv=5).mean().round(3))

Scaling targets#

For regression with neural networks, scaling the target (e.g. standardising, or log-transforming skewed targets) often stabilises training. Remember to invert the transformation for predictions — scikit-learn's TransformedTargetRegressor handles this.

Normalisation inside neural networks#

Deep networks extend the idea internally with batch normalisation, layer normalisation and related techniques that keep activations well scaled layer by layer — covered in the deep learning track.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Handling Missing Data: Mechanisms, Imputation and Indicators

Missing values are rarely random. We classify missingness as MCAR, MAR or MNAR, compare deletion and imputation strategies from simple to iterative, and show why a missingness indicator is often a feature in itself.

Intermediate⏱ 5 min#083
📈 Machine Learning

Encoding Categorical Variables: One-Hot, Ordinal, Target and Beyond

Models need numbers, but many features are categories. We compare one-hot, ordinal, frequency, target and hashing encoders, handle high cardinality and unseen categories, and avoid target-encoding leakage.

Beginner⏱ 5 min#084
📈 Machine Learning

Feature Engineering: Turning Raw Data into Signal

Better features beat better algorithms. We survey the craft — transformations, interactions, aggregations, date and text features, domain-driven ratios — and how to engineer features without leaking the target.

Intermediate⏱ 5 min#081