📈 Machine Learning · Lecture 9 of 47

The Bias–Variance Trade-off: Derivation and Intuition

Why do simple models underfit and complex models overfit? We derive the bias–variance decomposition of expected squared error, visualise it, and discuss how modern deep learning complicates the classical picture.

Every modelling decision — the degree of a polynomial, the depth of a tree, the strength of regularisation, the number of neighbours in k-NN — moves you along a single axis of trade-offs. On one side lies bias: systematic error from a model too simple to capture the truth. On the other lies variance: error from a model so flexible that it chases noise in the particular training set. Today we derive this decomposition precisely.

The setting#

Assume data is generated as $y = f(\mathbf{x}) + \epsilon$, where $f$ is the true function and $\epsilon$ is noise with $\mathbb{E}[\epsilon] = 0$ and $\text{Var}(\epsilon) = \sigma^2$. We draw a training set $\mathcal{D}$, fit a model $\hat{f}_{\mathcal{D}}$, and evaluate at a fixed test point $\mathbf{x}$.

The key conceptual step: the training set is random. Draw a different sample and you get a different fitted model. So $\hat{f}_{\mathcal{D}}(\mathbf{x})$ is a random variable. Let $\bar{f}(\mathbf{x}) = \mathbb{E}_{\mathcal{D}}[\hat{f}_{\mathcal{D}}(\mathbf{x})]$ be the average prediction over all possible training sets.

The decomposition#

The expected squared error at $\mathbf{x}$, averaging over training sets and noise, is

$$ \mathbb{E}\big[(y - \hat{f}_{\mathcal{D}}(\mathbf{x}))^2\big] = \underbrace{\big(f(\mathbf{x}) - \bar{f}(\mathbf{x})\big)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}_{\mathcal{D}}\big[(\hat{f}_{\mathcal{D}}(\mathbf{x}) - \bar{f}(\mathbf{x}))^2\big]}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Irreducible noise}} $$

Derivation sketch. Write $y - \hat{f} = (f - \bar{f}) + (\bar{f} - \hat{f}) + \epsilon$. Square and take expectations. The cross terms vanish because $\mathbb{E}[\epsilon] = 0$, $\epsilon$ is independent of $\mathcal{D}$, and $\mathbb{E}_{\mathcal{D}}[\bar{f} - \hat{f}] = 0$. $\blacksquare$

  • Bias measures how far the average model is from the truth — a property of the model family's assumptions.
  • Variance measures how much the model fluctuates across training sets — sensitivity to the particular sample.
  • Noise is the floor that no model can beat.

The archery analogy#

Imagine many archers, each trained on a different dataset, shooting at a target:

Low varianceHigh variance
Low biasTight cluster on the bullseye (ideal)Scattered around the bullseye
High biasTight cluster, off-centreScattered and off-centre

How complexity moves the balance#

Model / knobIncrease complexity by…Effect
Polynomial regressionhigher degreebias ↓, variance ↑
k-NNsmaller $k$bias ↓, variance ↑
Decision treegreater depthbias ↓, variance ↑
Ridge / lassosmaller $\lambda$bias ↓, variance ↑
Neural networkmore parameters, fewer regularisers(usually) bias ↓, variance ↑

Total error is minimised where the sum of bias² and variance is smallest — the bottom of the familiar U-shaped test-error curve.

Estimating bias and variance by simulation#

python
import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline

rng = np.random.default_rng(0)
f = lambda x: np.sin(2 * np.pi * x)
x_test = np.linspace(0.05, 0.95, 50)
sigma = 0.3

for degree in [1, 3, 9]:
    preds = []
    for trial in range(300):                              # 300 independent training sets
        x = rng.random(20); y = f(x) + rng.normal(0, sigma, 20)
        m = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(x[:, None], y)
        preds.append(m.predict(x_test[:, None]))
    preds = np.array(preds)
    bias2 = ((preds.mean(0) - f(x_test)) ** 2).mean()
    var = preds.var(0).mean()
    print(f"degree {degree}: bias^2={bias2:.3f}  variance={var:.3f}  "
          f"total={bias2 + var + sigma**2:.3f}")

Degree 1 shows high bias and low variance; degree 9 shows the reverse; degree 3 balances them.

Reducing each component#

To reduce bias: use a more flexible model, add informative features, reduce regularisation, train longer.

To reduce variance: collect more data, regularise, simplify the model, use feature selection, apply bagging/ensembling (averaging many high-variance models — the principle behind random forests), use data augmentation, stop early.

The modern twist: double descent#

Classical theory predicts that once a model is complex enough to fit the training data perfectly (the interpolation threshold), test error should explode. Yet very large neural networks interpolate their training data and still generalise well. Researchers observed double descent: as complexity increases past the interpolation threshold, test error first peaks and then decreases again. Explanations involve implicit regularisation by the optimiser — among the many solutions that fit the data, gradient descent tends to find smooth, low-norm ones. The bias–variance decomposition is still mathematically true; what changes is how variance behaves for heavily over-parameterised models. We examine this in the deep-learning track.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

What Is Machine Learning? The Learning Problem Formalised

Tom Mitchell's definition, the formal learning problem, empirical risk minimisation and the central goal of generalisation — the conceptual foundation for every model in this track.

Beginner⏱ 5 min#050
📈 Machine Learning

Principal Component Analysis (PCA): Theory and Practice

PCA finds the orthogonal directions of maximum variance. We derive it two ways — maximum variance and minimum reconstruction error — compute it via SVD, choose the number of components, and discuss its limits.

Intermediate⏱ 5 min#078
📈 Machine Learning

Softmax Regression and Multiclass Classification Strategies

How do we classify into more than two classes? We generalise logistic regression to softmax regression, derive its gradient, and compare it with one-vs-rest and one-vs-one strategies.

Beginner⏱ 5 min#057