Every modelling decision — the degree of a polynomial, the depth of a tree, the strength of regularisation, the number of neighbours in k-NN — moves you along a single axis of trade-offs. On one side lies bias: systematic error from a model too simple to capture the truth. On the other lies variance: error from a model so flexible that it chases noise in the particular training set. Today we derive this decomposition precisely.
The setting#
Assume data is generated as $y = f(\mathbf{x}) + \epsilon$, where $f$ is the true function and $\epsilon$ is noise with $\mathbb{E}[\epsilon] = 0$ and $\text{Var}(\epsilon) = \sigma^2$. We draw a training set $\mathcal{D}$, fit a model $\hat{f}_{\mathcal{D}}$, and evaluate at a fixed test point $\mathbf{x}$.
The key conceptual step: the training set is random. Draw a different sample and you get a different fitted model. So $\hat{f}_{\mathcal{D}}(\mathbf{x})$ is a random variable. Let $\bar{f}(\mathbf{x}) = \mathbb{E}_{\mathcal{D}}[\hat{f}_{\mathcal{D}}(\mathbf{x})]$ be the average prediction over all possible training sets.
The decomposition#
The expected squared error at $\mathbf{x}$, averaging over training sets and noise, is
Derivation sketch. Write $y - \hat{f} = (f - \bar{f}) + (\bar{f} - \hat{f}) + \epsilon$. Square and take expectations. The cross terms vanish because $\mathbb{E}[\epsilon] = 0$, $\epsilon$ is independent of $\mathcal{D}$, and $\mathbb{E}_{\mathcal{D}}[\bar{f} - \hat{f}] = 0$. $\blacksquare$
- Bias measures how far the average model is from the truth — a property of the model family's assumptions.
- Variance measures how much the model fluctuates across training sets — sensitivity to the particular sample.
- Noise is the floor that no model can beat.
The archery analogy#
Imagine many archers, each trained on a different dataset, shooting at a target:
| Low variance | High variance | |
|---|---|---|
| Low bias | Tight cluster on the bullseye (ideal) | Scattered around the bullseye |
| High bias | Tight cluster, off-centre | Scattered and off-centre |
How complexity moves the balance#
| Model / knob | Increase complexity by… | Effect |
|---|---|---|
| Polynomial regression | higher degree | bias ↓, variance ↑ |
| k-NN | smaller $k$ | bias ↓, variance ↑ |
| Decision tree | greater depth | bias ↓, variance ↑ |
| Ridge / lasso | smaller $\lambda$ | bias ↓, variance ↑ |
| Neural network | more parameters, fewer regularisers | (usually) bias ↓, variance ↑ |
Total error is minimised where the sum of bias² and variance is smallest — the bottom of the familiar U-shaped test-error curve.
Estimating bias and variance by simulation#
import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
rng = np.random.default_rng(0)
f = lambda x: np.sin(2 * np.pi * x)
x_test = np.linspace(0.05, 0.95, 50)
sigma = 0.3
for degree in [1, 3, 9]:
preds = []
for trial in range(300): # 300 independent training sets
x = rng.random(20); y = f(x) + rng.normal(0, sigma, 20)
m = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(x[:, None], y)
preds.append(m.predict(x_test[:, None]))
preds = np.array(preds)
bias2 = ((preds.mean(0) - f(x_test)) ** 2).mean()
var = preds.var(0).mean()
print(f"degree {degree}: bias^2={bias2:.3f} variance={var:.3f} "
f"total={bias2 + var + sigma**2:.3f}")Degree 1 shows high bias and low variance; degree 9 shows the reverse; degree 3 balances them.
Reducing each component#
To reduce bias: use a more flexible model, add informative features, reduce regularisation, train longer.
To reduce variance: collect more data, regularise, simplify the model, use feature selection, apply bagging/ensembling (averaging many high-variance models — the principle behind random forests), use data augmentation, stop early.
The modern twist: double descent#
Classical theory predicts that once a model is complex enough to fit the training data perfectly (the interpolation threshold), test error should explode. Yet very large neural networks interpolate their training data and still generalise well. Researchers observed double descent: as complexity increases past the interpolation threshold, test error first peaks and then decreases again. Explanations involve implicit regularisation by the optimiser — among the many solutions that fit the data, gradient descent tends to find smooth, low-norm ones. The bias–variance decomposition is still mathematically true; what changes is how variance behaves for heavily over-parameterised models. We examine this in the deep-learning track.