A multilayer perceptron (MLP) with even a single hidden layer is astonishingly expressive. This is formalised by the universal approximation theorem, one of the most quoted โ and most misunderstood โ results in deep learning. Today we state it, see why it is true, and then ask the more important question: if one layer is enough in principle, why do we use many?
The MLP#
A network with one hidden layer of $m$ units computes
โ a weighted sum of $m$ "ridge" functions, each a non-linearity applied to a projection of the input. Deeper MLPs stack such layers.
The theorem#
Universal Approximation Theorem (Cybenko 1989 for sigmoids; Hornik 1991; Leshno et al. 1993 for any non-polynomial activation, including ReLU). Let $\phi$ be continuous and not a polynomial. For any continuous function $g$ on a compact set $K \subset \mathbb{R}^d$ and any $\epsilon > 0$, there exist $m$ and parameters such that
In words: a single hidden layer, made wide enough, can approximate any continuous function on a bounded region as closely as we like.
Intuition: building functions from bumps#
With ReLU $\phi(z) = \max(0, z)$ in one dimension:
- A single ReLU is a hinge.
- The difference of two shifted ReLUs creates a ramp that flattens out: a step.
- The difference of two steps creates a bump โ a small tower over an interval.
- Sum many bumps of chosen heights and you can trace any continuous curve, like a histogram approximating a density.
import numpy as np
import torch, torch.nn as nn
x = torch.linspace(-3, 3, 400).unsqueeze(1)
target = torch.sin(2 * x) + 0.3 * x**2 # an arbitrary continuous function
for width in [2, 8, 64]:
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(1, width), nn.ReLU(), nn.Linear(width, 1))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
for _ in range(3000):
opt.zero_grad(); loss = ((net(x) - target) ** 2).mean(); loss.backward(); opt.step()
print(f"width {width:>3}: MSE = {loss.item():.5f}")Wider single-layer networks approximate the target ever more closely โ the theorem in action.
What the theorem does NOT say#
Why depth?#
If shallow networks are universal, why use deep ones? Because depth can be exponentially more efficient.
- Depth-separation results: there exist functions computable by small deep networks that require exponentially many units in shallow networks (e.g. Telgarsky 2016, using compositions of "sawtooth" functions; Eldan & Shamir 2016 for 3 vs 2 layers).
- Compositionality: a ReLU network partitions input space into linear regions. Each additional layer can "fold" space, so the number of linear regions can grow exponentially with depth but only polynomially with width.
- Hierarchical structure of real data: images are composed of objects, objects of parts, parts of edges; language of sentences, phrases and words. Deep networks mirror this compositional structure, reusing intermediate features across many higher-level concepts.
Width vs depth in practice#
- Very deep plain networks are hard to train (vanishing gradients) โ solved by residual connections and normalisation, discussed in later lectures.
- Very wide shallow networks tend to memorise rather than build reusable features.
- Modern architectures balance both, and scaling laws show performance improves predictably when depth, width and data grow together.
Designing an MLP#
For tabular data or as a component inside larger models:
- 2โ4 hidden layers; widths of 64โ1024; ReLU/GELU activations.
- Normalisation (batch or layer norm) and dropout for regularisation.
- Output layer: linear for regression; logits + softmax/sigmoid for classification.
- Tune learning rate first, then width/depth and regularisation.
In transformers, every block contains an MLP (the "feed-forward network") that holds a large share of the parameters and is believed to store much of the model's factual knowledge.