๐Ÿ”— Deep Learning ยท Lecture 3 of 38

Multilayer Perceptrons and the Universal Approximation Theorem

With one hidden layer, a network can approximate any continuous function โ€” so why go deep? We state the universal approximation theorem, build intuition with bumps, and explain the efficiency advantages of depth.

A multilayer perceptron (MLP) with even a single hidden layer is astonishingly expressive. This is formalised by the universal approximation theorem, one of the most quoted โ€” and most misunderstood โ€” results in deep learning. Today we state it, see why it is true, and then ask the more important question: if one layer is enough in principle, why do we use many?

The MLP#

A network with one hidden layer of $m$ units computes

$$ f(\mathbf{x}) = \sum_{j=1}^{m}v_j\,\phi(\mathbf{w}_j^\top\mathbf{x} + b_j) + c $$

โ€” a weighted sum of $m$ "ridge" functions, each a non-linearity applied to a projection of the input. Deeper MLPs stack such layers.

The theorem#

Universal Approximation Theorem (Cybenko 1989 for sigmoids; Hornik 1991; Leshno et al. 1993 for any non-polynomial activation, including ReLU). Let $\phi$ be continuous and not a polynomial. For any continuous function $g$ on a compact set $K \subset \mathbb{R}^d$ and any $\epsilon > 0$, there exist $m$ and parameters such that

$$ \sup_{\mathbf{x} \in K}|f(\mathbf{x}) - g(\mathbf{x})| < \epsilon $$

In words: a single hidden layer, made wide enough, can approximate any continuous function on a bounded region as closely as we like.

Intuition: building functions from bumps#

With ReLU $\phi(z) = \max(0, z)$ in one dimension:

  • A single ReLU is a hinge.
  • The difference of two shifted ReLUs creates a ramp that flattens out: a step.
  • The difference of two steps creates a bump โ€” a small tower over an interval.
  • Sum many bumps of chosen heights and you can trace any continuous curve, like a histogram approximating a density.
python
import numpy as np
import torch, torch.nn as nn

x = torch.linspace(-3, 3, 400).unsqueeze(1)
target = torch.sin(2 * x) + 0.3 * x**2                     # an arbitrary continuous function

for width in [2, 8, 64]:
    torch.manual_seed(0)
    net = nn.Sequential(nn.Linear(1, width), nn.ReLU(), nn.Linear(width, 1))
    opt = torch.optim.Adam(net.parameters(), lr=0.01)
    for _ in range(3000):
        opt.zero_grad(); loss = ((net(x) - target) ** 2).mean(); loss.backward(); opt.step()
    print(f"width {width:>3}: MSE = {loss.item():.5f}")

Wider single-layer networks approximate the target ever more closely โ€” the theorem in action.

What the theorem does NOT say#

Why depth?#

If shallow networks are universal, why use deep ones? Because depth can be exponentially more efficient.

  • Depth-separation results: there exist functions computable by small deep networks that require exponentially many units in shallow networks (e.g. Telgarsky 2016, using compositions of "sawtooth" functions; Eldan & Shamir 2016 for 3 vs 2 layers).
  • Compositionality: a ReLU network partitions input space into linear regions. Each additional layer can "fold" space, so the number of linear regions can grow exponentially with depth but only polynomially with width.
  • Hierarchical structure of real data: images are composed of objects, objects of parts, parts of edges; language of sentences, phrases and words. Deep networks mirror this compositional structure, reusing intermediate features across many higher-level concepts.

Width vs depth in practice#

  • Very deep plain networks are hard to train (vanishing gradients) โ€” solved by residual connections and normalisation, discussed in later lectures.
  • Very wide shallow networks tend to memorise rather than build reusable features.
  • Modern architectures balance both, and scaling laws show performance improves predictably when depth, width and data grow together.

Designing an MLP#

For tabular data or as a component inside larger models:

  • 2โ€“4 hidden layers; widths of 64โ€“1024; ReLU/GELU activations.
  • Normalisation (batch or layer norm) and dropout for regularisation.
  • Output layer: linear for regression; logits + softmax/sigmoid for classification.
  • Tune learning rate first, then width/depth and regularisation.

In transformers, every block contains an MLP (the "feed-forward network") that holds a large share of the parameters and is believed to store much of the model's factual knowledge.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

The Perceptron: The First Learning Machine

Rosenblatt's perceptron learned to classify by correcting its mistakes. We derive its learning rule, prove the convergence theorem, reveal its XOR limitation, and see how it foreshadowed modern networks.

Beginnerโฑ 5 min#098
๐Ÿ”— Deep Learning

Activation Functions: Sigmoid, Tanh, ReLU, GELU, SwiGLU and Softmax

The choice of non-linearity shapes how gradients flow and how networks learn. We compare the classical and modern activations, their derivatives and failure modes, and which to use where.

Beginnerโฑ 5 min#100
๐Ÿ”— Deep Learning

From Biological Neurons to Artificial Neural Networks

We open the Deep Learning track by tracing the path from biological neurons to artificial ones, defining a neural network precisely, and explaining why depth and learned representations changed AI.

Beginnerโฑ 4 min#097