Since around 2021, diffusion models have produced the most impressive generated images, and they now power text-to-image, video, audio and even molecule generation. The idea is almost paradoxical: to learn how to create data, learn how to remove noise from it. Start with pure noise, denoise step by step, and a realistic image emerges. This lecture derives the method carefully.
The forward (noising) process#
Take a data point $\mathbf{x}_0$ and gradually add Gaussian noise over $T$ steps (e.g. $T = 1000$), with a small variance schedule $\beta_1, \dots, \beta_T$:
After enough steps, $\mathbf{x}_T$ is essentially pure Gaussian noise. Crucially, because sums of Gaussians are Gaussian, we can jump to any step in closed form. Define $\alpha_t = 1 - \beta_t$ and $\bar{\alpha}_t = \prod_{s=1}^{t}\alpha_s$:
The reverse (denoising) process#
We want to run the process backwards: start from $\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$ and repeatedly sample $\mathbf{x}_{t-1}$ from $\mathbf{x}_t$. For small $\beta_t$, the true reverse step is approximately Gaussian, so we learn
DDPM: the simple loss#
Ho, Jain and Abbeel (2020) — Denoising Diffusion Probabilistic Models — showed that the variational bound on the likelihood simplifies beautifully if the network $\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)$ predicts the noise that was added. The training objective becomes:
Training loop: pick a random image, a random timestep and random noise; noise the image in one shot; ask the network to predict the noise; minimise MSE. No adversarial game, no mode collapse — just regression. The network is typically a U-Net with timestep embeddings (and attention layers), or a transformer (DiT).
Sampling#
Given the predicted noise, the reverse mean is
and we sample $\mathbf{x}_{t-1} = \boldsymbol{\mu}_\theta(\mathbf{x}_t, t) + \sigma_t\mathbf{z}$ (with $\mathbf{z} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$, and $\sigma_t^2 = \beta_t$ a common choice), for $t = T, \dots, 1$.
import torch
import torch.nn as nn
T = 1000
betas = torch.linspace(1e-4, 0.02, T)
alphas = 1 - betas
alpha_bar = torch.cumprod(alphas, dim=0)
class Denoiser(nn.Module): # a tiny MLP denoiser for 2-D toy data
def __init__(self, dim=2, hidden=256):
super().__init__()
self.t_emb = nn.Embedding(T, hidden)
self.net = nn.Sequential(nn.Linear(dim + hidden, hidden), nn.SiLU(),
nn.Linear(hidden, hidden), nn.SiLU(), nn.Linear(hidden, dim))
def forward(self, x, t):
return self.net(torch.cat([x, self.t_emb(t)], dim=1))
def training_loss(model, x0):
t = torch.randint(0, T, (x0.size(0),))
eps = torch.randn_like(x0)
ab = alpha_bar[t].unsqueeze(1)
xt = ab.sqrt() * x0 + (1 - ab).sqrt() * eps # jump straight to step t
return ((eps - model(xt, t)) ** 2).mean() # predict the noise
@torch.no_grad()
def sample(model, n=1000, dim=2):
x = torch.randn(n, dim)
for t in reversed(range(T)):
tt = torch.full((n,), t, dtype=torch.long)
eps = model(x, tt)
mean = (x - betas[t] / (1 - alpha_bar[t]).sqrt() * eps) / alphas[t].sqrt()
x = mean + (betas[t].sqrt() * torch.randn_like(x) if t > 0 else 0)
return x
from sklearn.datasets import make_moons
data = torch.tensor(make_moons(8000, noise=0.03)[0], dtype=torch.float32)
model = Denoiser(); opt = torch.optim.Adam(model.parameters(), 1e-3)
for step in range(3000):
loss = training_loss(model, data[torch.randint(0, len(data), (512,))])
opt.zero_grad(); loss.backward(); opt.step()
print("generated:", sample(model, 5))Plot the samples: they trace out the two moons, although the model never saw a formula for moons.
The score-based view#
Song and Ermon (2019) and Song et al. (2021) connected diffusion to score matching. The score is the gradient of the log-density, $\nabla_{\mathbf{x}}\log p(\mathbf{x})$ — it points towards regions of higher probability. The noise-prediction network estimates the score of the noised data distribution:
In continuous time, the forward process is a stochastic differential equation (SDE), and generation solves the reverse-time SDE — or an equivalent deterministic probability-flow ODE, which also enables exact likelihood computation. This unified view explains why diffusion models work: they learn the direction to move a noisy sample towards the data manifold at every noise level.
Faster sampling#
A thousand network evaluations per image is slow. Improvements:
- DDIM (Song, Meng & Ermon, 2021): a non-Markovian, deterministic sampler using the same trained model, producing good samples in 20–50 steps and enabling latent interpolation and inversion.
- Higher-order ODE solvers (DPM-Solver, Heun/EDM samplers): high quality in 10–25 steps.
- Distillation: progressive distillation, consistency models and adversarial distillation produce images in 1–4 steps.
Why diffusion won#
- Stable training — a simple regression loss, no adversarial dynamics.
- Excellent mode coverage and diversity, with sample quality that surpassed GANs on standard benchmarks (Dhariwal & Nichol, 2021, "Diffusion Models Beat GANs on Image Synthesis").
- Flexible conditioning — text, images, masks, poses — and guidance techniques (next lectures).
- Scalability to video, audio, 3-D and scientific data (protein structure generation, weather forecasting).