✨ Generative AI & LLMs · Lecture 6 of 30

Diffusion Models: Generating by Learning to Denoise

Diffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.

Since around 2021, diffusion models have produced the most impressive generated images, and they now power text-to-image, video, audio and even molecule generation. The idea is almost paradoxical: to learn how to create data, learn how to remove noise from it. Start with pure noise, denoise step by step, and a realistic image emerges. This lecture derives the method carefully.

The forward (noising) process#

Take a data point $\mathbf{x}_0$ and gradually add Gaussian noise over $T$ steps (e.g. $T = 1000$), with a small variance schedule $\beta_1, \dots, \beta_T$:

$$ q(\mathbf{x}_t \mid \mathbf{x}_{t-1}) = \mathcal{N}\left(\mathbf{x}_t;\; \sqrt{1 - \beta_t}\,\mathbf{x}_{t-1},\; \beta_t\mathbf{I}\right) $$

After enough steps, $\mathbf{x}_T$ is essentially pure Gaussian noise. Crucially, because sums of Gaussians are Gaussian, we can jump to any step in closed form. Define $\alpha_t = 1 - \beta_t$ and $\bar{\alpha}_t = \prod_{s=1}^{t}\alpha_s$:

$$ q(\mathbf{x}_t \mid \mathbf{x}_0) = \mathcal{N}\left(\mathbf{x}_t;\; \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0,\; (1 - \bar{\alpha}_t)\mathbf{I}\right) \quad\Longleftrightarrow\quad \mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\epsilon}, \;\; \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}) $$

The reverse (denoising) process#

We want to run the process backwards: start from $\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$ and repeatedly sample $\mathbf{x}_{t-1}$ from $\mathbf{x}_t$. For small $\beta_t$, the true reverse step is approximately Gaussian, so we learn

$$ p_\theta(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \mathcal{N}\left(\mathbf{x}_{t-1};\; \boldsymbol{\mu}_\theta(\mathbf{x}_t, t),\; \sigma_t^2\mathbf{I}\right) $$

DDPM: the simple loss#

Ho, Jain and Abbeel (2020) — Denoising Diffusion Probabilistic Models — showed that the variational bound on the likelihood simplifies beautifully if the network $\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)$ predicts the noise that was added. The training objective becomes:

$$ \mathcal{L}_{\text{simple}} = \mathbb{E}_{t, \mathbf{x}_0, \boldsymbol{\epsilon}}\left[\big\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta\big(\sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\epsilon},\; t\big)\big\|^2\right] $$

Training loop: pick a random image, a random timestep and random noise; noise the image in one shot; ask the network to predict the noise; minimise MSE. No adversarial game, no mode collapse — just regression. The network is typically a U-Net with timestep embeddings (and attention layers), or a transformer (DiT).

Sampling#

Given the predicted noise, the reverse mean is

$$ \boldsymbol{\mu}_\theta(\mathbf{x}_t, t) = \frac{1}{\sqrt{\alpha_t}}\left(\mathbf{x}_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)\right) $$

and we sample $\mathbf{x}_{t-1} = \boldsymbol{\mu}_\theta(\mathbf{x}_t, t) + \sigma_t\mathbf{z}$ (with $\mathbf{z} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$, and $\sigma_t^2 = \beta_t$ a common choice), for $t = T, \dots, 1$.

python
import torch
import torch.nn as nn

T = 1000
betas = torch.linspace(1e-4, 0.02, T)
alphas = 1 - betas
alpha_bar = torch.cumprod(alphas, dim=0)

class Denoiser(nn.Module):                          # a tiny MLP denoiser for 2-D toy data
    def __init__(self, dim=2, hidden=256):
        super().__init__()
        self.t_emb = nn.Embedding(T, hidden)
        self.net = nn.Sequential(nn.Linear(dim + hidden, hidden), nn.SiLU(),
                                 nn.Linear(hidden, hidden), nn.SiLU(), nn.Linear(hidden, dim))
    def forward(self, x, t):
        return self.net(torch.cat([x, self.t_emb(t)], dim=1))

def training_loss(model, x0):
    t = torch.randint(0, T, (x0.size(0),))
    eps = torch.randn_like(x0)
    ab = alpha_bar[t].unsqueeze(1)
    xt = ab.sqrt() * x0 + (1 - ab).sqrt() * eps     # jump straight to step t
    return ((eps - model(xt, t)) ** 2).mean()       # predict the noise

@torch.no_grad()
def sample(model, n=1000, dim=2):
    x = torch.randn(n, dim)
    for t in reversed(range(T)):
        tt = torch.full((n,), t, dtype=torch.long)
        eps = model(x, tt)
        mean = (x - betas[t] / (1 - alpha_bar[t]).sqrt() * eps) / alphas[t].sqrt()
        x = mean + (betas[t].sqrt() * torch.randn_like(x) if t > 0 else 0)
    return x

from sklearn.datasets import make_moons
data = torch.tensor(make_moons(8000, noise=0.03)[0], dtype=torch.float32)
model = Denoiser(); opt = torch.optim.Adam(model.parameters(), 1e-3)
for step in range(3000):
    loss = training_loss(model, data[torch.randint(0, len(data), (512,))])
    opt.zero_grad(); loss.backward(); opt.step()
print("generated:", sample(model, 5))

Plot the samples: they trace out the two moons, although the model never saw a formula for moons.

The score-based view#

Song and Ermon (2019) and Song et al. (2021) connected diffusion to score matching. The score is the gradient of the log-density, $\nabla_{\mathbf{x}}\log p(\mathbf{x})$ — it points towards regions of higher probability. The noise-prediction network estimates the score of the noised data distribution:

$$ \nabla_{\mathbf{x}_t}\log q(\mathbf{x}_t) \approx -\frac{\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)}{\sqrt{1 - \bar{\alpha}_t}} $$

In continuous time, the forward process is a stochastic differential equation (SDE), and generation solves the reverse-time SDE — or an equivalent deterministic probability-flow ODE, which also enables exact likelihood computation. This unified view explains why diffusion models work: they learn the direction to move a noisy sample towards the data manifold at every noise level.

Faster sampling#

A thousand network evaluations per image is slow. Improvements:

  • DDIM (Song, Meng & Ermon, 2021): a non-Markovian, deterministic sampler using the same trained model, producing good samples in 20–50 steps and enabling latent interpolation and inversion.
  • Higher-order ODE solvers (DPM-Solver, Heun/EDM samplers): high quality in 10–25 steps.
  • Distillation: progressive distillation, consistency models and adversarial distillation produce images in 1–4 steps.

Why diffusion won#

  • Stable training — a simple regression loss, no adversarial dynamics.
  • Excellent mode coverage and diversity, with sample quality that surpassed GANs on standard benchmarks (Dhariwal & Nichol, 2021, "Diffusion Models Beat GANs on Image Synthesis").
  • Flexible conditioning — text, images, masks, poses — and guidance techniques (next lectures).
  • Scalability to video, audio, 3-D and scientific data (protein structure generation, weather forecasting).
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Guidance in Diffusion Models: Classifier and Classifier-Free Guidance

Conditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.

Advanced⏱ 5 min#198
✨ Generative AI & LLMs

Generative Adversarial Networks: The Generator–Discriminator Game

GANs train a generator to fool a discriminator in a two-player game. We derive the minimax objective and its optimum, study training dynamics, mode collapse and the non-saturating loss, and build a DCGAN.

Advanced⏱ 5 min#193
✨ Generative AI & LLMs

Variational Autoencoders: Probabilistic Latent Spaces

VAEs turn autoencoders into generative models by learning a smooth, probabilistic latent space. We derive the evidence lower bound, the reparameterisation trick and the KL term, and discuss blurriness, posterior collapse and β-VAE.

Advanced⏱ 5 min#192