✨ Generative AI & LLMs · Lecture 8 of 30

Guidance in Diffusion Models: Classifier and Classifier-Free Guidance

Conditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.

A conditional diffusion model trained on image–caption pairs learns $p(\mathbf{x} \mid \mathbf{c})$. Sampled naively, its images are often only loosely related to the prompt and of mediocre quality. Guidance techniques steer sampling towards images that strongly match the condition. Classifier-free guidance (CFG) is arguably the single most important trick behind the quality of modern text-to-image and text-to-video systems.

Score functions and conditioning#

Recall that a diffusion model estimates the score $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ at each noise level. For conditional generation we want the conditional score. By Bayes' rule, $p(\mathbf{x}_t \mid \mathbf{c}) \propto p(\mathbf{c} \mid \mathbf{x}_t)\,p(\mathbf{x}_t)$, so

$$ \nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t \mid \mathbf{c}) = \nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t) + \nabla_{\mathbf{x}_t}\log p(\mathbf{c} \mid \mathbf{x}_t) $$

The conditional score is the unconditional score plus the gradient of a "classifier" that tells how well $\mathbf{x}_t$ matches the condition.

Classifier guidance#

Dhariwal and Nichol (2021) trained a separate classifier $p_\phi(\mathbf{c} \mid \mathbf{x}_t)$ on noisy images and scaled its gradient by a guidance weight $s$:

$$ \nabla\log p_s(\mathbf{x}_t \mid \mathbf{c}) = \nabla\log p(\mathbf{x}_t) + s\,\nabla\log p_\phi(\mathbf{c} \mid \mathbf{x}_t) $$

This corresponds to sampling from a sharpened distribution $\propto p(\mathbf{x})\,p(\mathbf{c} \mid \mathbf{x})^s$. Larger $s$ yields images that are more clearly of the class and higher in perceived quality, but less diverse. Drawbacks: you must train an extra classifier on noisy data, and its gradients can behave like adversarial perturbations.

Classifier-free guidance#

Ho and Salimans (2022) removed the classifier. Train one network to predict noise both with and without the condition, by randomly replacing the condition with a null token $\varnothing$ (e.g. an empty caption) for some fraction of training examples (often 10–20%). At sampling time, combine the two predictions:

$$ \tilde{\boldsymbol{\epsilon}}_\theta(\mathbf{x}_t, \mathbf{c}) = \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing) + w\,\big(\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \mathbf{c}) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing)\big) $$
  • $w = 0$: unconditional generation.
  • $w = 1$: plain conditional generation.
  • $w > 1$: extrapolate beyond the conditional prediction, in the direction that distinguishes "with prompt" from "without prompt".

Why does this work? Since noise predictions are proportional to (negative) scores, the difference $\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \mathbf{c}) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing)$ is proportional to $-\nabla\log p(\mathbf{x}_t \mid \mathbf{c}) + \nabla\log p(\mathbf{x}_t) = -\nabla\log p(\mathbf{c} \mid \mathbf{x}_t)$ — an implicit classifier gradient. CFG therefore approximates classifier guidance using the diffusion model itself.

python
import torch

def cfg_noise(model, x_t, t, cond_emb, null_emb, w=7.0):
    """One classifier-free-guided noise prediction (batched: unconditional + conditional)."""
    x_in = torch.cat([x_t, x_t])
    t_in = torch.cat([t, t])
    c_in = torch.cat([null_emb, cond_emb])
    eps_uncond, eps_cond = model(x_in, t_in, c_in).chunk(2)
    return eps_uncond + w * (eps_cond - eps_uncond)

# Training-side: condition dropout
def maybe_drop_condition(cond_emb, null_emb, p_drop=0.1):
    drop = torch.rand(cond_emb.size(0), *[1] * (cond_emb.dim() - 1)) < p_drop
    return torch.where(drop, null_emb.expand_as(cond_emb), cond_emb)

CFG doubles the cost of each sampling step (two forward passes, usually batched together).

The guidance scale trade-off#

Guidance scale $w$Effect
1Diverse but often weakly aligned, lower perceived quality
3–8Typical sweet spot for text-to-image
10–20Strong prompt adherence, less diversity, oversaturated colours and artefacts

Increasing $w$ moves along a fidelity–diversity curve: FID first improves, then worsens as diversity collapses, while CLIP-score (prompt alignment) keeps increasing. Practitioners tune $w$ per model and use case.

Negative prompts#

Replace the null condition with a negative prompt $\mathbf{c}_{\text{neg}}$:

$$ \tilde{\boldsymbol{\epsilon}} = \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \mathbf{c}_{\text{neg}}) + w\,\big(\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \mathbf{c}) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \mathbf{c}_{\text{neg}})\big) $$

Sampling moves towards the prompt and away from the negative prompt ("blurry, extra fingers, watermark").

Refinements#

  • Guidance rescaling / dynamic thresholding (Imagen) prevents oversaturation at high $w$ by clipping or rescaling predicted pixel values.
  • Guidance intervals: apply guidance only in the middle range of noise levels, improving diversity.
  • Guidance distillation: train a student to reproduce guided outputs in a single forward pass, halving cost.
  • Autoguidance: guide with a weaker version of the model instead of an unconditional one.
  • CFG applies beyond images — text-to-audio, text-to-video, and even language models (e.g. to strengthen adherence to instructions or context).
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Diffusion Models: Generating by Learning to Denoise

Diffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.

Advanced⏱ 6 min#196
✨ Generative AI & LLMs

Latent Diffusion and Stable Diffusion: Text-to-Image at Scale

Running diffusion in a compressed latent space made high-resolution text-to-image generation affordable. We dissect latent diffusion — autoencoder, U-Net with cross-attention, text encoder — and practical techniques like img2img, inpainting, ControlNet and fine-tuning.

Advanced⏱ 5 min#197
✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199