A conditional diffusion model trained on image–caption pairs learns $p(\mathbf{x} \mid \mathbf{c})$. Sampled naively, its images are often only loosely related to the prompt and of mediocre quality. Guidance techniques steer sampling towards images that strongly match the condition. Classifier-free guidance (CFG) is arguably the single most important trick behind the quality of modern text-to-image and text-to-video systems.
Score functions and conditioning#
Recall that a diffusion model estimates the score $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ at each noise level. For conditional generation we want the conditional score. By Bayes' rule, $p(\mathbf{x}_t \mid \mathbf{c}) \propto p(\mathbf{c} \mid \mathbf{x}_t)\,p(\mathbf{x}_t)$, so
The conditional score is the unconditional score plus the gradient of a "classifier" that tells how well $\mathbf{x}_t$ matches the condition.
Classifier guidance#
Dhariwal and Nichol (2021) trained a separate classifier $p_\phi(\mathbf{c} \mid \mathbf{x}_t)$ on noisy images and scaled its gradient by a guidance weight $s$:
This corresponds to sampling from a sharpened distribution $\propto p(\mathbf{x})\,p(\mathbf{c} \mid \mathbf{x})^s$. Larger $s$ yields images that are more clearly of the class and higher in perceived quality, but less diverse. Drawbacks: you must train an extra classifier on noisy data, and its gradients can behave like adversarial perturbations.
Classifier-free guidance#
Ho and Salimans (2022) removed the classifier. Train one network to predict noise both with and without the condition, by randomly replacing the condition with a null token $\varnothing$ (e.g. an empty caption) for some fraction of training examples (often 10–20%). At sampling time, combine the two predictions:
- $w = 0$: unconditional generation.
- $w = 1$: plain conditional generation.
- $w > 1$: extrapolate beyond the conditional prediction, in the direction that distinguishes "with prompt" from "without prompt".
Why does this work? Since noise predictions are proportional to (negative) scores, the difference $\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \mathbf{c}) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing)$ is proportional to $-\nabla\log p(\mathbf{x}_t \mid \mathbf{c}) + \nabla\log p(\mathbf{x}_t) = -\nabla\log p(\mathbf{c} \mid \mathbf{x}_t)$ — an implicit classifier gradient. CFG therefore approximates classifier guidance using the diffusion model itself.
import torch
def cfg_noise(model, x_t, t, cond_emb, null_emb, w=7.0):
"""One classifier-free-guided noise prediction (batched: unconditional + conditional)."""
x_in = torch.cat([x_t, x_t])
t_in = torch.cat([t, t])
c_in = torch.cat([null_emb, cond_emb])
eps_uncond, eps_cond = model(x_in, t_in, c_in).chunk(2)
return eps_uncond + w * (eps_cond - eps_uncond)
# Training-side: condition dropout
def maybe_drop_condition(cond_emb, null_emb, p_drop=0.1):
drop = torch.rand(cond_emb.size(0), *[1] * (cond_emb.dim() - 1)) < p_drop
return torch.where(drop, null_emb.expand_as(cond_emb), cond_emb)CFG doubles the cost of each sampling step (two forward passes, usually batched together).
The guidance scale trade-off#
| Guidance scale $w$ | Effect |
|---|---|
| 1 | Diverse but often weakly aligned, lower perceived quality |
| 3–8 | Typical sweet spot for text-to-image |
| 10–20 | Strong prompt adherence, less diversity, oversaturated colours and artefacts |
Increasing $w$ moves along a fidelity–diversity curve: FID first improves, then worsens as diversity collapses, while CLIP-score (prompt alignment) keeps increasing. Practitioners tune $w$ per model and use case.
Negative prompts#
Replace the null condition with a negative prompt $\mathbf{c}_{\text{neg}}$:
Sampling moves towards the prompt and away from the negative prompt ("blurry, extra fingers, watermark").
Refinements#
- Guidance rescaling / dynamic thresholding (Imagen) prevents oversaturation at high $w$ by clipping or rescaling predicted pixel values.
- Guidance intervals: apply guidance only in the middle range of noise levels, improving diversity.
- Guidance distillation: train a student to reproduce guided outputs in a single forward pass, halving cost.
- Autoguidance: guide with a weaker version of the model instead of an unconditional one.
- CFG applies beyond images — text-to-audio, text-to-video, and even language models (e.g. to strengthen adherence to instructions or context).