Pixel-space diffusion at high resolution is expensive: every denoising step processes hundreds of thousands of pixel values. In 2022, Rombach, Blattmann, Lorenz, Esser and Ommer published Latent Diffusion Models (LDMs), which run diffusion in the compact latent space of an autoencoder. Released openly as Stable Diffusion, it made high-quality text-to-image generation run on consumer GPUs and sparked an explosion of creative tools, research — and debate.
Architecture overview#
A latent diffusion model has three components:
- Autoencoder (VAE): the encoder $\mathcal{E}$ compresses an image $\mathbf{x}$ (e.g. $512 \times 512 \times 3$) into a latent $\mathbf{z} = \mathcal{E}(\mathbf{x})$ (e.g. $64 \times 64 \times 4$ — a factor of 48 fewer values); the decoder $\mathcal{D}$ maps latents back to images. It is trained once, with reconstruction, perceptual and adversarial losses plus a small KL penalty, so that latents keep perceptually important detail.
- Denoising network: a U-Net (or transformer) performing diffusion in latent space, with the usual noise-prediction objective:
- Conditioning encoder $\tau_\theta$: for text-to-image, a text encoder (a CLIP text encoder in Stable Diffusion 1.x; larger and multiple encoders in later versions) turns the prompt into a sequence of embeddings.
Conditioning through cross-attention#
The U-Net's intermediate feature maps attend to the text embeddings via cross-attention:
Each spatial location "looks up" the relevant words of the prompt. Visualising these cross-attention maps shows the word "dog" attending to the dog's region — the basis of attention-based editing methods like Prompt-to-Prompt.
Generation pipeline#
- Encode the prompt with the text encoder.
- Sample a random latent $\mathbf{z}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$.
- Denoise for 20–50 steps with a fast sampler, using classifier-free guidance (next lecture) to strengthen prompt adherence.
- Decode $\mathbf{z}_0$ with the VAE decoder into the final image.
# pip install diffusers transformers accelerate
import torch
from diffusers import StableDiffusionPipeline, DPMSolverMultistepScheduler
pipe = StableDiffusionPipeline.from_pretrained("stable-diffusion-v1-5/stable-diffusion-v1-5",
torch_dtype=torch.float16).to("cuda")
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config) # fast sampler
image = pipe("a watercolour illustration of children reading under a tree, soft morning light",
negative_prompt="blurry, distorted, text",
num_inference_steps=25, guidance_scale=7.0,
generator=torch.Generator("cuda").manual_seed(42)).images[0]
image.save("reading_tree.png")(Check the licence of any checkpoint you use; model licences often include use restrictions.)
Beyond text-to-image#
- Image-to-image (img2img / SDEdit): encode an input image, add noise to an intermediate step, then denoise with a new prompt — the noise strength trades faithfulness to the input against creativity.
- Inpainting: regenerate only a masked region, conditioned on the rest.
- ControlNet (Zhang et al., 2023): a trainable copy of the encoder, connected through zero-initialised convolutions, adds spatial conditions — edges, depth maps, human poses, segmentation maps — while keeping the base model frozen. It gives precise control over composition.
- Personalisation: DreamBooth fine-tunes on a few images of a subject bound to a rare token; Textual Inversion learns a new word embedding; LoRA adapters fine-tune efficiently and are easily shared.
- Super-resolution and upscaling with dedicated diffusion upscalers.
- Video: extending latents with a time dimension and temporal attention.
Architecture evolution#
Later models replaced parts of the recipe: larger and multiple text encoders (including T5-style encoders that improve text rendering and prompt understanding), higher-resolution training, and transformer backbones (Diffusion Transformers, DiT; Multimodal DiT) trained with flow matching / rectified flow. The core idea — generate in a learned latent space — has remained.