✨ Generative AI & LLMs · Lecture 7 of 30

Latent Diffusion and Stable Diffusion: Text-to-Image at Scale

Running diffusion in a compressed latent space made high-resolution text-to-image generation affordable. We dissect latent diffusion — autoencoder, U-Net with cross-attention, text encoder — and practical techniques like img2img, inpainting, ControlNet and fine-tuning.

Pixel-space diffusion at high resolution is expensive: every denoising step processes hundreds of thousands of pixel values. In 2022, Rombach, Blattmann, Lorenz, Esser and Ommer published Latent Diffusion Models (LDMs), which run diffusion in the compact latent space of an autoencoder. Released openly as Stable Diffusion, it made high-quality text-to-image generation run on consumer GPUs and sparked an explosion of creative tools, research — and debate.

Architecture overview#

A latent diffusion model has three components:

  1. Autoencoder (VAE): the encoder $\mathcal{E}$ compresses an image $\mathbf{x}$ (e.g. $512 \times 512 \times 3$) into a latent $\mathbf{z} = \mathcal{E}(\mathbf{x})$ (e.g. $64 \times 64 \times 4$ — a factor of 48 fewer values); the decoder $\mathcal{D}$ maps latents back to images. It is trained once, with reconstruction, perceptual and adversarial losses plus a small KL penalty, so that latents keep perceptually important detail.
  2. Denoising network: a U-Net (or transformer) performing diffusion in latent space, with the usual noise-prediction objective:
$$ \mathcal{L}_{\text{LDM}} = \mathbb{E}_{\mathcal{E}(\mathbf{x}), \mathbf{c}, \boldsymbol{\epsilon}, t}\left[\big\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta(\mathbf{z}_t, t, \tau_\theta(\mathbf{c}))\big\|^2\right] $$
  1. Conditioning encoder $\tau_\theta$: for text-to-image, a text encoder (a CLIP text encoder in Stable Diffusion 1.x; larger and multiple encoders in later versions) turns the prompt into a sequence of embeddings.

Conditioning through cross-attention#

The U-Net's intermediate feature maps attend to the text embeddings via cross-attention:

$$ \text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^\top}{\sqrt{d}}\right)\mathbf{V}, \quad \mathbf{Q} = \mathbf{W}_Q\,\varphi(\mathbf{z}_t), \; \mathbf{K} = \mathbf{W}_K\,\tau_\theta(\mathbf{c}), \; \mathbf{V} = \mathbf{W}_V\,\tau_\theta(\mathbf{c}) $$

Each spatial location "looks up" the relevant words of the prompt. Visualising these cross-attention maps shows the word "dog" attending to the dog's region — the basis of attention-based editing methods like Prompt-to-Prompt.

Generation pipeline#

  1. Encode the prompt with the text encoder.
  2. Sample a random latent $\mathbf{z}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$.
  3. Denoise for 20–50 steps with a fast sampler, using classifier-free guidance (next lecture) to strengthen prompt adherence.
  4. Decode $\mathbf{z}_0$ with the VAE decoder into the final image.
python
# pip install diffusers transformers accelerate
import torch
from diffusers import StableDiffusionPipeline, DPMSolverMultistepScheduler

pipe = StableDiffusionPipeline.from_pretrained("stable-diffusion-v1-5/stable-diffusion-v1-5",
                                               torch_dtype=torch.float16).to("cuda")
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)   # fast sampler
image = pipe("a watercolour illustration of children reading under a tree, soft morning light",
             negative_prompt="blurry, distorted, text",
             num_inference_steps=25, guidance_scale=7.0,
             generator=torch.Generator("cuda").manual_seed(42)).images[0]
image.save("reading_tree.png")

(Check the licence of any checkpoint you use; model licences often include use restrictions.)

Beyond text-to-image#

  • Image-to-image (img2img / SDEdit): encode an input image, add noise to an intermediate step, then denoise with a new prompt — the noise strength trades faithfulness to the input against creativity.
  • Inpainting: regenerate only a masked region, conditioned on the rest.
  • ControlNet (Zhang et al., 2023): a trainable copy of the encoder, connected through zero-initialised convolutions, adds spatial conditions — edges, depth maps, human poses, segmentation maps — while keeping the base model frozen. It gives precise control over composition.
  • Personalisation: DreamBooth fine-tunes on a few images of a subject bound to a rare token; Textual Inversion learns a new word embedding; LoRA adapters fine-tune efficiently and are easily shared.
  • Super-resolution and upscaling with dedicated diffusion upscalers.
  • Video: extending latents with a time dimension and temporal attention.

Architecture evolution#

Later models replaced parts of the recipe: larger and multiple text encoders (including T5-style encoders that improve text rendering and prompt understanding), higher-resolution training, and transformer backbones (Diffusion Transformers, DiT; Multimodal DiT) trained with flow matching / rectified flow. The core idea — generate in a learned latent space — has remained.

Limitations and responsibility#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Text-to-Image and Text-to-Video Generation: Systems, Control and Provenance

A systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.

Intermediate⏱ 5 min#219
✨ Generative AI & LLMs

Diffusion Models: Generating by Learning to Denoise

Diffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.

Advanced⏱ 6 min#196
✨ Generative AI & LLMs

Guidance in Diffusion Models: Classifier and Classifier-Free Guidance

Conditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.

Advanced⏱ 5 min#198