Type "a child reading beneath a banyan tree at sunset, watercolour style" and receive a finished illustration in seconds; describe a scene and receive a short video. Text-to-image and text-to-video models have moved from research demos to everyday creative tools. Building on the diffusion lectures, this lecture takes a systems view: how these models are designed, controlled, evaluated and governed.
Generations of text-to-image models#
| Approach | Examples | Idea |
|---|---|---|
| GAN-based | StackGAN, AttnGAN | Conditional GANs on text embeddings (limited quality) |
| Autoregressive over image tokens | DALL·E (2021), Parti | Tokenise images (VQ-VAE/VQGAN); generate tokens like text |
| Pixel diffusion with text encoders | GLIDE, Imagen | Diffusion guided by text; Imagen showed large text encoders (T5) matter |
| Latent diffusion | Stable Diffusion | Diffusion in compressed latent space; U-Net + cross-attention |
| Diffusion/flow transformers | DiT-based models, SD3, recent systems | Transformer backbones trained with flow matching, multiple text encoders |
A key finding from Imagen (Saharia et al., 2022): scaling the text encoder improved image–text alignment more than scaling the image model — understanding the prompt is half the battle. Improved caption data also matters: DALL·E 3 (2023) trained on highly descriptive synthetic recaptions of images, markedly improving prompt following.
From images to video#
Video adds a time dimension and demands temporal consistency (objects keep their identity, physics looks plausible). Approaches:
- Extend image diffusion models with temporal layers (temporal attention/convolution), often fine-tuning from image models.
- Spatio-temporal latent representations: compress video with a video autoencoder into latent "patches" in space and time, then generate with a diffusion transformer over these patches — the approach described publicly for several recent video models.
- Cascades: generate low-resolution keyframes, then interpolate frames and upscale.
Video generation is far more computationally expensive and still struggles with long durations, consistent object permanence, fine hand motion and physically accurate interactions.
Control beyond text#
Text alone is an imprecise specification. Practical tools add control:
- Image prompts and reference images (image-to-image, IP-Adapter-style conditioning).
- Spatial control: ControlNet with edges, depth, pose or segmentation maps; regional prompting.
- Inpainting and outpainting for editing parts of an image.
- Personalisation: DreamBooth, LoRA adapters for a subject or style.
- Instruction-based editing: "make it night-time" (e.g. InstructPix2Pix).
- Camera and motion control for video.
Evaluation#
- FID and related metrics for realism and diversity against reference image sets.
- CLIP score and VQA-based alignment metrics (does the image contain what the prompt asked for? — e.g. TIFA, compositional benchmarks such as GenEval testing counting, colours, positions).
- Text rendering accuracy.
- Human preference studies and learned preference models (e.g. trained on pairwise human choices).
- For video: temporal consistency, motion quality, and prompt adherence over time.
Compositional prompts ("a red cube on top of a blue sphere, to the left of a green cone") remain challenging — counting, spatial relations and attribute binding are common failure points.
Risks and responsible deployment#
Mitigations used across the industry:
- Training-data filtering and opt-out mechanisms.
- Prompt and output safety classifiers; restrictions on generating real, identifiable people.
- Provenance: C2PA content credentials attach cryptographically signed metadata describing how media was created or edited; invisible watermarks (e.g. SynthID) embed signals detectable by tools.
- Disclosure policies: label AI-generated media, particularly in news, advertising and humanitarian communication.
- Red-teaming before release.