✨ Generative AI & LLMs · Lecture 29 of 30

Text-to-Image and Text-to-Video Generation: Systems, Control and Provenance

A systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.

Type "a child reading beneath a banyan tree at sunset, watercolour style" and receive a finished illustration in seconds; describe a scene and receive a short video. Text-to-image and text-to-video models have moved from research demos to everyday creative tools. Building on the diffusion lectures, this lecture takes a systems view: how these models are designed, controlled, evaluated and governed.

Generations of text-to-image models#

ApproachExamplesIdea
GAN-basedStackGAN, AttnGANConditional GANs on text embeddings (limited quality)
Autoregressive over image tokensDALL·E (2021), PartiTokenise images (VQ-VAE/VQGAN); generate tokens like text
Pixel diffusion with text encodersGLIDE, ImagenDiffusion guided by text; Imagen showed large text encoders (T5) matter
Latent diffusionStable DiffusionDiffusion in compressed latent space; U-Net + cross-attention
Diffusion/flow transformersDiT-based models, SD3, recent systemsTransformer backbones trained with flow matching, multiple text encoders

A key finding from Imagen (Saharia et al., 2022): scaling the text encoder improved image–text alignment more than scaling the image model — understanding the prompt is half the battle. Improved caption data also matters: DALL·E 3 (2023) trained on highly descriptive synthetic recaptions of images, markedly improving prompt following.

From images to video#

Video adds a time dimension and demands temporal consistency (objects keep their identity, physics looks plausible). Approaches:

  • Extend image diffusion models with temporal layers (temporal attention/convolution), often fine-tuning from image models.
  • Spatio-temporal latent representations: compress video with a video autoencoder into latent "patches" in space and time, then generate with a diffusion transformer over these patches — the approach described publicly for several recent video models.
  • Cascades: generate low-resolution keyframes, then interpolate frames and upscale.

Video generation is far more computationally expensive and still struggles with long durations, consistent object permanence, fine hand motion and physically accurate interactions.

Control beyond text#

Text alone is an imprecise specification. Practical tools add control:

  • Image prompts and reference images (image-to-image, IP-Adapter-style conditioning).
  • Spatial control: ControlNet with edges, depth, pose or segmentation maps; regional prompting.
  • Inpainting and outpainting for editing parts of an image.
  • Personalisation: DreamBooth, LoRA adapters for a subject or style.
  • Instruction-based editing: "make it night-time" (e.g. InstructPix2Pix).
  • Camera and motion control for video.

Evaluation#

  • FID and related metrics for realism and diversity against reference image sets.
  • CLIP score and VQA-based alignment metrics (does the image contain what the prompt asked for? — e.g. TIFA, compositional benchmarks such as GenEval testing counting, colours, positions).
  • Text rendering accuracy.
  • Human preference studies and learned preference models (e.g. trained on pairwise human choices).
  • For video: temporal consistency, motion quality, and prompt adherence over time.

Compositional prompts ("a red cube on top of a blue sphere, to the left of a green cone") remain challenging — counting, spatial relations and attribute binding are common failure points.

Risks and responsible deployment#

Mitigations used across the industry:

  • Training-data filtering and opt-out mechanisms.
  • Prompt and output safety classifiers; restrictions on generating real, identifiable people.
  • Provenance: C2PA content credentials attach cryptographically signed metadata describing how media was created or edited; invisible watermarks (e.g. SynthID) embed signals detectable by tools.
  • Disclosure policies: label AI-generated media, particularly in news, advertising and humanitarian communication.
  • Red-teaming before release.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Latent Diffusion and Stable Diffusion: Text-to-Image at Scale

Running diffusion in a compressed latent space made high-resolution text-to-image generation affordable. We dissect latent diffusion — autoencoder, U-Net with cross-attention, text encoder — and practical techniques like img2img, inpainting, ControlNet and fine-tuning.

Advanced⏱ 5 min#197
✨ Generative AI & LLMs

Multimodal Models: Vision–Language and Beyond

Multimodal models understand and generate across text, images, audio and video. We cover fusion strategies, vision–language model architectures (LLaVA-style), training stages, capabilities like document understanding and VQA, and known failure modes.

Intermediate⏱ 5 min#218
✨ Generative AI & LLMs

Building an LLM Application End to End: From Idea to Production

A practical capstone for the Generative AI track — scoping a use case, choosing models, designing prompts, RAG and tools, evaluation, guardrails, cost and latency, deployment, monitoring and governance.

Intermediate⏱ 5 min#220