Humans experience the world through many senses at once. Multimodal AI systems likewise combine text with images, audio, video and other signals: they describe photos for blind users, read charts and scanned forms, answer questions about diagrams, transcribe and translate speech, and generate images from descriptions. Modern frontier assistants are natively multimodal. This lecture explains how such models are built.
Why multimodal?#
- Many tasks are inherently multimodal: document understanding (text + layout + images), medical reports with images, instructional videos.
- Modalities complement each other: images ground words in perception; text provides abstract knowledge.
- A single interface — "show it and ask" — is natural and accessible.
Fusion strategies#
| Strategy | How | Examples |
|---|---|---|
| Dual encoders (late fusion) | Separate encoders, aligned embedding space | CLIP, SigLIP — retrieval, zero-shot classification |
| Cross-attention fusion | Language model attends to visual features through inserted cross-attention layers | Flamingo |
| Projection into the LM (early fusion of tokens) | Visual features mapped into "visual tokens" fed to the LLM alongside text | LLaVA, many open VLMs |
| Natively multimodal | One model trained from the start on interleaved modalities, possibly with discrete image/audio tokens | Recent frontier models |
The LLaVA recipe#
A simple, influential open architecture (Liu et al., 2023):
- Vision encoder: a pretrained CLIP/SigLIP ViT turns an image into patch features.
- Projector: a linear layer or small MLP maps patch features into the LLM's embedding space → a sequence of visual tokens.
- Language model: a pretrained LLM processes
[visual tokens] + [text tokens]and generates the answer.
Training in two stages:
- Stage 1 — alignment: freeze the vision encoder and LLM; train only the projector on image–caption pairs so visual tokens "speak the LLM's language".
- Stage 2 — visual instruction tuning: train projector and LLM (fully or with LoRA) on instruction data about images — conversations, detailed descriptions, reasoning questions — much of it generated with the help of strong text models from image annotations.
Refinements include higher input resolution (tiling images into crops), better projectors, more diverse instruction data (charts, documents, OCR, multilingual), and video by sampling frames.
# Using an open vision-language model (sketch with Hugging Face)
from transformers import AutoProcessor, LlavaForConditionalGeneration
from PIL import Image
import torch
model_id = "llava-hf/llava-1.5-7b-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = LlavaForConditionalGeneration.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")
image = Image.open("water_point.jpg")
prompt = "USER: <image>\nDescribe any damage to the water point and whether it looks safe to use. ASSISTANT:"
inputs = processor(images=image, text=prompt, return_tensors="pt").to(model.device, torch.float16)
out = model.generate(**inputs, max_new_tokens=120, do_sample=False)
print(processor.decode(out[0], skip_special_tokens=True))Capabilities#
- Captioning and detailed description — accessibility (alt text), cataloguing.
- Visual question answering (VQA) — "How many people are in the queue?"
- Document, chart and screenshot understanding — reading forms, tables, infographics; OCR-free extraction.
- Visual reasoning — diagrams, maths figures, spatial relations.
- Grounding — pointing to regions (boxes) that correspond to phrases.
- Speech and audio — transcription, spoken dialogue, sound understanding in audio-capable models.
- Video — summarising and answering questions about clips.
Failure modes#
Evaluation#
VQA accuracy on benchmarks (VQAv2, TextVQA, DocVQA, ChartQA, MMMU for expert multimodal reasoning), hallucination rates, and — for real applications — task-specific test sets with images from your setting (lighting, devices, languages on signs and documents).
Multimodal generation#
The reverse direction — generating images, audio and video from text — uses diffusion and flow models (see the latent diffusion lecture), autoregressive models over discrete image tokens, and increasingly unified models that can both understand and generate multiple modalities within one system.