✨ Generative AI & LLMs · Lecture 28 of 30

Multimodal Models: Vision–Language and Beyond

Multimodal models understand and generate across text, images, audio and video. We cover fusion strategies, vision–language model architectures (LLaVA-style), training stages, capabilities like document understanding and VQA, and known failure modes.

Humans experience the world through many senses at once. Multimodal AI systems likewise combine text with images, audio, video and other signals: they describe photos for blind users, read charts and scanned forms, answer questions about diagrams, transcribe and translate speech, and generate images from descriptions. Modern frontier assistants are natively multimodal. This lecture explains how such models are built.

Why multimodal?#

  • Many tasks are inherently multimodal: document understanding (text + layout + images), medical reports with images, instructional videos.
  • Modalities complement each other: images ground words in perception; text provides abstract knowledge.
  • A single interface — "show it and ask" — is natural and accessible.

Fusion strategies#

StrategyHowExamples
Dual encoders (late fusion)Separate encoders, aligned embedding spaceCLIP, SigLIP — retrieval, zero-shot classification
Cross-attention fusionLanguage model attends to visual features through inserted cross-attention layersFlamingo
Projection into the LM (early fusion of tokens)Visual features mapped into "visual tokens" fed to the LLM alongside textLLaVA, many open VLMs
Natively multimodalOne model trained from the start on interleaved modalities, possibly with discrete image/audio tokensRecent frontier models

The LLaVA recipe#

A simple, influential open architecture (Liu et al., 2023):

  1. Vision encoder: a pretrained CLIP/SigLIP ViT turns an image into patch features.
  2. Projector: a linear layer or small MLP maps patch features into the LLM's embedding space → a sequence of visual tokens.
  3. Language model: a pretrained LLM processes [visual tokens] + [text tokens] and generates the answer.

Training in two stages:

  • Stage 1 — alignment: freeze the vision encoder and LLM; train only the projector on image–caption pairs so visual tokens "speak the LLM's language".
  • Stage 2 — visual instruction tuning: train projector and LLM (fully or with LoRA) on instruction data about images — conversations, detailed descriptions, reasoning questions — much of it generated with the help of strong text models from image annotations.

Refinements include higher input resolution (tiling images into crops), better projectors, more diverse instruction data (charts, documents, OCR, multilingual), and video by sampling frames.

python
# Using an open vision-language model (sketch with Hugging Face)
from transformers import AutoProcessor, LlavaForConditionalGeneration
from PIL import Image
import torch

model_id = "llava-hf/llava-1.5-7b-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = LlavaForConditionalGeneration.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")

image = Image.open("water_point.jpg")
prompt = "USER: <image>\nDescribe any damage to the water point and whether it looks safe to use. ASSISTANT:"
inputs = processor(images=image, text=prompt, return_tensors="pt").to(model.device, torch.float16)
out = model.generate(**inputs, max_new_tokens=120, do_sample=False)
print(processor.decode(out[0], skip_special_tokens=True))

Capabilities#

  • Captioning and detailed description — accessibility (alt text), cataloguing.
  • Visual question answering (VQA) — "How many people are in the queue?"
  • Document, chart and screenshot understanding — reading forms, tables, infographics; OCR-free extraction.
  • Visual reasoning — diagrams, maths figures, spatial relations.
  • Grounding — pointing to regions (boxes) that correspond to phrases.
  • Speech and audio — transcription, spoken dialogue, sound understanding in audio-capable models.
  • Video — summarising and answering questions about clips.

Failure modes#

Evaluation#

VQA accuracy on benchmarks (VQAv2, TextVQA, DocVQA, ChartQA, MMMU for expert multimodal reasoning), hallucination rates, and — for real applications — task-specific test sets with images from your setting (lighting, devices, languages on signs and documents).

Multimodal generation#

The reverse direction — generating images, audio and video from text — uses diffusion and flow models (see the latent diffusion lecture), autoregressive models over discrete image tokens, and increasingly unified models that can both understand and generate multiple modalities within one system.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

AI Agents and Tool Use: LLMs That Act

Agents let LLMs call tools, observe results and pursue multi-step goals. We cover function calling, the ReAct loop, planning and memory, multi-agent patterns, evaluation, and — critically — safety, permissions and human oversight.

Intermediate⏱ 5 min#217
✨ Generative AI & LLMs

Text-to-Image and Text-to-Video Generation: Systems, Control and Provenance

A systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.

Intermediate⏱ 5 min#219
✨ Generative AI & LLMs

Evaluating Large Language Models: Benchmarks, Arenas and Custom Evals

How good is an LLM — and for what? We survey capability benchmarks, human-preference arenas, safety evaluations, contamination and saturation problems, and how to build task-specific evaluation suites for your own applications.

Intermediate⏱ 5 min#216