A single frame shows a person with an arm raised. Are they waving, throwing or reaching? The answer lies in motion. Video understanding extends computer vision along the time axis for tasks such as action recognition, temporal localisation of events, video captioning, object tracking and anomaly detection in surveillance or industrial footage. Video is data-rich and computationally expensive, which shapes every design choice.
Representing video#
A video clip is a 4-D tensor $T \times H \times W \times C$ — for example, 16 frames of $224 \times 224$ RGB. Consecutive frames are highly redundant, so models usually sample frames sparsely (e.g. 8–32 frames spread across the clip).
Optical flow#
Optical flow estimates, for each pixel, its apparent motion between two frames. Classical methods rely on the brightness constancy assumption: a moving point keeps its intensity, giving
for flow $(u, v)$ — one equation in two unknowns (the aperture problem). Lucas–Kanade solves it by assuming constant flow in a small window; Horn–Schunck adds a global smoothness term. Deep networks (FlowNet, RAFT) now estimate flow far more accurately.
Two-stream networks#
Simonyan and Zisserman (2014) processed video with two CNNs:
- a spatial stream on single RGB frames (appearance);
- a temporal stream on stacks of optical-flow fields (motion);
and fused their predictions. Motion information substantially improved action recognition, but computing optical flow is expensive.
3-D convolutions#
A 3-D convolution extends kernels across time: a $3 \times 3 \times 3$ kernel slides over frames and pixels, learning spatio-temporal features directly.
- C3D (2015) showed 3-D CNNs learn useful video features.
- I3D (Carreira & Zisserman, 2017) "inflated" pretrained 2-D ImageNet filters into 3-D (repeating weights along time and rescaling), giving 3-D networks a strong initialisation. Trained on the large Kinetics dataset, I3D set new standards.
- (2+1)-D convolutions (R(2+1)D, 2018) factorise a 3-D conv into a 2-D spatial conv followed by a 1-D temporal conv — fewer parameters, an extra non-linearity, easier optimisation.
- SlowFast networks (2019) use a slow pathway (few frames, many channels) for appearance and a fast pathway (many frames, few channels) for motion — inspired by parallel processing streams in primate vision.
import torch
from torchvision.models.video import r2plus1d_18, R2Plus1D_18_Weights
weights = R2Plus1D_18_Weights.DEFAULT
model = r2plus1d_18(weights=weights).eval()
clip = torch.rand(1, 3, 16, 112, 112) # (batch, channels, frames, H, W)
with torch.no_grad():
probs = model(clip).softmax(-1)
top = probs.topk(3)
print([weights.meta["categories"][i] for i in top.indices[0]], top.values[0].round(decimals=3))Efficient temporal modelling#
- Temporal Segment Networks (TSN): sample one frame (or snippet) from each of several segments across the whole video and average predictions — cheap long-range coverage.
- Temporal Shift Module (TSM): shift a fraction of channels forward and backward in time within a 2-D CNN, exchanging information between neighbouring frames at almost zero cost.
Video transformers#
Treat video as a sequence of spatio-temporal patches ("tubelets"). Full attention over all patches in all frames is expensive, so architectures factorise it:
- TimeSformer: divided attention — temporal attention across frames at the same location, then spatial attention within a frame.
- ViViT: several factorised variants.
- Video Swin: local 3-D window attention.
- VideoMAE: masked autoencoding with very high mask ratios (~90%), exploiting video's redundancy for efficient self-supervised pretraining.
Large multimodal models now also process sampled video frames alongside text for video question answering and captioning.
Other video tasks#
- Temporal action localisation: find start and end times of actions in untrimmed videos.
- Multi-object tracking: detect objects in each frame and link them over time — "tracking by detection" with motion models (Kalman filters) and appearance embeddings (SORT, DeepSORT, ByteTrack).
- Video object segmentation and video anomaly detection.
Datasets and challenges#
Kinetics (hundreds of action classes), Something-Something (fine-grained actions requiring temporal reasoning, e.g. "pushing something from left to right" vs "right to left"), AVA (spatio-temporal action localisation), and Ego4D (egocentric video). A known pitfall: many actions in Kinetics can be recognised from scene context in a single frame (a swimming pool implies swimming), so strong scores do not always mean a model understands motion. Datasets like Something-Something test genuine temporal reasoning.