👁️ Computer Vision · Lecture 21 of 27

Video Understanding: Action Recognition and Temporal Modelling

Video adds time to vision. We study optical flow, two-stream networks, 3-D convolutions, (2+1)-D factorisation, video transformers and tracking, plus the computational tricks that make video models practical.

A single frame shows a person with an arm raised. Are they waving, throwing or reaching? The answer lies in motion. Video understanding extends computer vision along the time axis for tasks such as action recognition, temporal localisation of events, video captioning, object tracking and anomaly detection in surveillance or industrial footage. Video is data-rich and computationally expensive, which shapes every design choice.

Representing video#

A video clip is a 4-D tensor $T \times H \times W \times C$ — for example, 16 frames of $224 \times 224$ RGB. Consecutive frames are highly redundant, so models usually sample frames sparsely (e.g. 8–32 frames spread across the clip).

Optical flow#

Optical flow estimates, for each pixel, its apparent motion between two frames. Classical methods rely on the brightness constancy assumption: a moving point keeps its intensity, giving

$$ I_x u + I_y v + I_t = 0 $$

for flow $(u, v)$ — one equation in two unknowns (the aperture problem). Lucas–Kanade solves it by assuming constant flow in a small window; Horn–Schunck adds a global smoothness term. Deep networks (FlowNet, RAFT) now estimate flow far more accurately.

Two-stream networks#

Simonyan and Zisserman (2014) processed video with two CNNs:

  • a spatial stream on single RGB frames (appearance);
  • a temporal stream on stacks of optical-flow fields (motion);

and fused their predictions. Motion information substantially improved action recognition, but computing optical flow is expensive.

3-D convolutions#

A 3-D convolution extends kernels across time: a $3 \times 3 \times 3$ kernel slides over frames and pixels, learning spatio-temporal features directly.

  • C3D (2015) showed 3-D CNNs learn useful video features.
  • I3D (Carreira & Zisserman, 2017) "inflated" pretrained 2-D ImageNet filters into 3-D (repeating weights along time and rescaling), giving 3-D networks a strong initialisation. Trained on the large Kinetics dataset, I3D set new standards.
  • (2+1)-D convolutions (R(2+1)D, 2018) factorise a 3-D conv into a 2-D spatial conv followed by a 1-D temporal conv — fewer parameters, an extra non-linearity, easier optimisation.
  • SlowFast networks (2019) use a slow pathway (few frames, many channels) for appearance and a fast pathway (many frames, few channels) for motion — inspired by parallel processing streams in primate vision.
python
import torch
from torchvision.models.video import r2plus1d_18, R2Plus1D_18_Weights

weights = R2Plus1D_18_Weights.DEFAULT
model = r2plus1d_18(weights=weights).eval()
clip = torch.rand(1, 3, 16, 112, 112)               # (batch, channels, frames, H, W)
with torch.no_grad():
    probs = model(clip).softmax(-1)
top = probs.topk(3)
print([weights.meta["categories"][i] for i in top.indices[0]], top.values[0].round(decimals=3))

Efficient temporal modelling#

  • Temporal Segment Networks (TSN): sample one frame (or snippet) from each of several segments across the whole video and average predictions — cheap long-range coverage.
  • Temporal Shift Module (TSM): shift a fraction of channels forward and backward in time within a 2-D CNN, exchanging information between neighbouring frames at almost zero cost.

Video transformers#

Treat video as a sequence of spatio-temporal patches ("tubelets"). Full attention over all patches in all frames is expensive, so architectures factorise it:

  • TimeSformer: divided attention — temporal attention across frames at the same location, then spatial attention within a frame.
  • ViViT: several factorised variants.
  • Video Swin: local 3-D window attention.
  • VideoMAE: masked autoencoding with very high mask ratios (~90%), exploiting video's redundancy for efficient self-supervised pretraining.

Large multimodal models now also process sampled video frames alongside text for video question answering and captioning.

Other video tasks#

  • Temporal action localisation: find start and end times of actions in untrimmed videos.
  • Multi-object tracking: detect objects in each frame and link them over time — "tracking by detection" with motion models (Kalman filters) and appearance embeddings (SORT, DeepSORT, ByteTrack).
  • Video object segmentation and video anomaly detection.

Datasets and challenges#

Kinetics (hundreds of action classes), Something-Something (fine-grained actions requiring temporal reasoning, e.g. "pushing something from left to right" vs "right to left"), AVA (spatio-temporal action localisation), and Ego4D (egocentric video). A known pitfall: many actions in Kinetics can be recognised from scene context in a single frame (a swimming pool implies swimming), so strong scores do not always mean a model understands motion. Datasets like Something-Something test genuine temporal reasoning.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

Human Pose Estimation

Pose estimation locates body joints in images and video. We cover keypoint heatmap regression, top-down versus bottom-up approaches, part affinity fields, evaluation with OKS, 3-D pose, and applications from health to sport.

Intermediate⏱ 5 min#153
👁️ Computer Vision

Face Recognition: Metric Learning, ArcFace and Responsible Use

Face recognition maps faces to embeddings where the same person is close. We study the pipeline, triplet and angular-margin losses, verification versus identification, evaluation — and the serious ethical questions this technology raises.

Advanced⏱ 6 min#154
👁️ Computer Vision

3-D Vision: Stereo, Depth Estimation, Point Clouds and NeRF

Images are 2-D projections of a 3-D world. We cover camera geometry, stereo and monocular depth, structure from motion, point-cloud networks like PointNet, and neural scene representations such as NeRF and Gaussian splatting.

Advanced⏱ 5 min#156