👁️ Computer Vision · Lecture 19 of 27

Human Pose Estimation

Pose estimation locates body joints in images and video. We cover keypoint heatmap regression, top-down versus bottom-up approaches, part affinity fields, evaluation with OKS, 3-D pose, and applications from health to sport.

Where are a person's shoulders, elbows, wrists, hips, knees and ankles? Human pose estimation answers this by locating keypoints (joints) and connecting them into a skeleton. It supports physiotherapy and rehabilitation apps that check exercise form, fall detection for elderly care, sports analysis, animation and motion capture without suits, sign-language recognition and human–robot interaction.

Problem formulation#

For each person, predict $K$ keypoints (COCO uses 17: nose, eyes, ears, shoulders, elbows, wrists, hips, knees, ankles), each with coordinates $(x_k, y_k)$ and a visibility or confidence score.

Direct regression vs heatmaps#

Direct regression (DeepPose, 2014) predicts coordinates with a fully connected layer. It is simple but struggles to learn precise spatial mappings.

Heatmap regression became standard: for each keypoint, the network outputs a 2-D map $H_k$ whose values peak at the joint's location. The training target is a Gaussian centred on the ground truth:

$$ H_k^*(x, y) = \exp\left(-\frac{(x - x_k)^2 + (y - y_k)^2}{2\sigma^2}\right) $$

trained with per-pixel MSE. At inference, the keypoint is the argmax of the heatmap (with sub-pixel refinement). Heatmaps preserve spatial structure and let fully convolutional networks exploit locality.

Architectures that keep high resolution matter here: Stacked Hourglass networks (repeated down/up-sampling with skips), and HRNet, which maintains a high-resolution branch throughout while exchanging information with lower-resolution branches.

Multi-person pose: top-down vs bottom-up#

ApproachHowProsCons
Top-downDetect each person, crop, run single-person pose estimationHigh accuracyCost grows with number of people; depends on detector
Bottom-upDetect all keypoints in the image, then group them into peopleConstant cost regardless of crowd sizeGrouping is hard in crowds

OpenPose (Cao et al., 2017), a landmark bottom-up method, predicts keypoint heatmaps plus Part Affinity Fields (PAFs) — 2-D vector fields encoding the direction of limbs between joints. Grouping keypoints into skeletons becomes a matching problem scored by integrating the PAF along candidate limbs. It ran in real time on multiple people.

Modern one-stage methods (e.g. pose heads in YOLO-style detectors) predict boxes and keypoints together, and transformer-based models (e.g. ViTPose) achieve strong accuracy with plain ViT backbones.

Evaluation: Object Keypoint Similarity#

Accuracy is measured with OKS, analogous to IoU for keypoints:

$$ \text{OKS} = \frac{\sum_k\exp\left(-\frac{d_k^2}{2s^2\kappa_k^2}\right)\delta(v_k > 0)}{\sum_k\delta(v_k > 0)} $$

where $d_k$ is the distance between predicted and true keypoint $k$, $s$ is the object scale, and $\kappa_k$ is a per-keypoint constant reflecting annotation variability (hips are harder to place precisely than eyes). AP is then averaged over OKS thresholds, just as detection AP is averaged over IoU thresholds. Single-person benchmarks also use PCK (percentage of correct keypoints within a normalised distance).

Using a pretrained model#

python
import torch
from torchvision.models.detection import keypointrcnn_resnet50_fpn, KeypointRCNN_ResNet50_FPN_Weights
from torchvision.io import read_image
from torchvision.transforms.functional import convert_image_dtype

weights = KeypointRCNN_ResNet50_FPN_Weights.DEFAULT
model = keypointrcnn_resnet50_fpn(weights=weights).eval()
img = convert_image_dtype(read_image("people.jpg"), torch.float)
with torch.no_grad():
    out = model([img])[0]
keep = out["scores"] > 0.8
print("people:", int(keep.sum()), " keypoints tensor:", out["keypoints"][keep].shape)  # (P, 17, 3)
names = weights.meta["keypoint_names"]
print(dict(zip(names[:5], out["keypoints"][keep][0, :5, :2].round().tolist())))

From 2-D to 3-D and to action#

  • 3-D pose estimation lifts 2-D keypoints to 3-D, either from multiple calibrated cameras (triangulation) or from a single image by learned priors (a 2-D-to-3-D "lifting" network), often exploiting temporal consistency in video.
  • Body models such as SMPL represent full 3-D body shape and pose with a compact set of parameters.
  • Action recognition from skeletons: sequences of keypoints feed temporal models or graph convolutional networks treating the skeleton as a graph (ST-GCN), giving privacy-friendlier activity recognition than raw video.

Challenges#

Occlusion (people behind objects or each other), unusual poses, crowded scenes, motion blur, loose clothing (e.g. long robes hide leg joints), low resolution, and — importantly — dataset bias: training sets over-represent certain body types, clothing and activities. Evaluate on the population you intend to serve.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

Video Understanding: Action Recognition and Temporal Modelling

Video adds time to vision. We study optical flow, two-stream networks, 3-D convolutions, (2+1)-D factorisation, video transformers and tracking, plus the computational tricks that make video models practical.

Advanced⏱ 5 min#155
👁️ Computer Vision

CLIP: Connecting Images and Language

CLIP learns a shared embedding space for images and text from hundreds of millions of image–caption pairs. We explain its contrastive training, zero-shot classification with prompts, retrieval, limitations and its role in generative models.

Intermediate⏱ 5 min#152
👁️ Computer Vision

Face Recognition: Metric Learning, ArcFace and Responsible Use

Face recognition maps faces to embeddings where the same person is close. We study the pipeline, triplet and angular-margin losses, verification versus identification, evaluation — and the serious ethical questions this technology raises.

Advanced⏱ 6 min#154