Where are a person's shoulders, elbows, wrists, hips, knees and ankles? Human pose estimation answers this by locating keypoints (joints) and connecting them into a skeleton. It supports physiotherapy and rehabilitation apps that check exercise form, fall detection for elderly care, sports analysis, animation and motion capture without suits, sign-language recognition and human–robot interaction.
Problem formulation#
For each person, predict $K$ keypoints (COCO uses 17: nose, eyes, ears, shoulders, elbows, wrists, hips, knees, ankles), each with coordinates $(x_k, y_k)$ and a visibility or confidence score.
Direct regression vs heatmaps#
Direct regression (DeepPose, 2014) predicts coordinates with a fully connected layer. It is simple but struggles to learn precise spatial mappings.
Heatmap regression became standard: for each keypoint, the network outputs a 2-D map $H_k$ whose values peak at the joint's location. The training target is a Gaussian centred on the ground truth:
trained with per-pixel MSE. At inference, the keypoint is the argmax of the heatmap (with sub-pixel refinement). Heatmaps preserve spatial structure and let fully convolutional networks exploit locality.
Architectures that keep high resolution matter here: Stacked Hourglass networks (repeated down/up-sampling with skips), and HRNet, which maintains a high-resolution branch throughout while exchanging information with lower-resolution branches.
Multi-person pose: top-down vs bottom-up#
| Approach | How | Pros | Cons |
|---|---|---|---|
| Top-down | Detect each person, crop, run single-person pose estimation | High accuracy | Cost grows with number of people; depends on detector |
| Bottom-up | Detect all keypoints in the image, then group them into people | Constant cost regardless of crowd size | Grouping is hard in crowds |
OpenPose (Cao et al., 2017), a landmark bottom-up method, predicts keypoint heatmaps plus Part Affinity Fields (PAFs) — 2-D vector fields encoding the direction of limbs between joints. Grouping keypoints into skeletons becomes a matching problem scored by integrating the PAF along candidate limbs. It ran in real time on multiple people.
Modern one-stage methods (e.g. pose heads in YOLO-style detectors) predict boxes and keypoints together, and transformer-based models (e.g. ViTPose) achieve strong accuracy with plain ViT backbones.
Evaluation: Object Keypoint Similarity#
Accuracy is measured with OKS, analogous to IoU for keypoints:
where $d_k$ is the distance between predicted and true keypoint $k$, $s$ is the object scale, and $\kappa_k$ is a per-keypoint constant reflecting annotation variability (hips are harder to place precisely than eyes). AP is then averaged over OKS thresholds, just as detection AP is averaged over IoU thresholds. Single-person benchmarks also use PCK (percentage of correct keypoints within a normalised distance).
Using a pretrained model#
import torch
from torchvision.models.detection import keypointrcnn_resnet50_fpn, KeypointRCNN_ResNet50_FPN_Weights
from torchvision.io import read_image
from torchvision.transforms.functional import convert_image_dtype
weights = KeypointRCNN_ResNet50_FPN_Weights.DEFAULT
model = keypointrcnn_resnet50_fpn(weights=weights).eval()
img = convert_image_dtype(read_image("people.jpg"), torch.float)
with torch.no_grad():
out = model([img])[0]
keep = out["scores"] > 0.8
print("people:", int(keep.sum()), " keypoints tensor:", out["keypoints"][keep].shape) # (P, 17, 3)
names = weights.meta["keypoint_names"]
print(dict(zip(names[:5], out["keypoints"][keep][0, :5, :2].round().tolist())))From 2-D to 3-D and to action#
- 3-D pose estimation lifts 2-D keypoints to 3-D, either from multiple calibrated cameras (triangulation) or from a single image by learned priors (a 2-D-to-3-D "lifting" network), often exploiting temporal consistency in video.
- Body models such as SMPL represent full 3-D body shape and pose with a compact set of parameters.
- Action recognition from skeletons: sequences of keypoints feed temporal models or graph convolutional networks treating the skeleton as a graph (ST-GCN), giving privacy-friendlier activity recognition than raw video.
Challenges#
Occlusion (people behind objects or each other), unusual poses, crowded scenes, motion blur, loose clothing (e.g. long robes hide leg joints), low resolution, and — importantly — dataset bias: training sets over-represent certain body types, clothing and activities. Evaluate on the population you intend to serve.