Classification says "there is a person in this image". Object detection says "there are three people, here, here and here, and a bicycle there". It must output a variable number of bounding boxes, each with a class label and a confidence score. Detection underpins autonomous driving, retail analytics, wildlife counting from camera traps, and counting shelters or vehicles in satellite imagery after disasters.
Representing and evaluating boxes#
A box is typically $(x_{\min}, y_{\min}, x_{\max}, y_{\max})$ or (centre, width, height). Box overlap is measured by Intersection over Union:
A prediction is a true positive if its IoU with an unmatched ground-truth box of the same class exceeds a threshold (commonly 0.5).
Mean Average Precision (mAP): for each class, rank predictions by confidence, compute the precisionโrecall curve and its area (AP), then average over classes. The COCO benchmark reports mAP averaged over IoU thresholds 0.50 to 0.95 in steps of 0.05 (written mAP@[.5:.95]), rewarding precise localisation.
def iou(a, b):
x1, y1 = max(a[0], b[0]), max(a[1], b[1])
x2, y2 = min(a[2], b[2]), min(a[3], b[3])
inter = max(0, x2 - x1) * max(0, y2 - y1)
area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
return inter / (area(a) + area(b) - inter + 1e-9)
print(iou([10, 10, 50, 50], [30, 30, 70, 70])) # 400 / 2800 โ 0.143Non-maximum suppression (NMS)#
Detectors produce many overlapping boxes for the same object. NMS keeps the highest-scoring box and removes others that overlap it by more than a threshold (e.g. IoU > 0.5), repeating for the remaining boxes. Variants: Soft-NMS (decays scores instead of removing), class-aware NMS.
R-CNN (2014)#
Girshick et al.'s Regions with CNN features:
- Generate ~2,000 region proposals with selective search (a classical segmentation-based method).
- Warp each region to a fixed size and run a CNN on each one to extract features.
- Classify each region with SVMs; refine boxes with bounding-box regression.
It dramatically improved accuracy on PASCAL VOC but was extremely slow โ tens of seconds per image โ because the CNN ran thousands of times.
Fast R-CNN (2015)#
Run the CNN once on the whole image to get a feature map. For each proposal, RoI pooling extracts a fixed-size feature grid from the corresponding feature-map region. A single network then outputs class scores and box refinements, trained jointly with a multi-task loss (classification + smooth-L1 box regression). Much faster โ but selective search proposals remained the bottleneck.
Faster R-CNN (2015)#
Ren, He, Girshick and Sun replaced selective search with a learned Region Proposal Network (RPN) that shares the backbone features:
- At each feature-map location, place $k$ anchor boxes of different scales and aspect ratios (e.g. 3 ร 3 = 9).
- For each anchor, predict an objectness score and box offsets.
- Keep the top proposals after NMS; pass them to the Fast R-CNN head.
Box regression predicts offsets relative to an anchor $(x_a, y_a, w_a, h_a)$:
This parameterisation makes regression targets scale-invariant. Faster R-CNN became a near-real-time, end-to-end trainable detector and the standard two-stage design.
Feature Pyramid Networks (2017)#
Small objects are hard for detectors that use only the coarse final feature map. FPN (Lin et al.) builds a top-down pathway with lateral connections, producing semantically strong feature maps at multiple resolutions. Small objects are detected on high-resolution levels, large ones on coarse levels. Faster R-CNN + FPN became a very strong baseline.
import torch
from torchvision.models.detection import fasterrcnn_resnet50_fpn_v2, FasterRCNN_ResNet50_FPN_V2_Weights
from torchvision.models.detection.faster_rcnn import FastRCNNPredictor
weights = FasterRCNN_ResNet50_FPN_V2_Weights.DEFAULT
model = fasterrcnn_resnet50_fpn_v2(weights=weights).eval()
img = torch.rand(3, 480, 640)
with torch.no_grad():
out = model([img])[0]
print(out["boxes"].shape, out["labels"][:5], out["scores"][:5])
# Fine-tuning for your own classes (e.g. 3 classes + background)
in_feat = model.roi_heads.box_predictor.cls_score.in_features
model.roi_heads.box_predictor = FastRCNNPredictor(in_feat, num_classes=4)
# Training: model.train(); loss_dict = model(images, targets); sum(loss_dict.values()).backward()Targets are dictionaries with boxes (Nร4) and labels (N) per image.
Two-stage vs one-stage#
Two-stage detectors (propose, then classify and refine) are typically accurate, especially for small objects, but slower. One-stage detectors (YOLO, SSD, RetinaNet โ next lecture) predict boxes directly in one pass and are faster. The gap in accuracy has narrowed considerably over the years.