๐Ÿ‘๏ธ Computer Vision ยท Lecture 11 of 27

Object Detection I: R-CNN, Fast R-CNN and Faster R-CNN

Detection asks what objects are in an image and where. We define bounding boxes, IoU and mAP, then trace the two-stage R-CNN family from selective search to region proposal networks and feature pyramids.

Classification says "there is a person in this image". Object detection says "there are three people, here, here and here, and a bicycle there". It must output a variable number of bounding boxes, each with a class label and a confidence score. Detection underpins autonomous driving, retail analytics, wildlife counting from camera traps, and counting shelters or vehicles in satellite imagery after disasters.

Representing and evaluating boxes#

A box is typically $(x_{\min}, y_{\min}, x_{\max}, y_{\max})$ or (centre, width, height). Box overlap is measured by Intersection over Union:

$$ \text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} $$

A prediction is a true positive if its IoU with an unmatched ground-truth box of the same class exceeds a threshold (commonly 0.5).

Mean Average Precision (mAP): for each class, rank predictions by confidence, compute the precisionโ€“recall curve and its area (AP), then average over classes. The COCO benchmark reports mAP averaged over IoU thresholds 0.50 to 0.95 in steps of 0.05 (written mAP@[.5:.95]), rewarding precise localisation.

python
def iou(a, b):
    x1, y1 = max(a[0], b[0]), max(a[1], b[1])
    x2, y2 = min(a[2], b[2]), min(a[3], b[3])
    inter = max(0, x2 - x1) * max(0, y2 - y1)
    area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
    return inter / (area(a) + area(b) - inter + 1e-9)

print(iou([10, 10, 50, 50], [30, 30, 70, 70]))   # 400 / 2800 โ‰ˆ 0.143

Non-maximum suppression (NMS)#

Detectors produce many overlapping boxes for the same object. NMS keeps the highest-scoring box and removes others that overlap it by more than a threshold (e.g. IoU > 0.5), repeating for the remaining boxes. Variants: Soft-NMS (decays scores instead of removing), class-aware NMS.

R-CNN (2014)#

Girshick et al.'s Regions with CNN features:

  1. Generate ~2,000 region proposals with selective search (a classical segmentation-based method).
  2. Warp each region to a fixed size and run a CNN on each one to extract features.
  3. Classify each region with SVMs; refine boxes with bounding-box regression.

It dramatically improved accuracy on PASCAL VOC but was extremely slow โ€” tens of seconds per image โ€” because the CNN ran thousands of times.

Fast R-CNN (2015)#

Run the CNN once on the whole image to get a feature map. For each proposal, RoI pooling extracts a fixed-size feature grid from the corresponding feature-map region. A single network then outputs class scores and box refinements, trained jointly with a multi-task loss (classification + smooth-L1 box regression). Much faster โ€” but selective search proposals remained the bottleneck.

Faster R-CNN (2015)#

Ren, He, Girshick and Sun replaced selective search with a learned Region Proposal Network (RPN) that shares the backbone features:

  • At each feature-map location, place $k$ anchor boxes of different scales and aspect ratios (e.g. 3 ร— 3 = 9).
  • For each anchor, predict an objectness score and box offsets.
  • Keep the top proposals after NMS; pass them to the Fast R-CNN head.

Box regression predicts offsets relative to an anchor $(x_a, y_a, w_a, h_a)$:

$$ t_x = \frac{x - x_a}{w_a}, \quad t_y = \frac{y - y_a}{h_a}, \quad t_w = \log\frac{w}{w_a}, \quad t_h = \log\frac{h}{h_a} $$

This parameterisation makes regression targets scale-invariant. Faster R-CNN became a near-real-time, end-to-end trainable detector and the standard two-stage design.

Feature Pyramid Networks (2017)#

Small objects are hard for detectors that use only the coarse final feature map. FPN (Lin et al.) builds a top-down pathway with lateral connections, producing semantically strong feature maps at multiple resolutions. Small objects are detected on high-resolution levels, large ones on coarse levels. Faster R-CNN + FPN became a very strong baseline.

python
import torch
from torchvision.models.detection import fasterrcnn_resnet50_fpn_v2, FasterRCNN_ResNet50_FPN_V2_Weights
from torchvision.models.detection.faster_rcnn import FastRCNNPredictor

weights = FasterRCNN_ResNet50_FPN_V2_Weights.DEFAULT
model = fasterrcnn_resnet50_fpn_v2(weights=weights).eval()
img = torch.rand(3, 480, 640)
with torch.no_grad():
    out = model([img])[0]
print(out["boxes"].shape, out["labels"][:5], out["scores"][:5])

# Fine-tuning for your own classes (e.g. 3 classes + background)
in_feat = model.roi_heads.box_predictor.cls_score.in_features
model.roi_heads.box_predictor = FastRCNNPredictor(in_feat, num_classes=4)
# Training: model.train(); loss_dict = model(images, targets); sum(loss_dict.values()).backward()

Targets are dictionaries with boxes (Nร—4) and labels (N) per image.

Two-stage vs one-stage#

Two-stage detectors (propose, then classify and refine) are typically accurate, especially for small objects, but slower. One-stage detectors (YOLO, SSD, RetinaNet โ€” next lecture) predict boxes directly in one pass and are faster. The gap in accuracy has narrowed considerably over the years.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

Data Augmentation for Computer Vision

Augmentation multiplies your data by encoding known invariances. We survey geometric, photometric and occlusion augmentations, automated policies like RandAugment, mixing methods, and augmentation for detection and segmentation.

Intermediateโฑ 5 min#144
๐Ÿ‘๏ธ Computer Vision

Object Detection II: YOLO and Real-Time Detection

"You Only Look Once" reframed detection as a single regression problem, enabling real-time performance. We study the original YOLO grid formulation, its evolution, anchor-free heads, and practical training with modern tools.

Intermediateโฑ 5 min#146
๐Ÿ‘๏ธ Computer Vision

Building an Image Classification Pipeline End to End

A practical walkthrough of a real image classifier โ€” collecting and splitting data, preprocessing, choosing a pretrained backbone, training, evaluating per class, inspecting errors and exporting the model.

Beginnerโฑ 5 min#143