👁️ Computer Vision · Lecture 12 of 27

Object Detection II: YOLO and Real-Time Detection

"You Only Look Once" reframed detection as a single regression problem, enabling real-time performance. We study the original YOLO grid formulation, its evolution, anchor-free heads, and practical training with modern tools.

In 2016 Joseph Redmon and colleagues published YOLO — You Only Look Once. Instead of proposing regions and classifying each one, YOLO looks at the whole image once with a single network and directly predicts all boxes and classes. It ran at tens of frames per second on a GPU, bringing detection to real-time video. Successive YOLO versions have made it the most widely used family of practical detectors.

The original formulation#

YOLOv1 divides the image into an $S \times S$ grid (e.g. $7 \times 7$). The cell containing an object's centre is responsible for detecting it. Each cell predicts:

  • $B$ boxes, each with $(x, y, w, h)$ and a confidence $= P(\text{object}) \times \text{IoU}$;
  • $C$ class probabilities.

The output is a tensor of shape $S \times S \times (5B + C)$ — for PASCAL VOC with $S = 7$, $B = 2$, $C = 20$: $7 \times 7 \times 30$. Training uses a sum-of-squares loss with weights that emphasise box coordinates and down-weight cells without objects; width and height are predicted as square roots so errors on small boxes count more.

Strengths: very fast; sees the whole image, so it makes fewer background false positives than sliding-window methods. Weaknesses: each cell predicts few boxes, so crowds of small objects (flocks of birds) are hard; coarse localisation.

Evolution of YOLO#

VersionKey ideas
YOLOv2 / YOLO9000 (2017)Anchor boxes from k-means on training boxes, batch norm, higher resolution, multi-scale training
YOLOv3 (2018)Darknet-53 backbone, predictions at three scales (FPN-like), logistic class predictions
YOLOv4 (2020)"Bag of freebies" (mosaic augmentation, CIoU loss) and "bag of specials" (CSP backbone, PANet neck)
YOLOv5 onwards (community/industry)PyTorch implementations, auto-anchors, strong training recipes, easy tooling
YOLOX, YOLOv8 and laterAnchor-free heads, decoupled classification/regression heads, improved label assignment

(From version 4 onward, YOLO versions were developed by different groups, so the numbering does not indicate a single lineage.)

Anatomy of a modern one-stage detector#

  1. Backbone — extracts features (CSPDarknet, efficient CNNs).
  2. Neck — fuses multi-scale features (FPN + PAN: top-down and bottom-up paths).
  3. Head — at each location of each scale, predicts class scores and box geometry.

Anchor-free heads predict the distance from a location to the four box sides (as in FCOS) or the box centre and size directly, avoiding hand-tuned anchor shapes.

Label assignment — deciding which predictions are responsible for which ground-truth boxes — matters as much as architecture. Modern methods (SimOTA, task-aligned assignment) dynamically assign positives based on both classification and localisation quality.

Box losses based on IoU (GIoU, DIoU, CIoU) directly optimise overlap rather than coordinate differences, which aligns training with evaluation. GIoU, for instance, adds a penalty based on the smallest enclosing box $C$:

$$ \text{GIoU} = \text{IoU} - \frac{|C \setminus (A \cup B)|}{|C|} $$

so non-overlapping boxes still receive a useful gradient.

Training a YOLO model in practice#

Tools like the Ultralytics package make training straightforward. Labels use one text file per image with rows class x_center y_center width height, normalised to $[0, 1]$.

python
# pip install ultralytics
from ultralytics import YOLO

model = YOLO("yolov8n.pt")                  # small pretrained model (COCO)
model.train(data="shelters.yaml",           # paths to train/val images and class names
            epochs=100, imgsz=640, batch=16, patience=20)
metrics = model.val()                        # reports mAP50 and mAP50-95
results = model.predict("aerial_tile.jpg", conf=0.25, iou=0.5)
results[0].show()
model.export(format="onnx")                  # for deployment
yaml
# shelters.yaml
path: datasets/shelters
train: images/train
val: images/val
names: {0: tent, 1: building, 2: vehicle}

Speed–accuracy trade-offs#

YOLO families come in sizes (nano, small, medium, large, extra-large). Pick the smallest model meeting the accuracy requirement at your target frame rate on your hardware. Measure end-to-end latency including preprocessing and NMS, not just model FLOPs.

Beyond YOLO#

DETR (2020) reframed detection as set prediction with a transformer and bipartite matching, removing anchors and NMS entirely; later variants (Deformable DETR, DINO, RT-DETR) made transformer detectors fast and accurate. Open-vocabulary detectors such as OWL-ViT and Grounding DINO detect objects described by arbitrary text prompts — "detect water tanks" — without training on those classes.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

Object Detection III: SSD, RetinaNet and the Focal Loss

One-stage detectors face an extreme imbalance between background and objects. We study SSD's multi-scale default boxes, then derive RetinaNet's focal loss, which let one-stage detectors match two-stage accuracy.

Advanced⏱ 5 min#147
👁️ Computer Vision

Object Detection I: R-CNN, Fast R-CNN and Faster R-CNN

Detection asks what objects are in an image and where. We define bounding boxes, IoU and mAP, then trace the two-stage R-CNN family from selective search to region proposal networks and feature pyramids.

Intermediate⏱ 5 min#145
👁️ Computer Vision

Data Augmentation for Computer Vision

Augmentation multiplies your data by encoding known invariances. We survey geometric, photometric and occlusion augmentations, automated policies like RandAugment, mixing methods, and augmentation for detection and segmentation.

Intermediate⏱ 5 min#144