In 2016 Joseph Redmon and colleagues published YOLO — You Only Look Once. Instead of proposing regions and classifying each one, YOLO looks at the whole image once with a single network and directly predicts all boxes and classes. It ran at tens of frames per second on a GPU, bringing detection to real-time video. Successive YOLO versions have made it the most widely used family of practical detectors.
The original formulation#
YOLOv1 divides the image into an $S \times S$ grid (e.g. $7 \times 7$). The cell containing an object's centre is responsible for detecting it. Each cell predicts:
- $B$ boxes, each with $(x, y, w, h)$ and a confidence $= P(\text{object}) \times \text{IoU}$;
- $C$ class probabilities.
The output is a tensor of shape $S \times S \times (5B + C)$ — for PASCAL VOC with $S = 7$, $B = 2$, $C = 20$: $7 \times 7 \times 30$. Training uses a sum-of-squares loss with weights that emphasise box coordinates and down-weight cells without objects; width and height are predicted as square roots so errors on small boxes count more.
Strengths: very fast; sees the whole image, so it makes fewer background false positives than sliding-window methods. Weaknesses: each cell predicts few boxes, so crowds of small objects (flocks of birds) are hard; coarse localisation.
Evolution of YOLO#
| Version | Key ideas |
|---|---|
| YOLOv2 / YOLO9000 (2017) | Anchor boxes from k-means on training boxes, batch norm, higher resolution, multi-scale training |
| YOLOv3 (2018) | Darknet-53 backbone, predictions at three scales (FPN-like), logistic class predictions |
| YOLOv4 (2020) | "Bag of freebies" (mosaic augmentation, CIoU loss) and "bag of specials" (CSP backbone, PANet neck) |
| YOLOv5 onwards (community/industry) | PyTorch implementations, auto-anchors, strong training recipes, easy tooling |
| YOLOX, YOLOv8 and later | Anchor-free heads, decoupled classification/regression heads, improved label assignment |
(From version 4 onward, YOLO versions were developed by different groups, so the numbering does not indicate a single lineage.)
Anatomy of a modern one-stage detector#
- Backbone — extracts features (CSPDarknet, efficient CNNs).
- Neck — fuses multi-scale features (FPN + PAN: top-down and bottom-up paths).
- Head — at each location of each scale, predicts class scores and box geometry.
Anchor-free heads predict the distance from a location to the four box sides (as in FCOS) or the box centre and size directly, avoiding hand-tuned anchor shapes.
Label assignment — deciding which predictions are responsible for which ground-truth boxes — matters as much as architecture. Modern methods (SimOTA, task-aligned assignment) dynamically assign positives based on both classification and localisation quality.
Box losses based on IoU (GIoU, DIoU, CIoU) directly optimise overlap rather than coordinate differences, which aligns training with evaluation. GIoU, for instance, adds a penalty based on the smallest enclosing box $C$:
so non-overlapping boxes still receive a useful gradient.
Training a YOLO model in practice#
Tools like the Ultralytics package make training straightforward. Labels use one text file per image with rows class x_center y_center width height, normalised to $[0, 1]$.
# pip install ultralytics
from ultralytics import YOLO
model = YOLO("yolov8n.pt") # small pretrained model (COCO)
model.train(data="shelters.yaml", # paths to train/val images and class names
epochs=100, imgsz=640, batch=16, patience=20)
metrics = model.val() # reports mAP50 and mAP50-95
results = model.predict("aerial_tile.jpg", conf=0.25, iou=0.5)
results[0].show()
model.export(format="onnx") # for deployment# shelters.yaml
path: datasets/shelters
train: images/train
val: images/val
names: {0: tent, 1: building, 2: vehicle}Speed–accuracy trade-offs#
YOLO families come in sizes (nano, small, medium, large, extra-large). Pick the smallest model meeting the accuracy requirement at your target frame rate on your hardware. Measure end-to-end latency including preprocessing and NMS, not just model FLOPs.
Beyond YOLO#
DETR (2020) reframed detection as set prediction with a transformer and bipartite matching, removing anchors and NMS entirely; later variants (Deformable DETR, DINO, RT-DETR) made transformer detectors fast and accurate. Open-vocabulary detectors such as OWL-ViT and Grounding DINO detect objects described by arbitrary text prompts — "detect water tanks" — without training on those classes.