Semantic segmentation tells us which pixels are "person"; it does not tell us how many people there are or where one ends and another begins. Instance segmentation combines detection and segmentation: for every object instance it outputs a class, a confidence and a pixel mask. It is used to count and measure cells in microscopy, delineate individual buildings for damage assessment, separate overlapping fruit for robotic harvesting, and track individual animals.
Three flavours of segmentation#
| Task | Output | "Stuff" (sky, road) | "Things" (people, cars) |
|---|---|---|---|
| Semantic | Class per pixel | Yes | Class only, instances merged |
| Instance | Mask per object | Ignored | Separate instances |
| Panoptic | Class + instance ID per pixel | Yes | Separate instances |
Panoptic segmentation (Kirillov et al., 2019) unifies both: every pixel gets a class, and pixels of countable things also get an instance ID.
Mask R-CNN#
He, Gkioxari, Dollár and Girshick (2017) extended Faster R-CNN with a third, parallel branch:
- Backbone + FPN extract features.
- The RPN proposes regions.
- For each region, three heads run in parallel:
- classification,
- box regression,
- a mask head: a small fully convolutional network that predicts a $28 \times 28$ binary mask for each class (the mask for the predicted class is used).
The loss is the sum $\mathcal{L} = \mathcal{L}_{\text{cls}} + \mathcal{L}_{\text{box}} + \mathcal{L}_{\text{mask}}$, where the mask loss is per-pixel binary cross-entropy on the ground-truth class's mask only. Decoupling mask prediction from classification (no competition between classes at each pixel) improved results.
RoIAlign: the crucial detail#
Faster R-CNN's RoI pooling quantises region coordinates to the feature-map grid (rounding), then quantises again into pooling bins. For classification this misalignment is harmless, but for pixel-accurate masks, errors of even a feature-map cell (which may correspond to 16–32 image pixels) are significant.
RoIAlign removes all rounding: it samples feature values at exact, fractional locations using bilinear interpolation, then pools. This simple change improved mask accuracy substantially and also helped box detection.
import torch
from torchvision.ops import roi_align
features = torch.randn(1, 256, 50, 50) # feature map at stride 16
boxes = torch.tensor([[0, 103.7, 58.2, 331.9, 287.4]]) # (batch_idx, x1, y1, x2, y2) in image pixels
pooled = roi_align(features, boxes, output_size=(7, 7), spatial_scale=1 / 16, sampling_ratio=2, aligned=True)
print(pooled.shape) # (1, 256, 7, 7)Using and fine-tuning Mask R-CNN#
from torchvision.models.detection import maskrcnn_resnet50_fpn_v2, MaskRCNN_ResNet50_FPN_V2_Weights
from torchvision.models.detection.faster_rcnn import FastRCNNPredictor
from torchvision.models.detection.mask_rcnn import MaskRCNNPredictor
model = maskrcnn_resnet50_fpn_v2(weights=MaskRCNN_ResNet50_FPN_V2_Weights.DEFAULT)
num_classes = 2 # background + "building"
model.roi_heads.box_predictor = FastRCNNPredictor(model.roi_heads.box_predictor.cls_score.in_features, num_classes)
model.roi_heads.mask_predictor = MaskRCNNPredictor(256, 256, num_classes)
# Targets per image: {"boxes": (N,4), "labels": (N,), "masks": (N,H,W) uint8}At inference, each detection comes with a soft mask; threshold at 0.5 to obtain a binary mask in image coordinates.
Evaluation#
Instance segmentation uses mask AP: the same AP computation as detection, but IoU is computed between masks rather than boxes. Panoptic segmentation uses Panoptic Quality: $PQ = \text{SQ} \times \text{RQ}$ (segmentation quality × recognition quality).
Beyond Mask R-CNN#
- One-stage / bottom-up methods: YOLACT (prototype masks combined with per-instance coefficients, real-time), SOLO (predict masks by location).
- Query-based transformers: DETR-style models use learned object queries that each predict a class and mask, trained with bipartite matching — no anchors or NMS. MaskFormer and Mask2Former unified semantic, instance and panoptic segmentation in one architecture with state-of-the-art results.
- Promptable foundation models: Segment Anything (SAM) produces masks for any object given a point, box or (in later versions) text/concept prompt — transforming annotation workflows (see the lecture on vision foundation models).