👁️ Computer Vision · Lecture 15 of 27

Instance Segmentation: Mask R-CNN and Beyond

Instance segmentation separates each individual object with its own mask. We study Mask R-CNN's mask branch and RoIAlign, compare instance, semantic and panoptic segmentation, and survey query-based models like Mask2Former.

Semantic segmentation tells us which pixels are "person"; it does not tell us how many people there are or where one ends and another begins. Instance segmentation combines detection and segmentation: for every object instance it outputs a class, a confidence and a pixel mask. It is used to count and measure cells in microscopy, delineate individual buildings for damage assessment, separate overlapping fruit for robotic harvesting, and track individual animals.

Three flavours of segmentation#

TaskOutput"Stuff" (sky, road)"Things" (people, cars)
SemanticClass per pixelYesClass only, instances merged
InstanceMask per objectIgnoredSeparate instances
PanopticClass + instance ID per pixelYesSeparate instances

Panoptic segmentation (Kirillov et al., 2019) unifies both: every pixel gets a class, and pixels of countable things also get an instance ID.

Mask R-CNN#

He, Gkioxari, Dollár and Girshick (2017) extended Faster R-CNN with a third, parallel branch:

  1. Backbone + FPN extract features.
  2. The RPN proposes regions.
  3. For each region, three heads run in parallel:
    • classification,
    • box regression,
    • a mask head: a small fully convolutional network that predicts a $28 \times 28$ binary mask for each class (the mask for the predicted class is used).

The loss is the sum $\mathcal{L} = \mathcal{L}_{\text{cls}} + \mathcal{L}_{\text{box}} + \mathcal{L}_{\text{mask}}$, where the mask loss is per-pixel binary cross-entropy on the ground-truth class's mask only. Decoupling mask prediction from classification (no competition between classes at each pixel) improved results.

RoIAlign: the crucial detail#

Faster R-CNN's RoI pooling quantises region coordinates to the feature-map grid (rounding), then quantises again into pooling bins. For classification this misalignment is harmless, but for pixel-accurate masks, errors of even a feature-map cell (which may correspond to 16–32 image pixels) are significant.

RoIAlign removes all rounding: it samples feature values at exact, fractional locations using bilinear interpolation, then pools. This simple change improved mask accuracy substantially and also helped box detection.

python
import torch
from torchvision.ops import roi_align

features = torch.randn(1, 256, 50, 50)                 # feature map at stride 16
boxes = torch.tensor([[0, 103.7, 58.2, 331.9, 287.4]])  # (batch_idx, x1, y1, x2, y2) in image pixels
pooled = roi_align(features, boxes, output_size=(7, 7), spatial_scale=1 / 16, sampling_ratio=2, aligned=True)
print(pooled.shape)                                     # (1, 256, 7, 7)

Using and fine-tuning Mask R-CNN#

python
from torchvision.models.detection import maskrcnn_resnet50_fpn_v2, MaskRCNN_ResNet50_FPN_V2_Weights
from torchvision.models.detection.faster_rcnn import FastRCNNPredictor
from torchvision.models.detection.mask_rcnn import MaskRCNNPredictor

model = maskrcnn_resnet50_fpn_v2(weights=MaskRCNN_ResNet50_FPN_V2_Weights.DEFAULT)
num_classes = 2                                          # background + "building"
model.roi_heads.box_predictor = FastRCNNPredictor(model.roi_heads.box_predictor.cls_score.in_features, num_classes)
model.roi_heads.mask_predictor = MaskRCNNPredictor(256, 256, num_classes)
# Targets per image: {"boxes": (N,4), "labels": (N,), "masks": (N,H,W) uint8}

At inference, each detection comes with a soft mask; threshold at 0.5 to obtain a binary mask in image coordinates.

Evaluation#

Instance segmentation uses mask AP: the same AP computation as detection, but IoU is computed between masks rather than boxes. Panoptic segmentation uses Panoptic Quality: $PQ = \text{SQ} \times \text{RQ}$ (segmentation quality × recognition quality).

Beyond Mask R-CNN#

  • One-stage / bottom-up methods: YOLACT (prototype masks combined with per-instance coefficients, real-time), SOLO (predict masks by location).
  • Query-based transformers: DETR-style models use learned object queries that each predict a class and mask, trained with bipartite matching — no anchors or NMS. MaskFormer and Mask2Former unified semantic, instance and panoptic segmentation in one architecture with state-of-the-art results.
  • Promptable foundation models: Segment Anything (SAM) produces masks for any object given a point, box or (in later versions) text/concept prompt — transforming annotation workflows (see the lecture on vision foundation models).
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

Semantic Segmentation: FCN, U-Net and DeepLab

Segmentation labels every pixel. We cover fully convolutional networks, the encoder–decoder U-Net with skip connections, DeepLab's atrous convolutions, loss functions like Dice, and evaluation with IoU.

Intermediate⏱ 5 min#148
👁️ Computer Vision

Vision Transformers (ViT): Images as Sequences of Patches

Transformers conquered language, then vision. We dissect ViT's patch embeddings, class token and positional encodings, compare inductive biases with CNNs, and survey DeiT, Swin and hierarchical designs.

Advanced⏱ 5 min#150
👁️ Computer Vision

Object Detection III: SSD, RetinaNet and the Focal Loss

One-stage detectors face an extreme imbalance between background and objects. We study SSD's multi-scale default boxes, then derive RetinaNet's focal loss, which let one-stage detectors match two-stage accuracy.

Advanced⏱ 5 min#147