👁️ Computer Vision · Lecture 17 of 27

Self-Supervised Vision: SimCLR, MoCo, DINO and Masked Autoencoders

Labels are expensive, images are abundant. Self-supervised methods learn visual representations from unlabelled images through contrastive learning, self-distillation or masked reconstruction. We compare the main families and how to use them.

ImageNet's 1.3 million labelled images took years of human effort. Meanwhile, billions of unlabelled images exist. Self-supervised learning (SSL) designs "pretext" tasks whose labels come from the data itself, so a network can learn general-purpose visual features without human annotation. By around 2020–2021, self-supervised features rivalled and then, in many settings, surpassed supervised ImageNet pretraining for transfer learning.

Early pretext tasks#

Predicting image rotations (0°, 90°, 180°, 270°), solving jigsaw puzzles of shuffled patches, colourising grayscale images, predicting relative patch positions. These worked to a degree, but the learned features were tied to the pretext task.

Contrastive learning#

The key idea: two augmented views of the same image should have similar representations (positives), while views of different images should differ (negatives).

SimCLR#

Chen et al. (2020):

  1. Create two random augmentations of each image (crop + resize, colour distortion, blur).
  2. Encode with a backbone $f$ (e.g. ResNet-50) and a small projection head $g$ (an MLP): $\mathbf{z} = g(f(\mathbf{x}))$.
  3. Apply the NT-Xent (InfoNCE) loss: for a positive pair $(i, j)$ in a batch of $2N$ views,
$$ \ell_{i,j} = -\log\frac{\exp(\text{sim}(\mathbf{z}_i, \mathbf{z}_j)/\tau)}{\sum_{k \ne i}\exp(\text{sim}(\mathbf{z}_i, \mathbf{z}_k)/\tau)} $$

with cosine similarity and temperature $\tau$.

  1. After training, discard $g$ and use $f$'s features.

Findings: strong augmentation composition (especially crop + colour distortion) is crucial; the projection head improves the representation below it; large batches (thousands) provide many negatives.

python
import torch
import torch.nn.functional as F

def nt_xent(z1, z2, tau=0.2):
    """z1, z2: (N, d) projections of two views of the same N images."""
    z = F.normalize(torch.cat([z1, z2]), dim=1)             # (2N, d)
    sim = z @ z.T / tau
    sim.fill_diagonal_(float("-inf"))                        # exclude self-similarity
    n = z1.size(0)
    targets = torch.cat([torch.arange(n, 2 * n), torch.arange(0, n)])   # index of each positive
    return F.cross_entropy(sim, targets)

print(nt_xent(torch.randn(8, 128), torch.randn(8, 128)))

MoCo#

He et al. (2020) decoupled the number of negatives from batch size with a queue of past embeddings and a slowly updated momentum encoder (an exponential moving average of the online encoder), keeping queue embeddings consistent.

Non-contrastive methods#

Can we avoid negatives entirely? Naively, the network could collapse — map every image to the same vector. Methods prevent collapse differently:

  • BYOL (2020): an online network predicts the target network's representation of another view; the target is a momentum (EMA) copy, and an extra predictor head creates asymmetry.
  • SimSiam (2021): like BYOL without momentum; a stop-gradient on one branch is essential.
  • Barlow Twins / VICReg: make the cross-correlation matrix of two views' embeddings close to identity — decorrelating dimensions prevents collapse.
  • DINO (Caron et al., 2021): self-distillation with no labels. A student network matches the output distribution of a momentum teacher across different crops (the teacher sees large global crops, the student also sees small local crops). Centring and sharpening of teacher outputs prevent collapse. Remarkably, DINO-trained ViTs' attention maps segment objects without any supervision. DINOv2 (2023) scaled this with curated data and produced strong general-purpose features usable frozen for classification, segmentation and depth estimation.

Masked image modelling#

Inspired by BERT's masked language modelling:

  • MAE (He et al., 2022): mask a large fraction (~75%) of ViT patches, encode only the visible patches (cheap), and reconstruct the missing pixels with a lightweight decoder. The high mask ratio makes the task non-trivial (images are spatially redundant) and training efficient. MAE features fine-tune excellently.
  • BEiT, SimMIM: predict discrete tokens or pixels of masked patches.
  • I-JEPA (2023): predict the representations (not pixels) of masked regions from context, focusing on semantics rather than low-level detail.

Comparing the families#

FamilyExamplesSignalNeeds negativesNotes
ContrastiveSimCLR, MoCoInvariance across viewsYesStrong linear-probe features; batch/queue size matters
Self-distillationBYOL, DINOMatch EMA teacherNoDINO attention maps localise objects
Redundancy reductionBarlow Twins, VICRegDecorrelate dimensionsNoSimple, stable
Masked modellingMAE, BEiT, I-JEPAReconstruct/predict masked contentNoExcellent fine-tuning; efficient with ViTs

Evaluating representations#

  • Linear probing: freeze the backbone, train a linear classifier on top — measures how linearly separable the features are.
  • k-NN evaluation: classify by nearest neighbours in feature space — no training needed.
  • Fine-tuning and transfer to detection, segmentation and low-label regimes.

Why this matters in practice#

Self-supervised pretraining lets you exploit your own unlabelled domain data: thousands of unlabelled satellite tiles, X-rays, microscopy slides or field photos. Pretrain (or continue pretraining a public model) on them, then fine-tune on a small labelled set. In specialised domains this frequently beats ImageNet-pretrained initialisation.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

CLIP: Connecting Images and Language

CLIP learns a shared embedding space for images and text from hundreds of millions of image–caption pairs. We explain its contrastive training, zero-shot classification with prompts, retrieval, limitations and its role in generative models.

Intermediate⏱ 5 min#152
👁️ Computer Vision

Vision Transformers (ViT): Images as Sequences of Patches

Transformers conquered language, then vision. We dissect ViT's patch embeddings, class token and positional encodings, compare inductive biases with CNNs, and survey DeiT, Swin and hierarchical designs.

Advanced⏱ 5 min#150
👁️ Computer Vision

Instance Segmentation: Mask R-CNN and Beyond

Instance segmentation separates each individual object with its own mask. We study Mask R-CNN's mask branch and RoIAlign, compare instance, semantic and panoptic segmentation, and survey query-based models like Mask2Former.

Advanced⏱ 4 min#149