ImageNet's 1.3 million labelled images took years of human effort. Meanwhile, billions of unlabelled images exist. Self-supervised learning (SSL) designs "pretext" tasks whose labels come from the data itself, so a network can learn general-purpose visual features without human annotation. By around 2020–2021, self-supervised features rivalled and then, in many settings, surpassed supervised ImageNet pretraining for transfer learning.
Early pretext tasks#
Predicting image rotations (0°, 90°, 180°, 270°), solving jigsaw puzzles of shuffled patches, colourising grayscale images, predicting relative patch positions. These worked to a degree, but the learned features were tied to the pretext task.
Contrastive learning#
The key idea: two augmented views of the same image should have similar representations (positives), while views of different images should differ (negatives).
SimCLR#
Chen et al. (2020):
- Create two random augmentations of each image (crop + resize, colour distortion, blur).
- Encode with a backbone $f$ (e.g. ResNet-50) and a small projection head $g$ (an MLP): $\mathbf{z} = g(f(\mathbf{x}))$.
- Apply the NT-Xent (InfoNCE) loss: for a positive pair $(i, j)$ in a batch of $2N$ views,
with cosine similarity and temperature $\tau$.
- After training, discard $g$ and use $f$'s features.
Findings: strong augmentation composition (especially crop + colour distortion) is crucial; the projection head improves the representation below it; large batches (thousands) provide many negatives.
import torch
import torch.nn.functional as F
def nt_xent(z1, z2, tau=0.2):
"""z1, z2: (N, d) projections of two views of the same N images."""
z = F.normalize(torch.cat([z1, z2]), dim=1) # (2N, d)
sim = z @ z.T / tau
sim.fill_diagonal_(float("-inf")) # exclude self-similarity
n = z1.size(0)
targets = torch.cat([torch.arange(n, 2 * n), torch.arange(0, n)]) # index of each positive
return F.cross_entropy(sim, targets)
print(nt_xent(torch.randn(8, 128), torch.randn(8, 128)))MoCo#
He et al. (2020) decoupled the number of negatives from batch size with a queue of past embeddings and a slowly updated momentum encoder (an exponential moving average of the online encoder), keeping queue embeddings consistent.
Non-contrastive methods#
Can we avoid negatives entirely? Naively, the network could collapse — map every image to the same vector. Methods prevent collapse differently:
- BYOL (2020): an online network predicts the target network's representation of another view; the target is a momentum (EMA) copy, and an extra predictor head creates asymmetry.
- SimSiam (2021): like BYOL without momentum; a stop-gradient on one branch is essential.
- Barlow Twins / VICReg: make the cross-correlation matrix of two views' embeddings close to identity — decorrelating dimensions prevents collapse.
- DINO (Caron et al., 2021): self-distillation with no labels. A student network matches the output distribution of a momentum teacher across different crops (the teacher sees large global crops, the student also sees small local crops). Centring and sharpening of teacher outputs prevent collapse. Remarkably, DINO-trained ViTs' attention maps segment objects without any supervision. DINOv2 (2023) scaled this with curated data and produced strong general-purpose features usable frozen for classification, segmentation and depth estimation.
Masked image modelling#
Inspired by BERT's masked language modelling:
- MAE (He et al., 2022): mask a large fraction (~75%) of ViT patches, encode only the visible patches (cheap), and reconstruct the missing pixels with a lightweight decoder. The high mask ratio makes the task non-trivial (images are spatially redundant) and training efficient. MAE features fine-tune excellently.
- BEiT, SimMIM: predict discrete tokens or pixels of masked patches.
- I-JEPA (2023): predict the representations (not pixels) of masked regions from context, focusing on semantics rather than low-level detail.
Comparing the families#
| Family | Examples | Signal | Needs negatives | Notes |
|---|---|---|---|---|
| Contrastive | SimCLR, MoCo | Invariance across views | Yes | Strong linear-probe features; batch/queue size matters |
| Self-distillation | BYOL, DINO | Match EMA teacher | No | DINO attention maps localise objects |
| Redundancy reduction | Barlow Twins, VICReg | Decorrelate dimensions | No | Simple, stable |
| Masked modelling | MAE, BEiT, I-JEPA | Reconstruct/predict masked content | No | Excellent fine-tuning; efficient with ViTs |
Evaluating representations#
- Linear probing: freeze the backbone, train a linear classifier on top — measures how linearly separable the features are.
- k-NN evaluation: classify by nearest neighbours in feature space — no training needed.
- Fine-tuning and transfer to detection, segmentation and low-label regimes.
Why this matters in practice#
Self-supervised pretraining lets you exploit your own unlabelled domain data: thousands of unlabelled satellite tiles, X-rays, microscopy slides or field photos. Pretrain (or continue pretraining a public model) on them, then fine-tune on a small labelled set. In specialised domains this frequently beats ImageNet-pretrained initialisation.