In the Deep Learning track we met the idea of residual connections. Here we study ResNet as a concrete vision architecture โ the backbone behind countless detection, segmentation and medical-imaging systems, and still one of the first models practitioners try. ResNet won ILSVRC 2015 with a top-5 error of about 3.6% using an ensemble of networks up to 152 layers deep.
Building blocks#
Basic block (ResNet-18/34)#
Two $3 \times 3$ convolutions with BatchNorm and ReLU, plus an identity shortcut:
x โโโบ conv3ร3 โ BN โ ReLU โ conv3ร3 โ BN โโ(+)โโ ReLU โโโบ
โ โฒ
โโโโโโโโโโโโโโโโ identity โโโโโโโโโโโโโโโโโBottleneck block (ResNet-50/101/152)#
For deeper networks, a three-layer design keeps computation manageable:
- $1 \times 1$ conv reduces channels (e.g. 256 โ 64);
- $3 \times 3$ conv operates on the reduced channels;
- $1 \times 1$ conv restores channels (64 โ 256).
The expensive $3 \times 3$ operation runs on 4ร fewer channels. A bottleneck block with 256 input/output channels has about 70,000 weights, compared with about 1.2 million for two plain $3 \times 3$ layers at 256 channels.
When the spatial size halves (stride 2) and channels double, the shortcut uses a $1 \times 1$ convolution with stride 2 (a "projection shortcut").
The ResNet-50 layout#
| Stage | Output size (224 input) | Blocks |
|---|---|---|
| Stem: $7 \times 7$ conv, stride 2 + $3 \times 3$ max pool | $56 \times 56$ | โ |
| conv2_x | $56 \times 56$, 256 channels | 3 bottlenecks |
| conv3_x | $28 \times 28$, 512 | 4 |
| conv4_x | $14 \times 14$, 1024 | 6 |
| conv5_x | $7 \times 7$, 2048 | 3 |
| Global average pool + FC | 1000 classes | โ |
About 25.6 million parameters and ~4 GFLOPs per image โ fewer parameters than VGG-16 despite being far deeper and more accurate.
import torch
from torchvision import models
resnet = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)
print(sum(p.numel() for p in resnet.parameters()) / 1e6, "M parameters")
# Use as a feature extractor: drop the classification head
backbone = torch.nn.Sequential(*list(resnet.children())[:-1])
feats = backbone(torch.randn(2, 3, 224, 224)).flatten(1)
print(feats.shape) # (2, 2048) image embeddings
# Multi-scale feature maps for detection/segmentation
from torchvision.models.feature_extraction import create_feature_extractor
fx = create_feature_extractor(resnet, {"layer2": "c3", "layer3": "c4", "layer4": "c5"})
print({k: v.shape for k, v in fx(torch.randn(1, 3, 224, 224)).items()})Training recipe and "ResNet strikes back"#
The original recipe: SGD with momentum 0.9, weight decay $10^{-4}$, batch size 256, learning rate 0.1 divided by 10 when error plateaus, simple crop-and-flip augmentation, ~90 epochs.
Later work showed the same ResNet-50 can gain several points of ImageNet accuracy from a modern recipe alone โ longer training, cosine schedules, label smoothing, mixup/cutmix, RandAugment, stochastic depth, EMA of weights (Bello et al., 2021 "Revisiting ResNets"; Wightman et al., 2021 "ResNet strikes back"). A valuable lesson: compare architectures under equal training recipes.
Variants and descendants#
- Pre-activation ResNet (2016): BN โ ReLU โ conv ordering inside the branch, leaving a clean identity path; trained 1,000-layer networks.
- Wide ResNet (2016): fewer, wider layers can match or beat very deep thin ones.
- ResNeXt (2017): grouped convolutions in the bottleneck โ many parallel paths ("cardinality") โ a cleaner form of Inception's multi-branch idea; better accuracy at similar cost.
- SE-Net (2018): squeeze-and-excitation blocks reweight channels using global context:
A form of channel attention; won ILSVRC 2017.
- ResNet-D and bag of tricks (He et al., 2019): small stem and downsampling changes that add accuracy for free.
- ConvNeXt (2022): modernised ResNet design choices one by one towards Vision-Transformer-style components (larger $7 \times 7$ depthwise kernels, LayerNorm, GELU, inverted bottlenecks, fewer activations) and showed pure CNNs can match ViTs at similar scale.