๐Ÿ‘๏ธ Computer Vision ยท Lecture 6 of 27

ResNet in Depth: Architecture, Bottlenecks and Variants

ResNet's residual blocks enabled 152-layer networks and became the default vision backbone. We study basic and bottleneck blocks, the full ResNet-50 layout, training recipe, and descendants such as ResNeXt and ConvNeXt.

In the Deep Learning track we met the idea of residual connections. Here we study ResNet as a concrete vision architecture โ€” the backbone behind countless detection, segmentation and medical-imaging systems, and still one of the first models practitioners try. ResNet won ILSVRC 2015 with a top-5 error of about 3.6% using an ensemble of networks up to 152 layers deep.

Building blocks#

Basic block (ResNet-18/34)#

Two $3 \times 3$ convolutions with BatchNorm and ReLU, plus an identity shortcut:

text
x โ”€โ”€โ–บ conv3ร—3 โ”€ BN โ”€ ReLU โ”€ conv3ร—3 โ”€ BN โ”€โ”€(+)โ”€โ”€ ReLU โ”€โ”€โ–บ
โ”‚                                         โ–ฒ
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ identity โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Bottleneck block (ResNet-50/101/152)#

For deeper networks, a three-layer design keeps computation manageable:

  1. $1 \times 1$ conv reduces channels (e.g. 256 โ†’ 64);
  2. $3 \times 3$ conv operates on the reduced channels;
  3. $1 \times 1$ conv restores channels (64 โ†’ 256).

The expensive $3 \times 3$ operation runs on 4ร— fewer channels. A bottleneck block with 256 input/output channels has about 70,000 weights, compared with about 1.2 million for two plain $3 \times 3$ layers at 256 channels.

When the spatial size halves (stride 2) and channels double, the shortcut uses a $1 \times 1$ convolution with stride 2 (a "projection shortcut").

The ResNet-50 layout#

StageOutput size (224 input)Blocks
Stem: $7 \times 7$ conv, stride 2 + $3 \times 3$ max pool$56 \times 56$โ€”
conv2_x$56 \times 56$, 256 channels3 bottlenecks
conv3_x$28 \times 28$, 5124
conv4_x$14 \times 14$, 10246
conv5_x$7 \times 7$, 20483
Global average pool + FC1000 classesโ€”

About 25.6 million parameters and ~4 GFLOPs per image โ€” fewer parameters than VGG-16 despite being far deeper and more accurate.

python
import torch
from torchvision import models

resnet = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)
print(sum(p.numel() for p in resnet.parameters()) / 1e6, "M parameters")

# Use as a feature extractor: drop the classification head
backbone = torch.nn.Sequential(*list(resnet.children())[:-1])
feats = backbone(torch.randn(2, 3, 224, 224)).flatten(1)
print(feats.shape)                      # (2, 2048) image embeddings

# Multi-scale feature maps for detection/segmentation
from torchvision.models.feature_extraction import create_feature_extractor
fx = create_feature_extractor(resnet, {"layer2": "c3", "layer3": "c4", "layer4": "c5"})
print({k: v.shape for k, v in fx(torch.randn(1, 3, 224, 224)).items()})

Training recipe and "ResNet strikes back"#

The original recipe: SGD with momentum 0.9, weight decay $10^{-4}$, batch size 256, learning rate 0.1 divided by 10 when error plateaus, simple crop-and-flip augmentation, ~90 epochs.

Later work showed the same ResNet-50 can gain several points of ImageNet accuracy from a modern recipe alone โ€” longer training, cosine schedules, label smoothing, mixup/cutmix, RandAugment, stochastic depth, EMA of weights (Bello et al., 2021 "Revisiting ResNets"; Wightman et al., 2021 "ResNet strikes back"). A valuable lesson: compare architectures under equal training recipes.

Variants and descendants#

  • Pre-activation ResNet (2016): BN โ†’ ReLU โ†’ conv ordering inside the branch, leaving a clean identity path; trained 1,000-layer networks.
  • Wide ResNet (2016): fewer, wider layers can match or beat very deep thin ones.
  • ResNeXt (2017): grouped convolutions in the bottleneck โ€” many parallel paths ("cardinality") โ€” a cleaner form of Inception's multi-branch idea; better accuracy at similar cost.
  • SE-Net (2018): squeeze-and-excitation blocks reweight channels using global context:
$$ \mathbf{s} = \sigma\big(\mathbf{W}_2\,\text{ReLU}(\mathbf{W}_1\,\text{GAP}(\mathbf{X}))\big), \qquad \tilde{\mathbf{X}}_c = s_c\,\mathbf{X}_c $$

A form of channel attention; won ILSVRC 2017.

  • ResNet-D and bag of tricks (He et al., 2019): small stem and downsampling changes that add accuracy for free.
  • ConvNeXt (2022): modernised ResNet design choices one by one towards Vision-Transformer-style components (larger $7 \times 7$ depthwise kernels, LayerNorm, GELU, inverted bottlenecks, fewer activations) and showed pure CNNs can match ViTs at similar scale.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

EfficientNet and Principled Model Scaling

How should a network grow when you have more compute โ€” deeper, wider or higher resolution? EfficientNet's compound scaling answers "all three, in balance". We cover MBConv blocks, the B0โ€“B7 family and EfficientNetV2.

Intermediateโฑ 5 min#141
๐Ÿ‘๏ธ Computer Vision

LeNet and AlexNet: The Birth of Deep Vision

Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.

Beginnerโฑ 5 min#138
๐Ÿ‘๏ธ Computer Vision

VGG and GoogLeNet/Inception โ€” Depth and Multi-Scale Design

In 2014 two architectures pushed CNNs deeper in opposite styles: VGG with uniform stacks of 3ร—3 convolutions, GoogLeNet with parallel multi-scale Inception modules and 1ร—1 bottlenecks. We compare their designs and lessons.

Intermediateโฑ 5 min#139