The 2014 ImageNet challenge produced two architectures that shaped the next decade. VGGNet from Oxford's Visual Geometry Group showed that simply going deeper with a clean, uniform design works. GoogLeNet (Inception v1) from Google showed that careful design could achieve higher accuracy with far fewer parameters. Their ideas — small kernels, 1×1 bottlenecks, global pooling, multi-branch blocks — are everywhere today.
VGG: simplicity and depth#
Simonyan and Zisserman (2014) used a single design rule: only $3 \times 3$ convolutions (stride 1, padding 1) and $2 \times 2$ max pooling, with the number of channels doubling after each pool (64 → 128 → 256 → 512 → 512). VGG-16 has 13 convolutional layers and 3 fully connected layers.
Why $3 \times 3$? As we computed in the receptive-field lecture:
- Two stacked $3 \times 3$ layers see a $5 \times 5$ region; three see $7 \times 7$.
- Three $3 \times 3$ layers use $3 \times 9C^2 = 27C^2$ weights versus $49C^2$ for one $7 \times 7$ layer.
- Each extra layer adds a non-linearity, making the function more expressive.
VGG-16 reached about 7.3% top-5 error. Its weaknesses: roughly 138 million parameters (most in the first fully connected layer: $7 \times 7 \times 512 \times 4096 \approx 103$ million) and heavy computation (~15 GFLOPs per image). Yet its simplicity made it a favourite backbone for transfer learning, style transfer (VGG features define "perceptual loss") and early detection and segmentation systems.
GoogLeNet and the Inception module#
Szegedy et al. (2014) asked: which filter size should a layer use — $1\times1$, $3\times3$ or $5\times5$? Why not all of them in parallel? The Inception module runs several branches side by side and concatenates their outputs along the channel dimension:
┌── 1×1 conv ─────────────────────┐
├── 1×1 conv → 3×3 conv ──────────┤
input ──────┼── 1×1 conv → 5×5 conv ──────────┼── concatenate channels
└── 3×3 max pool → 1×1 conv ──────┘This captures features at multiple scales simultaneously.
The 1×1 bottleneck#
Naively, $5 \times 5$ convolutions on many channels are expensive. The key trick: first reduce channels with a cheap $1 \times 1$ convolution. For 256 input channels and 64 output channels:
- Direct $5 \times 5$: $256 \times 64 \times 25 \approx 410{,}000$ weights.
- $1 \times 1$ to 32 channels, then $5 \times 5$ to 64: $256 \times 32 + 32 \times 64 \times 25 \approx 59{,}000$ weights — about 7× fewer.
A $1 \times 1$ convolution is a small fully connected layer applied at every pixel across channels (the "network in network" idea of Lin et al., 2013).
Other GoogLeNet features#
- 22 layers deep yet only about 7 million parameters — roughly 20× fewer than VGG-16.
- Global average pooling instead of large fully connected layers.
- Auxiliary classifiers attached to intermediate layers during training to inject extra gradient into the middle of the network (later found to act mainly as regularisers).
- Top-5 error of about 6.7%, winning ILSVRC 2014.
import torch
import torch.nn as nn
class InceptionModule(nn.Module):
def __init__(self, cin, c1, c3r, c3, c5r, c5, cp):
super().__init__()
conv = lambda i, o, k: nn.Sequential(nn.Conv2d(i, o, k, padding=k // 2), nn.ReLU(inplace=True))
self.b1 = conv(cin, c1, 1)
self.b2 = nn.Sequential(conv(cin, c3r, 1), conv(c3r, c3, 3))
self.b3 = nn.Sequential(conv(cin, c5r, 1), conv(c5r, c5, 5))
self.b4 = nn.Sequential(nn.MaxPool2d(3, 1, 1), conv(cin, cp, 1))
def forward(self, x):
return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)
m = InceptionModule(192, 64, 96, 128, 16, 32, 32) # "inception 3a" configuration
print(m(torch.randn(1, 192, 28, 28)).shape) # (1, 256, 28, 28)Later Inception versions#
- Inception v2/v3 (2015): factorised $5 \times 5$ into two $3 \times 3$, and $n \times n$ into $1 \times n$ followed by $n \times 1$; added batch normalisation and label smoothing (which was introduced in this line of work).
- Inception v4 / Inception-ResNet (2016): combined Inception modules with residual connections.
- Xception (2017): pushed the idea to its extreme with depthwise separable convolutions — each channel's spatial pattern processed separately, then mixed with $1 \times 1$ convs.
Comparing the two philosophies#
| VGG-16 | GoogLeNet | |
|---|---|---|
| Depth | 16 | 22 |
| Parameters | ~138M | ~7M |
| Design | Uniform, simple | Multi-branch, engineered |
| Strength | Easy to understand and adapt | Efficient, accurate |
| Lasting ideas | Small kernels, doubling channels | 1×1 bottlenecks, multi-scale branches, global average pooling |