👁️ Computer Vision · Lecture 5 of 27

VGG and GoogLeNet/Inception — Depth and Multi-Scale Design

In 2014 two architectures pushed CNNs deeper in opposite styles: VGG with uniform stacks of 3×3 convolutions, GoogLeNet with parallel multi-scale Inception modules and 1×1 bottlenecks. We compare their designs and lessons.

The 2014 ImageNet challenge produced two architectures that shaped the next decade. VGGNet from Oxford's Visual Geometry Group showed that simply going deeper with a clean, uniform design works. GoogLeNet (Inception v1) from Google showed that careful design could achieve higher accuracy with far fewer parameters. Their ideas — small kernels, 1×1 bottlenecks, global pooling, multi-branch blocks — are everywhere today.

VGG: simplicity and depth#

Simonyan and Zisserman (2014) used a single design rule: only $3 \times 3$ convolutions (stride 1, padding 1) and $2 \times 2$ max pooling, with the number of channels doubling after each pool (64 → 128 → 256 → 512 → 512). VGG-16 has 13 convolutional layers and 3 fully connected layers.

Why $3 \times 3$? As we computed in the receptive-field lecture:

  • Two stacked $3 \times 3$ layers see a $5 \times 5$ region; three see $7 \times 7$.
  • Three $3 \times 3$ layers use $3 \times 9C^2 = 27C^2$ weights versus $49C^2$ for one $7 \times 7$ layer.
  • Each extra layer adds a non-linearity, making the function more expressive.

VGG-16 reached about 7.3% top-5 error. Its weaknesses: roughly 138 million parameters (most in the first fully connected layer: $7 \times 7 \times 512 \times 4096 \approx 103$ million) and heavy computation (~15 GFLOPs per image). Yet its simplicity made it a favourite backbone for transfer learning, style transfer (VGG features define "perceptual loss") and early detection and segmentation systems.

GoogLeNet and the Inception module#

Szegedy et al. (2014) asked: which filter size should a layer use — $1\times1$, $3\times3$ or $5\times5$? Why not all of them in parallel? The Inception module runs several branches side by side and concatenates their outputs along the channel dimension:

text
            ┌── 1×1 conv ─────────────────────┐
            ├── 1×1 conv → 3×3 conv ──────────┤
input ──────┼── 1×1 conv → 5×5 conv ──────────┼── concatenate channels
            └── 3×3 max pool → 1×1 conv ──────┘

This captures features at multiple scales simultaneously.

The 1×1 bottleneck#

Naively, $5 \times 5$ convolutions on many channels are expensive. The key trick: first reduce channels with a cheap $1 \times 1$ convolution. For 256 input channels and 64 output channels:

  • Direct $5 \times 5$: $256 \times 64 \times 25 \approx 410{,}000$ weights.
  • $1 \times 1$ to 32 channels, then $5 \times 5$ to 64: $256 \times 32 + 32 \times 64 \times 25 \approx 59{,}000$ weights — about 7× fewer.

A $1 \times 1$ convolution is a small fully connected layer applied at every pixel across channels (the "network in network" idea of Lin et al., 2013).

Other GoogLeNet features#

  • 22 layers deep yet only about 7 million parameters — roughly 20× fewer than VGG-16.
  • Global average pooling instead of large fully connected layers.
  • Auxiliary classifiers attached to intermediate layers during training to inject extra gradient into the middle of the network (later found to act mainly as regularisers).
  • Top-5 error of about 6.7%, winning ILSVRC 2014.
python
import torch
import torch.nn as nn

class InceptionModule(nn.Module):
    def __init__(self, cin, c1, c3r, c3, c5r, c5, cp):
        super().__init__()
        conv = lambda i, o, k: nn.Sequential(nn.Conv2d(i, o, k, padding=k // 2), nn.ReLU(inplace=True))
        self.b1 = conv(cin, c1, 1)
        self.b2 = nn.Sequential(conv(cin, c3r, 1), conv(c3r, c3, 3))
        self.b3 = nn.Sequential(conv(cin, c5r, 1), conv(c5r, c5, 5))
        self.b4 = nn.Sequential(nn.MaxPool2d(3, 1, 1), conv(cin, cp, 1))
    def forward(self, x):
        return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)

m = InceptionModule(192, 64, 96, 128, 16, 32, 32)       # "inception 3a" configuration
print(m(torch.randn(1, 192, 28, 28)).shape)              # (1, 256, 28, 28)

Later Inception versions#

  • Inception v2/v3 (2015): factorised $5 \times 5$ into two $3 \times 3$, and $n \times n$ into $1 \times n$ followed by $n \times 1$; added batch normalisation and label smoothing (which was introduced in this line of work).
  • Inception v4 / Inception-ResNet (2016): combined Inception modules with residual connections.
  • Xception (2017): pushed the idea to its extreme with depthwise separable convolutions — each channel's spatial pattern processed separately, then mixed with $1 \times 1$ convs.

Comparing the two philosophies#

VGG-16GoogLeNet
Depth1622
Parameters~138M~7M
DesignUniform, simpleMulti-branch, engineered
StrengthEasy to understand and adaptEfficient, accurate
Lasting ideasSmall kernels, doubling channels1×1 bottlenecks, multi-scale branches, global average pooling
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

LeNet and AlexNet: The Birth of Deep Vision

Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.

Beginner⏱ 5 min#138
👁️ Computer Vision

ResNet in Depth: Architecture, Bottlenecks and Variants

ResNet's residual blocks enabled 152-layer networks and became the default vision backbone. We study basic and bottleneck blocks, the full ResNet-50 layout, training recipe, and descendants such as ResNeXt and ConvNeXt.

Intermediate⏱ 5 min#140
👁️ Computer Vision

Classical Features: Harris Corners, SIFT, HOG and Bag of Visual Words

Before CNNs, vision relied on carefully engineered features. We study corner detection, SIFT keypoints and descriptors, HOG for pedestrian detection and the bag-of-visual-words model — ideas still used in geometry and robotics.

Intermediate⏱ 5 min#137