Every modern vision architecture descends from two landmark networks. LeNet-5 (1998) proved that convolutional networks trained with backpropagation could solve a real problem โ reading handwritten digits on cheques. AlexNet (2012) proved they could scale to a million natural images and beat every alternative by a wide margin. Studying them shows which ideas were present from the beginning and which innovations unlocked scale.
LeNet-5#
Yann LeCun and colleagues developed a series of convolutional networks from the late 1980s; LeNet-5 was described in the 1998 paper "Gradient-Based Learning Applied to Document Recognition". Systems based on this work were deployed to read a significant share of cheques in the United States.
Architecture for $32 \times 32$ grayscale input:
| Layer | Operation | Output |
|---|---|---|
| C1 | 6 filters $5\times5$ | $6 \times 28 \times 28$ |
| S2 | Subsampling (average pool) $2\times2$ | $6 \times 14 \times 14$ |
| C3 | 16 filters $5\times5$ | $16 \times 10 \times 10$ |
| S4 | Subsampling $2\times2$ | $16 \times 5 \times 5$ |
| C5 | 120 filters $5\times5$ | 120 |
| F6 | Fully connected | 84 |
| Output | 10 classes | 10 |
About 60,000 parameters, tanh-like activations. All the core ideas are already here: local receptive fields, weight sharing, subsampling, hierarchical features and end-to-end training by gradient descent.
import torch.nn as nn
class LeNet5(nn.Module):
def __init__(self, n_classes=10):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(1, 6, 5), nn.Tanh(), nn.AvgPool2d(2),
nn.Conv2d(6, 16, 5), nn.Tanh(), nn.AvgPool2d(2),
nn.Conv2d(16, 120, 5), nn.Tanh())
self.classifier = nn.Sequential(nn.Linear(120, 84), nn.Tanh(), nn.Linear(84, n_classes))
def forward(self, x): # x: (N, 1, 32, 32)
return self.classifier(self.features(x).flatten(1))So why did CNNs not take over computer vision in the 1990s? Datasets were small, computers were slow, and on the datasets of the time, SVMs with hand-crafted features were competitive and easier to use.
ImageNet#
Fei-Fei Li and colleagues built ImageNet, eventually containing over 14 million labelled images, labelled via crowdsourcing. The ImageNet Large Scale Visual Recognition Challenge (ILSVRC), run from 2010 to 2017, used a 1,000-class subset with about 1.2 million training images. It provided exactly what deep networks needed: scale.
AlexNet#
In 2012 Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton entered a deep CNN in ILSVRC and achieved a top-5 error of about 15.3%, versus about 26.2% for the next-best entry โ an unprecedented gap.
Architecture (for $224 \times 224 \times 3$ input): five convolutional layers (the first with $11 \times 11$ filters and stride 4), three max-pooling layers and three fully connected layers (4096, 4096, 1000), with about 60 million parameters โ a thousand times more than LeNet.
The innovations that mattered#
- ReLU activations โ trained several times faster than tanh units, avoiding saturation.
- GPU training โ trained on two NVIDIA GTX 580 GPUs (3 GB each) for about a week; the model was split across the two GPUs.
- Dropout (p = 0.5) in the fully connected layers โ crucial against overfitting the 60 million parameters.
- Data augmentation โ random $224 \times 224$ crops from $256 \times 256$ images, horizontal flips, and PCA-based colour perturbation, enlarging the effective dataset.
- Overlapping max pooling ($3 \times 3$ windows, stride 2).
- Local response normalisation โ later abandoned in favour of batch normalisation.
- SGD with momentum 0.9 and weight decay, with the learning rate reduced manually when validation error plateaued.
from torchvision import models
alexnet = models.alexnet(weights=models.AlexNet_Weights.IMAGENET1K_V1)
print(alexnet)
print("parameters:", sum(p.numel() for p in alexnet.parameters())) # ~61 millionImpact#
After 2012, almost every competitive ILSVRC entry used deep CNNs, and the error rate fell year after year (ZFNet 2013, VGG and GoogLeNet 2014, ResNet 2015). Industry invested heavily in GPUs and deep learning, and the approach spread to speech, language and beyond. The recipe โ big data + big compute + deep networks trained end-to-end โ became the paradigm of modern AI.