๐Ÿ‘๏ธ Computer Vision ยท Lecture 4 of 27

LeNet and AlexNet: The Birth of Deep Vision

Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.

Every modern vision architecture descends from two landmark networks. LeNet-5 (1998) proved that convolutional networks trained with backpropagation could solve a real problem โ€” reading handwritten digits on cheques. AlexNet (2012) proved they could scale to a million natural images and beat every alternative by a wide margin. Studying them shows which ideas were present from the beginning and which innovations unlocked scale.

LeNet-5#

Yann LeCun and colleagues developed a series of convolutional networks from the late 1980s; LeNet-5 was described in the 1998 paper "Gradient-Based Learning Applied to Document Recognition". Systems based on this work were deployed to read a significant share of cheques in the United States.

Architecture for $32 \times 32$ grayscale input:

LayerOperationOutput
C16 filters $5\times5$$6 \times 28 \times 28$
S2Subsampling (average pool) $2\times2$$6 \times 14 \times 14$
C316 filters $5\times5$$16 \times 10 \times 10$
S4Subsampling $2\times2$$16 \times 5 \times 5$
C5120 filters $5\times5$120
F6Fully connected84
Output10 classes10

About 60,000 parameters, tanh-like activations. All the core ideas are already here: local receptive fields, weight sharing, subsampling, hierarchical features and end-to-end training by gradient descent.

python
import torch.nn as nn

class LeNet5(nn.Module):
    def __init__(self, n_classes=10):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(1, 6, 5), nn.Tanh(), nn.AvgPool2d(2),
            nn.Conv2d(6, 16, 5), nn.Tanh(), nn.AvgPool2d(2),
            nn.Conv2d(16, 120, 5), nn.Tanh())
        self.classifier = nn.Sequential(nn.Linear(120, 84), nn.Tanh(), nn.Linear(84, n_classes))
    def forward(self, x):                       # x: (N, 1, 32, 32)
        return self.classifier(self.features(x).flatten(1))

So why did CNNs not take over computer vision in the 1990s? Datasets were small, computers were slow, and on the datasets of the time, SVMs with hand-crafted features were competitive and easier to use.

ImageNet#

Fei-Fei Li and colleagues built ImageNet, eventually containing over 14 million labelled images, labelled via crowdsourcing. The ImageNet Large Scale Visual Recognition Challenge (ILSVRC), run from 2010 to 2017, used a 1,000-class subset with about 1.2 million training images. It provided exactly what deep networks needed: scale.

AlexNet#

In 2012 Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton entered a deep CNN in ILSVRC and achieved a top-5 error of about 15.3%, versus about 26.2% for the next-best entry โ€” an unprecedented gap.

Architecture (for $224 \times 224 \times 3$ input): five convolutional layers (the first with $11 \times 11$ filters and stride 4), three max-pooling layers and three fully connected layers (4096, 4096, 1000), with about 60 million parameters โ€” a thousand times more than LeNet.

The innovations that mattered#

  1. ReLU activations โ€” trained several times faster than tanh units, avoiding saturation.
  2. GPU training โ€” trained on two NVIDIA GTX 580 GPUs (3 GB each) for about a week; the model was split across the two GPUs.
  3. Dropout (p = 0.5) in the fully connected layers โ€” crucial against overfitting the 60 million parameters.
  4. Data augmentation โ€” random $224 \times 224$ crops from $256 \times 256$ images, horizontal flips, and PCA-based colour perturbation, enlarging the effective dataset.
  5. Overlapping max pooling ($3 \times 3$ windows, stride 2).
  6. Local response normalisation โ€” later abandoned in favour of batch normalisation.
  7. SGD with momentum 0.9 and weight decay, with the learning rate reduced manually when validation error plateaued.
python
from torchvision import models
alexnet = models.alexnet(weights=models.AlexNet_Weights.IMAGENET1K_V1)
print(alexnet)
print("parameters:", sum(p.numel() for p in alexnet.parameters()))   # ~61 million

Impact#

After 2012, almost every competitive ILSVRC entry used deep CNNs, and the error rate fell year after year (ZFNet 2013, VGG and GoogLeNet 2014, ResNet 2015). Industry invested heavily in GPUs and deep learning, and the approach spread to speech, language and beyond. The recipe โ€” big data + big compute + deep networks trained end-to-end โ€” became the paradigm of modern AI.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

ResNet in Depth: Architecture, Bottlenecks and Variants

ResNet's residual blocks enabled 152-layer networks and became the default vision backbone. We study basic and bottleneck blocks, the full ResNet-50 layout, training recipe, and descendants such as ResNeXt and ConvNeXt.

Intermediateโฑ 5 min#140
๐Ÿ‘๏ธ Computer Vision

EfficientNet and Principled Model Scaling

How should a network grow when you have more compute โ€” deeper, wider or higher resolution? EfficientNet's compound scaling answers "all three, in balance". We cover MBConv blocks, the B0โ€“B7 family and EfficientNetV2.

Intermediateโฑ 5 min#141
๐Ÿ‘๏ธ Computer Vision

Classical Features: Harris Corners, SIFT, HOG and Bag of Visual Words

Before CNNs, vision relied on carefully engineered features. We study corner detection, SIFT keypoints and descriptors, HOG for pedestrian detection and the bag-of-visual-words model โ€” ideas still used in geometry and robotics.

Intermediateโฑ 5 min#137