👁️ Computer Vision · Lecture 7 of 27

EfficientNet and Principled Model Scaling

How should a network grow when you have more compute — deeper, wider or higher resolution? EfficientNet's compound scaling answers "all three, in balance". We cover MBConv blocks, the B0–B7 family and EfficientNetV2.

Given twice the compute, how should you enlarge a CNN? Before 2019, practitioners typically scaled one dimension: ResNet grew deeper (18 → 152 layers), Wide ResNets grew wider, and some models used higher-resolution input. Mingxing Tan and Quoc Le's EfficientNet (2019) showed that scaling all three dimensions together, in a fixed ratio, gives much better accuracy per unit of compute.

Three scaling dimensions#

  • Depth $d$ — more layers capture richer, more complex features but face diminishing returns and harder optimisation.
  • Width $w$ — more channels capture finer-grained features; very wide shallow networks struggle to learn high-level features.
  • Resolution $r$ — larger input images reveal finer detail; returns diminish at very high resolution.

Experiments showed that scaling any single dimension saturates quickly (e.g. accuracy stalls around 80% on ImageNet), and the dimensions interact: higher-resolution images benefit from deeper networks (bigger receptive fields) and wider networks (more fine-grained patterns).

Compound scaling#

EfficientNet scales all three with a single compound coefficient $\phi$:

$$ d = \alpha^\phi, \qquad w = \beta^\phi, \qquad r = \gamma^\phi, \qquad \text{subject to } \alpha\cdot\beta^2\cdot\gamma^2 \approx 2, \;\; \alpha, \beta, \gamma \ge 1 $$

FLOPs of a convolutional network scale roughly linearly with depth and quadratically with width and resolution, so this constraint makes total FLOPs grow by about $2^\phi$. A small grid search on the baseline found $\alpha = 1.2$, $\beta = 1.1$, $\gamma = 1.15$.

The baseline: EfficientNet-B0 and MBConv#

The baseline network was found by hardware-aware neural architecture search optimising accuracy and FLOPs. Its building block is the MBConv (mobile inverted bottleneck) from MobileNetV2, plus squeeze-and-excitation:

  1. $1 \times 1$ conv expands channels (e.g. 6×);
  2. depthwise $3 \times 3$ or $5 \times 5$ conv filters each channel spatially;
  3. squeeze-and-excitation reweights channels;
  4. $1 \times 1$ conv projects back to fewer channels;
  5. residual connection when shapes match (with stochastic depth during training).

It uses the SiLU (Swish) activation.

python
import torch
import torch.nn as nn

class MBConv(nn.Module):
    def __init__(self, cin, cout, expand=6, k=3, stride=1, se_ratio=0.25):
        super().__init__()
        mid = cin * expand
        self.use_res = stride == 1 and cin == cout
        self.block = nn.Sequential(
            nn.Conv2d(cin, mid, 1, bias=False), nn.BatchNorm2d(mid), nn.SiLU(),
            nn.Conv2d(mid, mid, k, stride, k // 2, groups=mid, bias=False), nn.BatchNorm2d(mid), nn.SiLU())
        se = max(1, int(cin * se_ratio))
        self.se = nn.Sequential(nn.AdaptiveAvgPool2d(1), nn.Conv2d(mid, se, 1), nn.SiLU(),
                                nn.Conv2d(se, mid, 1), nn.Sigmoid())
        self.project = nn.Sequential(nn.Conv2d(mid, cout, 1, bias=False), nn.BatchNorm2d(cout))
    def forward(self, x):
        h = self.block(x)
        h = self.project(h * self.se(h))
        return x + h if self.use_res else h

print(MBConv(40, 40)(torch.randn(1, 40, 28, 28)).shape)

The EfficientNet family#

Applying $\phi = 1, \dots, 7$ to B0 produced EfficientNet-B1 to B7. At publication, EfficientNet-B7 reached about 84.3% ImageNet top-1 accuracy while being 8.4× smaller and 6.1× faster at inference than the best existing CNN of similar accuracy. Smaller variants matched ResNet-50 with far fewer FLOPs and parameters, and the models transferred well to other datasets.

ModelInput resolutionParameters (approx.)
B02245.3M
B330012M
B545630M
B760066M

EfficientNetV2#

EfficientNet's FLOP efficiency did not always translate into fast training: depthwise convolutions underuse accelerators, and large images consume memory. EfficientNetV2 (Tan & Le, 2021) addressed this with:

  • Fused-MBConv blocks in early stages (a regular $3 \times 3$ conv replaces the expansion + depthwise pair, which runs faster on GPUs/TPUs);
  • training-aware NAS optimising accuracy, speed and parameter efficiency;
  • progressive learning: train with small images and weak regularisation first, then larger images with stronger regularisation.

Lessons#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

ResNet in Depth: Architecture, Bottlenecks and Variants

ResNet's residual blocks enabled 152-layer networks and became the default vision backbone. We study basic and bottleneck blocks, the full ResNet-50 layout, training recipe, and descendants such as ResNeXt and ConvNeXt.

Intermediate⏱ 5 min#140
👁️ Computer Vision

MobileNet and Efficient Architectures for Edge Devices

Phones, drones and microcontrollers need vision models that are small and fast. We derive the cost savings of depthwise separable convolutions and study MobileNet V1–V3, ShuffleNet and design principles for efficient inference.

Intermediate⏱ 5 min#142
👁️ Computer Vision

LeNet and AlexNet: The Birth of Deep Vision

Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.

Beginner⏱ 5 min#138