Given twice the compute, how should you enlarge a CNN? Before 2019, practitioners typically scaled one dimension: ResNet grew deeper (18 → 152 layers), Wide ResNets grew wider, and some models used higher-resolution input. Mingxing Tan and Quoc Le's EfficientNet (2019) showed that scaling all three dimensions together, in a fixed ratio, gives much better accuracy per unit of compute.
Three scaling dimensions#
- Depth $d$ — more layers capture richer, more complex features but face diminishing returns and harder optimisation.
- Width $w$ — more channels capture finer-grained features; very wide shallow networks struggle to learn high-level features.
- Resolution $r$ — larger input images reveal finer detail; returns diminish at very high resolution.
Experiments showed that scaling any single dimension saturates quickly (e.g. accuracy stalls around 80% on ImageNet), and the dimensions interact: higher-resolution images benefit from deeper networks (bigger receptive fields) and wider networks (more fine-grained patterns).
Compound scaling#
EfficientNet scales all three with a single compound coefficient $\phi$:
FLOPs of a convolutional network scale roughly linearly with depth and quadratically with width and resolution, so this constraint makes total FLOPs grow by about $2^\phi$. A small grid search on the baseline found $\alpha = 1.2$, $\beta = 1.1$, $\gamma = 1.15$.
The baseline: EfficientNet-B0 and MBConv#
The baseline network was found by hardware-aware neural architecture search optimising accuracy and FLOPs. Its building block is the MBConv (mobile inverted bottleneck) from MobileNetV2, plus squeeze-and-excitation:
- $1 \times 1$ conv expands channels (e.g. 6×);
- depthwise $3 \times 3$ or $5 \times 5$ conv filters each channel spatially;
- squeeze-and-excitation reweights channels;
- $1 \times 1$ conv projects back to fewer channels;
- residual connection when shapes match (with stochastic depth during training).
It uses the SiLU (Swish) activation.
import torch
import torch.nn as nn
class MBConv(nn.Module):
def __init__(self, cin, cout, expand=6, k=3, stride=1, se_ratio=0.25):
super().__init__()
mid = cin * expand
self.use_res = stride == 1 and cin == cout
self.block = nn.Sequential(
nn.Conv2d(cin, mid, 1, bias=False), nn.BatchNorm2d(mid), nn.SiLU(),
nn.Conv2d(mid, mid, k, stride, k // 2, groups=mid, bias=False), nn.BatchNorm2d(mid), nn.SiLU())
se = max(1, int(cin * se_ratio))
self.se = nn.Sequential(nn.AdaptiveAvgPool2d(1), nn.Conv2d(mid, se, 1), nn.SiLU(),
nn.Conv2d(se, mid, 1), nn.Sigmoid())
self.project = nn.Sequential(nn.Conv2d(mid, cout, 1, bias=False), nn.BatchNorm2d(cout))
def forward(self, x):
h = self.block(x)
h = self.project(h * self.se(h))
return x + h if self.use_res else h
print(MBConv(40, 40)(torch.randn(1, 40, 28, 28)).shape)The EfficientNet family#
Applying $\phi = 1, \dots, 7$ to B0 produced EfficientNet-B1 to B7. At publication, EfficientNet-B7 reached about 84.3% ImageNet top-1 accuracy while being 8.4× smaller and 6.1× faster at inference than the best existing CNN of similar accuracy. Smaller variants matched ResNet-50 with far fewer FLOPs and parameters, and the models transferred well to other datasets.
| Model | Input resolution | Parameters (approx.) |
|---|---|---|
| B0 | 224 | 5.3M |
| B3 | 300 | 12M |
| B5 | 456 | 30M |
| B7 | 600 | 66M |
EfficientNetV2#
EfficientNet's FLOP efficiency did not always translate into fast training: depthwise convolutions underuse accelerators, and large images consume memory. EfficientNetV2 (Tan & Le, 2021) addressed this with:
- Fused-MBConv blocks in early stages (a regular $3 \times 3$ conv replaces the expansion + depthwise pair, which runs faster on GPUs/TPUs);
- training-aware NAS optimising accuracy, speed and parameter efficiency;
- progressive learning: train with small images and weak regularisation first, then larger images with stronger regularisation.