A crop-disease detector used by farmers without reliable internet, a reading aid for the visually impaired, a camera that counts vehicles on a solar-powered pole — these applications must run on the device, with limited memory, compute and battery. Server-class models like ResNet-50 are too heavy. A family of architectures, led by MobileNet, was designed specifically for this setting.
The cost of a standard convolution#
For an input with $M$ channels, output $N$ channels, kernel $D_K \times D_K$ and output feature map $D_F \times D_F$, a standard convolution costs
multiply–accumulate operations: every output channel mixes every input channel at every spatial position.
Depthwise separable convolutions#
MobileNet (Howard et al., 2017) factorises this into two steps:
- Depthwise convolution — one $D_K \times D_K$ filter per input channel (spatial filtering only): cost $D_K^2 \cdot M \cdot D_F^2$.
- Pointwise convolution — a $1 \times 1$ convolution mixing channels: cost $M \cdot N \cdot D_F^2$.
The ratio of costs is
With $3 \times 3$ kernels, depthwise separable convolution is about 8–9× cheaper, with only a small accuracy drop.
import torch.nn as nn
def depthwise_separable(cin, cout, stride=1):
return nn.Sequential(
nn.Conv2d(cin, cin, 3, stride, 1, groups=cin, bias=False), # depthwise
nn.BatchNorm2d(cin), nn.ReLU6(inplace=True),
nn.Conv2d(cin, cout, 1, bias=False), # pointwise
nn.BatchNorm2d(cout), nn.ReLU6(inplace=True))
std = nn.Conv2d(256, 256, 3, padding=1, bias=False)
sep = depthwise_separable(256, 256)
count = lambda m: sum(p.numel() for p in m.parameters() if p.dim() > 1)
print("standard:", count(std), " separable:", count(sep)) # ~589k vs ~68k weightsMobileNetV1 also introduced two global knobs: a width multiplier $\alpha$ (thin every layer's channels) and a resolution multiplier $\rho$ (smaller inputs), letting one design family trade accuracy for speed across devices.
MobileNetV2: inverted residuals and linear bottlenecks#
Sandler et al. (2018) introduced the inverted residual block:
- Expand a thin input (e.g. 24 channels) with a $1 \times 1$ conv by a factor of 6;
- apply a depthwise $3 \times 3$ conv in the wide space;
- project back to a thin output with a $1 \times 1$ conv — with no activation (a linear bottleneck), because ReLU in a low-dimensional space destroys information;
- add a residual connection between the thin ends.
It is "inverted" relative to ResNet's bottleneck, which is wide at the ends and thin in the middle. Keeping the residual stream thin saves memory — critical on mobile devices.
MobileNetV3#
Howard et al. (2019) combined hardware-aware NAS (for block choices) with NetAdapt (for layer widths), added squeeze-and-excitation modules, and introduced the cheap hard-swish activation:
a piecewise-linear approximation of Swish that is fast and quantisation-friendly. It was released in Large and Small variants for different budgets.
Other efficient designs#
- ShuffleNet (2018): grouped $1 \times 1$ convolutions plus a channel shuffle so information flows between groups. ShuffleNetV2 proposed practical guidelines for real speed: equal input/output channel widths minimise memory access cost; excessive group convolution and network fragmentation slow things down; element-wise operations are not free.
- SqueezeNet (2016): AlexNet-level accuracy with 50× fewer parameters using "fire" modules.
- GhostNet: generate some feature maps with cheap linear operations.
- EfficientNet-Lite, MobileViT, EfficientFormer, MobileNetV4: efficient CNN and hybrid CNN–transformer designs tuned for mobile accelerators.
Deploying on the edge#
Architecture is only part of the story. The full recipe:
- Choose an efficient backbone matched to the hardware (CPU, mobile GPU, NPU, microcontroller).
- Quantise to INT8 (post-training or quantisation-aware).
- Prune or distil if needed.
- Convert to a mobile runtime: TensorFlow Lite / LiteRT, ONNX Runtime Mobile, Core ML, ExecuTorch, or TensorFlow Lite Micro for microcontrollers.
- Benchmark on the actual device — latency, memory, battery drain and accuracy on realistic, locally collected images.
import torch
from torchvision import models
model = models.mobilenet_v3_small(weights=models.MobileNet_V3_Small_Weights.IMAGENET1K_V1).eval()
print(sum(p.numel() for p in model.parameters()) / 1e6, "M parameters") # ~2.5M
example = torch.randn(1, 3, 224, 224)
torch.onnx.export(model, example, "mobilenet_v3_small.onnx", opset_version=17)