Designing neural architectures โ choosing layer types, widths, depths, connections โ was long a craft of human intuition and trial and error. Neural Architecture Search (NAS) automates this: an algorithm searches over a space of possible architectures for the one that performs best. NAS produced several state-of-the-art models, most famously the EfficientNet family, and it belongs to the broader field of AutoML, which also automates hyperparameter tuning and feature engineering.
The three components of NAS#
Elsken, Metzen and Hutter (2019) describe every NAS method by three choices:
- Search space โ which architectures can be expressed.
- Search strategy โ how to explore that space.
- Performance estimation โ how to evaluate a candidate cheaply.
Search spaces#
- Global / chain-structured: choose each layer's type and hyperparameters in sequence.
- Cell-based: search for a small repeated building block (a "cell") and stack it; this made search tractable and transferable (NASNet learned cells on CIFAR-10 and transferred them to ImageNet).
- Hierarchical / macro spaces: choose depths, widths and resolutions per stage.
The search space encodes strong human priors. A space made only of good building blocks makes even random search competitive โ an important and humbling lesson.
Search strategies#
Reinforcement learning#
Zoph and Le (2017) used an RNN controller to generate architecture descriptions, trained with policy gradients using validation accuracy as the reward. It produced excellent architectures but required enormous compute (hundreds of GPUs for weeks).
Evolutionary algorithms#
Maintain a population of architectures; mutate the good ones (add a layer, change a kernel); select by fitness. Regularized (aging) evolution (Real et al., 2019) โ removing the oldest rather than the worst individuals โ produced AmoebaNet, matching RL-based results.
Bayesian optimisation#
Model architecture performance with surrogates (e.g. with graph kernels or neural predictors) to choose promising candidates.
Differentiable NAS (DARTS)#
Liu et al. (2019) relaxed the discrete choice of operation on each edge into a softmax-weighted mixture:
Architecture parameters $\alpha$ and network weights are optimised jointly by gradient descent (a bilevel problem), and the strongest operation on each edge is kept at the end. This cut search cost to a few GPU-days, though DARTS can be unstable (e.g. collapsing towards parameter-free skip connections).
Cheap performance estimation#
Training every candidate to convergence is prohibitively expensive. Shortcuts:
- Low-fidelity estimates โ fewer epochs, smaller images, data subsets (combine with Hyperband).
- Learning-curve extrapolation โ stop unpromising runs early.
- Weight sharing / one-shot supernets โ train one large "supernet" containing all candidate architectures as sub-graphs; evaluate candidates by inheriting its weights (ENAS, Once-for-All). Fast, but inherited-weight rankings can correlate imperfectly with stand-alone performance.
- Zero-cost proxies โ score untrained networks by properties at initialisation (gradient statistics, activation patterns).
Hardware-aware NAS#
For deployment, accuracy is not the only objective. Multi-objective NAS includes latency, energy or memory on a specific device:
(MnasNet's reward). MobileNetV3 and EfficientNet's base network were found with such searches. Once-for-All trains one supernet and then extracts specialised sub-networks for many devices without retraining.
EfficientNet's compound scaling#
Tan and Le (2019) combined NAS with a principled scaling rule. After searching for a small baseline (EfficientNet-B0), they scaled depth $d$, width $w$ and resolution $r$ together:
so each increment of $\phi$ roughly doubles compute. The family B0โB7 achieved excellent accuracyโefficiency trade-offs.
Lessons from the NAS era#
AutoML in practice#
For most practitioners, AutoML means tools that automate model selection and tuning: Auto-sklearn, AutoGluon, H2O AutoML, FLAML, cloud AutoML services and Optuna-based pipelines. AutoGluon, for example, combines many model families with stacking and often provides a very strong tabular baseline with a few lines of code. Use AutoML as a baseline and accelerator, not a substitute for understanding the problem, the data and the evaluation.