Large neural networks can memorise training data. In 2012, Hinton and colleagues proposed a remarkably simple remedy: during training, randomly drop each neuron (set its output to zero) with some probability. This technique, dropout, was a key ingredient in AlexNet's ImageNet victory and became one of the most widely used regularisers in deep learning.
The mechanism#
During training, for each example and each unit, sample a mask $m_i \sim \text{Bernoulli}(1 - p)$, where $p$ is the drop probability:
At test time, all units are active. To keep the expected activation the same in both phases, we must scale. Modern implementations use inverted dropout, scaling during training so that inference needs no change:
import torch
def dropout(h, p=0.5, training=True):
if not training or p == 0:
return h
mask = (torch.rand_like(h) > p).float()
return h * mask / (1 - p)
h = torch.ones(8)
print(dropout(h, 0.5)) # some zeros, others scaled to 2.0
print(dropout(torch.ones(100000), 0.5).mean()) # ~1.0 in expectationWhy does it work?#
1. Preventing co-adaptation#
Without dropout, a neuron can rely on specific other neurons to correct its mistakes, forming fragile, co-adapted feature detectors. With dropout, any neuron may disappear, so each must be useful on its own and in many different combinations โ encouraging robust, redundant features.
2. An implicit ensemble#
A network with $n$ droppable units defines $2^n$ "thinned" sub-networks that share weights. Each training step trains one random sub-network. At test time, using all units with (inverted) scaling approximates the geometric average of the predictions of this exponentially large ensemble โ a cheap form of bagging.
3. Noise as regularisation#
For linear models, dropout on inputs is approximately equivalent to an adaptive L2 penalty that depends on feature variance. More generally, injecting noise during training discourages the model from depending on fine details of individual examples.
Where and how much?#
- Fully connected layers: $p = 0.5$ was the classic choice for large hidden layers (AlexNet, VGG).
- Input layer: small $p$ (0.1โ0.2) โ dropping too many inputs destroys information.
- Convolutional layers: standard dropout is less effective because neighbouring pixels are correlated; use spatial dropout (drop whole channels) or DropBlock (drop contiguous regions), or rely on data augmentation.
- Transformers: dropout of about 0.1 on attention weights, residual branches and embeddings is common for moderate-sized models; very large language models trained on huge datasets for a single epoch often use little or no dropout, because they are not in an overfitting regime.
- RNNs: apply the same dropout mask at every time step (variational dropout) rather than resampling per step.
Variants#
| Variant | Drops | Use |
|---|---|---|
| Standard dropout | individual activations | MLP layers |
| Spatial dropout | entire feature maps | CNNs |
| DropBlock | contiguous spatial regions | CNNs |
| DropConnect | individual weights | Regularising weights directly |
| Stochastic depth / DropPath | entire residual blocks | Deep ResNets, vision transformers |
| Attention dropout | attention probabilities | Transformers |
Stochastic depth deserves special mention: randomly skipping whole residual blocks during training regularises very deep networks and speeds training; it is standard in modern vision transformers and ConvNeXt.
Train/eval mode#
Dropout behaves differently in training and inference, so โ as with BatchNorm โ you must call model.train() and model.eval() correctly. Forgetting eval() makes predictions noisy and systematically worse.
Monte Carlo dropout for uncertainty#
Gal and Ghahramani (2016) showed that keeping dropout active at test time and averaging many stochastic forward passes approximates Bayesian inference over the weights. The spread of the predictions gives an estimate of model uncertainty.
import torch, torch.nn as nn
model = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Dropout(0.2),
nn.Linear(128, 128), nn.ReLU(), nn.Dropout(0.2), nn.Linear(128, 1))
# ... train on data x in [-2, 2] ...
def mc_predict(model, x, n=100):
model.train() # keep dropout ON
with torch.no_grad():
preds = torch.stack([model(x) for _ in range(n)])
return preds.mean(0), preds.std(0) # prediction and uncertainty
x_query = torch.tensor([[0.0], [5.0]]) # in-distribution vs far outside
mean, std = mc_predict(model, x_query)
print(mean.squeeze(), std.squeeze()) # after training, std should be larger at x = 5(If your model contains BatchNorm, switch only the dropout layers to training mode.) MC dropout is a practical, inexpensive way to flag inputs on which the model should not be trusted โ valuable in medical and humanitarian decision support.