๐Ÿ”— Deep Learning ยท Lecture 15 of 38

Dropout: Regularisation by Random Deletion

Randomly switching off neurons during training prevents co-adaptation and approximates an ensemble of exponentially many networks. We cover inverted dropout, where to apply it, its variants, and Monte Carlo dropout for uncertainty.

Large neural networks can memorise training data. In 2012, Hinton and colleagues proposed a remarkably simple remedy: during training, randomly drop each neuron (set its output to zero) with some probability. This technique, dropout, was a key ingredient in AlexNet's ImageNet victory and became one of the most widely used regularisers in deep learning.

The mechanism#

During training, for each example and each unit, sample a mask $m_i \sim \text{Bernoulli}(1 - p)$, where $p$ is the drop probability:

$$ \tilde{h}_i = m_i\,h_i $$

At test time, all units are active. To keep the expected activation the same in both phases, we must scale. Modern implementations use inverted dropout, scaling during training so that inference needs no change:

$$ \tilde{h}_i = \frac{m_i}{1 - p}\,h_i, \qquad \mathbb{E}[\tilde{h}_i] = h_i $$
python
import torch

def dropout(h, p=0.5, training=True):
    if not training or p == 0:
        return h
    mask = (torch.rand_like(h) > p).float()
    return h * mask / (1 - p)

h = torch.ones(8)
print(dropout(h, 0.5))                     # some zeros, others scaled to 2.0
print(dropout(torch.ones(100000), 0.5).mean())   # ~1.0 in expectation

Why does it work?#

1. Preventing co-adaptation#

Without dropout, a neuron can rely on specific other neurons to correct its mistakes, forming fragile, co-adapted feature detectors. With dropout, any neuron may disappear, so each must be useful on its own and in many different combinations โ€” encouraging robust, redundant features.

2. An implicit ensemble#

A network with $n$ droppable units defines $2^n$ "thinned" sub-networks that share weights. Each training step trains one random sub-network. At test time, using all units with (inverted) scaling approximates the geometric average of the predictions of this exponentially large ensemble โ€” a cheap form of bagging.

3. Noise as regularisation#

For linear models, dropout on inputs is approximately equivalent to an adaptive L2 penalty that depends on feature variance. More generally, injecting noise during training discourages the model from depending on fine details of individual examples.

Where and how much?#

  • Fully connected layers: $p = 0.5$ was the classic choice for large hidden layers (AlexNet, VGG).
  • Input layer: small $p$ (0.1โ€“0.2) โ€” dropping too many inputs destroys information.
  • Convolutional layers: standard dropout is less effective because neighbouring pixels are correlated; use spatial dropout (drop whole channels) or DropBlock (drop contiguous regions), or rely on data augmentation.
  • Transformers: dropout of about 0.1 on attention weights, residual branches and embeddings is common for moderate-sized models; very large language models trained on huge datasets for a single epoch often use little or no dropout, because they are not in an overfitting regime.
  • RNNs: apply the same dropout mask at every time step (variational dropout) rather than resampling per step.

Variants#

VariantDropsUse
Standard dropoutindividual activationsMLP layers
Spatial dropoutentire feature mapsCNNs
DropBlockcontiguous spatial regionsCNNs
DropConnectindividual weightsRegularising weights directly
Stochastic depth / DropPathentire residual blocksDeep ResNets, vision transformers
Attention dropoutattention probabilitiesTransformers

Stochastic depth deserves special mention: randomly skipping whole residual blocks during training regularises very deep networks and speeds training; it is standard in modern vision transformers and ConvNeXt.

Train/eval mode#

Dropout behaves differently in training and inference, so โ€” as with BatchNorm โ€” you must call model.train() and model.eval() correctly. Forgetting eval() makes predictions noisy and systematically worse.

Monte Carlo dropout for uncertainty#

Gal and Ghahramani (2016) showed that keeping dropout active at test time and averaging many stochastic forward passes approximates Bayesian inference over the weights. The spread of the predictions gives an estimate of model uncertainty.

python
import torch, torch.nn as nn

model = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Dropout(0.2),
                      nn.Linear(128, 128), nn.ReLU(), nn.Dropout(0.2), nn.Linear(128, 1))
# ... train on data x in [-2, 2] ...

def mc_predict(model, x, n=100):
    model.train()                             # keep dropout ON
    with torch.no_grad():
        preds = torch.stack([model(x) for _ in range(n)])
    return preds.mean(0), preds.std(0)        # prediction and uncertainty

x_query = torch.tensor([[0.0], [5.0]])        # in-distribution vs far outside
mean, std = mc_predict(model, x_query)
print(mean.squeeze(), std.squeeze())          # after training, std should be larger at x = 5

(If your model contains BatchNorm, switch only the dropout layers to training mode.) MC dropout is a practical, inexpensive way to flag inputs on which the model should not be trusted โ€” valuable in medical and humanitarian decision support.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Regularisation in Deep Learning: Weight Decay, Early Stopping, Augmentation and More

Deep networks can memorise anything, yet generalise well when regularised properly. We survey the toolkit โ€” weight decay, early stopping, data augmentation, mixup and cutmix, label smoothing โ€” and how to combine them.

Intermediateโฑ 5 min#112
๐Ÿ”— Deep Learning

Debugging Neural Network Training: A Systematic Recipe

Neural networks fail silently โ€” they train, but badly. We present a systematic recipe for finding bugs, from data inspection and overfitting a single batch to monitoring activations, gradients and learning curves.

Intermediateโฑ 6 min#126
๐Ÿ”— Deep Learning

Beyond BatchNorm: Layer, Group, Instance and RMS Normalisation

Normalisation layers differ only in which axes they average over โ€” yet that choice decides where they work. We compare LayerNorm, GroupNorm, InstanceNorm and RMSNorm, and the pre-norm versus post-norm debate in transformers.

Intermediateโฑ 4 min#110