Before a network learns anything, we must choose starting values for its weights. This seems like a detail, but it was one of the key obstacles that held back deep networks for years. Initialise too small and signals shrink to nothing as they pass through layers; too large and they explode. Today we derive the principled answer.
Why not zeros?#
If all weights in a layer start equal (e.g. all zero), every neuron in the layer computes the same output and receives the same gradient. They remain identical forever โ the layer behaves like a single neuron. Random initialisation breaks this symmetry. (Biases can safely start at zero.)
The variance argument#
Consider a layer $z_i = \sum_{j=1}^{n_{\text{in}}}w_{ij}x_j$ with independent zero-mean weights of variance $\text{Var}(w)$ and independent inputs with zero mean and variance $\text{Var}(x)$. Then
After $L$ layers, the variance is multiplied by $\left(n_{\text{in}}\text{Var}(w)\right)^L$ (ignoring activations). If $n_{\text{in}}\text{Var}(w) = 1.5$ and $L = 50$, activations grow by a factor of about $1.5^{50} \approx 6 \times 10^{8}$; if it equals 0.5, they shrink by $10^{15}$. The same argument applies to gradients flowing backwards. We want this factor to be exactly 1.
Xavier / Glorot initialisation#
For activations that are roughly linear around zero (tanh, sigmoid's central region), Glorot and Bengio (2010) required variance preservation in both the forward pass ($n_{\text{in}}\text{Var}(w) = 1$) and the backward pass ($n_{\text{out}}\text{Var}(w) = 1$), and compromised with the average:
Uniform version: $w \sim U\left[-\sqrt{\frac{6}{n_{\text{in}} + n_{\text{out}}}},\; \sqrt{\frac{6}{n_{\text{in}} + n_{\text{out}}}}\right]$.
He / Kaiming initialisation#
ReLU sets half its inputs to zero, which halves the second moment of the signal. He et al. (2015) compensated with a factor of 2:
This let them train very deep plain ReLU networks (30 layers) from scratch where Xavier initialisation stalled. It is the default for ReLU-family networks.
| Activation | Recommended initialisation |
|---|---|
| tanh, sigmoid, linear | Xavier / Glorot |
| ReLU, Leaky ReLU, GELU, SiLU | He / Kaiming (with the appropriate gain) |
| SELU | LeCun normal $\text{Var}(w) = 1/n_{\text{in}}$ |
Seeing it happen#
import torch
def activation_stats(init, act, depth=50, width=512):
x = torch.randn(1000, width)
for _ in range(depth):
W = torch.empty(width, width)
init(W)
x = act(x @ W.T)
return x.std().item()
inits = {
"N(0, 0.01)": lambda W: torch.nn.init.normal_(W, 0, 0.01),
"N(0, 1)": lambda W: torch.nn.init.normal_(W, 0, 1.0),
"Xavier": torch.nn.init.xavier_normal_,
"He (Kaiming)": lambda W: torch.nn.init.kaiming_normal_(W, nonlinearity="relu"),
}
for name, init in inits.items():
print(f"ReLU net, {name:<13}: std after 50 layers = {activation_stats(init, torch.relu):.3e}")Small fixed variance collapses activations to zero; unit variance explodes them; Xavier slowly shrinks them with ReLU; He keeps them stable.
Beyond variance: modern practice#
- Orthogonal initialisation โ initialise weight matrices as random orthogonal matrices (all singular values 1), which preserves norms exactly in linear networks and helps RNNs.
- Residual networks: each block adds to the input, $\mathbf{x} + F(\mathbf{x})$, so variance grows with depth. Common fixes: initialise the last layer of each residual branch to zero ("zero-init residual" / Fixup) so each block starts as the identity, or scale residual-branch initialisation by $1/\sqrt{2L}$ as in GPT-2.
- Normalisation layers (BatchNorm, LayerNorm) make networks much less sensitive to initialisation, but good initialisation still speeds training.
- Transformers: small normal initialisation (e.g. standard deviation 0.02) is common, with scaled residual projections.
- Maximal update parametrisation (ฮผP) prescribes how initialisation and learning rates should scale with width so that hyperparameters tuned on a small model transfer to a large one.