๐Ÿ”— Deep Learning ยท Lecture 11 of 38

Weight Initialisation: Xavier, He and Why It Matters

Bad initial weights make signals explode or vanish before training even starts. We derive variance-preserving initialisation for tanh (Xavier) and ReLU (He) networks and discuss modern practice for deep and residual models.

Before a network learns anything, we must choose starting values for its weights. This seems like a detail, but it was one of the key obstacles that held back deep networks for years. Initialise too small and signals shrink to nothing as they pass through layers; too large and they explode. Today we derive the principled answer.

Why not zeros?#

If all weights in a layer start equal (e.g. all zero), every neuron in the layer computes the same output and receives the same gradient. They remain identical forever โ€” the layer behaves like a single neuron. Random initialisation breaks this symmetry. (Biases can safely start at zero.)

The variance argument#

Consider a layer $z_i = \sum_{j=1}^{n_{\text{in}}}w_{ij}x_j$ with independent zero-mean weights of variance $\text{Var}(w)$ and independent inputs with zero mean and variance $\text{Var}(x)$. Then

$$ \text{Var}(z_i) = n_{\text{in}}\,\text{Var}(w)\,\text{Var}(x) $$

After $L$ layers, the variance is multiplied by $\left(n_{\text{in}}\text{Var}(w)\right)^L$ (ignoring activations). If $n_{\text{in}}\text{Var}(w) = 1.5$ and $L = 50$, activations grow by a factor of about $1.5^{50} \approx 6 \times 10^{8}$; if it equals 0.5, they shrink by $10^{15}$. The same argument applies to gradients flowing backwards. We want this factor to be exactly 1.

Xavier / Glorot initialisation#

For activations that are roughly linear around zero (tanh, sigmoid's central region), Glorot and Bengio (2010) required variance preservation in both the forward pass ($n_{\text{in}}\text{Var}(w) = 1$) and the backward pass ($n_{\text{out}}\text{Var}(w) = 1$), and compromised with the average:

$$ \text{Var}(w) = \frac{2}{n_{\text{in}} + n_{\text{out}}} $$

Uniform version: $w \sim U\left[-\sqrt{\frac{6}{n_{\text{in}} + n_{\text{out}}}},\; \sqrt{\frac{6}{n_{\text{in}} + n_{\text{out}}}}\right]$.

He / Kaiming initialisation#

ReLU sets half its inputs to zero, which halves the second moment of the signal. He et al. (2015) compensated with a factor of 2:

$$ \text{Var}(w) = \frac{2}{n_{\text{in}}}, \qquad w \sim \mathcal{N}\left(0, \frac{2}{n_{\text{in}}}\right) $$

This let them train very deep plain ReLU networks (30 layers) from scratch where Xavier initialisation stalled. It is the default for ReLU-family networks.

ActivationRecommended initialisation
tanh, sigmoid, linearXavier / Glorot
ReLU, Leaky ReLU, GELU, SiLUHe / Kaiming (with the appropriate gain)
SELULeCun normal $\text{Var}(w) = 1/n_{\text{in}}$

Seeing it happen#

python
import torch

def activation_stats(init, act, depth=50, width=512):
    x = torch.randn(1000, width)
    for _ in range(depth):
        W = torch.empty(width, width)
        init(W)
        x = act(x @ W.T)
    return x.std().item()

inits = {
    "N(0, 0.01)":  lambda W: torch.nn.init.normal_(W, 0, 0.01),
    "N(0, 1)":     lambda W: torch.nn.init.normal_(W, 0, 1.0),
    "Xavier":      torch.nn.init.xavier_normal_,
    "He (Kaiming)": lambda W: torch.nn.init.kaiming_normal_(W, nonlinearity="relu"),
}
for name, init in inits.items():
    print(f"ReLU net, {name:<13}: std after 50 layers = {activation_stats(init, torch.relu):.3e}")

Small fixed variance collapses activations to zero; unit variance explodes them; Xavier slowly shrinks them with ReLU; He keeps them stable.

Beyond variance: modern practice#

  • Orthogonal initialisation โ€” initialise weight matrices as random orthogonal matrices (all singular values 1), which preserves norms exactly in linear networks and helps RNNs.
  • Residual networks: each block adds to the input, $\mathbf{x} + F(\mathbf{x})$, so variance grows with depth. Common fixes: initialise the last layer of each residual branch to zero ("zero-init residual" / Fixup) so each block starts as the identity, or scale residual-branch initialisation by $1/\sqrt{2L}$ as in GPT-2.
  • Normalisation layers (BatchNorm, LayerNorm) make networks much less sensitive to initialisation, but good initialisation still speeds training.
  • Transformers: small normal initialisation (e.g. standard deviation 0.02) is common, with scaled residual projections.
  • Maximal update parametrisation (ฮผP) prescribes how initialisation and learning rates should scale with width so that hyperparameters tuned on a small model transfer to a large one.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Vanishing and Exploding Gradients โ€” Causes and Cures

Gradients are products of many Jacobians, so they can shrink or grow exponentially with depth. We analyse why, how to diagnose it, and the arsenal of fixes from ReLU and initialisation to residuals, normalisation, clipping and gating.

Intermediateโฑ 4 min#108
๐Ÿ”— Deep Learning

Residual Connections: Why Very Deep Networks Became Trainable

Deeper plain networks can train worse than shallower ones. Residual connections fix this by learning corrections to the identity. We explain the degradation problem, the gradient highway, and variants from ResNet to transformers.

Intermediateโฑ 5 min#120
๐Ÿ”— Deep Learning

Learning Rate Schedules, Warm-up and the LR Range Test

The learning rate is the most important hyperparameter, and it should change during training. We compare step, exponential, cosine and one-cycle schedules, explain why warm-up stabilises large models, and find good rates quickly.

Intermediateโฑ 5 min#106