๐Ÿ”— Deep Learning ยท Lecture 14 of 38

Beyond BatchNorm: Layer, Group, Instance and RMS Normalisation

Normalisation layers differ only in which axes they average over โ€” yet that choice decides where they work. We compare LayerNorm, GroupNorm, InstanceNorm and RMSNorm, and the pre-norm versus post-norm debate in transformers.

Batch normalisation computes statistics across the batch. That fails with small batches and is awkward for sequences. A family of alternatives normalises over other axes of the activation tensor, removing the dependence on batch size. One of them โ€” Layer Normalisation โ€” is in every transformer; its simplified cousin RMSNorm is in most recent large language models.

One formula, different axes#

All these methods compute

$$ \hat{x} = \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}}, \qquad y = \gamma\,\hat{x} + \beta $$

and differ only in which elements share $\mu$ and $\sigma$. For an image activation tensor of shape $(N, C, H, W)$:

MethodStatistics computed overDepends on batch?Typical use
BatchNorm$N, H, W$ (per channel)YesCNNs with moderate/large batches
LayerNorm$C, H, W$ (per example)NoTransformers, RNNs
InstanceNorm$H, W$ (per example, per channel)NoStyle transfer, image generation
GroupNormgroups of channels ร— $H, W$NoDetection/segmentation with small batches

For a transformer activation of shape $(N, T, D)$, LayerNorm normalises each token's $D$-dimensional vector independently.

Layer Normalisation#

Ba, Kiros and Hinton (2016) normalise across the features of each example (each token, in a transformer). Consequences:

  • identical computation in training and inference โ€” no running statistics;
  • works with batch size 1 and variable-length sequences;
  • well suited to recurrent and attention models.

Group Normalisation#

Wu and He (2018) divide channels into $G$ groups (e.g. 32) and normalise within each group per example. With $G = 1$ it equals LayerNorm; with $G = C$ it equals InstanceNorm. GroupNorm's accuracy is stable across batch sizes, making it the standard in detection and segmentation, where memory-hungry high-resolution images force tiny batches.

Instance Normalisation#

Normalises each channel of each image separately, removing instance-specific contrast and colour statistics. That is exactly what style transfer wants: the "style" of an image lives largely in these per-channel statistics. Adaptive Instance Normalisation (AdaIN) replaces $\gamma$ and $\beta$ with statistics of a style image โ€” the core idea behind StyleGAN's style modulation.

RMSNorm#

Zhang and Sennrich (2019) observed that re-centring (subtracting the mean) contributes little; re-scaling is what matters. RMSNorm drops the mean and the bias:

$$ y = \gamma\odot\frac{\mathbf{x}}{\text{RMS}(\mathbf{x})}, \qquad \text{RMS}(\mathbf{x}) = \sqrt{\frac{1}{D}\sum_{i=1}^{D}x_i^2 + \epsilon} $$

It is simpler and cheaper while matching LayerNorm's quality, and it has become the default in many modern LLMs (the LLaMA family, among others).

python
import torch
import torch.nn as nn

class RMSNorm(nn.Module):
    def __init__(self, dim, eps=1e-6):
        super().__init__()
        self.eps, self.weight = eps, nn.Parameter(torch.ones(dim))
    def forward(self, x):
        return self.weight * x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)

x = torch.randn(2, 5, 8) * 3 + 1                     # (batch, tokens, features)
ln, rms = nn.LayerNorm(8), RMSNorm(8)
print("LayerNorm per-token mean/std:", ln(x).mean(-1)[0, :3].detach(), ln(x).std(-1, unbiased=False)[0, :3].detach())
print("RMSNorm per-token RMS:", rms(x).pow(2).mean(-1).sqrt()[0, :3].detach())

Pre-norm versus post-norm transformers#

The original Transformer placed LayerNorm after the residual addition (post-norm):

$$ \mathbf{x}_{l+1} = \text{LN}\big(\mathbf{x}_l + F(\mathbf{x}_l)\big) $$

Most modern models use pre-norm, normalising the input to each sub-layer:

$$ \mathbf{x}_{l+1} = \mathbf{x}_l + F\big(\text{LN}(\mathbf{x}_l)\big) $$

Pre-norm keeps an untouched residual path from the output back to the input, so gradients flow more easily; it trains stably without long warm-up and scales to very deep models. Post-norm can achieve slightly better final quality when it trains successfully but is more fragile. Hybrids and extra normalisations (e.g. normalising queries and keys before attention, "QK-norm") are used to improve stability at scale.

Choosing a normalisation#

  • CNN, batch โ‰ฅ 16: BatchNorm.
  • CNN, small batches (detection, segmentation, medical 3-D): GroupNorm.
  • Transformers and RNNs: LayerNorm or RMSNorm, usually pre-norm.
  • Style transfer / image generation: InstanceNorm or AdaIN.
  • When in doubt for new sequence models: pre-norm RMSNorm.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Batch Normalisation: Faster, More Stable Training

BatchNorm normalises each feature using mini-batch statistics, then rescales it with learned parameters. We derive the forward pass, explain training-versus-inference behaviour, debate why it works, and list its pitfalls.

Intermediateโฑ 5 min#109
๐Ÿ”— Deep Learning

Dropout: Regularisation by Random Deletion

Randomly switching off neurons during training prevents co-adaptation and approximates an ensemble of exponentially many networks. We cover inverted dropout, where to apply it, its variants, and Monte Carlo dropout for uncertainty.

Beginnerโฑ 5 min#111
๐Ÿ”— Deep Learning

Vanishing and Exploding Gradients โ€” Causes and Cures

Gradients are products of many Jacobians, so they can shrink or grow exponentially with depth. We analyse why, how to diagnose it, and the arsenal of fixes from ReLU and initialisation to residuals, normalisation, clipping and gating.

Intermediateโฑ 4 min#108