Batch normalisation computes statistics across the batch. That fails with small batches and is awkward for sequences. A family of alternatives normalises over other axes of the activation tensor, removing the dependence on batch size. One of them โ Layer Normalisation โ is in every transformer; its simplified cousin RMSNorm is in most recent large language models.
One formula, different axes#
All these methods compute
and differ only in which elements share $\mu$ and $\sigma$. For an image activation tensor of shape $(N, C, H, W)$:
| Method | Statistics computed over | Depends on batch? | Typical use |
|---|---|---|---|
| BatchNorm | $N, H, W$ (per channel) | Yes | CNNs with moderate/large batches |
| LayerNorm | $C, H, W$ (per example) | No | Transformers, RNNs |
| InstanceNorm | $H, W$ (per example, per channel) | No | Style transfer, image generation |
| GroupNorm | groups of channels ร $H, W$ | No | Detection/segmentation with small batches |
For a transformer activation of shape $(N, T, D)$, LayerNorm normalises each token's $D$-dimensional vector independently.
Layer Normalisation#
Ba, Kiros and Hinton (2016) normalise across the features of each example (each token, in a transformer). Consequences:
- identical computation in training and inference โ no running statistics;
- works with batch size 1 and variable-length sequences;
- well suited to recurrent and attention models.
Group Normalisation#
Wu and He (2018) divide channels into $G$ groups (e.g. 32) and normalise within each group per example. With $G = 1$ it equals LayerNorm; with $G = C$ it equals InstanceNorm. GroupNorm's accuracy is stable across batch sizes, making it the standard in detection and segmentation, where memory-hungry high-resolution images force tiny batches.
Instance Normalisation#
Normalises each channel of each image separately, removing instance-specific contrast and colour statistics. That is exactly what style transfer wants: the "style" of an image lives largely in these per-channel statistics. Adaptive Instance Normalisation (AdaIN) replaces $\gamma$ and $\beta$ with statistics of a style image โ the core idea behind StyleGAN's style modulation.
RMSNorm#
Zhang and Sennrich (2019) observed that re-centring (subtracting the mean) contributes little; re-scaling is what matters. RMSNorm drops the mean and the bias:
It is simpler and cheaper while matching LayerNorm's quality, and it has become the default in many modern LLMs (the LLaMA family, among others).
import torch
import torch.nn as nn
class RMSNorm(nn.Module):
def __init__(self, dim, eps=1e-6):
super().__init__()
self.eps, self.weight = eps, nn.Parameter(torch.ones(dim))
def forward(self, x):
return self.weight * x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
x = torch.randn(2, 5, 8) * 3 + 1 # (batch, tokens, features)
ln, rms = nn.LayerNorm(8), RMSNorm(8)
print("LayerNorm per-token mean/std:", ln(x).mean(-1)[0, :3].detach(), ln(x).std(-1, unbiased=False)[0, :3].detach())
print("RMSNorm per-token RMS:", rms(x).pow(2).mean(-1).sqrt()[0, :3].detach())Pre-norm versus post-norm transformers#
The original Transformer placed LayerNorm after the residual addition (post-norm):
Most modern models use pre-norm, normalising the input to each sub-layer:
Pre-norm keeps an untouched residual path from the output back to the input, so gradients flow more easily; it trains stably without long warm-up and scales to very deep models. Post-norm can achieve slightly better final quality when it trains successfully but is more fragile. Hybrids and extra normalisations (e.g. normalising queries and keys before attention, "QK-norm") are used to improve stability at scale.
Choosing a normalisation#
- CNN, batch โฅ 16: BatchNorm.
- CNN, small batches (detection, segmentation, medical 3-D): GroupNorm.
- Transformers and RNNs: LayerNorm or RMSNorm, usually pre-norm.
- Style transfer / image generation: InstanceNorm or AdaIN.
- When in doubt for new sequence models: pre-norm RMSNorm.