Different parameters see very different gradients. The embedding of a rare word receives a gradient once in a thousand batches; a bias in the output layer receives one every step. A single global learning rate is too large for some parameters and too small for others. Adaptive optimisers scale each parameter's step by the history of its own gradients. Adam, the most famous of them, is probably the most widely used optimiser in deep learning.
AdaGrad#
Duchi, Hazan and Singer (2011) accumulate the sum of squared gradients per parameter:
(all operations element-wise). Parameters with large past gradients get smaller steps; rarely updated parameters get larger ones โ excellent for sparse features. Its flaw: $\mathbf{s}_t$ only grows, so the effective learning rate decays towards zero and training stalls in long non-convex runs.
RMSProp#
Hinton (in a 2012 lecture) replaced the sum with an exponential moving average, so old gradients are forgotten:
with $\rho \approx 0.9$โ$0.99$. Dividing by the root-mean-square gradient normalises step sizes across parameters.
Adam#
Kingma and Ba (2014) combined momentum (first moment) with RMSProp (second moment):
Defaults: $\beta_1 = 0.9$, $\beta_2 = 0.999$, $\epsilon = 10^{-8}$, $\eta = 10^{-3}$ (much smaller for large models).
Why bias correction?#
$\mathbf{m}_0 = \mathbf{v}_0 = \mathbf{0}$, so early averages are biased towards zero. At step 1, $\mathbf{v}_1 = 0.001\,\mathbf{g}_1^2$ โ a thousand times too small. Dividing by $1 - \beta^t$ corrects this exactly in expectation (if gradients were stationary). Without it, the first steps would be enormous because the denominator is tiny.
Properties#
- The step size per parameter is roughly bounded by $\eta$, since $|\hat{m}|/\sqrt{\hat{v}} \lesssim 1$ โ making Adam robust to gradient scale.
- It works well out of the box across many architectures, especially transformers, RNNs and GANs.
- It stores two extra values per parameter (memory = 3ร parameters, excluding gradients).
AdamW: decoupled weight decay#
With plain SGD, L2 regularisation (adding $\frac{\lambda}{2}\|\boldsymbol{\theta}\|^2$ to the loss) and weight decay (shrinking weights each step, $\boldsymbol{\theta} \leftarrow \boldsymbol{\theta} - \eta\lambda\boldsymbol{\theta}$) are equivalent. With Adam they are not: an L2 gradient term gets divided by $\sqrt{\hat{\mathbf{v}}}$, so parameters with large gradients are barely regularised. Loshchilov and Hutter (2017) proposed AdamW, which applies weight decay directly:
AdamW generalises better and is the standard optimiser for training transformers. Typical weight decay: 0.01โ0.1, usually not applied to biases and normalisation parameters.
import torch
model = torch.nn.Sequential(torch.nn.Linear(128, 256), torch.nn.GELU(),
torch.nn.LayerNorm(256), torch.nn.Linear(256, 10))
decay, no_decay = [], []
for name, p in model.named_parameters():
(no_decay if p.ndim == 1 else decay).append(p) # biases & norm weights are 1-D
opt = torch.optim.AdamW([{"params": decay, "weight_decay": 0.05},
{"params": no_decay, "weight_decay": 0.0}],
lr=3e-4, betas=(0.9, 0.999), eps=1e-8)Adam from scratch#
import numpy as np
def adam_step(theta, g, state, lr=1e-3, b1=0.9, b2=0.999, eps=1e-8, wd=0.0):
state["t"] += 1
state["m"] = b1 * state["m"] + (1 - b1) * g
state["v"] = b2 * state["v"] + (1 - b2) * g**2
m_hat = state["m"] / (1 - b1 ** state["t"])
v_hat = state["v"] / (1 - b2 ** state["t"])
return theta - lr * (m_hat / (np.sqrt(v_hat) + eps) + wd * theta)
# Minimise a badly scaled quadratic: 0.5*(x^2 + 1000*y^2)
theta = np.array([5.0, 5.0]); state = {"t": 0, "m": np.zeros(2), "v": np.zeros(2)}
for _ in range(2000):
g = np.array([1.0, 1000.0]) * theta
theta = adam_step(theta, g, state, lr=0.05)
print(theta.round(5))Adam handles the 1000ร difference in curvature gracefully because it normalises each coordinate's step.
Known issues and newer optimisers#
- Generalisation gap: on some vision tasks, Adam-trained models generalised slightly worse than SGD + momentum; AdamW and good schedules narrow this.
- Instability early in training for large models โ mitigated by learning-rate warm-up (next lecture).
- AMSGrad fixes a theoretical non-convergence example by keeping the maximum of past $\mathbf{v}_t$.
- LAMB/LARS add layer-wise trust ratios for very large batch training.
- Adafactor factorises the second-moment matrix to save memory.
- Lion (found by automated search) uses only the sign of a momentum-like update โ memory-efficient.
- Second-order-inspired methods such as Shampoo and SOAP precondition with approximate curvature matrices and have shown strong results in recent large-scale training.
| Optimiser | State per parameter | Typical use |
|---|---|---|
| SGD + momentum | 1 | CNNs for vision, well-tuned schedules |
| Adam / AdamW | 2 | Transformers, most new projects |
| Adafactor | < 1 (factored) | Memory-constrained large models |
| Lion | 1 | Memory-efficient alternative to AdamW |