๐Ÿ”— Deep Learning ยท Lecture 9 of 38

Optimisers II: AdaGrad, RMSProp, Adam and AdamW

Adaptive optimisers give each parameter its own learning rate. We derive AdaGrad, RMSProp and Adam including bias correction, explain why AdamW decouples weight decay, and survey newer optimisers.

Different parameters see very different gradients. The embedding of a rare word receives a gradient once in a thousand batches; a bias in the output layer receives one every step. A single global learning rate is too large for some parameters and too small for others. Adaptive optimisers scale each parameter's step by the history of its own gradients. Adam, the most famous of them, is probably the most widely used optimiser in deep learning.

AdaGrad#

Duchi, Hazan and Singer (2011) accumulate the sum of squared gradients per parameter:

$$ \mathbf{s}_t = \mathbf{s}_{t-1} + \mathbf{g}_t^2, \qquad \boldsymbol{\theta}_{t+1} = \boldsymbol{\theta}_t - \frac{\eta}{\sqrt{\mathbf{s}_t} + \epsilon}\odot\mathbf{g}_t $$

(all operations element-wise). Parameters with large past gradients get smaller steps; rarely updated parameters get larger ones โ€” excellent for sparse features. Its flaw: $\mathbf{s}_t$ only grows, so the effective learning rate decays towards zero and training stalls in long non-convex runs.

RMSProp#

Hinton (in a 2012 lecture) replaced the sum with an exponential moving average, so old gradients are forgotten:

$$ \mathbf{s}_t = \rho\,\mathbf{s}_{t-1} + (1 - \rho)\,\mathbf{g}_t^2, \qquad \boldsymbol{\theta}_{t+1} = \boldsymbol{\theta}_t - \frac{\eta}{\sqrt{\mathbf{s}_t} + \epsilon}\odot\mathbf{g}_t $$

with $\rho \approx 0.9$โ€“$0.99$. Dividing by the root-mean-square gradient normalises step sizes across parameters.

Adam#

Kingma and Ba (2014) combined momentum (first moment) with RMSProp (second moment):

$$ \begin{aligned} \mathbf{m}_t &= \beta_1\mathbf{m}_{t-1} + (1 - \beta_1)\,\mathbf{g}_t \\ \mathbf{v}_t &= \beta_2\mathbf{v}_{t-1} + (1 - \beta_2)\,\mathbf{g}_t^2 \\ \hat{\mathbf{m}}_t &= \frac{\mathbf{m}_t}{1 - \beta_1^t}, \qquad \hat{\mathbf{v}}_t = \frac{\mathbf{v}_t}{1 - \beta_2^t} \\ \boldsymbol{\theta}_{t+1} &= \boldsymbol{\theta}_t - \eta\,\frac{\hat{\mathbf{m}}_t}{\sqrt{\hat{\mathbf{v}}_t} + \epsilon} \end{aligned} $$

Defaults: $\beta_1 = 0.9$, $\beta_2 = 0.999$, $\epsilon = 10^{-8}$, $\eta = 10^{-3}$ (much smaller for large models).

Why bias correction?#

$\mathbf{m}_0 = \mathbf{v}_0 = \mathbf{0}$, so early averages are biased towards zero. At step 1, $\mathbf{v}_1 = 0.001\,\mathbf{g}_1^2$ โ€” a thousand times too small. Dividing by $1 - \beta^t$ corrects this exactly in expectation (if gradients were stationary). Without it, the first steps would be enormous because the denominator is tiny.

Properties#

  • The step size per parameter is roughly bounded by $\eta$, since $|\hat{m}|/\sqrt{\hat{v}} \lesssim 1$ โ€” making Adam robust to gradient scale.
  • It works well out of the box across many architectures, especially transformers, RNNs and GANs.
  • It stores two extra values per parameter (memory = 3ร— parameters, excluding gradients).

AdamW: decoupled weight decay#

With plain SGD, L2 regularisation (adding $\frac{\lambda}{2}\|\boldsymbol{\theta}\|^2$ to the loss) and weight decay (shrinking weights each step, $\boldsymbol{\theta} \leftarrow \boldsymbol{\theta} - \eta\lambda\boldsymbol{\theta}$) are equivalent. With Adam they are not: an L2 gradient term gets divided by $\sqrt{\hat{\mathbf{v}}}$, so parameters with large gradients are barely regularised. Loshchilov and Hutter (2017) proposed AdamW, which applies weight decay directly:

$$ \boldsymbol{\theta}_{t+1} = \boldsymbol{\theta}_t - \eta\left(\frac{\hat{\mathbf{m}}_t}{\sqrt{\hat{\mathbf{v}}_t} + \epsilon} + \lambda\,\boldsymbol{\theta}_t\right) $$

AdamW generalises better and is the standard optimiser for training transformers. Typical weight decay: 0.01โ€“0.1, usually not applied to biases and normalisation parameters.

python
import torch

model = torch.nn.Sequential(torch.nn.Linear(128, 256), torch.nn.GELU(),
                            torch.nn.LayerNorm(256), torch.nn.Linear(256, 10))
decay, no_decay = [], []
for name, p in model.named_parameters():
    (no_decay if p.ndim == 1 else decay).append(p)       # biases & norm weights are 1-D
opt = torch.optim.AdamW([{"params": decay, "weight_decay": 0.05},
                         {"params": no_decay, "weight_decay": 0.0}],
                        lr=3e-4, betas=(0.9, 0.999), eps=1e-8)

Adam from scratch#

python
import numpy as np

def adam_step(theta, g, state, lr=1e-3, b1=0.9, b2=0.999, eps=1e-8, wd=0.0):
    state["t"] += 1
    state["m"] = b1 * state["m"] + (1 - b1) * g
    state["v"] = b2 * state["v"] + (1 - b2) * g**2
    m_hat = state["m"] / (1 - b1 ** state["t"])
    v_hat = state["v"] / (1 - b2 ** state["t"])
    return theta - lr * (m_hat / (np.sqrt(v_hat) + eps) + wd * theta)

# Minimise a badly scaled quadratic: 0.5*(x^2 + 1000*y^2)
theta = np.array([5.0, 5.0]); state = {"t": 0, "m": np.zeros(2), "v": np.zeros(2)}
for _ in range(2000):
    g = np.array([1.0, 1000.0]) * theta
    theta = adam_step(theta, g, state, lr=0.05)
print(theta.round(5))

Adam handles the 1000ร— difference in curvature gracefully because it normalises each coordinate's step.

Known issues and newer optimisers#

  • Generalisation gap: on some vision tasks, Adam-trained models generalised slightly worse than SGD + momentum; AdamW and good schedules narrow this.
  • Instability early in training for large models โ€” mitigated by learning-rate warm-up (next lecture).
  • AMSGrad fixes a theoretical non-convergence example by keeping the maximum of past $\mathbf{v}_t$.
  • LAMB/LARS add layer-wise trust ratios for very large batch training.
  • Adafactor factorises the second-moment matrix to save memory.
  • Lion (found by automated search) uses only the sign of a momentum-like update โ€” memory-efficient.
  • Second-order-inspired methods such as Shampoo and SOAP precondition with approximate curvature matrices and have shown strong results in recent large-scale training.
OptimiserState per parameterTypical use
SGD + momentum1CNNs for vision, well-tuned schedules
Adam / AdamW2Transformers, most new projects
Adafactor< 1 (factored)Memory-constrained large models
Lion1Memory-efficient alternative to AdamW
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Optimisers I: SGD, Momentum and Nesterov Acceleration

Plain SGD zig-zags through ravines and crawls across plateaus. Momentum accumulates velocity to fix both. We derive heavy-ball and Nesterov momentum, analyse their effect on ill-conditioned problems, and give tuning advice.

Intermediateโฑ 5 min#104
๐Ÿ”— Deep Learning

Loss Landscapes, Saddle Points and Flat Minima

What does the surface that SGD descends actually look like? We study critical points in high dimensions, visualise loss landscapes, discuss sharp versus flat minima, mode connectivity and why architecture shapes trainability.

Advancedโฑ 6 min#134
๐Ÿ”— Deep Learning

Learning Rate Schedules, Warm-up and the LR Range Test

The learning rate is the most important hyperparameter, and it should change during training. We compare step, exponential, cosine and one-cycle schedules, explain why warm-up stabilises large models, and find good rates quickly.

Intermediateโฑ 5 min#106