One of the most consequential empirical discoveries in modern AI is that language-model performance improves predictably as we scale up. Loss decreases as a smooth power law in the number of parameters, the amount of training data and the compute used. These scaling laws turned model development into something closer to engineering: organisations could forecast the benefit of a larger training run before spending tens of millions of dollars on it.
Power laws#
Kaplan et al. (OpenAI, 2020) trained many transformer language models across seven orders of magnitude of scale and found that test loss $L$ follows
where $N$ is the number of (non-embedding) parameters, $D$ the number of training tokens, $C$ the compute, and the exponents were small (roughly 0.05–0.1). On a log–log plot these are straight lines. Other architectural details (depth vs width, number of heads) mattered much less than scale, within reasonable ranges.
A useful rule of thumb for transformer training compute:
(about 2 FLOPs per parameter per token for the forward pass and 4 for the backward pass).
Compute-optimal training: Chinchilla#
Given a fixed compute budget, how should we split it between model size and data? Kaplan et al. suggested growing parameters faster than data, and many large models of 2020–2021 (e.g. GPT-3 with 175B parameters trained on about 300B tokens) followed that advice.
Hoffmann et al. (DeepMind, 2022) revisited the question more carefully, fitting
where $E$ is the irreducible loss. Minimising $L$ subject to $C = 6ND$, they found that parameters and tokens should grow roughly in equal proportion — about 20 tokens per parameter for compute-optimal training. Their 70B-parameter Chinchilla, trained on 1.4 trillion tokens with the same compute as the 280B-parameter Gopher, outperformed Gopher on most benchmarks. Many earlier large models had been significantly undertrained.
import numpy as np
def chinchilla_optimal(C, tokens_per_param=20):
"""Compute-optimal N and D for budget C under C = 6 N D and D = k N."""
N = np.sqrt(C / (6 * tokens_per_param))
return N, tokens_per_param * N
for C in [1e21, 1e23, 1e25]:
N, D = chinchilla_optimal(C)
print(f"C={C:.0e} FLOPs -> N ≈ {N / 1e9:,.1f}B params, D ≈ {D / 1e9:,.0f}B tokens")Beyond compute-optimal: inference-aware scaling#
Chinchilla optimises training compute. But a model that is served to millions of users incurs enormous inference cost, which scales with model size. It is often better to train a smaller model on far more data than Chinchilla-optimal (hundreds or thousands of tokens per parameter) — more training compute, but cheaper and faster at inference. Many modern open models (e.g. the LLaMA series) followed this "overtraining" strategy.
Emergent abilities?#
Wei et al. (2022) reported emergent abilities — tasks (e.g. multi-digit arithmetic, some reasoning benchmarks) where performance stays near chance for smaller models and then jumps sharply beyond a certain scale. Schaeffer et al. (2023) argued that many such jumps are artefacts of discontinuous metrics (like exact-match accuracy): with continuous metrics (e.g. per-token probability of the correct answer), improvement is often smooth and predictable. The debate continues, but it highlights a key lesson: how you measure shapes what you conclude about capabilities.
Data constraints#
Scaling laws assume ever more high-quality data. Estimates suggest the stock of high-quality public human-written text is finite and being consumed rapidly. Responses include: training for multiple epochs (with diminishing returns — Muennighoff et al. found up to about four epochs of repeated data behave almost like fresh data), better data filtering and curation, synthetic data generated by models (with care to avoid degradation), multimodal data, and code.
Other scaling dimensions#
- Test-time compute: letting models "think longer" (sampling more reasoning steps, search, self-verification) improves performance on hard problems — a new axis of scaling explored by recent reasoning models.
- Mixture-of-experts: increasing parameters without proportional compute.
- Scaling laws for downstream tasks, fine-tuning, and other modalities (vision, speech, protein models) have also been observed, usually with different exponents.