✨ Generative AI & LLMs · Lecture 10 of 30

Scaling Laws: How Performance Grows with Compute, Data and Parameters

Language-model loss follows smooth power laws in model size, data and compute. We examine the Kaplan and Chinchilla scaling laws, compute-optimal training, the debate on emergent abilities, and what scaling means for the field.

One of the most consequential empirical discoveries in modern AI is that language-model performance improves predictably as we scale up. Loss decreases as a smooth power law in the number of parameters, the amount of training data and the compute used. These scaling laws turned model development into something closer to engineering: organisations could forecast the benefit of a larger training run before spending tens of millions of dollars on it.

Power laws#

Kaplan et al. (OpenAI, 2020) trained many transformer language models across seven orders of magnitude of scale and found that test loss $L$ follows

$$ L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}, \qquad L(D) \approx \left(\frac{D_c}{D}\right)^{\alpha_D}, \qquad L(C) \approx \left(\frac{C_c}{C}\right)^{\alpha_C} $$

where $N$ is the number of (non-embedding) parameters, $D$ the number of training tokens, $C$ the compute, and the exponents were small (roughly 0.05–0.1). On a log–log plot these are straight lines. Other architectural details (depth vs width, number of heads) mattered much less than scale, within reasonable ranges.

A useful rule of thumb for transformer training compute:

$$ C \approx 6ND \quad \text{FLOPs} $$

(about 2 FLOPs per parameter per token for the forward pass and 4 for the backward pass).

Compute-optimal training: Chinchilla#

Given a fixed compute budget, how should we split it between model size and data? Kaplan et al. suggested growing parameters faster than data, and many large models of 2020–2021 (e.g. GPT-3 with 175B parameters trained on about 300B tokens) followed that advice.

Hoffmann et al. (DeepMind, 2022) revisited the question more carefully, fitting

$$ L(N, D) = E + \frac{A}{N^\alpha} + \frac{B}{D^\beta} $$

where $E$ is the irreducible loss. Minimising $L$ subject to $C = 6ND$, they found that parameters and tokens should grow roughly in equal proportion — about 20 tokens per parameter for compute-optimal training. Their 70B-parameter Chinchilla, trained on 1.4 trillion tokens with the same compute as the 280B-parameter Gopher, outperformed Gopher on most benchmarks. Many earlier large models had been significantly undertrained.

python
import numpy as np

def chinchilla_optimal(C, tokens_per_param=20):
    """Compute-optimal N and D for budget C under C = 6 N D and D = k N."""
    N = np.sqrt(C / (6 * tokens_per_param))
    return N, tokens_per_param * N

for C in [1e21, 1e23, 1e25]:
    N, D = chinchilla_optimal(C)
    print(f"C={C:.0e} FLOPs -> N ≈ {N / 1e9:,.1f}B params, D ≈ {D / 1e9:,.0f}B tokens")

Beyond compute-optimal: inference-aware scaling#

Chinchilla optimises training compute. But a model that is served to millions of users incurs enormous inference cost, which scales with model size. It is often better to train a smaller model on far more data than Chinchilla-optimal (hundreds or thousands of tokens per parameter) — more training compute, but cheaper and faster at inference. Many modern open models (e.g. the LLaMA series) followed this "overtraining" strategy.

Emergent abilities?#

Wei et al. (2022) reported emergent abilities — tasks (e.g. multi-digit arithmetic, some reasoning benchmarks) where performance stays near chance for smaller models and then jumps sharply beyond a certain scale. Schaeffer et al. (2023) argued that many such jumps are artefacts of discontinuous metrics (like exact-match accuracy): with continuous metrics (e.g. per-token probability of the correct answer), improvement is often smooth and predictable. The debate continues, but it highlights a key lesson: how you measure shapes what you conclude about capabilities.

Data constraints#

Scaling laws assume ever more high-quality data. Estimates suggest the stock of high-quality public human-written text is finite and being consumed rapidly. Responses include: training for multiple epochs (with diminishing returns — Muennighoff et al. found up to about four epochs of repeated data behave almost like fresh data), better data filtering and curation, synthetic data generated by models (with care to avoid degradation), multimodal data, and code.

Other scaling dimensions#

  • Test-time compute: letting models "think longer" (sampling more reasoning steps, search, self-verification) improves performance on hard problems — a new axis of scaling explored by recent reasoning models.
  • Mixture-of-experts: increasing parameters without proportional compute.
  • Scaling laws for downstream tasks, fine-tuning, and other modalities (vision, speech, protein models) have also been observed, usually with different exponents.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199
✨ Generative AI & LLMs

Pretraining LLMs: Data Pipelines, Objectives and Infrastructure

What actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.

Advanced⏱ 5 min#201
✨ Generative AI & LLMs

Guidance in Diffusion Models: Classifier and Classifier-Free Guidance

Conditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.

Advanced⏱ 5 min#198