The architecture of a modern LLM fits on a single page; the data and engineering behind pretraining fill entire teams. Studies repeatedly find that data quality and composition matter as much as model size. This lecture opens the black box of pretraining.
Stage 1: sourcing data#
Typical sources:
- Web crawls (e.g. Common Crawl) — enormous but noisy.
- Curated text: Wikipedia, books, academic papers, news.
- Code from public repositories — improves coding and, reportedly, structured reasoning.
- Multilingual text — often under-represented; deliberate up-sampling improves other languages.
- Mathematical and scientific text, dialogue data, and increasingly synthetic data.
Legal and ethical questions — copyright, consent, personal data, robots.txt opt-outs — are active areas of litigation and policy. Responsible data practice documents sources and respects opt-outs.
Stage 2: extraction and filtering#
- Text extraction from HTML (removing menus, boilerplate, ads).
- Language identification (e.g. fastText classifiers).
- Heuristic quality filters — as in the Gopher and C4 pipelines: minimum length, ratio of alphabetic characters, repeated lines, "lorem ipsum", excessive symbols, very long words, pages with too many bullet points.
- Model-based quality filtering — classifiers trained to recognise high-quality or educational text. FineWeb-Edu, for example, used an LLM-annotated "educational value" classifier, and training on the filtered subset improved knowledge and reasoning benchmarks notably.
- Toxicity and personal-information filtering — remove hateful content and scrub emails, phone numbers and IDs (trade-off: aggressive toxicity filters can remove dialects and discussions by marginalised communities).
- Decontamination — remove text overlapping with evaluation benchmarks.
Stage 3: deduplication#
The web is massively duplicated (templates, mirrors, quoted text). Deduplication improves quality, reduces memorisation of training examples and saves compute (Lee et al., 2022).
- Exact deduplication: hashing documents or substrings (suffix arrays find repeated spans).
- Near-duplicate detection: MinHash with locality-sensitive hashing over character or word n-gram shingles estimates Jaccard similarity efficiently at web scale.
import hashlib, re
def shingles(text, n=5):
words = re.findall(r"\w+", text.lower())
return {" ".join(words[i:i + n]) for i in range(len(words) - n + 1)}
def minhash(sh, num_perm=64):
return [min(int(hashlib.md5(f"{seed}:{s}".encode()).hexdigest(), 16) for s in sh)
for seed in range(num_perm)]
a = "Birth registration is free at the civil registry office for all children born in the country."
b = "Birth registration is free at the civil registry office for all children born in this country."
ma, mb = minhash(shingles(a)), minhash(shingles(b))
print("estimated Jaccard:", sum(x == y for x, y in zip(ma, mb)) / len(ma))Stage 4: data mixture#
The final corpus mixes sources with chosen weights — e.g. web, code, books, papers, multilingual — and may up-sample high-quality sources for several epochs. Mixture weights are tuned with small-scale proxy experiments (methods such as DoReMi optimise them automatically). Many recipes also use a final annealing phase: near the end of training, with a decaying learning rate, emphasise the highest-quality data (maths, code, curated text) for disproportionate gains.
Stage 5: tokenisation#
Train a byte-level BPE or SentencePiece tokeniser on a representative sample, with vocabulary sizes of roughly 32K–256K. Larger vocabularies reduce sequence length (especially for multilingual text) at the cost of a bigger embedding matrix.
The objective and training#
Standard causal language modelling: minimise cross-entropy of each next token. Documents are packed into fixed-length sequences (e.g. 4,096–8,192 tokens) separated by end-of-text tokens, often with attention masks preventing cross-document attention. Some models add fill-in-the-middle training (rearranging a document so the middle is predicted after prefix and suffix) for code infilling. Long-context capability is usually added in a later stage with longer sequences and adjusted positional encodings.
Stability at scale#
Large training runs can suffer loss spikes and divergence. Common measures: pre-norm or RMSNorm, careful initialisation, learning-rate warm-up and decay, gradient clipping, bf16 mixed precision, AdamW with tuned $\beta_2$ and epsilon, QK-normalisation or logit soft-capping, z-loss on output logits, and — when a spike occurs — rewinding to an earlier checkpoint and skipping the offending data batches.
Infrastructure#
Thousands of accelerators with fast interconnects; data, tensor, pipeline and sequence parallelism with sharded optimisers (see the distributed training lecture); high-throughput data loaders; frequent checkpointing; automatic failure detection and restart. Model FLOPs utilisation (MFU) — the fraction of theoretical peak compute actually achieved — is a key efficiency metric; good large runs reach roughly 40–60%.
Monitoring during pretraining#
- Training and validation loss on held-out data by source.
- Downstream evaluations at intermediate checkpoints (few-shot benchmarks) to catch problems early.
- Small proxy models trained on candidate data mixtures to predict large-model behaviour.