✨ Generative AI & LLMs · Lecture 11 of 30

Pretraining LLMs: Data Pipelines, Objectives and Infrastructure

What actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.

The architecture of a modern LLM fits on a single page; the data and engineering behind pretraining fill entire teams. Studies repeatedly find that data quality and composition matter as much as model size. This lecture opens the black box of pretraining.

Stage 1: sourcing data#

Typical sources:

  • Web crawls (e.g. Common Crawl) — enormous but noisy.
  • Curated text: Wikipedia, books, academic papers, news.
  • Code from public repositories — improves coding and, reportedly, structured reasoning.
  • Multilingual text — often under-represented; deliberate up-sampling improves other languages.
  • Mathematical and scientific text, dialogue data, and increasingly synthetic data.

Legal and ethical questions — copyright, consent, personal data, robots.txt opt-outs — are active areas of litigation and policy. Responsible data practice documents sources and respects opt-outs.

Stage 2: extraction and filtering#

  1. Text extraction from HTML (removing menus, boilerplate, ads).
  2. Language identification (e.g. fastText classifiers).
  3. Heuristic quality filters — as in the Gopher and C4 pipelines: minimum length, ratio of alphabetic characters, repeated lines, "lorem ipsum", excessive symbols, very long words, pages with too many bullet points.
  4. Model-based quality filtering — classifiers trained to recognise high-quality or educational text. FineWeb-Edu, for example, used an LLM-annotated "educational value" classifier, and training on the filtered subset improved knowledge and reasoning benchmarks notably.
  5. Toxicity and personal-information filtering — remove hateful content and scrub emails, phone numbers and IDs (trade-off: aggressive toxicity filters can remove dialects and discussions by marginalised communities).
  6. Decontamination — remove text overlapping with evaluation benchmarks.

Stage 3: deduplication#

The web is massively duplicated (templates, mirrors, quoted text). Deduplication improves quality, reduces memorisation of training examples and saves compute (Lee et al., 2022).

  • Exact deduplication: hashing documents or substrings (suffix arrays find repeated spans).
  • Near-duplicate detection: MinHash with locality-sensitive hashing over character or word n-gram shingles estimates Jaccard similarity efficiently at web scale.
python
import hashlib, re

def shingles(text, n=5):
    words = re.findall(r"\w+", text.lower())
    return {" ".join(words[i:i + n]) for i in range(len(words) - n + 1)}

def minhash(sh, num_perm=64):
    return [min(int(hashlib.md5(f"{seed}:{s}".encode()).hexdigest(), 16) for s in sh)
            for seed in range(num_perm)]

a = "Birth registration is free at the civil registry office for all children born in the country."
b = "Birth registration is free at the civil registry office for all children born in this country."
ma, mb = minhash(shingles(a)), minhash(shingles(b))
print("estimated Jaccard:", sum(x == y for x, y in zip(ma, mb)) / len(ma))

Stage 4: data mixture#

The final corpus mixes sources with chosen weights — e.g. web, code, books, papers, multilingual — and may up-sample high-quality sources for several epochs. Mixture weights are tuned with small-scale proxy experiments (methods such as DoReMi optimise them automatically). Many recipes also use a final annealing phase: near the end of training, with a decaying learning rate, emphasise the highest-quality data (maths, code, curated text) for disproportionate gains.

Stage 5: tokenisation#

Train a byte-level BPE or SentencePiece tokeniser on a representative sample, with vocabulary sizes of roughly 32K–256K. Larger vocabularies reduce sequence length (especially for multilingual text) at the cost of a bigger embedding matrix.

The objective and training#

Standard causal language modelling: minimise cross-entropy of each next token. Documents are packed into fixed-length sequences (e.g. 4,096–8,192 tokens) separated by end-of-text tokens, often with attention masks preventing cross-document attention. Some models add fill-in-the-middle training (rearranging a document so the middle is predicted after prefix and suffix) for code infilling. Long-context capability is usually added in a later stage with longer sequences and adjusted positional encodings.

Stability at scale#

Large training runs can suffer loss spikes and divergence. Common measures: pre-norm or RMSNorm, careful initialisation, learning-rate warm-up and decay, gradient clipping, bf16 mixed precision, AdamW with tuned $\beta_2$ and epsilon, QK-normalisation or logit soft-capping, z-loss on output logits, and — when a spike occurs — rewinding to an earlier checkpoint and skipping the offending data batches.

Infrastructure#

Thousands of accelerators with fast interconnects; data, tensor, pipeline and sequence parallelism with sharded optimisers (see the distributed training lecture); high-throughput data loaders; frequent checkpointing; automatic failure detection and restart. Model FLOPs utilisation (MFU) — the fraction of theoretical peak compute actually achieved — is a key efficiency metric; good large runs reach roughly 40–60%.

Monitoring during pretraining#

  • Training and validation loss on held-out data by source.
  • Downstream evaluations at intermediate checkpoints (few-shot benchmarks) to catch problems early.
  • Small proxy models trained on candidate data mixtures to predict large-model behaviour.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199
✨ Generative AI & LLMs

Scaling Laws: How Performance Grows with Compute, Data and Parameters

Language-model loss follows smooth power laws in model size, data and compute. We examine the Kaplan and Chinchilla scaling laws, compute-optimal training, the debate on emergent abilities, and what scaling means for the field.

Advanced⏱ 5 min#200
✨ Generative AI & LLMs

Instruction Tuning: Teaching Language Models to Follow Directions

A base model continues text; an instruction-tuned model answers requests. We cover supervised fine-tuning data (human-written, templated and synthetic), chat formats, loss masking, what instruction tuning changes, and practical recipes.

Intermediate⏱ 5 min#202