Large Language Models (LLMs) are the most visible AI technology of our time. They draft emails, explain physics, write code, translate, summarise reports and hold conversations in dozens of languages. This lecture gives you a clear map of how they are built and what they can and cannot do; subsequent lectures examine each stage in depth.
What is an LLM?#
An LLM is (almost always) a decoder-only transformer trained to predict the next token, with billions to hundreds of billions of parameters, trained on trillions of tokens of text (and often code, images and other modalities). Given a sequence of tokens, it outputs a probability distribution over the next token; repeated sampling generates text.
The training pipeline#
| Stage | Data | Objective | Result |
|---|---|---|---|
| 1. Pretraining | Trillions of tokens: web, books, code, papers, multilingual text | Next-token prediction | Base model: broad knowledge and skills, but just continues text |
| 2. Supervised fine-tuning (SFT) | Thousands to millions of instruction–response demonstrations | Next-token prediction on responses | Follows instructions, adopts assistant format |
| 3. Preference optimisation | Human (or AI) comparisons of responses | RLHF, DPO or similar | More helpful, honest, harmless; better style |
| 4. Reasoning / RL training (in recent models) | Problems with verifiable answers (maths, code) | Reinforcement learning on outcomes | Longer, more reliable step-by-step reasoning |
| 5. Safety evaluation and red-teaming | Adversarial prompts, expert testing | — | Identified risks, mitigations, usage policies |
Pretraining dominates compute (often thousands of GPUs for weeks or months); post-training stages are far cheaper but shape behaviour dramatically.
Capabilities#
LLMs can:
- follow instructions and hold multi-turn conversations;
- perform in-context learning from examples in the prompt;
- write and debug code;
- summarise, translate, classify and extract information;
- reason step by step on many problems, especially when prompted or trained to do so;
- use tools — search, calculators, code interpreters, APIs — when given function-calling interfaces;
- process images, audio and documents (multimodal models).
Limitations#
- Hallucination: fluent, confident statements that are false or unsupported.
- Knowledge cut-off: knowledge is frozen at training time unless connected to retrieval or tools.
- Reasoning brittleness: errors on problems requiring long exact computation, careful counting, or novel structures; sensitivity to phrasing.
- Context limits: finite context windows and imperfect use of long contexts.
- Bias and representational harms absorbed from training data.
- Security issues: prompt injection, jailbreaks, data leakage.
- Cost and latency of large models; environmental footprint.
How LLMs are used#
- Prompting — zero-shot or few-shot instructions (prompt engineering).
- Retrieval-augmented generation (RAG) — ground answers in your documents.
- Fine-tuning — adapt to a domain, format or task, often with parameter-efficient methods (LoRA).
- Agents — LLMs that plan and call tools in loops to complete tasks.
- Embedding models — related encoders for search and clustering.
Open and closed models#
- Closed/proprietary models are accessed via APIs; typically strongest capabilities; data handling governed by provider terms.
- Open-weight models can be downloaded and run locally or on private servers; important for privacy, customisation, cost control and research. They vary widely in licence terms.
For organisations handling sensitive personal data — health records, refugee case files — the ability to run models on controlled infrastructure can be decisive.
A first API-style example#
# Using an open model locally with Hugging Face transformers
from transformers import pipeline
chat = pipeline("text-generation", model="Qwen/Qwen2.5-0.5B-Instruct")
messages = [
{"role": "system", "content": "You are a concise tutor for university students."},
{"role": "user", "content": "Explain overfitting in two sentences with an everyday analogy."},
]
out = chat(messages, max_new_tokens=120, do_sample=False)
print(out[0]["generated_text"][-1]["content"])The chat template converts role-tagged messages into the special token format the model was fine-tuned on — using the wrong template degrades performance.