While BERT showed the power of bidirectional encoders for understanding, OpenAI pursued a different path: generative pretraining with a decoder-only transformer trained simply to predict the next token. Each generation of the GPT (Generative Pre-trained Transformer) family scaled this recipe up, and the results reshaped the field — culminating in conversational assistants used by hundreds of millions of people.
The objective: next-token prediction#
Given text $x_1, \dots, x_T$, a GPT model maximises
using a stack of transformer blocks with causal (masked) self-attention, so each position attends only to earlier positions. Every position in every training sequence provides a training signal, and no labels are required — any text will do.
GPT-1 (2018): pretrain, then fine-tune#
Radford et al. pretrained a 12-layer, ~117M-parameter decoder on BooksCorpus, then fine-tuned it on downstream tasks with task-specific input formatting (e.g. concatenating premise and hypothesis with delimiters). It improved the state of the art on many benchmarks, establishing generative pretraining as effective for transfer — just months before BERT.
GPT-2 (2019): zero-shot task transfer#
GPT-2 scaled to 1.5 billion parameters trained on WebText (about 40 GB of text from outbound links on Reddit with some engagement). The paper's thesis: a sufficiently large language model trained on diverse text learns to perform tasks without explicit supervision, because the tasks appear naturally in text. Writing "TL;DR:" after an article elicits a summary; "English: … French:" elicits translation. GPT-2's fluent generations led OpenAI to stage its release, citing misuse concerns — an early public debate about responsible release of language models.
GPT-3 (2020): in-context learning#
GPT-3 scaled to 175 billion parameters trained on hundreds of billions of tokens. Its headline finding was in-context learning: without any gradient updates, the model performs a new task from a description and a few examples placed in the prompt.
Translate English to French:
sea otter => loutre de mer
cheese => fromage
peppermint =>Performance improved dramatically with scale and with the number of examples (zero-shot < one-shot < few-shot). Some abilities appeared only in the largest models. GPT-3 also demonstrated limitations: factual errors, inconsistent reasoning, sensitivity to prompt wording, and biases absorbed from web data.
From language model to assistant#
A raw pretrained model continues text; it does not reliably follow instructions or behave helpfully and safely. InstructGPT (Ouyang et al., 2022) added:
- Supervised fine-tuning (SFT) on demonstrations of good responses to instructions;
- Reinforcement learning from human feedback (RLHF): a reward model trained on human preference comparisons, then policy optimisation (PPO) against it.
Human evaluators preferred outputs of a 1.3B-parameter InstructGPT model over those of the 175B GPT-3 base model — alignment training mattered more than raw size for usefulness. ChatGPT (late 2022) applied this recipe to dialogue and brought LLMs to the mainstream. Later GPT-4-class models added multimodal inputs and much stronger reasoning, alongside many competing model families (Claude, Gemini, LLaMA, Mistral, Qwen, DeepSeek and others). (The Generative AI track covers alignment and modern LLMs in depth.)
Why does next-token prediction yield broad capabilities?#
To predict the next token well across the whole internet, a model benefits from modelling grammar, facts, the structure of arguments, the behaviour of code, the conventions of dialogue, and patterns of reasoning found in text. Prediction is a form of compression, and good compression requires modelling the regularities that generated the data. Scale provides the capacity; diverse data provides the curriculum.
Generating text#
At each step the model outputs a distribution over the vocabulary; a decoding strategy picks the next token (greedy, beam, temperature sampling, top-k, nucleus). Covered in detail in the Generative AI track.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tok = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2").eval()
prompt = "Machine learning can help teachers by"
ids = tok(prompt, return_tensors="pt").input_ids
with torch.no_grad():
out = model.generate(ids, max_new_tokens=40, do_sample=True, temperature=0.8, top_p=0.9,
pad_token_id=tok.eos_token_id)
print(tok.decode(out[0], skip_special_tokens=True))
# Next-token probabilities
with torch.no_grad():
probs = model(ids).logits[0, -1].softmax(-1)
top = probs.topk(5)
print([(tok.decode(int(i)), round(float(p), 3)) for p, i in zip(top.values, top.indices)])Encoder vs decoder, revisited#
| BERT-style encoders | GPT-style decoders | |
|---|---|---|
| Attention | Bidirectional | Causal |
| Objective | Masked token prediction | Next-token prediction |
| Strength | Understanding, embeddings, efficient fine-tuning | Generation, in-context learning, general assistants |
| Typical size in practice | 100M–1B | 1B to hundreds of billions |
Decoder-only models became the dominant architecture for general-purpose LLMs because generation subsumes many tasks and the objective scales cleanly.