💬 NLP & Transformers · Lecture 17 of 29

The GPT Family: Autoregressive Language Models from GPT-1 to Today

GPT models are decoder-only transformers trained to predict the next token. We trace GPT-1 through GPT-3's few-shot learning to instruction-tuned assistants, and explain why next-token prediction at scale produces broad capabilities.

While BERT showed the power of bidirectional encoders for understanding, OpenAI pursued a different path: generative pretraining with a decoder-only transformer trained simply to predict the next token. Each generation of the GPT (Generative Pre-trained Transformer) family scaled this recipe up, and the results reshaped the field — culminating in conversational assistants used by hundreds of millions of people.

The objective: next-token prediction#

Given text $x_1, \dots, x_T$, a GPT model maximises

$$ \mathcal{L}(\theta) = \sum_{t=1}^{T}\log P_\theta(x_t \mid x_1, \dots, x_{t-1}) $$

using a stack of transformer blocks with causal (masked) self-attention, so each position attends only to earlier positions. Every position in every training sequence provides a training signal, and no labels are required — any text will do.

GPT-1 (2018): pretrain, then fine-tune#

Radford et al. pretrained a 12-layer, ~117M-parameter decoder on BooksCorpus, then fine-tuned it on downstream tasks with task-specific input formatting (e.g. concatenating premise and hypothesis with delimiters). It improved the state of the art on many benchmarks, establishing generative pretraining as effective for transfer — just months before BERT.

GPT-2 (2019): zero-shot task transfer#

GPT-2 scaled to 1.5 billion parameters trained on WebText (about 40 GB of text from outbound links on Reddit with some engagement). The paper's thesis: a sufficiently large language model trained on diverse text learns to perform tasks without explicit supervision, because the tasks appear naturally in text. Writing "TL;DR:" after an article elicits a summary; "English: … French:" elicits translation. GPT-2's fluent generations led OpenAI to stage its release, citing misuse concerns — an early public debate about responsible release of language models.

GPT-3 (2020): in-context learning#

GPT-3 scaled to 175 billion parameters trained on hundreds of billions of tokens. Its headline finding was in-context learning: without any gradient updates, the model performs a new task from a description and a few examples placed in the prompt.

text
Translate English to French:
sea otter => loutre de mer
cheese => fromage
peppermint =>

Performance improved dramatically with scale and with the number of examples (zero-shot < one-shot < few-shot). Some abilities appeared only in the largest models. GPT-3 also demonstrated limitations: factual errors, inconsistent reasoning, sensitivity to prompt wording, and biases absorbed from web data.

From language model to assistant#

A raw pretrained model continues text; it does not reliably follow instructions or behave helpfully and safely. InstructGPT (Ouyang et al., 2022) added:

  1. Supervised fine-tuning (SFT) on demonstrations of good responses to instructions;
  2. Reinforcement learning from human feedback (RLHF): a reward model trained on human preference comparisons, then policy optimisation (PPO) against it.

Human evaluators preferred outputs of a 1.3B-parameter InstructGPT model over those of the 175B GPT-3 base model — alignment training mattered more than raw size for usefulness. ChatGPT (late 2022) applied this recipe to dialogue and brought LLMs to the mainstream. Later GPT-4-class models added multimodal inputs and much stronger reasoning, alongside many competing model families (Claude, Gemini, LLaMA, Mistral, Qwen, DeepSeek and others). (The Generative AI track covers alignment and modern LLMs in depth.)

Why does next-token prediction yield broad capabilities?#

To predict the next token well across the whole internet, a model benefits from modelling grammar, facts, the structure of arguments, the behaviour of code, the conventions of dialogue, and patterns of reasoning found in text. Prediction is a form of compression, and good compression requires modelling the regularities that generated the data. Scale provides the capacity; diverse data provides the curriculum.

Generating text#

At each step the model outputs a distribution over the vocabulary; a decoding strategy picks the next token (greedy, beam, temperature sampling, top-k, nucleus). Covered in detail in the Generative AI track.

python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

tok = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2").eval()
prompt = "Machine learning can help teachers by"
ids = tok(prompt, return_tensors="pt").input_ids
with torch.no_grad():
    out = model.generate(ids, max_new_tokens=40, do_sample=True, temperature=0.8, top_p=0.9,
                         pad_token_id=tok.eos_token_id)
print(tok.decode(out[0], skip_special_tokens=True))

# Next-token probabilities
with torch.no_grad():
    probs = model(ids).logits[0, -1].softmax(-1)
top = probs.topk(5)
print([(tok.decode(int(i)), round(float(p), 3)) for p, i in zip(top.values, top.indices)])

Encoder vs decoder, revisited#

BERT-style encodersGPT-style decoders
AttentionBidirectionalCausal
ObjectiveMasked token predictionNext-token prediction
StrengthUnderstanding, embeddings, efficient fine-tuningGeneration, in-context learning, general assistants
Typical size in practice100M–1B1B to hundreds of billions

Decoder-only models became the dominant architecture for general-purpose LLMs because generation subsumes many tasks and the objective scales cleanly.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

N-gram Language Models, Smoothing and Perplexity

A language model assigns probabilities to sequences of words. We derive n-gram models from the chain rule and Markov assumption, fix zero probabilities with smoothing, generate text, and evaluate with perplexity.

Intermediate⏱ 5 min#165
💬 NLP & Transformers

BERT: Bidirectional Encoder Representations from Transformers

BERT showed that a bidirectional transformer pretrained on unlabelled text could be fine-tuned to beat task-specific models across NLP. We cover masked language modelling, input format, fine-tuning patterns, and successors such as RoBERTa, DeBERTa and multilingual encoders.

Intermediate⏱ 5 min#177
💬 NLP & Transformers

T5 and BART: Encoder–Decoder Pretraining and Text-to-Text Learning

T5 casts every NLP task as text in, text out; BART pretrains as a denoising autoencoder. We cover span corruption, the text-to-text framework, the lessons of T5's systematic study, and when encoder–decoders are the right choice.

Intermediate⏱ 5 min#179