In June 2017, eight researchers at Google published "Attention Is All You Need". Their Transformer dispensed with recurrence entirely, relying on attention to relate every position in a sequence to every other. It trained faster than RNNs, achieved state-of-the-art translation, and — within a few years — became the architecture behind BERT, GPT, T5, Vision Transformers, speech models, protein models and today's large language models. This is the most important architecture in modern AI, so we will study it carefully.
The big picture#
The original Transformer is an encoder–decoder model for translation:
- The encoder maps the source tokens to contextual representations using a stack of $N = 6$ identical layers.
- The decoder generates the target one token at a time using a stack of $N = 6$ layers that attend both to previously generated tokens and to the encoder output.
Base model dimensions: $d_{\text{model}} = 512$, 8 attention heads, feed-forward inner size 2048.
Step 1: token embeddings#
Tokens (subwords) are mapped to vectors of dimension $d_{\text{model}}$ by a learned embedding table (scaled by $\sqrt{d_{\text{model}}}$ in the original paper).
Step 2: positional encoding#
Self-attention treats its input as a set — it has no notion of order. The Transformer adds a positional encoding to each embedding. The original used fixed sinusoids:
Different dimensions oscillate at different frequencies, giving each position a unique signature, and relative offsets correspond to linear transformations. (Modern models often use learned or rotary encodings — see the positional-encoding lecture.)
Step 3: scaled dot-product attention#
Each token's vector is projected into a query $\mathbf{q}$, key $\mathbf{k}$ and value $\mathbf{v}$. For all tokens at once, with matrices $\mathbf{Q}, \mathbf{K}, \mathbf{V}$:
Every token computes a weighted average of all tokens' values, with weights determined by query–key similarity. "The animal didn't cross the street because it was too tired": the representation of "it" can draw heavily on "animal".
Step 4: multi-head attention#
One attention pattern is limiting. Multi-head attention runs $h$ attention operations in parallel, each with its own learned projections into a smaller subspace ($d_k = d_{\text{model}}/h$), then concatenates and projects the results:
Different heads can specialise — some track syntax, some coreference, some adjacent tokens.
Step 5: position-wise feed-forward network#
After attention, each position passes independently through the same two-layer MLP:
Attention mixes information across positions; the FFN processes each position. The FFN holds about two-thirds of a layer's parameters.
Step 6: residual connections and layer normalisation#
Each sub-layer (attention, FFN) is wrapped with a residual connection and LayerNorm. The original used post-norm, $\text{LN}(\mathbf{x} + \text{Sublayer}(\mathbf{x}))$; most modern models use pre-norm, $\mathbf{x} + \text{Sublayer}(\text{LN}(\mathbf{x}))$, which trains more stably.
Step 7: the decoder and masking#
Each decoder layer has three sub-layers:
- Masked self-attention over the target prefix. A causal mask sets scores for future positions to $-\infty$ before the softmax, so position $t$ cannot peek at tokens $> t$ — preserving the autoregressive property while allowing all positions to be trained in parallel.
- Cross-attention: queries come from the decoder; keys and values come from the encoder output — this is how the decoder "looks at" the source sentence.
- Feed-forward network.
A final linear layer and softmax produce next-token probabilities. Padding masks also prevent attention to padding tokens in batches of different-length sentences.
A compact implementation#
import math, torch
import torch.nn as nn
import torch.nn.functional as F
class MultiHeadAttention(nn.Module):
def __init__(self, d, h, dropout=0.1):
super().__init__()
self.h, self.dk = h, d // h
self.qkv, self.out, self.drop = nn.Linear(d, 3 * d), nn.Linear(d, d), nn.Dropout(dropout)
def forward(self, x, mask=None): # x: (B, T, d)
B, T, d = x.shape
q, k, v = self.qkv(x).view(B, T, 3, self.h, self.dk).permute(2, 0, 3, 1, 4) # each (B, h, T, dk)
scores = q @ k.transpose(-2, -1) / math.sqrt(self.dk)
if mask is not None:
scores = scores.masked_fill(mask == 0, float("-inf"))
att = self.drop(F.softmax(scores, dim=-1))
return self.out((att @ v).transpose(1, 2).reshape(B, T, d))
class Block(nn.Module): # pre-norm transformer block
def __init__(self, d=256, h=8, ff=1024, dropout=0.1):
super().__init__()
self.ln1, self.ln2 = nn.LayerNorm(d), nn.LayerNorm(d)
self.attn = MultiHeadAttention(d, h, dropout)
self.ffn = nn.Sequential(nn.Linear(d, ff), nn.GELU(), nn.Linear(ff, d), nn.Dropout(dropout))
def forward(self, x, mask=None):
x = x + self.attn(self.ln1(x), mask)
return x + self.ffn(self.ln2(x))
T = 10
causal = torch.tril(torch.ones(T, T)).view(1, 1, T, T) # lower-triangular mask
x = torch.randn(2, T, 256)
print(Block()(x, causal).shape) # (2, 10, 256)Why the Transformer won#
- Parallelism: all positions are processed simultaneously during training — ideal for GPUs — unlike sequential RNNs.
- Short paths: any two positions interact in one layer, making long-range dependencies easier to learn.
- Scalability: performance keeps improving with more data, parameters and compute (scaling laws).
- Generality: the same block works for text, images (patches), audio, proteins and more.
The main cost: self-attention is $O(T^2)$ in sequence length in time and memory — addressed by efficient attention methods covered later.
Three families#
| Family | Architecture | Pretraining | Examples | Best for |
|---|---|---|---|---|
| Encoder-only | Bidirectional encoder | Masked language modelling | BERT, RoBERTa | Classification, extraction, embeddings |
| Decoder-only | Causal decoder | Next-token prediction | GPT family, LLaMA | Generation, general-purpose LLMs |
| Encoder–decoder | Both + cross-attention | Denoising / span corruption | T5, BART, original Transformer | Translation, summarisation |