Self-attention is permutation-equivariant: shuffle the input tokens and the outputs are shuffled the same way, with no other change. Without positional information, "dog bites man" and "man bites dog" would look identical to a transformer. How we inject position has a large effect on how well models handle long contexts and whether they can extrapolate to sequences longer than those seen in training.
Absolute positional encodings#
Sinusoidal (original Transformer)#
added to token embeddings. Each pair of dimensions is a rotating 2-D vector with a different frequency — like the hands of a clock running at many speeds. Key property: $PE_{p+k}$ is a linear function (a rotation) of $PE_p$, which in principle lets the model reason about relative offsets. No parameters, and defined for any position.
Learned absolute embeddings#
A trainable vector per position (BERT, GPT-2). Flexible, but undefined beyond the maximum trained length (e.g. 512 or 1024 positions) — the model cannot extrapolate.
Relative position methods#
What usually matters linguistically is relative distance ("the word two positions back"), not absolute index. Shaw et al. (2018) added learned relative-position terms to attention scores; T5 uses a learned scalar bias per relative-distance bucket added to attention logits.
Rotary Position Embedding (RoPE)#
RoPE (Su et al., 2021) is used by LLaMA, Mistral, Qwen, and many other modern LLMs. Instead of adding a vector to embeddings, it rotates query and key vectors by an angle proportional to their position. Split a $d$-dimensional vector into $d/2$ pairs; rotate pair $i$ at position $p$ by angle $p\theta_i$ with $\theta_i = 10000^{-2i/d}$:
Because rotations compose, the dot product between a query at position $m$ and a key at position $n$ depends only on their relative offset $m - n$:
RoPE thus encodes absolute position in the representation while making attention relative — with no extra parameters, compatible with efficient attention kernels and the KV cache.
import torch
def rope(x, base=10000.0):
"""x: (..., T, d) with even d. Rotates pairs of dimensions by position-dependent angles."""
T, d = x.shape[-2], x.shape[-1]
theta = base ** (-torch.arange(0, d, 2, dtype=torch.float32) / d) # (d/2,)
angles = torch.arange(T, dtype=torch.float32)[:, None] * theta[None] # (T, d/2)
cos, sin = angles.cos(), angles.sin()
x1, x2 = x[..., 0::2], x[..., 1::2]
out = torch.empty_like(x)
out[..., 0::2] = x1 * cos - x2 * sin
out[..., 1::2] = x1 * sin + x2 * cos
return out
# Check the relative property: score depends only on the offset
q, k = torch.randn(1, 64), torch.randn(1, 64)
seq_q = rope(q.repeat(50, 1)); seq_k = rope(k.repeat(50, 1))
print((seq_q[10] @ seq_k[7]).item(), (seq_q[30] @ seq_k[27]).item()) # equal: both offset 3ALiBi: attention with linear biases#
ALiBi (Press, Smith & Lewis, 2022) uses no positional embeddings at all. It adds a fixed, head-specific linear penalty to attention scores proportional to distance:
with slopes $m_h$ forming a geometric sequence across heads (e.g. $1/2, 1/4, \dots$). Distant tokens are penalised more; different heads attend at different ranges. ALiBi showed strong length extrapolation: models trained on short sequences performed well on longer ones.
Comparison#
| Method | Type | Parameters | Extrapolation | Used in |
|---|---|---|---|---|
| Sinusoidal | Absolute, added | None | Limited | Original Transformer |
| Learned | Absolute, added | $T_{\max} \times d$ | None | BERT, GPT-2 |
| T5 relative bias | Relative, score bias | Few | Moderate | T5 |
| RoPE | Relative via rotation | None | Moderate; extendable | LLaMA, many LLMs |
| ALiBi | Relative, linear score bias | None | Good | BLOOM, MPT |
Extending context length#
Training on very long sequences is expensive, so models are often trained at moderate length and then extended:
- Position interpolation (Chen et al., 2023): rescale positions so a longer sequence maps into the trained range (e.g. divide positions by 4), then fine-tune briefly.
- NTK-aware scaling and YaRN: rescale RoPE frequencies non-uniformly — preserving high-frequency (local) information while stretching low-frequency (long-range) components.
- Continued pretraining on long documents with the adjusted encoding.
Even with long context windows, models may use information in the middle of long inputs less reliably than at the beginning or end ("lost in the middle", Liu et al., 2023). Evaluate long-context behaviour with retrieval-style tests, not only perplexity.
NoPE: no positional encoding?#
Causal decoder models can partially infer position without explicit encodings, because the causal mask itself breaks symmetry (a token can count how many tokens precede it). Some studies found such models generalise reasonably to longer lengths, though explicit encodings remain standard.