💬 NLP & Transformers · Lecture 15 of 29

Positional Encodings: Sinusoidal, Learned, RoPE and ALiBi

Attention is order-blind, so transformers need positional information. We compare absolute sinusoidal and learned encodings with relative methods — rotary embeddings (RoPE) and ALiBi — and discuss extending context length.

Self-attention is permutation-equivariant: shuffle the input tokens and the outputs are shuffled the same way, with no other change. Without positional information, "dog bites man" and "man bites dog" would look identical to a transformer. How we inject position has a large effect on how well models handle long contexts and whether they can extrapolate to sequences longer than those seen in training.

Absolute positional encodings#

Sinusoidal (original Transformer)#

$$ PE_{(p, 2i)} = \sin\left(\frac{p}{10000^{2i/d}}\right), \qquad PE_{(p, 2i+1)} = \cos\left(\frac{p}{10000^{2i/d}}\right) $$

added to token embeddings. Each pair of dimensions is a rotating 2-D vector with a different frequency — like the hands of a clock running at many speeds. Key property: $PE_{p+k}$ is a linear function (a rotation) of $PE_p$, which in principle lets the model reason about relative offsets. No parameters, and defined for any position.

Learned absolute embeddings#

A trainable vector per position (BERT, GPT-2). Flexible, but undefined beyond the maximum trained length (e.g. 512 or 1024 positions) — the model cannot extrapolate.

Relative position methods#

What usually matters linguistically is relative distance ("the word two positions back"), not absolute index. Shaw et al. (2018) added learned relative-position terms to attention scores; T5 uses a learned scalar bias per relative-distance bucket added to attention logits.

Rotary Position Embedding (RoPE)#

RoPE (Su et al., 2021) is used by LLaMA, Mistral, Qwen, and many other modern LLMs. Instead of adding a vector to embeddings, it rotates query and key vectors by an angle proportional to their position. Split a $d$-dimensional vector into $d/2$ pairs; rotate pair $i$ at position $p$ by angle $p\theta_i$ with $\theta_i = 10000^{-2i/d}$:

$$ \begin{bmatrix} q'_{2i} \\ q'_{2i+1} \end{bmatrix} = \begin{bmatrix} \cos p\theta_i & -\sin p\theta_i \\ \sin p\theta_i & \cos p\theta_i \end{bmatrix}\begin{bmatrix} q_{2i} \\ q_{2i+1} \end{bmatrix} $$

Because rotations compose, the dot product between a query at position $m$ and a key at position $n$ depends only on their relative offset $m - n$:

$$ \langle R_m\mathbf{q},\, R_n\mathbf{k}\rangle = \langle\mathbf{q},\, R_{n-m}\mathbf{k}\rangle $$

RoPE thus encodes absolute position in the representation while making attention relative — with no extra parameters, compatible with efficient attention kernels and the KV cache.

python
import torch

def rope(x, base=10000.0):
    """x: (..., T, d) with even d. Rotates pairs of dimensions by position-dependent angles."""
    T, d = x.shape[-2], x.shape[-1]
    theta = base ** (-torch.arange(0, d, 2, dtype=torch.float32) / d)        # (d/2,)
    angles = torch.arange(T, dtype=torch.float32)[:, None] * theta[None]      # (T, d/2)
    cos, sin = angles.cos(), angles.sin()
    x1, x2 = x[..., 0::2], x[..., 1::2]
    out = torch.empty_like(x)
    out[..., 0::2] = x1 * cos - x2 * sin
    out[..., 1::2] = x1 * sin + x2 * cos
    return out

# Check the relative property: score depends only on the offset
q, k = torch.randn(1, 64), torch.randn(1, 64)
seq_q = rope(q.repeat(50, 1)); seq_k = rope(k.repeat(50, 1))
print((seq_q[10] @ seq_k[7]).item(), (seq_q[30] @ seq_k[27]).item())   # equal: both offset 3

ALiBi: attention with linear biases#

ALiBi (Press, Smith & Lewis, 2022) uses no positional embeddings at all. It adds a fixed, head-specific linear penalty to attention scores proportional to distance:

$$ \text{score}(i, j) = \mathbf{q}_i^\top\mathbf{k}_j - m_h\cdot(i - j) $$

with slopes $m_h$ forming a geometric sequence across heads (e.g. $1/2, 1/4, \dots$). Distant tokens are penalised more; different heads attend at different ranges. ALiBi showed strong length extrapolation: models trained on short sequences performed well on longer ones.

Comparison#

MethodTypeParametersExtrapolationUsed in
SinusoidalAbsolute, addedNoneLimitedOriginal Transformer
LearnedAbsolute, added$T_{\max} \times d$NoneBERT, GPT-2
T5 relative biasRelative, score biasFewModerateT5
RoPERelative via rotationNoneModerate; extendableLLaMA, many LLMs
ALiBiRelative, linear score biasNoneGoodBLOOM, MPT

Extending context length#

Training on very long sequences is expensive, so models are often trained at moderate length and then extended:

  • Position interpolation (Chen et al., 2023): rescale positions so a longer sequence maps into the trained range (e.g. divide positions by 4), then fine-tune briefly.
  • NTK-aware scaling and YaRN: rescale RoPE frequencies non-uniformly — preserving high-frequency (local) information while stretching low-frequency (long-range) components.
  • Continued pretraining on long documents with the adjusted encoding.

Even with long context windows, models may use information in the middle of long inputs less reliably than at the beginning or end ("lost in the middle", Liu et al., 2023). Evaluate long-context behaviour with retrieval-style tests, not only perplexity.

NoPE: no positional encoding?#

Causal decoder models can partially infer position without explicit encodings, because the causal mask itself breaks symmetry (a token can count how many tokens precede it). Some studies found such models generalise reasonably to longer lengths, though explicit encodings remain standard.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Self-Attention in Depth: Intuition, Complexity and Variants

A deeper look at self-attention — what attention heads learn, the geometry of queries and keys, computational complexity, causal masking, KV caching, and efficient variants such as multi-query and grouped-query attention.

Advanced⏱ 6 min#175
💬 NLP & Transformers

BERT: Bidirectional Encoder Representations from Transformers

BERT showed that a bidirectional transformer pretrained on unlabelled text could be fine-tuned to beat task-specific models across NLP. We cover masked language modelling, input format, fine-tuning patterns, and successors such as RoBERTa, DeBERTa and multilingual encoders.

Intermediate⏱ 5 min#177
💬 NLP & Transformers

The Transformer Architecture Explained, Block by Block

The 2017 Transformer replaced recurrence with attention and became the foundation of modern AI. We walk through embeddings, positional encoding, multi-head self-attention, feed-forward layers, residuals, normalisation, masking and the encoder–decoder design.

Intermediate⏱ 6 min#174