A language model does not produce text directly. At each step it outputs a probability distribution over tens of thousands of tokens, and a decoding strategy decides which token comes next. The same model can be repetitive and dull, creative and coherent, or incoherent, depending on this choice. Understanding decoding is essential for anyone deploying LLMs.
Deterministic decoding#
Greedy decoding#
Always choose the most probable token: $y_t = \arg\max_w P(w \mid y_{<t})$. Fast and deterministic. For open-ended generation it often becomes repetitive and bland; locally optimal choices may not produce globally likely sequences.
Beam search#
Keep the $B$ highest-probability partial sequences at each step; expand each and keep the best $B$ overall. Finds higher-likelihood sequences than greedy decoding, with length normalisation to avoid favouring short outputs. Standard for translation, summarisation and speech recognition, where there is roughly one correct output. For open-ended text, Holtzman et al. (2020) showed that maximising likelihood leads to degenerate, repetitive text — human text is not the most probable text; it often contains surprising word choices.
Stochastic decoding (sampling)#
Sample the next token from the distribution. To control randomness:
Temperature#
Rescale logits $z_i$ before the softmax:
- $T < 1$: sharper distribution → more focused, conservative, deterministic.
- $T = 1$: the model's distribution.
- $T > 1$: flatter → more diverse and creative, but more errors.
- $T \to 0$: greedy decoding.
Top-k sampling#
Sample only among the $k$ most probable tokens (renormalised). Fixed $k$ is a problem: when the distribution is flat (many plausible continuations), $k$ may be too small; when it is peaked, $k$ may admit nonsense.
Nucleus (top-p) sampling#
Holtzman et al. (2020): choose the smallest set of tokens whose cumulative probability exceeds $p$ (e.g. 0.9), and sample from it. The candidate set adapts to the model's confidence — small when the model is sure, large when many continuations are plausible.
Min-p sampling#
Keep tokens whose probability is at least a fraction (e.g. 0.05–0.1) of the top token's probability — scales naturally with confidence and works well at higher temperatures.
import torch
def sample_next(logits, temperature=1.0, top_k=None, top_p=None, min_p=None):
logits = logits / max(temperature, 1e-6)
probs = torch.softmax(logits, dim=-1)
if top_k is not None:
kth = torch.topk(probs, top_k).values[-1]
probs = torch.where(probs >= kth, probs, torch.zeros_like(probs))
if top_p is not None:
sorted_p, idx = torch.sort(probs, descending=True)
cum = torch.cumsum(sorted_p, dim=0)
keep = cum - sorted_p < top_p # keep tokens until mass reaches p
mask = torch.zeros_like(probs, dtype=torch.bool); mask[idx[keep]] = True
probs = torch.where(mask, probs, torch.zeros_like(probs))
if min_p is not None:
probs = torch.where(probs >= min_p * probs.max(), probs, torch.zeros_like(probs))
probs = probs / probs.sum()
return int(torch.multinomial(probs, 1))
logits = torch.tensor([3.0, 2.5, 1.0, 0.2, -1.0, -3.0])
print([sample_next(logits, temperature=0.7, top_p=0.9) for _ in range(10)])Controlling repetition and content#
- Repetition penalty: reduce logits of tokens already generated.
- Frequency / presence penalties: penalise tokens by how often (or whether) they appeared.
- No-repeat n-gram constraints: forbid repeating any n-gram (common in summarisation).
- Stop sequences and maximum length.
- Logit bias: increase or forbid specific tokens.
Constrained and structured decoding#
When outputs must follow a format — valid JSON, a regular expression, a grammar, a label from a fixed set — constrained decoding masks out tokens that would violate the constraint at each step, guaranteeing syntactically valid output (libraries such as Outlines, guidance-style tools and API "structured output" modes). This is far more reliable than asking nicely in the prompt.
Choosing settings#
| Use case | Suggested decoding |
|---|---|
| Classification, extraction, factual QA | Greedy / temperature 0 (plus constrained output) |
| Code generation | Low temperature (0–0.3); sample several and test |
| Translation, summarisation | Beam search (4–5) or low temperature |
| Conversational assistant | Temperature ~0.6–0.8 with top-p 0.9 |
| Brainstorming, creative writing | Temperature ~0.9–1.2, top-p 0.95 or min-p |
| Reasoning with self-consistency | Temperature ~0.6–0.8, many samples, majority vote |
Decoding and quality#
Sampling settings interact with truthfulness: higher temperature increases hallucination risk; greedy decoding can loop. For factual tasks, prefer low temperature combined with grounding (RAG). For creativity, allow diversity but add review. Always evaluate the full system — model plus decoding — on your own test set.