A language model is asked for a court case that supports an argument, and it invents one — complete with a plausible name, citation and quotation. This actually happened: in 2023, lawyers in a US case were sanctioned for submitting a brief containing non-existent cases generated by a chatbot. Hallucination — generating content that is false, fabricated or unsupported by the provided sources — is the most important reliability problem of LLMs. As builders, we must understand why it happens and design systems that contain it.
Types of hallucination#
- Factuality errors (extrinsic): statements contradicting world knowledge — wrong dates, invented statistics, fictitious references.
- Faithfulness errors (intrinsic): outputs that contradict or go beyond the provided context — a summary adding details not in the source; a RAG answer citing a passage that does not support it.
- Fabricated entities: non-existent papers, URLs, people, legal cases, medicine names.
- Reasoning errors: plausible-looking steps with an incorrect conclusion.
- Instruction inconsistency: ignoring constraints ("answer only from the document").
Why do LLMs hallucinate?#
- The training objective rewards plausibility, not truth. Next-token prediction learns to produce text that looks like the training data. A fluent fake citation looks much like a real one.
- Knowledge is stored imperfectly. Facts seen rarely during pretraining (long-tail entities, less-documented regions and languages) are recalled poorly; models blend similar facts.
- Knowledge cut-off. Events after training are unknown, yet the model may answer anyway.
- Pressure to answer. Instruction tuning and preference optimisation often reward confident, complete answers over "I don't know", and evaluation benchmarks typically give no credit for abstaining — incentivising guessing.
- Decoding. Sampling can pick low-probability tokens that commit the model to a false path; once written, the model tends to stay consistent with its own errors ("snowballing").
- Context problems. Long or noisy contexts, conflicting sources, or retrieval failures in RAG.
Detecting hallucinations#
- Grounding checks: for each claim, check entailment against retrieved sources with an NLI model or an LLM judge; flag unsupported claims.
- Self-consistency / sampling divergence: sample several answers; if they disagree on facts, the model is likely unsure (SelfCheckGPT uses this idea).
- Token-level uncertainty: low probability or high entropy on key tokens (names, numbers) signals risk; semantic entropy (Farquhar et al., 2024) clusters sampled answers by meaning before computing entropy, detecting confabulations more reliably than lexical variation.
- Verification with tools: look up claims in search engines, databases or APIs; execute code; recompute numbers.
- Citation validation: confirm that cited documents exist and actually support the statement.
from transformers import pipeline
nli = pipeline("text-classification", model="MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli", top_k=None)
def grounded(source, claim, threshold=0.6):
scores = {r["label"].lower(): r["score"] for r in nli({"text": source, "text_pair": claim})[0]}
return scores.get("entailment", 0) >= threshold, scores
source = "Birth registration is free at the civil registry office within 45 days of birth."
for claim in ["Registration is free within 45 days.", "Registration costs 500 taka after 30 days."]:
ok, s = grounded(source, claim)
print(("SUPPORTED " if ok else "UNSUPPORTED"), claim, {k: round(v, 2) for k, v in s.items()})Mitigation strategies#
At the system level (most effective):
- Retrieval-augmented generation — ground answers in trusted, current documents.
- Require citations and verify them automatically.
- Allow and encourage abstention: "If the sources do not contain the answer, say so." Reward abstention in evaluation.
- Use tools for facts and computation: calculators, code, databases, search.
- Constrain outputs: structured formats, allowed values, validation against schemas and business rules.
- Human review for consequential outputs; clear UI signals of uncertainty and source links.
At the model level:
- Fine-tune with examples of correct refusals and "I don't know" responses.
- Preference optimisation that rewards factual accuracy (with verified labels), not just style.
- Decoding: lower temperature for factual tasks.
- Training on data with fewer errors; improving long-tail knowledge.
Designing for residual risk#
Measuring it#
Build a test set with questions whose answers are known, including unanswerable and adversarial questions, and measure: accuracy, hallucination rate (confident wrong answers), abstention rate, and faithfulness to provided sources. Benchmarks such as TruthfulQA (questions designed around common misconceptions) and factuality evaluations of long-form text (e.g. FActScore, which checks atomic facts in biographies) illustrate approaches you can adapt.