✨ Generative AI & LLMs · Lecture 25 of 30

Hallucination in LLMs: Causes, Detection and Mitigation

LLMs sometimes produce fluent but false or unsupported content. We classify hallucinations, explain why they arise from training and decoding, and survey detection methods and mitigation — from retrieval and citations to calibrated abstention.

A language model is asked for a court case that supports an argument, and it invents one — complete with a plausible name, citation and quotation. This actually happened: in 2023, lawyers in a US case were sanctioned for submitting a brief containing non-existent cases generated by a chatbot. Hallucination — generating content that is false, fabricated or unsupported by the provided sources — is the most important reliability problem of LLMs. As builders, we must understand why it happens and design systems that contain it.

Types of hallucination#

  • Factuality errors (extrinsic): statements contradicting world knowledge — wrong dates, invented statistics, fictitious references.
  • Faithfulness errors (intrinsic): outputs that contradict or go beyond the provided context — a summary adding details not in the source; a RAG answer citing a passage that does not support it.
  • Fabricated entities: non-existent papers, URLs, people, legal cases, medicine names.
  • Reasoning errors: plausible-looking steps with an incorrect conclusion.
  • Instruction inconsistency: ignoring constraints ("answer only from the document").

Why do LLMs hallucinate?#

  1. The training objective rewards plausibility, not truth. Next-token prediction learns to produce text that looks like the training data. A fluent fake citation looks much like a real one.
  2. Knowledge is stored imperfectly. Facts seen rarely during pretraining (long-tail entities, less-documented regions and languages) are recalled poorly; models blend similar facts.
  3. Knowledge cut-off. Events after training are unknown, yet the model may answer anyway.
  4. Pressure to answer. Instruction tuning and preference optimisation often reward confident, complete answers over "I don't know", and evaluation benchmarks typically give no credit for abstaining — incentivising guessing.
  5. Decoding. Sampling can pick low-probability tokens that commit the model to a false path; once written, the model tends to stay consistent with its own errors ("snowballing").
  6. Context problems. Long or noisy contexts, conflicting sources, or retrieval failures in RAG.

Detecting hallucinations#

  • Grounding checks: for each claim, check entailment against retrieved sources with an NLI model or an LLM judge; flag unsupported claims.
  • Self-consistency / sampling divergence: sample several answers; if they disagree on facts, the model is likely unsure (SelfCheckGPT uses this idea).
  • Token-level uncertainty: low probability or high entropy on key tokens (names, numbers) signals risk; semantic entropy (Farquhar et al., 2024) clusters sampled answers by meaning before computing entropy, detecting confabulations more reliably than lexical variation.
  • Verification with tools: look up claims in search engines, databases or APIs; execute code; recompute numbers.
  • Citation validation: confirm that cited documents exist and actually support the statement.
python
from transformers import pipeline

nli = pipeline("text-classification", model="MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli", top_k=None)

def grounded(source, claim, threshold=0.6):
    scores = {r["label"].lower(): r["score"] for r in nli({"text": source, "text_pair": claim})[0]}
    return scores.get("entailment", 0) >= threshold, scores

source = "Birth registration is free at the civil registry office within 45 days of birth."
for claim in ["Registration is free within 45 days.", "Registration costs 500 taka after 30 days."]:
    ok, s = grounded(source, claim)
    print(("SUPPORTED  " if ok else "UNSUPPORTED"), claim, {k: round(v, 2) for k, v in s.items()})

Mitigation strategies#

At the system level (most effective):

  1. Retrieval-augmented generation — ground answers in trusted, current documents.
  2. Require citations and verify them automatically.
  3. Allow and encourage abstention: "If the sources do not contain the answer, say so." Reward abstention in evaluation.
  4. Use tools for facts and computation: calculators, code, databases, search.
  5. Constrain outputs: structured formats, allowed values, validation against schemas and business rules.
  6. Human review for consequential outputs; clear UI signals of uncertainty and source links.

At the model level:

  • Fine-tune with examples of correct refusals and "I don't know" responses.
  • Preference optimisation that rewards factual accuracy (with verified labels), not just style.
  • Decoding: lower temperature for factual tasks.
  • Training on data with fewer errors; improving long-tail knowledge.

Designing for residual risk#

Measuring it#

Build a test set with questions whose answers are known, including unanswerable and adversarial questions, and measure: accuracy, hallucination rate (confident wrong answers), abstention rate, and faithfulness to provided sources. Benchmarks such as TruthfulQA (questions designed around common misconceptions) and factuality evaluations of long-form text (e.g. FActScore, which checks atomic facts in biographies) illustrate approaches you can adapt.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Retrieval-Augmented Generation (RAG): Grounding LLMs in Your Documents

RAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.

Intermediate⏱ 6 min#208
✨ Generative AI & LLMs

Decoding Strategies: Greedy, Beam Search, Temperature, Top-k and Top-p

A language model outputs probabilities; a decoding strategy turns them into text. We compare greedy and beam search with temperature, top-k, nucleus and min-p sampling, repetition penalties and constrained decoding, and when to use each.

Intermediate⏱ 5 min#214
✨ Generative AI & LLMs

Evaluating Large Language Models: Benchmarks, Arenas and Custom Evals

How good is an LLM — and for what? We survey capability benchmarks, human-preference arenas, safety evaluations, contamination and saturation problems, and how to build task-specific evaluation suites for your own applications.

Intermediate⏱ 5 min#216