✨ Generative AI & LLMs · Lecture 18 of 30

Retrieval-Augmented Generation (RAG): Grounding LLMs in Your Documents

RAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.

An LLM on its own knows only what it absorbed during training — possibly outdated, never including your organisation's private documents, and mixed with hallucinations. Retrieval-Augmented Generation (RAG) addresses this by retrieving relevant passages from a knowledge base at query time and asking the model to answer using those passages. Proposed by Lewis et al. (2020), RAG has become the most common architecture for building LLM applications over organisational knowledge: policy handbooks, FAQs, research reports, technical manuals.

Why RAG?#

  • Fresh knowledge: update the document index, not the model.
  • Private and domain knowledge without fine-tuning.
  • Grounding and citations: answers can point to sources, enabling verification.
  • Reduced hallucination (not eliminated).
  • Access control: retrieve only documents the user is allowed to see.

The pipeline#

1. Ingestion#

Collect documents (PDFs, web pages, spreadsheets), extract text (OCR for scans), clean it, and attach metadata: title, section, date, language, source URL, access permissions.

2. Chunking#

Split documents into passages small enough to be specific but large enough to be self-contained — often 200–800 tokens with some overlap. Better: split along structure (headings, sections, list items), keep tables intact, and prepend the document title and section heading to each chunk so it makes sense in isolation.

3. Embedding and indexing#

Embed each chunk with a text-embedding model (multilingual if your content or users are multilingual) and store vectors in a vector index. Also index the text for keyword (BM25) search.

4. Retrieval#

Embed the user's question; retrieve the top-$k$ chunks by vector similarity, combined with keyword search (hybrid retrieval), filtered by metadata (language, date, permissions).

5. Re-ranking#

Re-score the top candidates (e.g. top 30) with a cross-encoder re-ranker and keep the best few (e.g. 5).

6. Generation#

Insert the passages into a prompt that instructs the model to answer only from them, cite sources, and say when the answer is not present.

7. Post-processing#

Validate citations, format the answer, log the interaction (with privacy safeguards) for evaluation.

A minimal RAG system#

python
import numpy as np
from sentence_transformers import SentenceTransformer, CrossEncoder

chunks = [
    {"id": "reg-1", "text": "Birth registration is free at the civil registry office within 45 days of birth."},
    {"id": "reg-2", "text": "Late birth registration after 45 days requires a sworn statement and two witnesses."},
    {"id": "hlth-1", "text": "Children under five receive free vaccinations every Tuesday at the health post."},
    {"id": "cash-1", "text": "Cash assistance applications are reviewed within 30 days; decisions are sent by SMS."},
]
embedder = SentenceTransformer("intfloat/multilingual-e5-small")
E = embedder.encode(["passage: " + c["text"] for c in chunks], normalize_embeddings=True)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")

def retrieve(question, k=3, keep=2):
    q = embedder.encode(["query: " + question], normalize_embeddings=True)[0]
    cand = np.argsort(-(E @ q))[:k]
    scores = reranker.predict([(question, chunks[i]["text"]) for i in cand])
    return [chunks[cand[i]] for i in np.argsort(-scores)[:keep]]

def build_prompt(question, passages):
    context = "\n".join(f"[{p['id']}] {p['text']}" for p in passages)
    return (f"Answer the question using ONLY the sources below. Cite source ids in brackets. "
            f"If the sources do not contain the answer, say you don't know.\n\n"
            f"Sources:\n{context}\n\nQuestion: {question}\nAnswer:")

q = "My baby was born two months ago. Can I still register the birth?"
print(build_prompt(q, retrieve(q)))
# answer = llm(build_prompt(q, retrieve(q)))  # send to any chat model

Evaluating RAG#

Evaluate the two halves separately and together:

  • Retrieval: recall@k and MRR on a labelled set of questions with their relevant chunks. If retrieval fails, generation cannot succeed.
  • Generation:
    • Faithfulness / groundedness — is every claim supported by the retrieved passages?
    • Answer relevance — does it address the question?
    • Correctness — against reference answers written by experts.
    • Citation accuracy — do cited sources support the statements?
    • Abstention — does it say "I don't know" for unanswerable questions?

Frameworks such as RAGAS and LLM-as-judge rubrics help scale evaluation; validate them against expert judgements on a sample.

Common failure modes and fixes#

FailureFix
Relevant chunk not retrievedHybrid search, better chunking, query rewriting, multilingual embeddings
Retrieved but ignored or misreadRe-ranking, fewer/better passages, clearer prompt, put key passages first
Outdated or conflicting sourcesMetadata filters by date/version; show dates; resolve conflicts in the index
Question needs several documentsMulti-query retrieval, iterative/agentic retrieval
Hallucinated detailsStricter grounding instructions, citation verification, abstention
Wrong languageDetect language; multilingual embeddings; answer in the user's language

Advanced RAG patterns#

  • Query rewriting / expansion: reformulate ambiguous or conversational questions ("what about for adults?") into standalone queries.
  • HyDE: generate a hypothetical answer, embed it, and retrieve with it.
  • Parent–child chunking: retrieve small chunks for precision, pass their larger parent sections for context.
  • Agentic RAG: the model decides when and what to retrieve, iterating until it has enough evidence.
  • GraphRAG: build a knowledge graph of entities and relations to answer global or multi-hop questions.

Security and privacy#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Vector Databases and Approximate Nearest Neighbour Search

Embedding-based applications need fast similarity search over millions of vectors. We explain exact vs approximate search, HNSW graphs, IVF and product quantisation, filtering, and how to choose and operate a vector store.

Intermediate⏱ 6 min#209
✨ Generative AI & LLMs

Hallucination in LLMs: Causes, Detection and Mitigation

LLMs sometimes produce fluent but false or unsupported content. We classify hallucinations, explain why they arise from training and decoding, and survey detection methods and mitigation — from retrieval and citations to calibrated abstention.

Intermediate⏱ 6 min#215
✨ Generative AI & LLMs

Building an LLM Application End to End: From Idea to Production

A practical capstone for the Generative AI track — scoping a use case, choosing models, designing prompts, RAG and tools, evaluation, guardrails, cost and latency, deployment, monitoring and governance.

Intermediate⏱ 5 min#220