An LLM on its own knows only what it absorbed during training — possibly outdated, never including your organisation's private documents, and mixed with hallucinations. Retrieval-Augmented Generation (RAG) addresses this by retrieving relevant passages from a knowledge base at query time and asking the model to answer using those passages. Proposed by Lewis et al. (2020), RAG has become the most common architecture for building LLM applications over organisational knowledge: policy handbooks, FAQs, research reports, technical manuals.
Why RAG?#
- Fresh knowledge: update the document index, not the model.
- Private and domain knowledge without fine-tuning.
- Grounding and citations: answers can point to sources, enabling verification.
- Reduced hallucination (not eliminated).
- Access control: retrieve only documents the user is allowed to see.
The pipeline#
1. Ingestion#
Collect documents (PDFs, web pages, spreadsheets), extract text (OCR for scans), clean it, and attach metadata: title, section, date, language, source URL, access permissions.
2. Chunking#
Split documents into passages small enough to be specific but large enough to be self-contained — often 200–800 tokens with some overlap. Better: split along structure (headings, sections, list items), keep tables intact, and prepend the document title and section heading to each chunk so it makes sense in isolation.
3. Embedding and indexing#
Embed each chunk with a text-embedding model (multilingual if your content or users are multilingual) and store vectors in a vector index. Also index the text for keyword (BM25) search.
4. Retrieval#
Embed the user's question; retrieve the top-$k$ chunks by vector similarity, combined with keyword search (hybrid retrieval), filtered by metadata (language, date, permissions).
5. Re-ranking#
Re-score the top candidates (e.g. top 30) with a cross-encoder re-ranker and keep the best few (e.g. 5).
6. Generation#
Insert the passages into a prompt that instructs the model to answer only from them, cite sources, and say when the answer is not present.
7. Post-processing#
Validate citations, format the answer, log the interaction (with privacy safeguards) for evaluation.
A minimal RAG system#
import numpy as np
from sentence_transformers import SentenceTransformer, CrossEncoder
chunks = [
{"id": "reg-1", "text": "Birth registration is free at the civil registry office within 45 days of birth."},
{"id": "reg-2", "text": "Late birth registration after 45 days requires a sworn statement and two witnesses."},
{"id": "hlth-1", "text": "Children under five receive free vaccinations every Tuesday at the health post."},
{"id": "cash-1", "text": "Cash assistance applications are reviewed within 30 days; decisions are sent by SMS."},
]
embedder = SentenceTransformer("intfloat/multilingual-e5-small")
E = embedder.encode(["passage: " + c["text"] for c in chunks], normalize_embeddings=True)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
def retrieve(question, k=3, keep=2):
q = embedder.encode(["query: " + question], normalize_embeddings=True)[0]
cand = np.argsort(-(E @ q))[:k]
scores = reranker.predict([(question, chunks[i]["text"]) for i in cand])
return [chunks[cand[i]] for i in np.argsort(-scores)[:keep]]
def build_prompt(question, passages):
context = "\n".join(f"[{p['id']}] {p['text']}" for p in passages)
return (f"Answer the question using ONLY the sources below. Cite source ids in brackets. "
f"If the sources do not contain the answer, say you don't know.\n\n"
f"Sources:\n{context}\n\nQuestion: {question}\nAnswer:")
q = "My baby was born two months ago. Can I still register the birth?"
print(build_prompt(q, retrieve(q)))
# answer = llm(build_prompt(q, retrieve(q))) # send to any chat modelEvaluating RAG#
Evaluate the two halves separately and together:
- Retrieval: recall@k and MRR on a labelled set of questions with their relevant chunks. If retrieval fails, generation cannot succeed.
- Generation:
- Faithfulness / groundedness — is every claim supported by the retrieved passages?
- Answer relevance — does it address the question?
- Correctness — against reference answers written by experts.
- Citation accuracy — do cited sources support the statements?
- Abstention — does it say "I don't know" for unanswerable questions?
Frameworks such as RAGAS and LLM-as-judge rubrics help scale evaluation; validate them against expert judgements on a sample.
Common failure modes and fixes#
| Failure | Fix |
|---|---|
| Relevant chunk not retrieved | Hybrid search, better chunking, query rewriting, multilingual embeddings |
| Retrieved but ignored or misread | Re-ranking, fewer/better passages, clearer prompt, put key passages first |
| Outdated or conflicting sources | Metadata filters by date/version; show dates; resolve conflicts in the index |
| Question needs several documents | Multi-query retrieval, iterative/agentic retrieval |
| Hallucinated details | Stricter grounding instructions, citation verification, abstention |
| Wrong language | Detect language; multilingual embeddings; answer in the user's language |
Advanced RAG patterns#
- Query rewriting / expansion: reformulate ambiguous or conversational questions ("what about for adults?") into standalone queries.
- HyDE: generate a hypothetical answer, embed it, and retrieve with it.
- Parent–child chunking: retrieve small chunks for precision, pass their larger parent sections for context.
- Agentic RAG: the model decides when and what to retrieve, iterating until it has enough evidence.
- GraphRAG: build a knowledge graph of entities and relations to answer global or multi-hop questions.