You have studied transformers, prompting, RAG, fine-tuning, agents and evaluation. This capstone lecture puts them together into a disciplined process for building an LLM application that actually works for real users. Our running example: a multilingual assistant that answers staff and community questions using an organisation's service guidelines.
Step 1: define the problem and success#
- Users and context: who asks what, in which languages, through which channel (web, WhatsApp, voice)? What literacy and connectivity constraints exist?
- Scope: which questions are in scope; which must be escalated to humans (medical advice, legal status, protection concerns)?
- Success metrics: answer correctness, faithfulness to guidelines, correct escalation, user satisfaction, time saved, cost per conversation.
- Risks: what is the worst thing the system could say or do, and to whom?
Write this down before writing code. A clear non-goals list is as valuable as the goals.
Step 2: start simple#
Build the simplest thing that might work — often a good prompt plus RAG over the guidelines — and test it on 30 real questions. Only add complexity (fine-tuning, agents, multi-step pipelines) when evaluation shows a need.
Step 3: choose models#
Consider: quality on your evaluation set (not only leaderboards), language support, latency, cost, context length, data-handling terms, licence, and whether you need to self-host for privacy. A common pattern is routing: a small, cheap model for simple requests and a larger model for complex ones.
Step 4: design the architecture#
User message
→ input checks (language detection, PII redaction, safety filter, prompt-injection heuristics)
→ intent routing (FAQ / needs tool / out-of-scope / urgent → human)
→ retrieval (hybrid search over approved guidelines, filtered by language & permissions)
→ generation (system prompt + retrieved passages + conversation state)
→ output checks (grounding/citation verification, safety filter, format validation)
→ response with sources, or escalation to a human
→ logging (with privacy safeguards) for evaluation & monitoringKeep components separable and testable: retrieval, generation and guardrails should each have their own tests and metrics.
Step 5: prompts and data as code#
Version prompts, retrieval configurations and evaluation sets in source control. Store model identifiers and parameters with every logged response, so behaviour changes can be traced.
Step 6: evaluate continuously#
- Offline: a golden set covering typical questions, edge cases, every language, adversarial inputs, and questions that must be refused or escalated. Grade with programmatic checks, validated LLM judges and expert review.
- Regression tests: run the suite on every change.
- Online: user feedback (thumbs up/down with reasons), escalation rates, resolution rates, A/B tests where appropriate.
from dataclasses import dataclass, field
import time
@dataclass
class Trace:
request_id: str
model: str
prompt_version: str
retrieved_ids: list = field(default_factory=list)
latency_ms: float = 0.0
input_tokens: int = 0
output_tokens: int = 0
grounded: bool | None = None
escalated: bool = False
def answer(question, retriever, llm, verifier, prompt_version="v7"):
t0 = time.time()
passages = retriever(question)
reply = llm(question, passages)
ok = verifier(reply["text"], passages) # citation / entailment check
trace = Trace(reply["id"], reply["model"], prompt_version, [p["id"] for p in passages],
(time.time() - t0) * 1000, reply["usage"]["in"], reply["usage"]["out"], ok,
escalated=not ok)
text = reply["text"] if ok else "I'm not certain about this. I've passed your question to a staff member."
return text, trace # log the trace (never raw personal data)Step 7: guardrails and safety#
- Input and output safety classifiers; topic restrictions.
- PII handling: redact or avoid storing personal data; comply with data-protection law and organisational policy.
- Prompt-injection defences: treat retrieved and user content as data; restrict tool permissions.
- Escalation paths to trained humans for sensitive topics, distress signals and low-confidence answers.
- Transparency: tell users they are talking to an AI and how their data is used.
Step 8: cost and latency#
Estimate cost = (input tokens + output tokens) × price × volume. Reduce it with shorter prompts, fewer and better retrieved passages, response-length limits, caching of frequent answers, prompt/prefix caching, smaller models for easy requests, and batching for offline jobs. Measure p50 and p95 latency; stream tokens to improve perceived speed.
Step 9: deploy and monitor#
- Gradual rollout (internal users → pilot group → wider).
- Dashboards for quality signals, escalation rates, latency, cost, error rates and safety flags.
- Watch for drift: new topics, changed policies (update the index!), model version changes by providers.
- Incident response: a way to quickly disable features, roll back prompts or models, and notify users.
Step 10: governance#
Document the system (purpose, data sources, models, evaluations, known limitations, human oversight) in a model/system card; conduct a risk and impact assessment, especially where the system affects access to services; review regularly with domain experts and, where possible, representatives of affected communities.