✨ Generative AI & LLMs · Lecture 30 of 30

Building an LLM Application End to End: From Idea to Production

A practical capstone for the Generative AI track — scoping a use case, choosing models, designing prompts, RAG and tools, evaluation, guardrails, cost and latency, deployment, monitoring and governance.

You have studied transformers, prompting, RAG, fine-tuning, agents and evaluation. This capstone lecture puts them together into a disciplined process for building an LLM application that actually works for real users. Our running example: a multilingual assistant that answers staff and community questions using an organisation's service guidelines.

Step 1: define the problem and success#

  • Users and context: who asks what, in which languages, through which channel (web, WhatsApp, voice)? What literacy and connectivity constraints exist?
  • Scope: which questions are in scope; which must be escalated to humans (medical advice, legal status, protection concerns)?
  • Success metrics: answer correctness, faithfulness to guidelines, correct escalation, user satisfaction, time saved, cost per conversation.
  • Risks: what is the worst thing the system could say or do, and to whom?

Write this down before writing code. A clear non-goals list is as valuable as the goals.

Step 2: start simple#

Build the simplest thing that might work — often a good prompt plus RAG over the guidelines — and test it on 30 real questions. Only add complexity (fine-tuning, agents, multi-step pipelines) when evaluation shows a need.

Step 3: choose models#

Consider: quality on your evaluation set (not only leaderboards), language support, latency, cost, context length, data-handling terms, licence, and whether you need to self-host for privacy. A common pattern is routing: a small, cheap model for simple requests and a larger model for complex ones.

Step 4: design the architecture#

text
User message
  → input checks (language detection, PII redaction, safety filter, prompt-injection heuristics)
  → intent routing (FAQ / needs tool / out-of-scope / urgent → human)
  → retrieval (hybrid search over approved guidelines, filtered by language & permissions)
  → generation (system prompt + retrieved passages + conversation state)
  → output checks (grounding/citation verification, safety filter, format validation)
  → response with sources, or escalation to a human
  → logging (with privacy safeguards) for evaluation & monitoring

Keep components separable and testable: retrieval, generation and guardrails should each have their own tests and metrics.

Step 5: prompts and data as code#

Version prompts, retrieval configurations and evaluation sets in source control. Store model identifiers and parameters with every logged response, so behaviour changes can be traced.

Step 6: evaluate continuously#

  • Offline: a golden set covering typical questions, edge cases, every language, adversarial inputs, and questions that must be refused or escalated. Grade with programmatic checks, validated LLM judges and expert review.
  • Regression tests: run the suite on every change.
  • Online: user feedback (thumbs up/down with reasons), escalation rates, resolution rates, A/B tests where appropriate.
python
from dataclasses import dataclass, field
import time

@dataclass
class Trace:
    request_id: str
    model: str
    prompt_version: str
    retrieved_ids: list = field(default_factory=list)
    latency_ms: float = 0.0
    input_tokens: int = 0
    output_tokens: int = 0
    grounded: bool | None = None
    escalated: bool = False

def answer(question, retriever, llm, verifier, prompt_version="v7"):
    t0 = time.time()
    passages = retriever(question)
    reply = llm(question, passages)
    ok = verifier(reply["text"], passages)                  # citation / entailment check
    trace = Trace(reply["id"], reply["model"], prompt_version, [p["id"] for p in passages],
                  (time.time() - t0) * 1000, reply["usage"]["in"], reply["usage"]["out"], ok,
                  escalated=not ok)
    text = reply["text"] if ok else "I'm not certain about this. I've passed your question to a staff member."
    return text, trace                                       # log the trace (never raw personal data)

Step 7: guardrails and safety#

  • Input and output safety classifiers; topic restrictions.
  • PII handling: redact or avoid storing personal data; comply with data-protection law and organisational policy.
  • Prompt-injection defences: treat retrieved and user content as data; restrict tool permissions.
  • Escalation paths to trained humans for sensitive topics, distress signals and low-confidence answers.
  • Transparency: tell users they are talking to an AI and how their data is used.

Step 8: cost and latency#

Estimate cost = (input tokens + output tokens) × price × volume. Reduce it with shorter prompts, fewer and better retrieved passages, response-length limits, caching of frequent answers, prompt/prefix caching, smaller models for easy requests, and batching for offline jobs. Measure p50 and p95 latency; stream tokens to improve perceived speed.

Step 9: deploy and monitor#

  • Gradual rollout (internal users → pilot group → wider).
  • Dashboards for quality signals, escalation rates, latency, cost, error rates and safety flags.
  • Watch for drift: new topics, changed policies (update the index!), model version changes by providers.
  • Incident response: a way to quickly disable features, roll back prompts or models, and notify users.

Step 10: governance#

Document the system (purpose, data sources, models, evaluations, known limitations, human oversight) in a model/system card; conduct a risk and impact assessment, especially where the system affects access to services; review regularly with domain experts and, where possible, representatives of affected communities.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Retrieval-Augmented Generation (RAG): Grounding LLMs in Your Documents

RAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.

Intermediate⏱ 6 min#208
✨ Generative AI & LLMs

Text-to-Image and Text-to-Video Generation: Systems, Control and Provenance

A systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.

Intermediate⏱ 5 min#219
✨ Generative AI & LLMs

Multimodal Models: Vision–Language and Beyond

Multimodal models understand and generate across text, images, audio and video. We cover fusion strategies, vision–language model architectures (LLaVA-style), training stages, capabilities like document understanding and VQA, and known failure modes.

Intermediate⏱ 5 min#218