๐Ÿ’ฌ NLP & Transformers ยท Lecture 26 of 29

Dialogue Systems and Chatbots: From Rules to LLM Assistants

We compare rule-based, task-oriented and open-domain dialogue systems; cover intent detection, slot filling and dialogue state tracking; and show how LLM-based assistants with retrieval and tools are designed, evaluated and deployed safely.

From ELIZA in 1966 to today's AI assistants, conversation has been a grand challenge of AI. Chatbots now answer customer questions, help people book appointments, guide users through complex procedures and provide information in many languages. Designing a good conversational system requires more than a powerful model: it requires understanding the users, the tasks, the failure modes and the right hand-offs to humans.

Types of dialogue systems#

TypeGoalExampleTypical technology
Rule-based / scriptedFollow predefined flowsMenu-driven SMS botDecision trees, keyword rules
Task-orientedComplete a specific taskBook an appointment, check case statusNLU + state tracking + policy + templates; or LLM with tools
Question answeringAnswer information requestsFAQ botRetrieval + reader/LLM (RAG)
Open-domain / socialEngaging general conversationCompanion chatbotsLarge language models

The classic task-oriented pipeline#

  1. Natural Language Understanding (NLU):
    • Intent detection โ€” what the user wants (book_appointment, check_status) โ€” a text classification problem.
    • Slot filling โ€” extract parameters (date=Tuesday, clinic=Camp 4) โ€” a sequence-labelling problem, often solved jointly with intent detection.
  2. Dialogue State Tracking (DST): maintain a structured record of what is known so far, updated every turn (the user may change their mind).
  3. Dialogue Policy: decide the next action โ€” ask for a missing slot, confirm, query a database, hand over to a human.
  4. Natural Language Generation (NLG): turn the action into text (templates or a generator).
text
User:   I need to see a doctor for my son on Tuesday.
NLU:    intent=book_appointment, slots={patient: son, date: Tuesday}
State:  {intent: book_appointment, patient: son, date: Tuesday, clinic: ?}
Policy: request(clinic)
NLG:    "Which clinic would you like to visit? Camp 4 or Camp 7?"

This modular design is controllable, testable and auditable โ€” valuable when conversations have real consequences โ€” but brittle to phrasings and flows its designers did not anticipate.

End-to-end and LLM-based assistants#

Large language models can conduct fluent, flexible conversation directly. Production assistants combine an LLM with:

  • a system prompt defining role, tone, scope and rules;
  • retrieval over approved documents (RAG) for accurate, current information;
  • tools/function calling to query databases or take actions (check a case status, book a slot) through well-defined, permission-checked APIs;
  • memory of the conversation (and, with consent, of user preferences);
  • guardrails: input/output filters, topic restrictions, refusal of unsafe requests, and escalation to humans.
python
# A minimal intent + slot baseline for a task-oriented bot
import re

INTENTS = {
    "book_appointment": ["appointment", "see a doctor", "book", "schedule"],
    "check_status": ["status", "my case", "application", "update"],
    "opening_hours": ["open", "hours", "close", "time"],
}
DAYS = r"(monday|tuesday|wednesday|thursday|friday|saturday|sunday)"

def nlu(text):
    t = text.lower()
    intent = max(INTENTS, key=lambda k: sum(kw in t for kw in INTENTS[k]))
    slots = {}
    if m := re.search(DAYS, t):
        slots["date"] = m.group(1)
    if m := re.search(r"camp\s*(\d+)", t):
        slots["clinic"] = f"Camp {m.group(1)}"
    return intent, slots

print(nlu("Can I book a doctor appointment at camp 4 on Tuesday?"))

In practice, replace the keyword rules with a fine-tuned classifier or an LLM with structured (JSON) outputs โ€” but keep the explicit state and policy where reliability matters.

Designing good conversations#

  • Scope clearly: tell users what the bot can and cannot do.
  • Handle failure gracefully: when unsure, ask a clarifying question or offer a human hand-off rather than guessing.
  • Confirm consequential actions before executing them.
  • Accessibility and language: support the languages, literacy levels and channels (SMS, WhatsApp, voice) your users actually use.
  • Privacy: collect the minimum data needed; explain how it is used; avoid storing sensitive details unnecessarily.
  • Transparency: users should know they are talking to an AI.

Evaluation#

  • Task-oriented: task success rate, turns to completion, slot accuracy, joint goal accuracy for state tracking.
  • QA bots: answer correctness and faithfulness against sources, containment rate (resolved without human) balanced against correct escalation.
  • Conversation quality: human ratings of helpfulness, coherence, empathy and safety; A/B tests with real users.
  • Red-teaming: deliberately attempt to make the bot give harmful, false or out-of-scope answers, including prompt-injection attacks hidden in retrieved documents.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Multilingual and Low-Resource NLP (with a Focus on Bangla)

Most of the world's languages have little digital data. We examine why this matters, how multilingual models enable cross-lingual transfer, the specific challenges of languages like Bangla, and practical strategies for building NLP in low-resource settings.

Intermediateโฑ 5 min#186
๐Ÿ’ฌ NLP & Transformers

Automatic Speech Recognition: From HMMs to Whisper

Speech recognition converts audio into text. We cover audio features and spectrograms, the classical HMMโ€“GMM pipeline, end-to-end neural models with CTC and attention, self-supervised wav2vec 2.0, Whisper, and evaluation with word error rate.

Intermediateโฑ 5 min#188
๐Ÿ’ฌ NLP & Transformers

Topic Modelling: LDA, NMF and Neural Topic Models

Topic models discover themes in large document collections without labels. We derive Latent Dirichlet Allocation's generative story, compare it with NMF and embedding-based BERTopic, and discuss evaluating and interpreting topics.

Intermediateโฑ 5 min#185