💬 NLP & Transformers · Lecture 20 of 29

Question Answering: Extractive, Open-Domain and Generative

Question answering systems return answers, not documents. We cover extractive reading comprehension with span prediction, open-domain retriever–reader pipelines, generative QA, evaluation metrics and the problem of unanswerable questions.

Search engines return lists of documents; people usually want answers. "What documents do I need to register a birth?" "When does the vaccination clinic open?" Question answering (QA) systems read text and respond directly. QA is a central NLP task, a benchmark of machine reading comprehension, and — through retrieval-augmented generation — the basis of many modern assistant systems.

Types of QA#

TypeSettingOutput
Extractive (reading comprehension)Question + a given passageA span copied from the passage
Open-domainQuestion only; search a large corpusSpan or generated answer
Generative / abstractiveQuestion (+ context)Free-form text
Knowledge-base QAQuestion over a structured knowledge graph/databaseEntity, value or query result
Multi-hopAnswer requires combining facts from several documentsSpan or text
ConversationalQuestions depend on earlier turnsSpan or text

Extractive QA#

SQuAD (Stanford Question Answering Dataset, 2016) contains 100,000+ questions on Wikipedia paragraphs, each answered by a span. BERT-style models solve it by encoding [CLS] question [SEP] passage [SEP] and predicting, for each passage token, the probability of being the answer start and end:

$$ P_{\text{start}}(i) = \text{softmax}_i(\mathbf{w}_s^\top\mathbf{h}_i), \qquad P_{\text{end}}(j) = \text{softmax}_j(\mathbf{w}_e^\top\mathbf{h}_j) $$

The predicted span maximises $P_{\text{start}}(i)\,P_{\text{end}}(j)$ subject to $i \le j$ and a maximum length. Within a couple of years of SQuAD's release, models exceeded the reported human performance on it — though this reflected the dataset's limits as much as true comprehension.

Unanswerable questions#

SQuAD 2.0 added questions whose answer is not in the passage. The model must abstain (predict the [CLS] position as a "no answer" span) when its best span score is below a threshold. Knowing when not to answer is essential for trustworthy systems.

python
from transformers import pipeline
qa = pipeline("question-answering", model="deepset/roberta-base-squad2")
context = ("Birth registration is free of charge at the civil registry office. Parents should bring "
           "the hospital birth notification and their identity documents. The office is open "
           "Sunday to Thursday from 9 am to 4 pm.")
for q in ["What should parents bring?", "When is the office open?", "How much does a passport cost?"]:
    r = qa(question=q, context=context, handle_impossible_answer=True)
    print(f"{q} -> {r['answer']!r} ({r['score']:.2f})")

The last question should yield an empty answer — the passage does not contain it.

Evaluation#

  • Exact Match (EM): prediction equals a gold answer after normalisation (lowercase, remove punctuation and articles).
  • Token F1: overlap between predicted and gold answer tokens — partial credit.
  • For generative QA: human judgement, reference-based metrics, and increasingly LLM-based judges (with caution); faithfulness to sources is critical.

Open-domain QA: retriever + reader#

When no passage is given, the system must first find relevant text in a large corpus:

  1. Retriever: select the top-$k$ passages. Classical sparse retrieval (BM25) or dense retrieval — encode questions and passages into vectors with dual encoders (e.g. DPR, Karpukhin et al. 2020) and retrieve by maximum inner product.
  2. Reader: extract or generate the answer from retrieved passages.

Dense retrievers are trained contrastively: a question should be closer to its gold passage than to other passages in the batch (in-batch negatives) and to "hard negatives" retrieved by BM25. Combining sparse and dense retrieval (hybrid search) and re-ranking the top results with a cross-encoder (which reads question and passage together) further improves accuracy.

Generative QA and RAG#

Sequence-to-sequence models (T5, BART) and LLMs generate answers conditioned on retrieved passages. Fusion-in-Decoder encodes each passage separately and lets the decoder attend to all of them. Retrieval-Augmented Generation (RAG) with LLMs is now the dominant pattern for question answering over organisational documents (covered in depth in the Generative AI track).

Challenges#

  • Lexical gap between question and answer text (addressed by dense retrieval).
  • Multi-hop reasoning across documents.
  • Temporal questions — answers change over time; the knowledge source must be current.
  • Ambiguous questions — clarify rather than guess.
  • Multilingual QA — questions and documents in different languages; multilingual encoders enable cross-lingual retrieval.
  • Dataset artefacts — models can exploit superficial cues (answer type, lexical overlap) rather than truly reading; adversarial evaluation reveals this.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Fine-Tuning Pretrained Language Models: A Practical Guide

How to adapt a pretrained language model to your task reliably — choosing a model, preparing data, hyperparameters, handling small data and instability, continued pretraining, and evaluating properly.

Intermediate⏱ 5 min#180
💬 NLP & Transformers

Text Summarisation: Extractive and Abstractive Methods

Summarisation condenses documents while preserving key information. We compare extractive methods (TextRank) with abstractive neural models, evaluate with ROUGE and factual-consistency checks, and discuss long documents and hallucination.

Intermediate⏱ 5 min#182
💬 NLP & Transformers

T5 and BART: Encoder–Decoder Pretraining and Text-to-Text Learning

T5 casts every NLP task as text in, text out; BART pretrains as a denoising autoencoder. We cover span corruption, the text-to-text framework, the lessons of T5's systematic study, and when encoder–decoders are the right choice.

Intermediate⏱ 5 min#179