Search engines return lists of documents; people usually want answers. "What documents do I need to register a birth?" "When does the vaccination clinic open?" Question answering (QA) systems read text and respond directly. QA is a central NLP task, a benchmark of machine reading comprehension, and — through retrieval-augmented generation — the basis of many modern assistant systems.
Types of QA#
| Type | Setting | Output |
|---|---|---|
| Extractive (reading comprehension) | Question + a given passage | A span copied from the passage |
| Open-domain | Question only; search a large corpus | Span or generated answer |
| Generative / abstractive | Question (+ context) | Free-form text |
| Knowledge-base QA | Question over a structured knowledge graph/database | Entity, value or query result |
| Multi-hop | Answer requires combining facts from several documents | Span or text |
| Conversational | Questions depend on earlier turns | Span or text |
Extractive QA#
SQuAD (Stanford Question Answering Dataset, 2016) contains 100,000+ questions on Wikipedia paragraphs, each answered by a span. BERT-style models solve it by encoding [CLS] question [SEP] passage [SEP] and predicting, for each passage token, the probability of being the answer start and end:
The predicted span maximises $P_{\text{start}}(i)\,P_{\text{end}}(j)$ subject to $i \le j$ and a maximum length. Within a couple of years of SQuAD's release, models exceeded the reported human performance on it — though this reflected the dataset's limits as much as true comprehension.
Unanswerable questions#
SQuAD 2.0 added questions whose answer is not in the passage. The model must abstain (predict the [CLS] position as a "no answer" span) when its best span score is below a threshold. Knowing when not to answer is essential for trustworthy systems.
from transformers import pipeline
qa = pipeline("question-answering", model="deepset/roberta-base-squad2")
context = ("Birth registration is free of charge at the civil registry office. Parents should bring "
"the hospital birth notification and their identity documents. The office is open "
"Sunday to Thursday from 9 am to 4 pm.")
for q in ["What should parents bring?", "When is the office open?", "How much does a passport cost?"]:
r = qa(question=q, context=context, handle_impossible_answer=True)
print(f"{q} -> {r['answer']!r} ({r['score']:.2f})")The last question should yield an empty answer — the passage does not contain it.
Evaluation#
- Exact Match (EM): prediction equals a gold answer after normalisation (lowercase, remove punctuation and articles).
- Token F1: overlap between predicted and gold answer tokens — partial credit.
- For generative QA: human judgement, reference-based metrics, and increasingly LLM-based judges (with caution); faithfulness to sources is critical.
Open-domain QA: retriever + reader#
When no passage is given, the system must first find relevant text in a large corpus:
- Retriever: select the top-$k$ passages. Classical sparse retrieval (BM25) or dense retrieval — encode questions and passages into vectors with dual encoders (e.g. DPR, Karpukhin et al. 2020) and retrieve by maximum inner product.
- Reader: extract or generate the answer from retrieved passages.
Dense retrievers are trained contrastively: a question should be closer to its gold passage than to other passages in the batch (in-batch negatives) and to "hard negatives" retrieved by BM25. Combining sparse and dense retrieval (hybrid search) and re-ranking the top results with a cross-encoder (which reads question and passage together) further improves accuracy.
Generative QA and RAG#
Sequence-to-sequence models (T5, BART) and LLMs generate answers conditioned on retrieved passages. Fusion-in-Decoder encodes each passage separately and lets the decoder attend to all of them. Retrieval-Augmented Generation (RAG) with LLMs is now the dominant pattern for question answering over organisational documents (covered in depth in the Generative AI track).
Challenges#
- Lexical gap between question and answer text (addressed by dense retrieval).
- Multi-hop reasoning across documents.
- Temporal questions — answers change over time; the knowledge source must be current.
- Ambiguous questions — clarify rather than guess.
- Multilingual QA — questions and documents in different languages; multilingual encoders enable cross-lingual retrieval.
- Dataset artefacts — models can exploit superficial cues (answer type, lexical overlap) rather than truly reading; adversarial evaluation reveals this.