💬 NLP & Transformers · Lecture 21 of 29

Text Summarisation: Extractive and Abstractive Methods

Summarisation condenses documents while preserving key information. We compare extractive methods (TextRank) with abstractive neural models, evaluate with ROUGE and factual-consistency checks, and discuss long documents and hallucination.

Analysts face more reports, articles, meeting transcripts and case notes than anyone can read. Automatic summarisation produces a shorter version that preserves the most important information. It is among the most useful — and most error-prone — NLP applications, because a fluent summary that states something false can be worse than no summary at all.

Two paradigms#

  • Extractive: select the most important sentences from the source and concatenate them. Faithful by construction (every sentence appears in the source), but can be choppy and redundant.
  • Abstractive: generate new text that paraphrases and condenses. More fluent and concise, but risks hallucination — content not supported by the source.

Summaries can also be single-document or multi-document, generic or query-focused ("summarise what this report says about water supply"), and targeted to different lengths and audiences.

Extractive methods#

Frequency-based (Luhn, 1958): score sentences by the frequency of their important words.

TextRank (Mihalcea & Tarau, 2004): build a graph whose nodes are sentences and whose edge weights are sentence similarities; run PageRank; pick the top-ranked sentences. Sentences similar to many others are "central".

python
import numpy as np, re
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

def textrank_summary(text, k=2, d=0.85, iters=50):
    sents = [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]
    X = TfidfVectorizer().fit_transform(sents)
    S = cosine_similarity(X); np.fill_diagonal(S, 0)
    S = S / (S.sum(1, keepdims=True) + 1e-9)          # row-normalise into transition probabilities
    r = np.ones(len(sents)) / len(sents)
    for _ in range(iters):
        r = (1 - d) / len(sents) + d * S.T @ r       # PageRank iteration
    top = sorted(np.argsort(-r)[:k])                  # keep original order
    return " ".join(sents[i] for i in top)

doc = ("Floods affected five districts this week. Around 20,000 families were displaced. "
       "Authorities opened 40 temporary shelters in schools. Clean water is the most urgent need. "
       "Health workers warned of possible cholera outbreaks. Schools will remain closed until Sunday.")
print(textrank_summary(doc))

Supervised extractive models (e.g. BERTSum) classify each sentence as "include" or not, trained on labels derived from reference summaries.

Abstractive methods#

Encoder–decoder transformers fine-tuned on (document, summary) pairs: BART, PEGASUS (pretrained with "gap sentence generation" — masking whole important sentences and generating them, closely matching the summarisation task), T5, and — increasingly — instruction-tuned LLMs prompted to summarise.

Common datasets: CNN/DailyMail (news highlights, fairly extractive), XSum (single-sentence, highly abstractive BBC summaries), arXiv/PubMed (long scientific papers), and meeting and dialogue summarisation sets.

Evaluation#

ROUGE (Lin, 2004) measures n-gram overlap with reference summaries, emphasising recall:

  • ROUGE-1 / ROUGE-2: unigram / bigram overlap.
  • ROUGE-L: longest common subsequence.
$$ \text{ROUGE-N recall} = \frac{\sum_{\text{gram}_n \in \text{ref}}\text{Count}_{\text{match}}(\text{gram}_n)}{\sum_{\text{gram}_n \in \text{ref}}\text{Count}(\text{gram}_n)} $$

ROUGE is cheap but rewards word overlap, not correctness: a summary can score well while stating something false, or score poorly while being an excellent paraphrase.

python
# pip install rouge-score
from rouge_score import rouge_scorer
sc = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL"], use_stemmer=True)
ref = "About 20,000 families were displaced by floods; clean water is the most urgent need."
hyp = "Floods displaced 20,000 families, and clean water is urgently needed."
print({k: round(v.fmeasure, 3) for k, v in sc.score(ref, hyp).items()})

Factual consistency (faithfulness) requires separate evaluation:

  • Entailment-based checks: does the source entail each summary sentence (using an NLI model)? (e.g. SummaC)
  • QA-based checks: generate questions from the summary, answer them from the source, and compare (QAGS, QuestEval).
  • LLM-as-judge evaluations with explicit rubrics — useful at scale, but validate against human judgements.
  • Human evaluation: coverage, faithfulness, fluency, conciseness.

Studies found that a substantial fraction of abstractive summaries from earlier neural models contained unsupported information (Maynez et al., 2020). Modern LLMs are better, but not immune.

Long documents#

Transformers have limited context windows. Strategies:

  • Truncation (loses information);
  • Chunk-and-merge / map-reduce: summarise chunks, then summarise the summaries;
  • Extract-then-abstract: select important passages first;
  • Long-context models (Longformer/LED, long-context LLMs) — but check whether information from the middle of long inputs is used.

Responsible summarisation#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Evaluating NLP Systems: Perplexity, BLEU, ROUGE, BERTScore and Human Judgement

How do we know if a language system is good? We survey intrinsic and extrinsic evaluation, overlap metrics, embedding-based metrics, learned metrics, LLM judges, human evaluation, benchmarks and their pitfalls.

Intermediate⏱ 5 min#183
💬 NLP & Transformers

Question Answering: Extractive, Open-Domain and Generative

Question answering systems return answers, not documents. We cover extractive reading comprehension with span prediction, open-domain retriever–reader pipelines, generative QA, evaluation metrics and the problem of unanswerable questions.

Intermediate⏱ 5 min#181
💬 NLP & Transformers

Fine-Tuning Pretrained Language Models: A Practical Guide

How to adapt a pretrained language model to your task reliably — choosing a model, preparing data, hyperparameters, handling small data and instability, continued pretraining, and evaluating properly.

Intermediate⏱ 5 min#180