Analysts face more reports, articles, meeting transcripts and case notes than anyone can read. Automatic summarisation produces a shorter version that preserves the most important information. It is among the most useful — and most error-prone — NLP applications, because a fluent summary that states something false can be worse than no summary at all.
Two paradigms#
- Extractive: select the most important sentences from the source and concatenate them. Faithful by construction (every sentence appears in the source), but can be choppy and redundant.
- Abstractive: generate new text that paraphrases and condenses. More fluent and concise, but risks hallucination — content not supported by the source.
Summaries can also be single-document or multi-document, generic or query-focused ("summarise what this report says about water supply"), and targeted to different lengths and audiences.
Extractive methods#
Frequency-based (Luhn, 1958): score sentences by the frequency of their important words.
TextRank (Mihalcea & Tarau, 2004): build a graph whose nodes are sentences and whose edge weights are sentence similarities; run PageRank; pick the top-ranked sentences. Sentences similar to many others are "central".
import numpy as np, re
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
def textrank_summary(text, k=2, d=0.85, iters=50):
sents = [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]
X = TfidfVectorizer().fit_transform(sents)
S = cosine_similarity(X); np.fill_diagonal(S, 0)
S = S / (S.sum(1, keepdims=True) + 1e-9) # row-normalise into transition probabilities
r = np.ones(len(sents)) / len(sents)
for _ in range(iters):
r = (1 - d) / len(sents) + d * S.T @ r # PageRank iteration
top = sorted(np.argsort(-r)[:k]) # keep original order
return " ".join(sents[i] for i in top)
doc = ("Floods affected five districts this week. Around 20,000 families were displaced. "
"Authorities opened 40 temporary shelters in schools. Clean water is the most urgent need. "
"Health workers warned of possible cholera outbreaks. Schools will remain closed until Sunday.")
print(textrank_summary(doc))Supervised extractive models (e.g. BERTSum) classify each sentence as "include" or not, trained on labels derived from reference summaries.
Abstractive methods#
Encoder–decoder transformers fine-tuned on (document, summary) pairs: BART, PEGASUS (pretrained with "gap sentence generation" — masking whole important sentences and generating them, closely matching the summarisation task), T5, and — increasingly — instruction-tuned LLMs prompted to summarise.
Common datasets: CNN/DailyMail (news highlights, fairly extractive), XSum (single-sentence, highly abstractive BBC summaries), arXiv/PubMed (long scientific papers), and meeting and dialogue summarisation sets.
Evaluation#
ROUGE (Lin, 2004) measures n-gram overlap with reference summaries, emphasising recall:
- ROUGE-1 / ROUGE-2: unigram / bigram overlap.
- ROUGE-L: longest common subsequence.
ROUGE is cheap but rewards word overlap, not correctness: a summary can score well while stating something false, or score poorly while being an excellent paraphrase.
# pip install rouge-score
from rouge_score import rouge_scorer
sc = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL"], use_stemmer=True)
ref = "About 20,000 families were displaced by floods; clean water is the most urgent need."
hyp = "Floods displaced 20,000 families, and clean water is urgently needed."
print({k: round(v.fmeasure, 3) for k, v in sc.score(ref, hyp).items()})Factual consistency (faithfulness) requires separate evaluation:
- Entailment-based checks: does the source entail each summary sentence (using an NLI model)? (e.g. SummaC)
- QA-based checks: generate questions from the summary, answer them from the source, and compare (QAGS, QuestEval).
- LLM-as-judge evaluations with explicit rubrics — useful at scale, but validate against human judgements.
- Human evaluation: coverage, faithfulness, fluency, conciseness.
Studies found that a substantial fraction of abstractive summaries from earlier neural models contained unsupported information (Maynez et al., 2020). Modern LLMs are better, but not immune.
Long documents#
Transformers have limited context windows. Strategies:
- Truncation (loses information);
- Chunk-and-merge / map-reduce: summarise chunks, then summarise the summaries;
- Extract-then-abstract: select important passages first;
- Long-context models (Longformer/LED, long-context LLMs) — but check whether information from the middle of long inputs is used.