Evaluation is the compass of NLP research and engineering. A model can only improve on what we measure โ and it will happily exploit any weakness in the measurement. Generated text is especially hard to evaluate: there are many correct translations of a sentence and many good summaries of a document. This lecture surveys the toolbox and its limits.
Intrinsic vs extrinsic evaluation#
- Intrinsic: measures a component in isolation (perplexity of a language model, word-similarity correlation of embeddings).
- Extrinsic: measures impact on a downstream task or real user outcomes (does better perplexity improve speech recognition? do users resolve their question faster?).
Intrinsic metrics are cheap and fast; extrinsic ones are what ultimately matter.
Classification-style tasks#
Accuracy, precision, recall, F1 (macro for imbalance), entity-level F1 for NER, exact match and token F1 for extractive QA โ covered in earlier lectures. Always report confidence intervals or variance across seeds.
Perplexity#
For language models:
Lower is better. Comparable only for models with the same tokeniser (or when normalised per character/byte, e.g. bits per byte). Perplexity measures predictive fit, not helpfulness, truthfulness or safety.
Overlap-based generation metrics#
| Metric | Measures | Typical use |
|---|---|---|
| BLEU | Modified n-gram precision + brevity penalty | Machine translation |
| chrF | Character n-gram F-score | MT, morphologically rich languages |
| ROUGE-1/2/L | n-gram / LCS overlap (recall-oriented) | Summarisation |
| METEOR | Unigram matching with stems and synonyms | MT, captioning |
| CIDEr | TF-IDF-weighted n-gram similarity to many references | Image captioning |
They are cheap and reproducible, but reward surface overlap: they penalise valid paraphrases and cannot detect factual errors. They correlate reasonably with human judgement when comparing systems of very different quality, and poorly when comparing strong systems.
Embedding-based and learned metrics#
- BERTScore (Zhang et al., 2020): match each token of the candidate to its most similar token in the reference using contextual embeddings; compute precision, recall and F1 of these similarities. Captures paraphrases.
- COMET, BLEURT: neural models trained to predict human quality ratings โ substantially better correlation with human judgements in machine translation.
- Faithfulness metrics: NLI-based and QA-based consistency checks for summarisation.
# pip install bert-score sacrebleu
from bert_score import score
import sacrebleu
cands = ["The clinic will open at 9 am tomorrow."]
refs = ["Tomorrow the health centre opens at nine in the morning."]
P, R, F = score(cands, refs, lang="en")
print("BERTScore F1:", round(F.item(), 3), " BLEU:", round(sacrebleu.corpus_bleu(cands, [refs]).score, 1))BLEU is low (little exact overlap) while BERTScore is high (the meaning matches).
LLM-as-a-judge#
Large language models can grade outputs against rubrics or compare two responses pairwise. This scales evaluation of open-ended generation and correlates well with human preferences on many tasks (e.g. studies accompanying MT-Bench and Chatbot Arena). Known biases:
- Position bias โ preferring the first (or second) answer; mitigate by swapping order.
- Verbosity bias โ preferring longer answers.
- Self-preference โ favouring outputs from the same model family.
- Limited reliability on specialised or factual content the judge itself gets wrong.
Validate LLM judges against human labels on a sample before trusting them.
Human evaluation#
The gold standard, but it must be designed carefully:
- Clear criteria (adequacy, fluency, faithfulness, helpfulness, harmlessness) with rubrics.
- Trained annotators; measure inter-annotator agreement.
- Pairwise comparisons are often more reliable than absolute scores.
- Blind the evaluators to which system produced which output; randomise order.
- Sufficient sample size for statistical significance.
- Domain experts for specialised content; native speakers for each language.
Benchmarks and their pitfalls#
Benchmarks (GLUE, SuperGLUE, SQuAD, MMLU, HellaSwag, BIG-bench, HELM and many more) drive progress but have recurring problems:
- Saturation: models reach or exceed human baselines quickly, often before the underlying capability is solved.
- Annotation artefacts and shortcuts: e.g. in natural-language inference datasets, hypothesis-only models performed far above chance because words like "not" correlated with contradiction.
- Contamination: test items appear in web-scale training data, inflating scores.
- Narrowness: benchmarks cover a small slice of real use, often English-centric.
- Goodhart's law: "When a measure becomes a target, it ceases to be a good measure."
Responses include adversarial and dynamic benchmarks (Adversarial NLI, Dynabench), held-out private test sets, contamination checks, holistic evaluations across many metrics (HELM), and โ most important for practitioners โ custom evaluation sets built from your own use case.