๐Ÿ’ฌ NLP & Transformers ยท Lecture 22 of 29

Evaluating NLP Systems: Perplexity, BLEU, ROUGE, BERTScore and Human Judgement

How do we know if a language system is good? We survey intrinsic and extrinsic evaluation, overlap metrics, embedding-based metrics, learned metrics, LLM judges, human evaluation, benchmarks and their pitfalls.

Evaluation is the compass of NLP research and engineering. A model can only improve on what we measure โ€” and it will happily exploit any weakness in the measurement. Generated text is especially hard to evaluate: there are many correct translations of a sentence and many good summaries of a document. This lecture surveys the toolbox and its limits.

Intrinsic vs extrinsic evaluation#

  • Intrinsic: measures a component in isolation (perplexity of a language model, word-similarity correlation of embeddings).
  • Extrinsic: measures impact on a downstream task or real user outcomes (does better perplexity improve speech recognition? do users resolve their question faster?).

Intrinsic metrics are cheap and fast; extrinsic ones are what ultimately matter.

Classification-style tasks#

Accuracy, precision, recall, F1 (macro for imbalance), entity-level F1 for NER, exact match and token F1 for extractive QA โ€” covered in earlier lectures. Always report confidence intervals or variance across seeds.

Perplexity#

For language models:

$$ \text{PPL} = \exp\left(-\frac{1}{N}\sum_{i=1}^{N}\log P(w_i \mid w_{<i})\right) $$

Lower is better. Comparable only for models with the same tokeniser (or when normalised per character/byte, e.g. bits per byte). Perplexity measures predictive fit, not helpfulness, truthfulness or safety.

Overlap-based generation metrics#

MetricMeasuresTypical use
BLEUModified n-gram precision + brevity penaltyMachine translation
chrFCharacter n-gram F-scoreMT, morphologically rich languages
ROUGE-1/2/Ln-gram / LCS overlap (recall-oriented)Summarisation
METEORUnigram matching with stems and synonymsMT, captioning
CIDErTF-IDF-weighted n-gram similarity to many referencesImage captioning

They are cheap and reproducible, but reward surface overlap: they penalise valid paraphrases and cannot detect factual errors. They correlate reasonably with human judgement when comparing systems of very different quality, and poorly when comparing strong systems.

Embedding-based and learned metrics#

  • BERTScore (Zhang et al., 2020): match each token of the candidate to its most similar token in the reference using contextual embeddings; compute precision, recall and F1 of these similarities. Captures paraphrases.
  • COMET, BLEURT: neural models trained to predict human quality ratings โ€” substantially better correlation with human judgements in machine translation.
  • Faithfulness metrics: NLI-based and QA-based consistency checks for summarisation.
python
# pip install bert-score sacrebleu
from bert_score import score
import sacrebleu
cands = ["The clinic will open at 9 am tomorrow."]
refs = ["Tomorrow the health centre opens at nine in the morning."]
P, R, F = score(cands, refs, lang="en")
print("BERTScore F1:", round(F.item(), 3), " BLEU:", round(sacrebleu.corpus_bleu(cands, [refs]).score, 1))

BLEU is low (little exact overlap) while BERTScore is high (the meaning matches).

LLM-as-a-judge#

Large language models can grade outputs against rubrics or compare two responses pairwise. This scales evaluation of open-ended generation and correlates well with human preferences on many tasks (e.g. studies accompanying MT-Bench and Chatbot Arena). Known biases:

  • Position bias โ€” preferring the first (or second) answer; mitigate by swapping order.
  • Verbosity bias โ€” preferring longer answers.
  • Self-preference โ€” favouring outputs from the same model family.
  • Limited reliability on specialised or factual content the judge itself gets wrong.

Validate LLM judges against human labels on a sample before trusting them.

Human evaluation#

The gold standard, but it must be designed carefully:

  • Clear criteria (adequacy, fluency, faithfulness, helpfulness, harmlessness) with rubrics.
  • Trained annotators; measure inter-annotator agreement.
  • Pairwise comparisons are often more reliable than absolute scores.
  • Blind the evaluators to which system produced which output; randomise order.
  • Sufficient sample size for statistical significance.
  • Domain experts for specialised content; native speakers for each language.

Benchmarks and their pitfalls#

Benchmarks (GLUE, SuperGLUE, SQuAD, MMLU, HellaSwag, BIG-bench, HELM and many more) drive progress but have recurring problems:

  • Saturation: models reach or exceed human baselines quickly, often before the underlying capability is solved.
  • Annotation artefacts and shortcuts: e.g. in natural-language inference datasets, hypothesis-only models performed far above chance because words like "not" correlated with contradiction.
  • Contamination: test items appear in web-scale training data, inflating scores.
  • Narrowness: benchmarks cover a small slice of real use, often English-centric.
  • Goodhart's law: "When a measure becomes a target, it ceases to be a good measure."

Responses include adversarial and dynamic benchmarks (Adversarial NLI, Dynabench), held-out private test sets, contamination checks, holistic evaluations across many metrics (HELM), and โ€” most important for practitioners โ€” custom evaluation sets built from your own use case.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Text Summarisation: Extractive and Abstractive Methods

Summarisation condenses documents while preserving key information. We compare extractive methods (TextRank) with abstractive neural models, evaluate with ROUGE and factual-consistency checks, and discuss long documents and hallucination.

Intermediateโฑ 5 min#182
๐Ÿ’ฌ NLP & Transformers

Neural Machine Translation: From Seq2Seq to Multilingual Transformers

Machine translation is one of NLP's oldest and most impactful tasks. We trace its evolution to neural systems, cover training data, subword vocabularies, back-translation, evaluation with BLEU and COMET, and the challenges of low-resource languages.

Intermediateโฑ 5 min#173
๐Ÿ’ฌ NLP & Transformers

Text Classification: From Linear Models to Fine-Tuned Transformers

Text classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.

Beginnerโฑ 4 min#169