A disease affects 1% of patients. A model that always predicts "healthy" achieves 99% accuracy โ and is completely useless. This is why professionals never report accuracy alone. Today we learn the vocabulary of classification metrics and, more importantly, how to choose the metric that matches the real costs of mistakes.
The confusion matrix#
For binary classification with a "positive" class (the thing we are trying to detect):
| Predicted positive | Predicted negative | |
|---|---|---|
| Actually positive | True Positive (TP) | False Negative (FN) |
| Actually negative | False Positive (FP) | True Negative (TN) |
Every metric is a function of these four numbers.
The core metrics#
$F_1$ is the harmonic mean of precision and recall. Unlike the arithmetic mean, it is low if either component is low, so a model cannot hide terrible recall behind high precision.
The generalised $F_\beta$ weights recall $\beta$ times as important as precision:
Use $F_2$ when missing positives is worse; $F_{0.5}$ when false alarms are worse.
Precision versus recall: choosing by consequences#
| Application | Costlier error | Emphasise |
|---|---|---|
| Cancer screening | Missing a cancer (FN) | Recall |
| Spam filtering | Losing a real email (FP) | Precision |
| Fraud detection | Depends: investigation cost vs fraud loss | Precision at a fixed recall, or cost-based |
| Identifying vulnerable families for urgent follow-up | Missing a family in need (FN) | Recall, with capacity constraints on FP |
| Automated content removal | Removing legitimate speech (FP) | Precision |
Metrics for imbalanced data#
- Balanced accuracy = (Recall + Specificity)/2 โ the average recall over classes. The "always healthy" model scores 0.5.
- Matthews Correlation Coefficient (MCC):
ranges from โ1 to +1, uses all four cells and is informative even with severe imbalance.
- Cohen's kappa โ agreement corrected for chance.
Multiclass averaging#
For $K$ classes, compute per-class precision/recall/F1 (each class vs the rest), then average:
- Macro: unweighted mean over classes โ every class matters equally; exposes poor minority-class performance.
- Weighted: mean weighted by class frequency.
- Micro: pool all TP/FP/FN globally โ for single-label multiclass it equals accuracy.
Computing everything#
import numpy as np
from sklearn.metrics import (confusion_matrix, accuracy_score, precision_score, recall_score,
f1_score, balanced_accuracy_score, matthews_corrcoef,
classification_report)
rng = np.random.default_rng(0)
y_true = (rng.random(1000) < 0.05).astype(int) # 5% positives
scores = np.where(y_true == 1, rng.beta(5, 2, 1000), rng.beta(2, 5, 1000))
y_pred = (scores >= 0.5).astype(int)
print(confusion_matrix(y_true, y_pred))
print("accuracy ", round(accuracy_score(y_true, y_pred), 3))
print("precision ", round(precision_score(y_true, y_pred), 3))
print("recall ", round(recall_score(y_true, y_pred), 3))
print("F1 ", round(f1_score(y_true, y_pred), 3))
print("balanced accuracy ", round(balanced_accuracy_score(y_true, y_pred), 3))
print("MCC ", round(matthews_corrcoef(y_true, y_pred), 3))
print("always-negative accuracy:", round(accuracy_score(y_true, np.zeros_like(y_true)), 3))
print(classification_report(y_true, y_pred, digits=3))Notice that the trivial always-negative model has accuracy of about 95% here โ higher than you might expect from a real model โ while its recall, F1 and MCC are zero.
Cost-sensitive evaluation#
When you know the costs, use them directly. If a false negative costs $c_{FN}$ and a false positive $c_{FP}$:
For a well-calibrated probabilistic classifier, the cost-minimising threshold is
If missing a positive is nine times worse than a false alarm, predict positive whenever $p \ge 0.1$.
Beyond aggregate metrics#
- Slice metrics by subgroup (region, gender, language) to detect unequal performance.
- Error analysis: read the actual false positives and negatives.
- Confidence intervals: metrics from small test sets are noisy (see the hypothesis-testing lecture).