๐Ÿ“ˆ Machine Learning ยท Lecture 12 of 47

Evaluation Metrics for Classification: Accuracy, Precision, Recall and F1

Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.

A disease affects 1% of patients. A model that always predicts "healthy" achieves 99% accuracy โ€” and is completely useless. This is why professionals never report accuracy alone. Today we learn the vocabulary of classification metrics and, more importantly, how to choose the metric that matches the real costs of mistakes.

The confusion matrix#

For binary classification with a "positive" class (the thing we are trying to detect):

Predicted positivePredicted negative
Actually positiveTrue Positive (TP)False Negative (FN)
Actually negativeFalse Positive (FP)True Negative (TN)

Every metric is a function of these four numbers.

The core metrics#

$$ \text{Accuracy} = \frac{TP + TN}{TP + TN + FP + FN} $$
$$ \text{Precision} = \frac{TP}{TP + FP} \qquad \text{"When the model says positive, how often is it right?"} $$
$$ \text{Recall (Sensitivity, TPR)} = \frac{TP}{TP + FN} \qquad \text{"Of all real positives, how many did we find?"} $$
$$ \text{Specificity (TNR)} = \frac{TN}{TN + FP} \qquad \text{False Positive Rate} = 1 - \text{Specificity} $$
$$ F_1 = 2\cdot\frac{\text{Precision}\cdot\text{Recall}}{\text{Precision} + \text{Recall}} $$

$F_1$ is the harmonic mean of precision and recall. Unlike the arithmetic mean, it is low if either component is low, so a model cannot hide terrible recall behind high precision.

The generalised $F_\beta$ weights recall $\beta$ times as important as precision:

$$ F_\beta = (1 + \beta^2)\frac{\text{Precision}\cdot\text{Recall}}{\beta^2\,\text{Precision} + \text{Recall}} $$

Use $F_2$ when missing positives is worse; $F_{0.5}$ when false alarms are worse.

Precision versus recall: choosing by consequences#

ApplicationCostlier errorEmphasise
Cancer screeningMissing a cancer (FN)Recall
Spam filteringLosing a real email (FP)Precision
Fraud detectionDepends: investigation cost vs fraud lossPrecision at a fixed recall, or cost-based
Identifying vulnerable families for urgent follow-upMissing a family in need (FN)Recall, with capacity constraints on FP
Automated content removalRemoving legitimate speech (FP)Precision

Metrics for imbalanced data#

  • Balanced accuracy = (Recall + Specificity)/2 โ€” the average recall over classes. The "always healthy" model scores 0.5.
  • Matthews Correlation Coefficient (MCC):
$$ \text{MCC} = \frac{TP\cdot TN - FP\cdot FN}{\sqrt{(TP + FP)(TP + FN)(TN + FP)(TN + FN)}} $$

ranges from โˆ’1 to +1, uses all four cells and is informative even with severe imbalance.

  • Cohen's kappa โ€” agreement corrected for chance.

Multiclass averaging#

For $K$ classes, compute per-class precision/recall/F1 (each class vs the rest), then average:

  • Macro: unweighted mean over classes โ€” every class matters equally; exposes poor minority-class performance.
  • Weighted: mean weighted by class frequency.
  • Micro: pool all TP/FP/FN globally โ€” for single-label multiclass it equals accuracy.

Computing everything#

python
import numpy as np
from sklearn.metrics import (confusion_matrix, accuracy_score, precision_score, recall_score,
                             f1_score, balanced_accuracy_score, matthews_corrcoef,
                             classification_report)

rng = np.random.default_rng(0)
y_true = (rng.random(1000) < 0.05).astype(int)            # 5% positives
scores = np.where(y_true == 1, rng.beta(5, 2, 1000), rng.beta(2, 5, 1000))
y_pred = (scores >= 0.5).astype(int)

print(confusion_matrix(y_true, y_pred))
print("accuracy          ", round(accuracy_score(y_true, y_pred), 3))
print("precision         ", round(precision_score(y_true, y_pred), 3))
print("recall            ", round(recall_score(y_true, y_pred), 3))
print("F1                ", round(f1_score(y_true, y_pred), 3))
print("balanced accuracy ", round(balanced_accuracy_score(y_true, y_pred), 3))
print("MCC               ", round(matthews_corrcoef(y_true, y_pred), 3))
print("always-negative accuracy:", round(accuracy_score(y_true, np.zeros_like(y_true)), 3))
print(classification_report(y_true, y_pred, digits=3))

Notice that the trivial always-negative model has accuracy of about 95% here โ€” higher than you might expect from a real model โ€” while its recall, F1 and MCC are zero.

Cost-sensitive evaluation#

When you know the costs, use them directly. If a false negative costs $c_{FN}$ and a false positive $c_{FP}$:

$$ \text{Expected cost} = c_{FN}\cdot FN + c_{FP}\cdot FP $$

For a well-calibrated probabilistic classifier, the cost-minimising threshold is

$$ t^* = \frac{c_{FP}}{c_{FP} + c_{FN}} $$

If missing a positive is nine times worse than a false alarm, predict positive whenever $p \ge 0.1$.

Beyond aggregate metrics#

  • Slice metrics by subgroup (region, gender, language) to detect unequal performance.
  • Error analysis: read the actual false positives and negatives.
  • Confidence intervals: metrics from small test sets are noisy (see the hypothesis-testing lecture).
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Regression Metrics: MSE, RMSE, MAE, Rยฒ and Beyond

How good is a numeric prediction? We compare MSE, RMSE, MAE, Rยฒ, MAPE and quantile loss, explain how each responds to outliers and scale, and match metrics to real decisions.

Beginnerโฑ 5 min#063
๐Ÿ“ˆ Machine Learning

The Machine Learning Workflow: From Problem to Deployed Model

Successful ML projects follow a disciplined process โ€” problem framing, data collection, exploration, baselines, iteration, evaluation and deployment. We walk through it with a complete scikit-learn example.

Beginnerโฑ 5 min#052
๐Ÿ“ˆ Machine Learning

Train, Validation and Test Splits โ€” and Cross-Validation Done Right

Trustworthy evaluation starts with correct data splits. We cover hold-out validation, k-fold, stratified, group and time-series cross-validation, nested CV for tuning, and the leakage traps that invalidate results.

Beginnerโฑ 5 min#060