Most classifiers output a score or probability, and we choose a threshold to turn it into a decision. Different thresholds give different confusion matrices. Rather than evaluating one arbitrary threshold, we can evaluate the ranking quality of the scores across all thresholds. That is what ROC and precision–recall curves do.
The ROC curve#
The Receiver Operating Characteristic curve (originally from radar signal detection in World War II) plots, for every threshold:
- the True Positive Rate (recall) $TPR = \frac{TP}{TP + FN}$ on the y-axis,
- against the False Positive Rate $FPR = \frac{FP}{FP + TN}$ on the x-axis.
As the threshold decreases from $+\infty$ to $-\infty$, the curve moves from $(0, 0)$ (predict nothing positive) to $(1, 1)$ (predict everything positive).
- A perfect classifier passes through $(0, 1)$.
- A random classifier follows the diagonal $TPR = FPR$.
- A curve below the diagonal indicates a classifier whose scores are inverted.
AUC: area under the ROC curve#
The AUC (or ROC-AUC) summarises the curve with one number between 0 and 1. It has a beautiful probabilistic interpretation:
— the probability that a randomly chosen positive example receives a higher score than a randomly chosen negative one. It is equivalent to the normalised Mann–Whitney U statistic. AUC = 0.5 means random ranking; AUC = 1.0 means perfect ranking.
Properties:
- Threshold-independent — evaluates ranking, not a particular decision.
- Invariant to class imbalance in the sense that TPR and FPR are each computed within one class.
- Invariant to monotone transformations of scores — it ignores calibration entirely.
Why ROC can mislead under heavy imbalance#
Suppose 10 positives and 100,000 negatives. An FPR of 1% sounds tiny — but it means 1,000 false positives against at most 10 true positives. Precision is below 1%. The ROC curve looks excellent while the model is operationally useless, because FPR's denominator (all negatives) is huge.
The precision–recall curve#
The PR curve plots precision against recall for every threshold. Because precision's denominator includes false positives directly, the PR curve is sensitive to imbalance and shows the practical cost of catching more positives.
- The baseline for a random classifier is a horizontal line at the positive prevalence $\pi = P(y = 1)$ — not 0.5.
- The area under the PR curve is summarised as Average Precision (AP).
Computing and plotting#
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_curve, roc_auc_score, precision_recall_curve, average_precision_score
X, y = make_classification(n_samples=20_000, n_features=20, weights=[0.98, 0.02],
class_sep=0.8, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, test_size=0.3, random_state=0)
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(11, 4.5))
for name, m in [("Logistic", LogisticRegression(max_iter=1000)),
("Random forest", RandomForestClassifier(n_estimators=300, random_state=0))]:
s = m.fit(X_tr, y_tr).predict_proba(X_te)[:, 1]
fpr, tpr, _ = roc_curve(y_te, s)
prec, rec, _ = precision_recall_curve(y_te, s)
ax1.plot(fpr, tpr, label=f"{name} (AUC={roc_auc_score(y_te, s):.3f})")
ax2.plot(rec, prec, label=f"{name} (AP={average_precision_score(y_te, s):.3f})")
ax1.plot([0, 1], [0, 1], "k--", lw=1); ax1.set_xlabel("FPR"); ax1.set_ylabel("TPR"); ax1.legend()
ax2.axhline(y_te.mean(), color="k", ls="--", lw=1, label="prevalence")
ax2.set_xlabel("Recall"); ax2.set_ylabel("Precision"); ax2.legend()
plt.tight_layout(); plt.show()Run it: both models may have high ROC-AUC, yet their AP values are far lower and more clearly separated — the PR curve reveals the difficulty of the rare class.
Choosing an operating point#
A curve is not a decision. To deploy, pick a threshold using:
- A constraint: "recall must be at least 90%" → choose the threshold achieving that with the highest precision.
- Capacity: "our team can review 200 cases per day" → choose the threshold that flags 200 cases (precision at top-k).
- Costs: minimise expected cost as shown in the previous lecture.
- Youden's J statistic $\max(TPR - FPR)$ — a common default in medicine.
Always choose the threshold on the validation set, then report test performance at that fixed threshold.
Calibration is a separate question#
A model can rank perfectly (AUC = 1) while its probabilities are badly miscalibrated (predicting 0.6 for every positive and 0.4 for every negative). If decisions depend on the value of the probability — expected-cost thresholds, risk communication — check calibration with reliability diagrams and the Brier score $\frac{1}{n}\sum_i(p_i - y_i)^2$, and recalibrate with Platt scaling or isotonic regression if needed.
Multiclass extensions#
For $K$ classes, compute one-vs-rest curves per class and average (macro or weighted), or report per-class curves. Scikit-learn's roc_auc_score(..., multi_class="ovr") implements this.