📈 Machine Learning · Lecture 13 of 47

ROC Curves, AUC and Precision–Recall Curves

A classifier's score supports many thresholds. ROC and precision–recall curves summarise all of them. We construct both, interpret AUC probabilistically, and learn when each curve tells the truth.

Most classifiers output a score or probability, and we choose a threshold to turn it into a decision. Different thresholds give different confusion matrices. Rather than evaluating one arbitrary threshold, we can evaluate the ranking quality of the scores across all thresholds. That is what ROC and precision–recall curves do.

The ROC curve#

The Receiver Operating Characteristic curve (originally from radar signal detection in World War II) plots, for every threshold:

  • the True Positive Rate (recall) $TPR = \frac{TP}{TP + FN}$ on the y-axis,
  • against the False Positive Rate $FPR = \frac{FP}{FP + TN}$ on the x-axis.

As the threshold decreases from $+\infty$ to $-\infty$, the curve moves from $(0, 0)$ (predict nothing positive) to $(1, 1)$ (predict everything positive).

  • A perfect classifier passes through $(0, 1)$.
  • A random classifier follows the diagonal $TPR = FPR$.
  • A curve below the diagonal indicates a classifier whose scores are inverted.

AUC: area under the ROC curve#

The AUC (or ROC-AUC) summarises the curve with one number between 0 and 1. It has a beautiful probabilistic interpretation:

$$ \text{AUC} = P\big(s(\mathbf{x}^+) > s(\mathbf{x}^-)\big) $$

— the probability that a randomly chosen positive example receives a higher score than a randomly chosen negative one. It is equivalent to the normalised Mann–Whitney U statistic. AUC = 0.5 means random ranking; AUC = 1.0 means perfect ranking.

Properties:

  • Threshold-independent — evaluates ranking, not a particular decision.
  • Invariant to class imbalance in the sense that TPR and FPR are each computed within one class.
  • Invariant to monotone transformations of scores — it ignores calibration entirely.

Why ROC can mislead under heavy imbalance#

Suppose 10 positives and 100,000 negatives. An FPR of 1% sounds tiny — but it means 1,000 false positives against at most 10 true positives. Precision is below 1%. The ROC curve looks excellent while the model is operationally useless, because FPR's denominator (all negatives) is huge.

The precision–recall curve#

The PR curve plots precision against recall for every threshold. Because precision's denominator includes false positives directly, the PR curve is sensitive to imbalance and shows the practical cost of catching more positives.

  • The baseline for a random classifier is a horizontal line at the positive prevalence $\pi = P(y = 1)$ — not 0.5.
  • The area under the PR curve is summarised as Average Precision (AP).

Computing and plotting#

python
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_curve, roc_auc_score, precision_recall_curve, average_precision_score

X, y = make_classification(n_samples=20_000, n_features=20, weights=[0.98, 0.02],
                           class_sep=0.8, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, test_size=0.3, random_state=0)

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(11, 4.5))
for name, m in [("Logistic", LogisticRegression(max_iter=1000)),
                ("Random forest", RandomForestClassifier(n_estimators=300, random_state=0))]:
    s = m.fit(X_tr, y_tr).predict_proba(X_te)[:, 1]
    fpr, tpr, _ = roc_curve(y_te, s)
    prec, rec, _ = precision_recall_curve(y_te, s)
    ax1.plot(fpr, tpr, label=f"{name} (AUC={roc_auc_score(y_te, s):.3f})")
    ax2.plot(rec, prec, label=f"{name} (AP={average_precision_score(y_te, s):.3f})")
ax1.plot([0, 1], [0, 1], "k--", lw=1); ax1.set_xlabel("FPR"); ax1.set_ylabel("TPR"); ax1.legend()
ax2.axhline(y_te.mean(), color="k", ls="--", lw=1, label="prevalence")
ax2.set_xlabel("Recall"); ax2.set_ylabel("Precision"); ax2.legend()
plt.tight_layout(); plt.show()

Run it: both models may have high ROC-AUC, yet their AP values are far lower and more clearly separated — the PR curve reveals the difficulty of the rare class.

Choosing an operating point#

A curve is not a decision. To deploy, pick a threshold using:

  • A constraint: "recall must be at least 90%" → choose the threshold achieving that with the highest precision.
  • Capacity: "our team can review 200 cases per day" → choose the threshold that flags 200 cases (precision at top-k).
  • Costs: minimise expected cost as shown in the previous lecture.
  • Youden's J statistic $\max(TPR - FPR)$ — a common default in medicine.

Always choose the threshold on the validation set, then report test performance at that fixed threshold.

Calibration is a separate question#

A model can rank perfectly (AUC = 1) while its probabilities are badly miscalibrated (predicting 0.6 for every positive and 0.4 for every negative). If decisions depend on the value of the probability — expected-cost thresholds, risk communication — check calibration with reliability diagrams and the Brier score $\frac{1}{n}\sum_i(p_i - y_i)^2$, and recalibrate with Platt scaling or isotonic regression if needed.

Multiclass extensions#

For $K$ classes, compute one-vs-rest curves per class and average (macro or weighted), or report per-class curves. Scikit-learn's roc_auc_score(..., multi_class="ovr") implements this.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Evaluation Metrics for Classification: Accuracy, Precision, Recall and F1

Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.

Beginner⏱ 5 min#061
📈 Machine Learning

Regression Metrics: MSE, RMSE, MAE, R² and Beyond

How good is a numeric prediction? We compare MSE, RMSE, MAE, R², MAPE and quantile loss, explain how each responds to outliers and scale, and match metrics to real decisions.

Beginner⏱ 5 min#063
📈 Machine Learning

Train, Validation and Test Splits — and Cross-Validation Done Right

Trustworthy evaluation starts with correct data splits. We cover hold-out validation, k-fold, stratified, group and time-series cross-validation, nested CV for tuning, and the leakage traps that invalidate results.

Beginner⏱ 5 min#060