Your new model scores 87.4% accuracy; the baseline scores 86.9%. Is your model better — or did you get lucky with this test set, this random seed, this data split? Many published ML "improvements" fail to replicate because authors never asked this question. Today we learn the statistical tools to answer it honestly.
The logic of hypothesis testing#
- State a null hypothesis $H_0$: "the two models have the same true accuracy."
- Choose a test statistic that measures the observed difference.
- Compute the p-value: the probability, assuming $H_0$ is true, of observing a difference at least as extreme as the one you saw.
- If $p < \alpha$ (commonly 0.05), reject $H_0$.
Two error types:
| $H_0$ true | $H_0$ false | |
|---|---|---|
| Reject $H_0$ | Type I error (false positive), rate $\alpha$ | Correct (power $= 1 - \beta$) |
| Fail to reject | Correct | Type II error (false negative), rate $\beta$ |
Confidence intervals for accuracy#
Accuracy on $n$ independent test examples is a binomial proportion. A normal-approximation 95% confidence interval is
For $\hat{p} = 0.87$ and $n = 1000$: $0.87 \pm 0.021$. A 0.5% difference between two models is well inside this uncertainty. (For small $n$ or extreme $\hat{p}$, use the Wilson interval.)
Paired comparisons: McNemar's test#
When two classifiers are evaluated on the same test set, their errors are correlated. Comparing their confidence intervals independently wastes information. McNemar's test focuses on disagreements:
| B correct | B wrong | |
|---|---|---|
| A correct | $n_{11}$ | $n_{10}$ |
| A wrong | $n_{01}$ | $n_{00}$ |
Under $H_0$, disagreements are equally likely in both directions. The statistic
is approximately $\chi^2$ with one degree of freedom. For small counts, use the exact binomial test on $n_{10}$ out of $n_{10} + n_{01}$.
Comparing across data splits and seeds#
Performance also varies with random initialisation and data splits. Good practice:
- Run each method with multiple seeds (e.g. 5–10) and report mean ± standard deviation.
- For cross-validation, use a paired t-test across folds — but note that folds share training data, which violates independence and inflates false positives. The 5×2 cross-validation paired t-test (Dietterich) and the corrected resampled t-test (Nadeau & Bengio) address this.
- For many datasets, compare methods with the Wilcoxon signed-rank test, which does not assume normality.
The bootstrap: a universal tool#
When no formula exists (for F1, AUC, BLEU…), use the bootstrap: resample the test set with replacement many times, recompute the metric, and read confidence intervals from the empirical distribution.
import numpy as np
rng = np.random.default_rng(0)
n = 1000
y_true = rng.integers(0, 2, n)
pred_a = np.where(rng.random(n) < 0.87, y_true, 1 - y_true) # ~87% accurate
pred_b = np.where(rng.random(n) < 0.86, y_true, 1 - y_true) # ~86% accurate
diffs = []
for _ in range(10_000):
idx = rng.integers(0, n, n) # paired resampling: same indices for both
diffs.append((pred_a[idx] == y_true[idx]).mean() - (pred_b[idx] == y_true[idx]).mean())
lo, hi = np.percentile(diffs, [2.5, 97.5])
print(f"accuracy difference: {np.mean(diffs):.4f}, 95% CI [{lo:.4f}, {hi:.4f}]")
# McNemar exact test
from scipy.stats import binomtest
a_ok, b_ok = pred_a == y_true, pred_b == y_true
n10, n01 = int((a_ok & ~b_ok).sum()), int((~a_ok & b_ok).sum())
print("McNemar p-value:", binomtest(n10, n10 + n01, 0.5).pvalue)If the interval for the difference includes zero, you cannot claim an improvement.
Multiple comparisons and the garden of forking paths#
Test 20 hyperparameter settings at $\alpha = 0.05$ and, on average, one will look "significant" by chance. Corrections:
- Bonferroni: use $\alpha/m$ for $m$ tests (conservative).
- Holm–Bonferroni: a uniformly more powerful step-down version.
- Benjamini–Hochberg: controls the false discovery rate.
More insidious is test-set overfitting: repeatedly evaluating on the test set and keeping what works turns the test set into a training set. Use a validation set for all decisions and touch the test set once at the end.
A checklist for claiming "Model A beats Model B"#
- Same data splits, same preprocessing, comparable tuning budgets.
- Multiple seeds; report mean, standard deviation and confidence intervals.
- Paired test (McNemar, bootstrap on differences, Wilcoxon).
- Correct for multiple comparisons.
- Report the effect size and discuss practical significance.
- The test set was used only once.