∑ Mathematics for ML · Lecture 20 of 25

Statistical Hypothesis Testing for ML: Is Model A Really Better?

A 1% accuracy gain may be noise. We cover confidence intervals, p-values, paired tests, McNemar's test, the bootstrap and multiple-comparison pitfalls so your experimental claims hold up.

Your new model scores 87.4% accuracy; the baseline scores 86.9%. Is your model better — or did you get lucky with this test set, this random seed, this data split? Many published ML "improvements" fail to replicate because authors never asked this question. Today we learn the statistical tools to answer it honestly.

The logic of hypothesis testing#

  1. State a null hypothesis $H_0$: "the two models have the same true accuracy."
  2. Choose a test statistic that measures the observed difference.
  3. Compute the p-value: the probability, assuming $H_0$ is true, of observing a difference at least as extreme as the one you saw.
  4. If $p < \alpha$ (commonly 0.05), reject $H_0$.

Two error types:

$H_0$ true$H_0$ false
Reject $H_0$Type I error (false positive), rate $\alpha$Correct (power $= 1 - \beta$)
Fail to rejectCorrectType II error (false negative), rate $\beta$

Confidence intervals for accuracy#

Accuracy on $n$ independent test examples is a binomial proportion. A normal-approximation 95% confidence interval is

$$ \hat{p} \pm 1.96\sqrt{\frac{\hat{p}(1 - \hat{p})}{n}} $$

For $\hat{p} = 0.87$ and $n = 1000$: $0.87 \pm 0.021$. A 0.5% difference between two models is well inside this uncertainty. (For small $n$ or extreme $\hat{p}$, use the Wilson interval.)

Paired comparisons: McNemar's test#

When two classifiers are evaluated on the same test set, their errors are correlated. Comparing their confidence intervals independently wastes information. McNemar's test focuses on disagreements:

B correctB wrong
A correct$n_{11}$$n_{10}$
A wrong$n_{01}$$n_{00}$

Under $H_0$, disagreements are equally likely in both directions. The statistic

$$ \chi^2 = \frac{(|n_{10} - n_{01}| - 1)^2}{n_{10} + n_{01}} $$

is approximately $\chi^2$ with one degree of freedom. For small counts, use the exact binomial test on $n_{10}$ out of $n_{10} + n_{01}$.

Comparing across data splits and seeds#

Performance also varies with random initialisation and data splits. Good practice:

  • Run each method with multiple seeds (e.g. 5–10) and report mean ± standard deviation.
  • For cross-validation, use a paired t-test across folds — but note that folds share training data, which violates independence and inflates false positives. The 5×2 cross-validation paired t-test (Dietterich) and the corrected resampled t-test (Nadeau & Bengio) address this.
  • For many datasets, compare methods with the Wilcoxon signed-rank test, which does not assume normality.

The bootstrap: a universal tool#

When no formula exists (for F1, AUC, BLEU…), use the bootstrap: resample the test set with replacement many times, recompute the metric, and read confidence intervals from the empirical distribution.

python
import numpy as np

rng = np.random.default_rng(0)
n = 1000
y_true = rng.integers(0, 2, n)
pred_a = np.where(rng.random(n) < 0.87, y_true, 1 - y_true)   # ~87% accurate
pred_b = np.where(rng.random(n) < 0.86, y_true, 1 - y_true)   # ~86% accurate

diffs = []
for _ in range(10_000):
    idx = rng.integers(0, n, n)                # paired resampling: same indices for both
    diffs.append((pred_a[idx] == y_true[idx]).mean() - (pred_b[idx] == y_true[idx]).mean())
lo, hi = np.percentile(diffs, [2.5, 97.5])
print(f"accuracy difference: {np.mean(diffs):.4f}, 95% CI [{lo:.4f}, {hi:.4f}]")

# McNemar exact test
from scipy.stats import binomtest
a_ok, b_ok = pred_a == y_true, pred_b == y_true
n10, n01 = int((a_ok & ~b_ok).sum()), int((~a_ok & b_ok).sum())
print("McNemar p-value:", binomtest(n10, n10 + n01, 0.5).pvalue)

If the interval for the difference includes zero, you cannot claim an improvement.

Multiple comparisons and the garden of forking paths#

Test 20 hyperparameter settings at $\alpha = 0.05$ and, on average, one will look "significant" by chance. Corrections:

  • Bonferroni: use $\alpha/m$ for $m$ tests (conservative).
  • Holm–Bonferroni: a uniformly more powerful step-down version.
  • Benjamini–Hochberg: controls the false discovery rate.

More insidious is test-set overfitting: repeatedly evaluating on the test set and keeping what works turns the test set into a training set. Use a validation set for all decisions and touch the test set once at the end.

A checklist for claiming "Model A beats Model B"#

  1. Same data splits, same preprocessing, comparable tuning budgets.
  2. Multiple seeds; report mean, standard deviation and confidence intervals.
  3. Paired test (McNemar, bootstrap on differences, Wilcoxon).
  4. Correct for multiple comparisons.
  5. Report the effect size and discuss practical significance.
  6. The test set was used only once.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

MAP Estimation, Priors and the Bayesian View of Regularisation

Adding a prior to maximum likelihood gives MAP estimation — and reveals that L2 and L1 regularisation are Gaussian and Laplace priors. We also meet conjugate priors and full Bayesian inference.

Intermediate⏱ 5 min#039
∑ Mathematics for ML

Maximum Likelihood Estimation: How Models Learn from Data

Most loss functions in ML are negative log-likelihoods in disguise. We define MLE, derive estimators for Bernoulli and Gaussian models, and prove that MSE and cross-entropy arise from maximum likelihood.

Intermediate⏱ 5 min#038
∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036