A new recommendation model improves offline accuracy by 3%. Will users actually find what they need faster? A new triage model has better AUC. Will caseworkers resolve urgent cases sooner? Offline metrics are proxies; the real question is causal: what happens to outcomes when we deploy this model? Randomised online experiments — A/B tests — answer it with the same logic as clinical trials.
Why offline evaluation is not enough#
- Offline data reflects the old system's behaviour (recommendations shape clicks), creating feedback loops.
- Proxy metrics (accuracy, AUC) may not map to outcomes (user satisfaction, time saved, harm avoided).
- Real-world effects include user adaptation, latency, UI interaction and workflow changes.
The basic design#
- Randomise units (users, sessions, households, clinics) into control (current model, A) and treatment (new model, B).
- Run both concurrently for a pre-specified period.
- Compare a pre-registered primary metric between groups.
Randomisation makes the groups comparable in expectation, so a significant difference can be attributed to the model.
Choosing metrics#
- Primary (decision) metric: the outcome you care about — e.g. proportion of urgent cases handled within 48 hours.
- Guardrail metrics: things that must not get worse — latency, error rates, complaint rates, fairness across groups, cost.
- Diagnostic metrics: help explain results (clicks, model confidence distributions).
Specify metrics, analysis method and stopping rules before starting (pre-registration) to avoid cherry-picking.
Sample size and power#
To detect a difference between two proportions $p_A$ and $p_B$ with significance level $\alpha$ and power $1 - \beta$, the required sample size per group is approximately
Small effects require large samples: halving the detectable effect quadruples $n$.
from scipy.stats import norm
import numpy as np
from statsmodels.stats.proportion import proportions_ztest
def sample_size(p_a, p_b, alpha=0.05, power=0.8):
z = norm.ppf(1 - alpha / 2) + norm.ppf(power)
return int(np.ceil(z**2 * (p_a * (1 - p_a) + p_b * (1 - p_b)) / (p_b - p_a) ** 2))
print("per-group n to detect 10% -> 11%:", sample_size(0.10, 0.11))
print("per-group n to detect 10% -> 12%:", sample_size(0.10, 0.12))
# Analysis after the experiment
successes = np.array([1_130, 1_020]) # treatment, control
totals = np.array([10_000, 10_000])
stat, p = proportions_ztest(successes, totals)
diff = successes[0] / totals[0] - successes[1] / totals[1]
se = np.sqrt(sum(s / t * (1 - s / t) / t for s, t in zip(successes, totals)))
print(f"difference = {diff:.4f}, 95% CI [{diff - 1.96 * se:.4f}, {diff + 1.96 * se:.4f}], p = {p:.4f}")Common pitfalls#
- Peeking: checking results repeatedly and stopping when $p < 0.05$ inflates false positives dramatically. Use a fixed horizon, or sequential testing methods designed for continuous monitoring (e.g. alpha-spending, always-valid confidence sequences).
- Multiple comparisons: testing many metrics or segments guarantees some false "wins"; correct for it or pre-specify.
- Wrong randomisation unit: randomising by request when users see both variants causes contamination; analyse at the level you randomise.
- Interference / network effects: treating one unit affects others (e.g. shared resources, marketplaces, social networks). Use cluster randomisation (by region, clinic or camp) when needed.
- Novelty and learning effects: users react differently at first; run long enough.
- Sample ratio mismatch: if groups are not the expected sizes, something is broken in assignment or logging — investigate before trusting results.
- Simpson's paradox and heterogeneity: an overall improvement can hide harm to a subgroup — examine pre-specified segments.
Alternatives and complements#
- Interleaving (for ranking/search): mix results from both rankers in one list and see which gets more engagement — very sensitive with small samples.
- Multi-armed bandits: shift traffic adaptively towards better variants (less regret, harder inference).
- Quasi-experimental methods when randomisation is impossible: difference-in-differences, regression discontinuity, synthetic controls.
- Off-policy evaluation from logs before any live test.