⚙️ MLOps & Engineering · Lecture 13 of 15

A/B Testing and Online Evaluation of ML Models

Offline metrics do not guarantee real-world impact. Online experiments measure what a model actually changes. We design randomised A/B tests for ML, compute sample sizes, avoid common pitfalls, and discuss ethics of experimenting with people.

A new recommendation model improves offline accuracy by 3%. Will users actually find what they need faster? A new triage model has better AUC. Will caseworkers resolve urgent cases sooner? Offline metrics are proxies; the real question is causal: what happens to outcomes when we deploy this model? Randomised online experiments — A/B tests — answer it with the same logic as clinical trials.

Why offline evaluation is not enough#

  • Offline data reflects the old system's behaviour (recommendations shape clicks), creating feedback loops.
  • Proxy metrics (accuracy, AUC) may not map to outcomes (user satisfaction, time saved, harm avoided).
  • Real-world effects include user adaptation, latency, UI interaction and workflow changes.

The basic design#

  1. Randomise units (users, sessions, households, clinics) into control (current model, A) and treatment (new model, B).
  2. Run both concurrently for a pre-specified period.
  3. Compare a pre-registered primary metric between groups.

Randomisation makes the groups comparable in expectation, so a significant difference can be attributed to the model.

Choosing metrics#

  • Primary (decision) metric: the outcome you care about — e.g. proportion of urgent cases handled within 48 hours.
  • Guardrail metrics: things that must not get worse — latency, error rates, complaint rates, fairness across groups, cost.
  • Diagnostic metrics: help explain results (clicks, model confidence distributions).

Specify metrics, analysis method and stopping rules before starting (pre-registration) to avoid cherry-picking.

Sample size and power#

To detect a difference between two proportions $p_A$ and $p_B$ with significance level $\alpha$ and power $1 - \beta$, the required sample size per group is approximately

$$ n \approx \frac{\left(z_{1-\alpha/2} + z_{1-\beta}\right)^2\,\big[p_A(1 - p_A) + p_B(1 - p_B)\big]}{(p_B - p_A)^2} $$

Small effects require large samples: halving the detectable effect quadruples $n$.

python
from scipy.stats import norm
import numpy as np
from statsmodels.stats.proportion import proportions_ztest

def sample_size(p_a, p_b, alpha=0.05, power=0.8):
    z = norm.ppf(1 - alpha / 2) + norm.ppf(power)
    return int(np.ceil(z**2 * (p_a * (1 - p_a) + p_b * (1 - p_b)) / (p_b - p_a) ** 2))

print("per-group n to detect 10% -> 11%:", sample_size(0.10, 0.11))
print("per-group n to detect 10% -> 12%:", sample_size(0.10, 0.12))

# Analysis after the experiment
successes = np.array([1_130, 1_020])          # treatment, control
totals = np.array([10_000, 10_000])
stat, p = proportions_ztest(successes, totals)
diff = successes[0] / totals[0] - successes[1] / totals[1]
se = np.sqrt(sum(s / t * (1 - s / t) / t for s, t in zip(successes, totals)))
print(f"difference = {diff:.4f}, 95% CI [{diff - 1.96 * se:.4f}, {diff + 1.96 * se:.4f}], p = {p:.4f}")

Common pitfalls#

  1. Peeking: checking results repeatedly and stopping when $p < 0.05$ inflates false positives dramatically. Use a fixed horizon, or sequential testing methods designed for continuous monitoring (e.g. alpha-spending, always-valid confidence sequences).
  2. Multiple comparisons: testing many metrics or segments guarantees some false "wins"; correct for it or pre-specify.
  3. Wrong randomisation unit: randomising by request when users see both variants causes contamination; analyse at the level you randomise.
  4. Interference / network effects: treating one unit affects others (e.g. shared resources, marketplaces, social networks). Use cluster randomisation (by region, clinic or camp) when needed.
  5. Novelty and learning effects: users react differently at first; run long enough.
  6. Sample ratio mismatch: if groups are not the expected sizes, something is broken in assignment or logging — investigate before trusting results.
  7. Simpson's paradox and heterogeneity: an overall improvement can hide harm to a subgroup — examine pre-specified segments.

Alternatives and complements#

  • Interleaving (for ranking/search): mix results from both rankers in one list and see which gets more engagement — very sensitive with small samples.
  • Multi-armed bandits: shift traffic adaptively towards better variants (less regret, harder inference).
  • Quasi-experimental methods when randomisation is impossible: difference-in-differences, regression discontinuity, synthetic controls.
  • Off-policy evaluation from logs before any live test.

Ethics of experimentation#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Data Labelling and Annotation: Building High-Quality Datasets

Labels are the foundation of supervised learning, yet labelling is often rushed. We cover annotation guidelines, workflows and tools, measuring agreement, handling label noise, model-assisted labelling, and fair treatment of annotators.

Beginner⏱ 5 min#253
⚙️ MLOps & Engineering

From Notebook to Production Code: Structuring Clean ML Projects

Notebooks are great for exploration and terrible for production. We cover a clean project layout, configuration, modular code, typing and testing, logging, packaging, and a workflow for moving from exploration to maintainable software.

Beginner⏱ 5 min#255
⚙️ MLOps & Engineering

GPUs and Hardware for Machine Learning

Understanding hardware helps you train faster and cheaper. We explain why GPUs suit deep learning, the roles of memory capacity and bandwidth, precision and tensor cores, estimating requirements, and choosing between local, cloud and free resources.

Intermediate⏱ 5 min#252