A model is a snapshot of the world at training time. The world keeps moving: a new policy changes who applies for services, a pandemic changes behaviour, a new phone model changes image colours, an upstream form changes a field's meaning. Without monitoring, a model can become inaccurate — or unfair — for months before anyone notices. Monitoring is how we detect such problems early.
What can go wrong#
- Data (covariate) drift: the input distribution changes, $P(\mathbf{x})$ shifts. Example: a larger share of applicants from a new region.
- Concept drift: the relationship between inputs and outputs changes, $P(y \mid \mathbf{x})$ shifts. Example: the same household characteristics now imply different needs after price inflation.
- Label (prior) drift: the class balance changes, $P(y)$ shifts. Example: a disease outbreak increases positive cases.
- Prediction drift: the distribution of model outputs changes — often the first visible symptom.
- Data quality issues: missing values surge, units change, a feature is silently set to a default — technically a pipeline bug rather than real-world drift, but the most common cause of sudden degradation.
- Upstream/downstream changes: new software versions, changed feature definitions, new consumers using outputs differently.
Monitoring layers#
- Operational: uptime, latency (p50/p95/p99), error rates, throughput, resource usage, cost.
- Data quality: schema violations, missing-value rates, range violations, new categories, volume.
- Distribution drift: per-feature and prediction distributions versus a reference window.
- Performance: accuracy, precision/recall, calibration — when (delayed) ground-truth labels arrive.
- Fairness and slices: performance and outcome rates per subgroup.
- Business/outcome metrics: is the system still achieving its purpose?
Measuring drift#
Population Stability Index (PSI) — widely used in credit scoring. Bin a feature (using reference quantiles), then compare the proportions of reference ($p_i$) and current ($q_i$) data per bin:
Common rules of thumb: PSI < 0.1 little change; 0.1–0.25 moderate; > 0.25 significant shift (treat thresholds as starting points, not laws).
Other measures: Kolmogorov–Smirnov test for continuous features, chi-squared test for categorical features, Jensen–Shannon or Wasserstein distances, and domain classifiers — train a model to distinguish reference from current data; if it succeeds (AUC well above 0.5), the distributions differ, and its feature importances show where.
import numpy as np
from scipy.stats import ks_2samp
def psi(reference, current, bins=10, eps=1e-6):
edges = np.quantile(reference, np.linspace(0, 1, bins + 1))
edges[0], edges[-1] = -np.inf, np.inf
p = np.histogram(reference, edges)[0] / len(reference) + eps
q = np.histogram(current, edges)[0] / len(current) + eps
return float(np.sum((q - p) * np.log(q / p)))
rng = np.random.default_rng(0)
ref = rng.lognormal(9.0, 0.6, 20_000) # income at training time
cur_same = rng.lognormal(9.0, 0.6, 5_000)
cur_shift = rng.lognormal(9.25, 0.7, 5_000) # inflation + more variance
for name, cur in [("same", cur_same), ("shifted", cur_shift)]:
print(f"{name:<8} PSI={psi(ref, cur):.3f} KS p-value={ks_2samp(ref, cur).pvalue:.2e}")Monitoring without labels#
Ground truth often arrives late (did the applicant's situation actually worsen?) or never. Proxies:
- input and prediction drift;
- changes in the model's confidence distribution;
- agreement between the production model and a reference or shadow model;
- performance estimation methods that predict accuracy from confidence under covariate shift (e.g. confidence-based estimators), with caution;
- targeted human review of a random sample of predictions — often the most reliable signal.
Alerts and responses#
Design alerting deliberately: too many alerts cause fatigue; too few miss problems.
- Define thresholds per metric, windows (hourly, daily), and minimum sample sizes.
- Route alerts to owners with runbooks: what to check, how to roll back, when to retrain.
- Responses range from investigating a data pipeline bug, to recalibrating thresholds, retraining on recent data, rolling back, or pausing automated decisions in favour of manual processing.
Retraining strategies#
- Scheduled (e.g. monthly) — simple, predictable.
- Triggered by drift or performance degradation.
- Continuous/online learning — powerful but risky (feedback loops, poisoning); use with safeguards.
Every retrained model must pass the same validation gates as the original — including fairness checks — before deployment.
Tools: Evidently, NannyML, WhyLabs/whylogs, Arize, cloud monitoring services, and Prometheus/Grafana for operational metrics.