⚙️ MLOps & Engineering · Lecture 7 of 15

Model Monitoring: Detecting Data Drift and Concept Drift

Deployed models degrade silently as the world changes. We define data drift, concept drift and prediction drift, measure drift with PSI and statistical tests, monitor without labels, and design alerts and retraining triggers.

A model is a snapshot of the world at training time. The world keeps moving: a new policy changes who applies for services, a pandemic changes behaviour, a new phone model changes image colours, an upstream form changes a field's meaning. Without monitoring, a model can become inaccurate — or unfair — for months before anyone notices. Monitoring is how we detect such problems early.

What can go wrong#

  • Data (covariate) drift: the input distribution changes, $P(\mathbf{x})$ shifts. Example: a larger share of applicants from a new region.
  • Concept drift: the relationship between inputs and outputs changes, $P(y \mid \mathbf{x})$ shifts. Example: the same household characteristics now imply different needs after price inflation.
  • Label (prior) drift: the class balance changes, $P(y)$ shifts. Example: a disease outbreak increases positive cases.
  • Prediction drift: the distribution of model outputs changes — often the first visible symptom.
  • Data quality issues: missing values surge, units change, a feature is silently set to a default — technically a pipeline bug rather than real-world drift, but the most common cause of sudden degradation.
  • Upstream/downstream changes: new software versions, changed feature definitions, new consumers using outputs differently.

Monitoring layers#

  1. Operational: uptime, latency (p50/p95/p99), error rates, throughput, resource usage, cost.
  2. Data quality: schema violations, missing-value rates, range violations, new categories, volume.
  3. Distribution drift: per-feature and prediction distributions versus a reference window.
  4. Performance: accuracy, precision/recall, calibration — when (delayed) ground-truth labels arrive.
  5. Fairness and slices: performance and outcome rates per subgroup.
  6. Business/outcome metrics: is the system still achieving its purpose?

Measuring drift#

Population Stability Index (PSI) — widely used in credit scoring. Bin a feature (using reference quantiles), then compare the proportions of reference ($p_i$) and current ($q_i$) data per bin:

$$ \text{PSI} = \sum_i(q_i - p_i)\ln\frac{q_i}{p_i} $$

Common rules of thumb: PSI < 0.1 little change; 0.1–0.25 moderate; > 0.25 significant shift (treat thresholds as starting points, not laws).

Other measures: Kolmogorov–Smirnov test for continuous features, chi-squared test for categorical features, Jensen–Shannon or Wasserstein distances, and domain classifiers — train a model to distinguish reference from current data; if it succeeds (AUC well above 0.5), the distributions differ, and its feature importances show where.

python
import numpy as np
from scipy.stats import ks_2samp

def psi(reference, current, bins=10, eps=1e-6):
    edges = np.quantile(reference, np.linspace(0, 1, bins + 1))
    edges[0], edges[-1] = -np.inf, np.inf
    p = np.histogram(reference, edges)[0] / len(reference) + eps
    q = np.histogram(current, edges)[0] / len(current) + eps
    return float(np.sum((q - p) * np.log(q / p)))

rng = np.random.default_rng(0)
ref = rng.lognormal(9.0, 0.6, 20_000)                 # income at training time
cur_same = rng.lognormal(9.0, 0.6, 5_000)
cur_shift = rng.lognormal(9.25, 0.7, 5_000)            # inflation + more variance
for name, cur in [("same", cur_same), ("shifted", cur_shift)]:
    print(f"{name:<8} PSI={psi(ref, cur):.3f}  KS p-value={ks_2samp(ref, cur).pvalue:.2e}")

Monitoring without labels#

Ground truth often arrives late (did the applicant's situation actually worsen?) or never. Proxies:

  • input and prediction drift;
  • changes in the model's confidence distribution;
  • agreement between the production model and a reference or shadow model;
  • performance estimation methods that predict accuracy from confidence under covariate shift (e.g. confidence-based estimators), with caution;
  • targeted human review of a random sample of predictions — often the most reliable signal.

Alerts and responses#

Design alerting deliberately: too many alerts cause fatigue; too few miss problems.

  • Define thresholds per metric, windows (hourly, daily), and minimum sample sizes.
  • Route alerts to owners with runbooks: what to check, how to roll back, when to retrain.
  • Responses range from investigating a data pipeline bug, to recalibrating thresholds, retraining on recent data, rolling back, or pausing automated decisions in favour of manual processing.

Retraining strategies#

  • Scheduled (e.g. monthly) — simple, predictable.
  • Triggered by drift or performance degradation.
  • Continuous/online learning — powerful but risky (feedback loops, poisoning); use with safeguards.

Every retrained model must pass the same validation gates as the original — including fairness checks — before deployment.

Tools: Evidently, NannyML, WhyLabs/whylogs, Arize, cloud monitoring services, and Prometheus/Grafana for operational metrics.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Docker and Containers for Machine Learning

"It works on my machine" is not a deployment strategy. We explain containers, write efficient Dockerfiles for ML training and serving, handle GPUs, and cover image size, security and orchestration basics.

Beginner⏱ 5 min#247
⚙️ MLOps & Engineering

CI/CD for Machine Learning: Testing and Automating ML Systems

Continuous integration and delivery bring software-engineering discipline to ML. We cover the testing pyramid for ML — code, data and model tests — quality gates, continuous training, and a practical GitHub Actions workflow.

Intermediate⏱ 5 min#249
⚙️ MLOps & Engineering

Model Serving: Batch, Real-Time APIs and Streaming

A model creates value only when its predictions reach people and systems. We compare batch, online and streaming serving, build a FastAPI prediction service with validation, and cover latency, scaling and safe rollout.

Intermediate⏱ 5 min#246