In software engineering, continuous integration (CI) automatically builds and tests every code change, and continuous delivery/deployment (CD) automatically releases changes that pass. ML systems need the same discipline — plus more, because behaviour depends on data and trained artefacts, not just code. This lecture shows how to test ML systems and automate their release safely.
What changes in an ML system?#
A new model can result from changes to:
- code (feature logic, training script, serving code);
- data (new training data, schema changes);
- configuration (hyperparameters, thresholds);
- dependencies (library versions).
CI/CD for ML (sometimes called CI/CD/CT — adding continuous training) must test all of them.
The testing pyramid for ML#
1. Code tests (fast, run on every commit)#
- Unit tests for feature functions, preprocessing and utilities.
- Edge cases: missing values, empty inputs, unseen categories, extreme values.
- API contract tests for the serving interface.
2. Data tests#
- Schema and constraint validation (types, ranges, uniqueness).
- Distribution checks against reference data.
- Leakage checks: no target-derived features; no overlap between train and test by ID.
3. Model tests#
- Performance gates: metrics must exceed a threshold and not regress versus the current production model beyond a tolerance.
- Slice tests: minimum performance per subgroup/language/region.
- Behavioural tests (inspired by CheckList, Ribeiro et al., 2020): invariance ("changing a name should not change the prediction"), directional expectations ("more household members should not decrease need score, all else equal"), and minimum functionality tests.
- Robustness to noise and typos; calibration checks.
- Training smoke test: train on a tiny sample for a few steps and confirm loss decreases.
4. Integration and end-to-end tests#
- The full pipeline runs from raw data to a served prediction in a staging environment.
- The packaged model loads and produces the same predictions as in training (training–serving parity).
# tests/test_model_behaviour.py (run with pytest)
import joblib, pandas as pd, pytest
model = joblib.load("artifacts/model.joblib")
BASE = {"members": 5, "children_under_5": 1, "monthly_income": 9000.0, "has_disability": False, "district": "Dhaka"}
def score(**overrides):
return model.predict_proba(pd.DataFrame([{**BASE, **overrides}]))[0, 1]
def test_directional_income():
# Lower income should not reduce the priority score (monotonic expectation)
assert score(monthly_income=3000.0) >= score(monthly_income=15000.0) - 1e-6
def test_deterministic_predictions():
# The same input must always produce the same score (no hidden randomness at inference)
assert abs(score() - score()) < 1e-12
@pytest.mark.parametrize("children", [0, 3, 6])
def test_valid_probability(children):
s = score(children_under_5=children)
assert 0.0 <= s <= 1.0Quality gates and model promotion#
A trained candidate is promoted only if it passes gates, for example:
- validation AUC ≥ 0.85 and not worse than production by more than 0.005;
- recall for the high-priority class ≥ 0.80 in every region;
- calibration error below a threshold;
- all behavioural tests pass;
- latency and model size within budget;
- human sign-off for high-impact models.
Results are recorded in the model registry with the run that produced them.
A GitHub Actions workflow#
# .github/workflows/ml-ci.yml
name: ml-ci
on:
pull_request:
push:
branches: [main]
jobs:
test-and-train:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: {python-version: "3.11"}
- run: pip install -r requirements.txt
- name: Lint and unit tests
run: |
ruff check src tests
pytest tests/unit -q
- name: Data validation
run: python src/validate.py data/sample.parquet
- name: Smoke-train on a sample
run: python src/train.py --config configs/smoke.yaml
- name: Model quality gates and behavioural tests
run: |
python src/evaluate.py --gates configs/gates.yaml
pytest tests/model -q
- name: Build container
run: docker build -t triage-api:${{ github.sha }} .Full training on large data usually runs on dedicated infrastructure (triggered by the pipeline or schedule), not inside a CI runner; CI runs fast checks and smoke tests.
Continuous training and delivery#
- Continuous training (CT): pipelines retrain automatically on new data (scheduled or drift-triggered), then run the same gates.
- Continuous delivery: passing models are packaged and deployed to staging automatically; production promotion may be automatic (low-risk models) or require approval (high-risk models).
- Rollback: keep the previous model ready; automate rollback on alert.