⚙️ MLOps & Engineering · Lecture 8 of 15

CI/CD for Machine Learning: Testing and Automating ML Systems

Continuous integration and delivery bring software-engineering discipline to ML. We cover the testing pyramid for ML — code, data and model tests — quality gates, continuous training, and a practical GitHub Actions workflow.

In software engineering, continuous integration (CI) automatically builds and tests every code change, and continuous delivery/deployment (CD) automatically releases changes that pass. ML systems need the same discipline — plus more, because behaviour depends on data and trained artefacts, not just code. This lecture shows how to test ML systems and automate their release safely.

What changes in an ML system?#

A new model can result from changes to:

  • code (feature logic, training script, serving code);
  • data (new training data, schema changes);
  • configuration (hyperparameters, thresholds);
  • dependencies (library versions).

CI/CD for ML (sometimes called CI/CD/CT — adding continuous training) must test all of them.

The testing pyramid for ML#

1. Code tests (fast, run on every commit)#

  • Unit tests for feature functions, preprocessing and utilities.
  • Edge cases: missing values, empty inputs, unseen categories, extreme values.
  • API contract tests for the serving interface.

2. Data tests#

  • Schema and constraint validation (types, ranges, uniqueness).
  • Distribution checks against reference data.
  • Leakage checks: no target-derived features; no overlap between train and test by ID.

3. Model tests#

  • Performance gates: metrics must exceed a threshold and not regress versus the current production model beyond a tolerance.
  • Slice tests: minimum performance per subgroup/language/region.
  • Behavioural tests (inspired by CheckList, Ribeiro et al., 2020): invariance ("changing a name should not change the prediction"), directional expectations ("more household members should not decrease need score, all else equal"), and minimum functionality tests.
  • Robustness to noise and typos; calibration checks.
  • Training smoke test: train on a tiny sample for a few steps and confirm loss decreases.

4. Integration and end-to-end tests#

  • The full pipeline runs from raw data to a served prediction in a staging environment.
  • The packaged model loads and produces the same predictions as in training (training–serving parity).
python
# tests/test_model_behaviour.py  (run with pytest)
import joblib, pandas as pd, pytest

model = joblib.load("artifacts/model.joblib")
BASE = {"members": 5, "children_under_5": 1, "monthly_income": 9000.0, "has_disability": False, "district": "Dhaka"}

def score(**overrides):
    return model.predict_proba(pd.DataFrame([{**BASE, **overrides}]))[0, 1]

def test_directional_income():
    # Lower income should not reduce the priority score (monotonic expectation)
    assert score(monthly_income=3000.0) >= score(monthly_income=15000.0) - 1e-6

def test_deterministic_predictions():
    # The same input must always produce the same score (no hidden randomness at inference)
    assert abs(score() - score()) < 1e-12

@pytest.mark.parametrize("children", [0, 3, 6])
def test_valid_probability(children):
    s = score(children_under_5=children)
    assert 0.0 <= s <= 1.0

Quality gates and model promotion#

A trained candidate is promoted only if it passes gates, for example:

  • validation AUC ≥ 0.85 and not worse than production by more than 0.005;
  • recall for the high-priority class ≥ 0.80 in every region;
  • calibration error below a threshold;
  • all behavioural tests pass;
  • latency and model size within budget;
  • human sign-off for high-impact models.

Results are recorded in the model registry with the run that produced them.

A GitHub Actions workflow#

yaml
# .github/workflows/ml-ci.yml
name: ml-ci
on:
  pull_request:
  push:
    branches: [main]
jobs:
  test-and-train:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: {python-version: "3.11"}
      - run: pip install -r requirements.txt
      - name: Lint and unit tests
        run: |
          ruff check src tests
          pytest tests/unit -q
      - name: Data validation
        run: python src/validate.py data/sample.parquet
      - name: Smoke-train on a sample
        run: python src/train.py --config configs/smoke.yaml
      - name: Model quality gates and behavioural tests
        run: |
          python src/evaluate.py --gates configs/gates.yaml
          pytest tests/model -q
      - name: Build container
        run: docker build -t triage-api:${{ github.sha }} .

Full training on large data usually runs on dedicated infrastructure (triggered by the pipeline or schedule), not inside a CI runner; CI runs fast checks and smoke tests.

Continuous training and delivery#

  • Continuous training (CT): pipelines retrain automatically on new data (scheduled or drift-triggered), then run the same gates.
  • Continuous delivery: passing models are packaged and deployed to staging automatically; production promotion may be automatic (low-risk models) or require approval (high-risk models).
  • Rollback: keep the previous model ready; automate rollback on alert.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

From Notebook to Production Code: Structuring Clean ML Projects

Notebooks are great for exploration and terrible for production. We cover a clean project layout, configuration, modular code, typing and testing, logging, packaging, and a workflow for moving from exploration to maintainable software.

Beginner⏱ 5 min#255
⚙️ MLOps & Engineering

Model Monitoring: Detecting Data Drift and Concept Drift

Deployed models degrade silently as the world changes. We define data drift, concept drift and prediction drift, measure drift with PSI and statistical tests, monitor without labels, and design alerts and retraining triggers.

Intermediate⏱ 5 min#248
⚙️ MLOps & Engineering

Feature Stores and Training–Serving Consistency

Feature stores manage features as shared, versioned assets available offline for training and online for serving. We explain training–serving skew, point-in-time correctness, online vs offline stores, and when a feature store is worth it.

Advanced⏱ 4 min#250