A student trains a model in a notebook, reaches 94% accuracy, and considers the project done. In an organisation, that is where the real work begins. The model must receive fresh data reliably, produce predictions for real users with acceptable latency, be monitored as the world changes, be retrained and redeployed safely, and be auditable months later. MLOps — machine learning operations — is the set of practices that makes this possible.
Why ML in production is hard#
In their influential paper "Hidden Technical Debt in Machine Learning Systems" (Sculley et al., NeurIPS 2015), Google engineers showed a now-famous diagram: the ML code is a tiny box surrounded by much larger boxes — data collection, data verification, feature extraction, configuration, serving infrastructure, monitoring, analysis tools, process management. ML systems also accumulate unique forms of technical debt:
- Entanglement: "changing anything changes everything" — adding or modifying one feature changes how the model uses all others.
- Data dependencies: upstream data sources change silently (a sensor is recalibrated; a form field changes meaning).
- Feedback loops: model predictions influence the data used to retrain it (recommendations shape clicks).
- Glue code and pipeline jungles: ad-hoc scripts that nobody fully understands.
- Configuration debt: hundreds of settings with poor documentation.
- Undeclared consumers: other systems quietly depend on your model's outputs.
The ML lifecycle#
Problem framing → Data collection & labelling → Data validation → Feature engineering
→ Training & experiment tracking → Evaluation & validation → Packaging
→ Deployment (batch / online / edge) → Monitoring → Retraining → (repeat)Unlike traditional software, ML systems can degrade without any code change, because the data changes. MLOps treats data, models and code as first-class, versioned artefacts.
Core principles#
- Versioning everything: code (Git), data (dataset versions/snapshots), models (registry), configurations and environments.
- Automation: reproducible pipelines for training, evaluation and deployment instead of manual notebook steps.
- Continuous testing: code tests, data tests (schemas, distributions), model tests (performance thresholds, fairness slices, robustness).
- Continuous delivery: safe, repeatable deployment with rollbacks.
- Monitoring: of inputs, predictions, performance, latency, cost — and of harm.
- Reproducibility and auditability: be able to answer "which data, code and settings produced this prediction?"
- Collaboration: shared tools and conventions between data scientists, engineers, domain experts and risk/compliance teams.
MLOps maturity levels#
A widely cited framework (Google Cloud's MLOps levels) describes:
- Level 0 — manual: notebooks, manual hand-off of a model file to engineers, rare releases, no monitoring.
- Level 1 — ML pipeline automation: automated, reproducible training pipelines; continuous training triggered by new data; model registry; monitoring.
- Level 2 — CI/CD pipeline automation: automated testing and deployment of the pipelines themselves, enabling rapid, reliable iteration.
Most organisations — including many NGOs and public agencies — are at level 0. Moving to level 1 for important models brings most of the benefit.
A minimal production-minded project#
project/
├── data/ # raw data is NOT committed; use versioned storage + a manifest
├── src/
│ ├── data.py # loading + validation
│ ├── features.py # feature engineering (shared by training AND serving)
│ ├── train.py # training entry point, reads config
│ ├── evaluate.py # metrics, slices, thresholds
│ └── serve.py # prediction API
├── configs/train.yaml # hyperparameters, paths, seeds
├── tests/ # unit, data and model tests
├── Dockerfile # reproducible environment
└── pipeline.yaml # orchestration definition (e.g. DVC, Airflow, Kubeflow, Prefect)# src/train.py — a pipeline step that is reproducible and self-documenting
import json, hashlib, yaml, joblib, pandas as pd
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import roc_auc_score
def file_hash(path):
return hashlib.sha256(open(path, "rb").read()).hexdigest()[:12]
def main(cfg_path="configs/train.yaml"):
cfg = yaml.safe_load(open(cfg_path))
train, valid = pd.read_parquet(cfg["train_path"]), pd.read_parquet(cfg["valid_path"])
X, y = train.drop(columns=cfg["target"]), train[cfg["target"]]
model = HistGradientBoostingClassifier(**cfg["params"], random_state=cfg["seed"]).fit(X, y)
auc = roc_auc_score(valid[cfg["target"]], model.predict_proba(valid.drop(columns=cfg["target"]))[:, 1])
assert auc >= cfg["min_auc"], f"Model below quality gate: {auc:.3f}"
joblib.dump(model, "artifacts/model.joblib")
json.dump({"auc": auc, "train_data": file_hash(cfg["train_path"]), "config": cfg},
open("artifacts/metadata.json", "w"), indent=2)
if __name__ == "__main__":
main()Note the quality gate (the model is not saved if it underperforms) and the metadata linking the model to its data and configuration.
Roles#
Data scientists, ML engineers, data engineers, platform/DevOps engineers, product owners, domain experts and governance teams all contribute. In small teams one person wears many hats — which makes simple, well-documented practices even more valuable.