⚙️ MLOps & Engineering · Lecture 1 of 15

What Is MLOps? From Notebook to Reliable Production System

Most ML projects fail not because of the model but because of everything around it. We define MLOps, walk through the ML lifecycle, describe maturity levels, and examine the hidden technical debt of machine learning systems.

A student trains a model in a notebook, reaches 94% accuracy, and considers the project done. In an organisation, that is where the real work begins. The model must receive fresh data reliably, produce predictions for real users with acceptable latency, be monitored as the world changes, be retrained and redeployed safely, and be auditable months later. MLOps — machine learning operations — is the set of practices that makes this possible.

Why ML in production is hard#

In their influential paper "Hidden Technical Debt in Machine Learning Systems" (Sculley et al., NeurIPS 2015), Google engineers showed a now-famous diagram: the ML code is a tiny box surrounded by much larger boxes — data collection, data verification, feature extraction, configuration, serving infrastructure, monitoring, analysis tools, process management. ML systems also accumulate unique forms of technical debt:

  • Entanglement: "changing anything changes everything" — adding or modifying one feature changes how the model uses all others.
  • Data dependencies: upstream data sources change silently (a sensor is recalibrated; a form field changes meaning).
  • Feedback loops: model predictions influence the data used to retrain it (recommendations shape clicks).
  • Glue code and pipeline jungles: ad-hoc scripts that nobody fully understands.
  • Configuration debt: hundreds of settings with poor documentation.
  • Undeclared consumers: other systems quietly depend on your model's outputs.

The ML lifecycle#

text
Problem framing → Data collection & labelling → Data validation → Feature engineering
      → Training & experiment tracking → Evaluation & validation → Packaging
      → Deployment (batch / online / edge) → Monitoring → Retraining → (repeat)

Unlike traditional software, ML systems can degrade without any code change, because the data changes. MLOps treats data, models and code as first-class, versioned artefacts.

Core principles#

  1. Versioning everything: code (Git), data (dataset versions/snapshots), models (registry), configurations and environments.
  2. Automation: reproducible pipelines for training, evaluation and deployment instead of manual notebook steps.
  3. Continuous testing: code tests, data tests (schemas, distributions), model tests (performance thresholds, fairness slices, robustness).
  4. Continuous delivery: safe, repeatable deployment with rollbacks.
  5. Monitoring: of inputs, predictions, performance, latency, cost — and of harm.
  6. Reproducibility and auditability: be able to answer "which data, code and settings produced this prediction?"
  7. Collaboration: shared tools and conventions between data scientists, engineers, domain experts and risk/compliance teams.

MLOps maturity levels#

A widely cited framework (Google Cloud's MLOps levels) describes:

  • Level 0 — manual: notebooks, manual hand-off of a model file to engineers, rare releases, no monitoring.
  • Level 1 — ML pipeline automation: automated, reproducible training pipelines; continuous training triggered by new data; model registry; monitoring.
  • Level 2 — CI/CD pipeline automation: automated testing and deployment of the pipelines themselves, enabling rapid, reliable iteration.

Most organisations — including many NGOs and public agencies — are at level 0. Moving to level 1 for important models brings most of the benefit.

A minimal production-minded project#

text
project/
├── data/                 # raw data is NOT committed; use versioned storage + a manifest
├── src/
│   ├── data.py           # loading + validation
│   ├── features.py       # feature engineering (shared by training AND serving)
│   ├── train.py          # training entry point, reads config
│   ├── evaluate.py       # metrics, slices, thresholds
│   └── serve.py          # prediction API
├── configs/train.yaml    # hyperparameters, paths, seeds
├── tests/                # unit, data and model tests
├── Dockerfile            # reproducible environment
└── pipeline.yaml         # orchestration definition (e.g. DVC, Airflow, Kubeflow, Prefect)
python
# src/train.py — a pipeline step that is reproducible and self-documenting
import json, hashlib, yaml, joblib, pandas as pd
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import roc_auc_score

def file_hash(path):
    return hashlib.sha256(open(path, "rb").read()).hexdigest()[:12]

def main(cfg_path="configs/train.yaml"):
    cfg = yaml.safe_load(open(cfg_path))
    train, valid = pd.read_parquet(cfg["train_path"]), pd.read_parquet(cfg["valid_path"])
    X, y = train.drop(columns=cfg["target"]), train[cfg["target"]]
    model = HistGradientBoostingClassifier(**cfg["params"], random_state=cfg["seed"]).fit(X, y)
    auc = roc_auc_score(valid[cfg["target"]], model.predict_proba(valid.drop(columns=cfg["target"]))[:, 1])
    assert auc >= cfg["min_auc"], f"Model below quality gate: {auc:.3f}"
    joblib.dump(model, "artifacts/model.joblib")
    json.dump({"auc": auc, "train_data": file_hash(cfg["train_path"]), "config": cfg},
              open("artifacts/metadata.json", "w"), indent=2)

if __name__ == "__main__":
    main()

Note the quality gate (the model is not saved if it underperforms) and the metadata linking the model to its data and configuration.

Roles#

Data scientists, ML engineers, data engineers, platform/DevOps engineers, product owners, domain experts and governance teams all contribute. In small teams one person wears many hats — which makes simple, well-documented practices even more valuable.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Data Versioning, Validation and Pipelines

Models are only as good as their data — and data changes. We cover versioning datasets, validating schemas and distributions, building reproducible data pipelines, and orchestration tools.

Intermediate⏱ 4 min#243
⚙️ MLOps & Engineering

Experiment Tracking: Never Lose a Result Again

ML development involves hundreds of runs with different data, code and hyperparameters. We cover what to track, how to use MLflow for runs, metrics and artefacts, the model registry, and good experiment hygiene.

Beginner⏱ 4 min#244
⚙️ MLOps & Engineering

Reproducibility in Machine Learning

Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.

Intermediate⏱ 4 min#245