⚙️ MLOps & Engineering · Lecture 3 of 15

Experiment Tracking: Never Lose a Result Again

ML development involves hundreds of runs with different data, code and hyperparameters. We cover what to track, how to use MLflow for runs, metrics and artefacts, the model registry, and good experiment hygiene.

"Which settings produced that 91% model from two weeks ago?" If you cannot answer this in seconds, you need experiment tracking. ML development is empirical: you try many ideas, most fail, and the few that succeed must be reproducible. Spreadsheets and filenames like model_final_v3_really_final.pkl do not scale. Tracking tools record every run systematically.

What to track for every run#

  • Code version: Git commit hash (and whether the working tree had uncommitted changes).
  • Data version: dataset identifier or hash, split definitions.
  • Configuration: all hyperparameters, feature lists, preprocessing options, random seeds.
  • Environment: library versions, hardware (GPU type).
  • Metrics: training/validation curves, final metrics, per-slice metrics (by language, region, class).
  • Artefacts: the model file, plots (confusion matrix, calibration), sample predictions, evaluation reports.
  • Notes and tags: the hypothesis being tested, the conclusion.

Tools#

  • MLflow: open source; tracking server, UI, model registry, model packaging. Can run locally or self-hosted — useful for sensitive environments.
  • Weights & Biases, Neptune, Comet: hosted platforms with rich visualisation and collaboration.
  • TensorBoard: training curves, especially for deep learning.
  • Aim, ClearML, DVC experiments: other open-source options.

MLflow in practice#

python
import mlflow, mlflow.sklearn, subprocess
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import f1_score, roc_auc_score, ConfusionMatrixDisplay
import matplotlib.pyplot as plt

mlflow.set_tracking_uri("file:./mlruns")                 # or a shared tracking server URL
mlflow.set_experiment("triage-classifier")

X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_tr, X_va, y_tr, y_va = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)

for n_estimators in [100, 300]:
    for max_depth in [4, None]:
        with mlflow.start_run(run_name=f"rf-{n_estimators}-{max_depth}"):
            commit = subprocess.run(["git", "rev-parse", "HEAD"], capture_output=True, text=True).stdout.strip()
            mlflow.set_tags({"git_commit": commit or "not-a-git-repo", "data_version": "bc-sklearn-v1",
                             "hypothesis": "deeper trees improve recall"})
            params = {"n_estimators": n_estimators, "max_depth": max_depth, "random_state": 0}
            mlflow.log_params(params)
            model = RandomForestClassifier(**params).fit(X_tr, y_tr)
            proba = model.predict_proba(X_va)[:, 1]
            mlflow.log_metrics({"val_auc": roc_auc_score(y_va, proba),
                                "val_f1": f1_score(y_va, proba > 0.5)})
            ConfusionMatrixDisplay.from_predictions(y_va, proba > 0.5)
            mlflow.log_figure(plt.gcf(), "confusion_matrix.png"); plt.close()
            mlflow.sklearn.log_model(model, name="model", input_example=X_va.head(3))
# Then run `mlflow ui` and compare runs side by side.

Autologging (mlflow.autolog()) captures parameters and metrics automatically for many libraries.

The model registry#

A model registry is the catalogue of models that might be deployed:

  • each registered model has versions, linked to the runs (and thus data and code) that produced them;
  • versions carry aliases or stages (e.g. "candidate", "champion", "archived");
  • promotion requires review — metrics, fairness checks, sign-off;
  • deployment systems load "the current champion" rather than a file path, enabling clean rollbacks.
python
from mlflow import MlflowClient
client = MlflowClient()
best = mlflow.search_runs(experiment_names=["triage-classifier"], order_by=["metrics.val_auc DESC"]).iloc[0]
mv = mlflow.register_model(f"runs:/{best.run_id}/model", "triage-classifier")
client.set_registered_model_alias("triage-classifier", "candidate", mv.version)

Experiment hygiene#

Beyond metrics: comparing models properly#

Leaderboards of runs invite chasing tiny improvements. Compare candidates with confidence intervals, paired tests and slice metrics (see the hypothesis-testing lecture). A model 0.3% better on average but 5% worse for a minority language may be the wrong choice.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Reproducibility in Machine Learning

Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.

Intermediate⏱ 4 min#245
⚙️ MLOps & Engineering

Data Versioning, Validation and Pipelines

Models are only as good as their data — and data changes. We cover versioning datasets, validating schemas and distributions, building reproducible data pipelines, and orchestration tools.

Intermediate⏱ 4 min#243
⚙️ MLOps & Engineering

What Is MLOps? From Notebook to Reliable Production System

Most ML projects fail not because of the model but because of everything around it. We define MLOps, walk through the ML lifecycle, describe maturity levels, and examine the hidden technical debt of machine learning systems.

Beginner⏱ 5 min#242