"Which settings produced that 91% model from two weeks ago?" If you cannot answer this in seconds, you need experiment tracking. ML development is empirical: you try many ideas, most fail, and the few that succeed must be reproducible. Spreadsheets and filenames like model_final_v3_really_final.pkl do not scale. Tracking tools record every run systematically.
What to track for every run#
- Code version: Git commit hash (and whether the working tree had uncommitted changes).
- Data version: dataset identifier or hash, split definitions.
- Configuration: all hyperparameters, feature lists, preprocessing options, random seeds.
- Environment: library versions, hardware (GPU type).
- Metrics: training/validation curves, final metrics, per-slice metrics (by language, region, class).
- Artefacts: the model file, plots (confusion matrix, calibration), sample predictions, evaluation reports.
- Notes and tags: the hypothesis being tested, the conclusion.
Tools#
- MLflow: open source; tracking server, UI, model registry, model packaging. Can run locally or self-hosted — useful for sensitive environments.
- Weights & Biases, Neptune, Comet: hosted platforms with rich visualisation and collaboration.
- TensorBoard: training curves, especially for deep learning.
- Aim, ClearML, DVC experiments: other open-source options.
MLflow in practice#
import mlflow, mlflow.sklearn, subprocess
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import f1_score, roc_auc_score, ConfusionMatrixDisplay
import matplotlib.pyplot as plt
mlflow.set_tracking_uri("file:./mlruns") # or a shared tracking server URL
mlflow.set_experiment("triage-classifier")
X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_tr, X_va, y_tr, y_va = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
for n_estimators in [100, 300]:
for max_depth in [4, None]:
with mlflow.start_run(run_name=f"rf-{n_estimators}-{max_depth}"):
commit = subprocess.run(["git", "rev-parse", "HEAD"], capture_output=True, text=True).stdout.strip()
mlflow.set_tags({"git_commit": commit or "not-a-git-repo", "data_version": "bc-sklearn-v1",
"hypothesis": "deeper trees improve recall"})
params = {"n_estimators": n_estimators, "max_depth": max_depth, "random_state": 0}
mlflow.log_params(params)
model = RandomForestClassifier(**params).fit(X_tr, y_tr)
proba = model.predict_proba(X_va)[:, 1]
mlflow.log_metrics({"val_auc": roc_auc_score(y_va, proba),
"val_f1": f1_score(y_va, proba > 0.5)})
ConfusionMatrixDisplay.from_predictions(y_va, proba > 0.5)
mlflow.log_figure(plt.gcf(), "confusion_matrix.png"); plt.close()
mlflow.sklearn.log_model(model, name="model", input_example=X_va.head(3))
# Then run `mlflow ui` and compare runs side by side.Autologging (mlflow.autolog()) captures parameters and metrics automatically for many libraries.
The model registry#
A model registry is the catalogue of models that might be deployed:
- each registered model has versions, linked to the runs (and thus data and code) that produced them;
- versions carry aliases or stages (e.g. "candidate", "champion", "archived");
- promotion requires review — metrics, fairness checks, sign-off;
- deployment systems load "the current champion" rather than a file path, enabling clean rollbacks.
from mlflow import MlflowClient
client = MlflowClient()
best = mlflow.search_runs(experiment_names=["triage-classifier"], order_by=["metrics.val_auc DESC"]).iloc[0]
mv = mlflow.register_model(f"runs:/{best.run_id}/model", "triage-classifier")
client.set_registered_model_alias("triage-classifier", "candidate", mv.version)Experiment hygiene#
Beyond metrics: comparing models properly#
Leaderboards of runs invite chasing tiny improvements. Compare candidates with confidence intervals, paired tests and slice metrics (see the hypothesis-testing lecture). A model 0.3% better on average but 5% worse for a minority language may be the wrong choice.