Reproducibility is a cornerstone of science and a practical necessity in engineering. Yet surveys and replication studies across ML have found that many published results are hard to reproduce: code is missing, hyperparameters are unreported, baselines are under-tuned, and results vary with random seeds. In production, irreproducibility means you cannot debug a model, explain a decision, or safely retrain. This lecture covers the causes and cures.
Levels of reproducibility#
- Repeatability: the same team, same setup, gets the same result.
- Reproducibility: a different team, using the original code and data, gets the same result.
- Replicability: a different team, with independent implementation or data, reaches the same conclusion.
Aim for the first two in every project; the third is the gold standard of science.
Sources of non-reproducibility#
- Randomness: weight initialisation, data shuffling, dropout, augmentation, train/test splits, sampling.
- Non-deterministic hardware operations: some GPU kernels (e.g. atomic additions in certain convolution backward passes and scatter operations) produce slightly different results run to run; different GPU models or library versions change floating-point results.
- Software environment: library versions (a new scikit-learn default can change results), CUDA/cuDNN versions, operating system.
- Data: unversioned datasets, changing upstream sources, undocumented filtering, time-dependent queries ("last 30 days").
- Hidden configuration: unrecorded hyperparameters, manual notebook steps executed out of order.
- Evaluation variance: small test sets and single-seed results make "improvements" indistinguishable from noise.
Controlling randomness#
import os, random
import numpy as np
import torch
def set_seed(seed: int = 42, deterministic: bool = True):
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
torch.cuda.manual_seed_all(seed)
os.environ["PYTHONHASHSEED"] = str(seed)
if deterministic:
torch.backends.cudnn.benchmark = False # avoid auto-tuned, variable algorithms
torch.use_deterministic_algorithms(True, warn_only=True)
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8" # needed for deterministic cuBLAS ops
set_seed(42)
# DataLoader workers need their own seeding
def worker_init_fn(worker_id):
s = torch.initial_seed() % 2**32
np.random.seed(s); random.seed(s)
g = torch.Generator().manual_seed(42)
# loader = DataLoader(ds, shuffle=True, generator=g, worker_init_fn=worker_init_fn, num_workers=4)Deterministic mode can slow training; for many projects, fixing seeds (without full determinism) plus reporting variance across seeds is the pragmatic choice.
Controlling the environment#
- Pin dependencies:
requirements.txtwith exact versions, or lock files (pip-tools,poetry.lock,uv.lock,conda-lock). - Containers: Docker images capture the operating system, libraries and CUDA stack (next lectures).
- Record hardware and driver versions with each run.
# requirements.txt (pinned)
numpy==2.1.3
pandas==2.2.3
scikit-learn==1.5.2
torch==2.5.1Controlling data and configuration#
- Version datasets (DVC, snapshots) and record their hashes with each run.
- Store all hyperparameters in config files; log them automatically.
- Replace notebooks with scripts for final pipelines — or use tools that execute notebooks top-to-bottom in CI.
- Save split indices, not just the splitting code.
Reporting for research#
The NeurIPS and ICML reproducibility checklists ask authors to report, among other things: dataset details and preprocessing; hyperparameters and how they were chosen; number of runs, seeds and variance; compute used; and code availability. A good paper or thesis includes:
- Code and instructions to reproduce each table and figure.
- Exact data versions (or instructions to obtain them).
- Hyperparameter search spaces and budgets for baselines too.
- Means and confidence intervals over multiple seeds.
- Compute requirements.
- Known limitations.
Reproducibility in production#
For deployed models, reproducibility supports accountability: for any prediction, you should be able to identify the model version, its training data version, the code and configuration, and the input features used. Log prediction requests with model version identifiers (respecting privacy rules), and keep trained artefacts immutable in a registry.