⚙️ MLOps & Engineering · Lecture 4 of 15

Reproducibility in Machine Learning

Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.

Reproducibility is a cornerstone of science and a practical necessity in engineering. Yet surveys and replication studies across ML have found that many published results are hard to reproduce: code is missing, hyperparameters are unreported, baselines are under-tuned, and results vary with random seeds. In production, irreproducibility means you cannot debug a model, explain a decision, or safely retrain. This lecture covers the causes and cures.

Levels of reproducibility#

  • Repeatability: the same team, same setup, gets the same result.
  • Reproducibility: a different team, using the original code and data, gets the same result.
  • Replicability: a different team, with independent implementation or data, reaches the same conclusion.

Aim for the first two in every project; the third is the gold standard of science.

Sources of non-reproducibility#

  1. Randomness: weight initialisation, data shuffling, dropout, augmentation, train/test splits, sampling.
  2. Non-deterministic hardware operations: some GPU kernels (e.g. atomic additions in certain convolution backward passes and scatter operations) produce slightly different results run to run; different GPU models or library versions change floating-point results.
  3. Software environment: library versions (a new scikit-learn default can change results), CUDA/cuDNN versions, operating system.
  4. Data: unversioned datasets, changing upstream sources, undocumented filtering, time-dependent queries ("last 30 days").
  5. Hidden configuration: unrecorded hyperparameters, manual notebook steps executed out of order.
  6. Evaluation variance: small test sets and single-seed results make "improvements" indistinguishable from noise.

Controlling randomness#

python
import os, random
import numpy as np
import torch

def set_seed(seed: int = 42, deterministic: bool = True):
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    torch.cuda.manual_seed_all(seed)
    os.environ["PYTHONHASHSEED"] = str(seed)
    if deterministic:
        torch.backends.cudnn.benchmark = False                 # avoid auto-tuned, variable algorithms
        torch.use_deterministic_algorithms(True, warn_only=True)
        os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"      # needed for deterministic cuBLAS ops

set_seed(42)

# DataLoader workers need their own seeding
def worker_init_fn(worker_id):
    s = torch.initial_seed() % 2**32
    np.random.seed(s); random.seed(s)

g = torch.Generator().manual_seed(42)
# loader = DataLoader(ds, shuffle=True, generator=g, worker_init_fn=worker_init_fn, num_workers=4)

Deterministic mode can slow training; for many projects, fixing seeds (without full determinism) plus reporting variance across seeds is the pragmatic choice.

Controlling the environment#

  • Pin dependencies: requirements.txt with exact versions, or lock files (pip-tools, poetry.lock, uv.lock, conda-lock).
  • Containers: Docker images capture the operating system, libraries and CUDA stack (next lectures).
  • Record hardware and driver versions with each run.
text
# requirements.txt (pinned)
numpy==2.1.3
pandas==2.2.3
scikit-learn==1.5.2
torch==2.5.1

Controlling data and configuration#

  • Version datasets (DVC, snapshots) and record their hashes with each run.
  • Store all hyperparameters in config files; log them automatically.
  • Replace notebooks with scripts for final pipelines — or use tools that execute notebooks top-to-bottom in CI.
  • Save split indices, not just the splitting code.

Reporting for research#

The NeurIPS and ICML reproducibility checklists ask authors to report, among other things: dataset details and preprocessing; hyperparameters and how they were chosen; number of runs, seeds and variance; compute used; and code availability. A good paper or thesis includes:

  1. Code and instructions to reproduce each table and figure.
  2. Exact data versions (or instructions to obtain them).
  3. Hyperparameter search spaces and budgets for baselines too.
  4. Means and confidence intervals over multiple seeds.
  5. Compute requirements.
  6. Known limitations.

Reproducibility in production#

For deployed models, reproducibility supports accountability: for any prediction, you should be able to identify the model version, its training data version, the code and configuration, and the input features used. Log prediction requests with model version identifiers (respecting privacy rules), and keep trained artefacts immutable in a registry.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Experiment Tracking: Never Lose a Result Again

ML development involves hundreds of runs with different data, code and hyperparameters. We cover what to track, how to use MLflow for runs, metrics and artefacts, the model registry, and good experiment hygiene.

Beginner⏱ 4 min#244
⚙️ MLOps & Engineering

Docker and Containers for Machine Learning

"It works on my machine" is not a deployment strategy. We explain containers, write efficient Dockerfiles for ML training and serving, handle GPUs, and cover image size, security and orchestration basics.

Beginner⏱ 5 min#247
⚙️ MLOps & Engineering

Model Serving: Batch, Real-Time APIs and Streaming

A model creates value only when its predictions reach people and systems. We compare batch, online and streaming serving, build a FastAPI prediction service with validation, and cover latency, scaling and safe rollout.

Intermediate⏱ 5 min#246