⚙️ MLOps & Engineering · Lecture 14 of 15

From Notebook to Production Code: Structuring Clean ML Projects

Notebooks are great for exploration and terrible for production. We cover a clean project layout, configuration, modular code, typing and testing, logging, packaging, and a workflow for moving from exploration to maintainable software.

Jupyter notebooks are wonderful for exploring data and prototyping models: immediate feedback, inline plots, narrative. But notebooks that grow into production systems become fragile: hidden state from out-of-order cell execution, copy-pasted code, no tests, hard-to-review diffs. Most of your career's ML work will be maintained by others — or by you in a year, having forgotten everything. Writing clean, well-structured ML code is a professional skill that makes your work reproducible, reviewable and reliable.

Notebooks: use them for what they are good at#

  • ✅ Exploration, visualisation, quick experiments, teaching, reports.
  • ❌ Core logic that must be reused, tested or scheduled.

A good workflow: explore in a notebook → move stable functions into a Python package → import them back into notebooks → run training and inference as scripts or pipelines. Restart the kernel and "run all" regularly to catch hidden-state bugs. Tools like nbstripout remove outputs before committing; jupytext pairs notebooks with plain .py files for readable diffs.

A clean project layout#

text
aid-triage/
├── pyproject.toml            # package metadata + dependencies + tool config
├── README.md                 # purpose, setup, how to train/evaluate/serve
├── configs/
│   ├── train.yaml
│   └── gates.yaml
├── src/aid_triage/
│   ├── __init__.py
│   ├── data.py               # loading, validation
│   ├── features.py           # feature engineering (shared by train and serve)
│   ├── model.py              # model definition / pipeline construction
│   ├── train.py              # CLI entry point
│   ├── evaluate.py           # metrics, slices, reports
│   └── serve.py              # API
├── notebooks/                # exploration only; import from src
├── tests/
│   ├── test_features.py
│   └── test_model.py
└── .github/workflows/ci.yml

Templates such as Cookiecutter Data Science provide similar starting structures.

Principles of clean ML code#

1. Small, pure functions#

Functions that take inputs and return outputs without hidden side effects are easy to test and reuse.

python
# src/aid_triage/features.py
from __future__ import annotations
import pandas as pd

def dependency_ratio(members: pd.Series, working_age: pd.Series) -> pd.Series:
    """Number of dependants per working-age member; 0 working-age members -> NaN."""
    dependants = members - working_age
    return dependants / working_age.where(working_age > 0)

def build_features(df: pd.DataFrame) -> pd.DataFrame:
    out = pd.DataFrame(index=df.index)
    out["dependency_ratio"] = dependency_ratio(df["members"], df["working_age_members"])
    out["income_per_member"] = df["monthly_income"] / df["members"].clip(lower=1)
    out["has_child_under_5"] = (df["children_under_5"] > 0).astype(int)
    return out

2. Configuration, not constants#

Put paths, hyperparameters, thresholds and seeds in config files (YAML/TOML) or typed config objects — not scattered through code.

python
from dataclasses import dataclass
import yaml

@dataclass(frozen=True)
class TrainConfig:
    train_path: str
    valid_path: str
    target: str
    learning_rate: float = 0.05
    max_leaf_nodes: int = 31
    seed: int = 42

cfg = TrainConfig(**yaml.safe_load(open("configs/train.yaml")))

3. Type hints and docstrings#

They document intent, enable editor support and catch bugs with type checkers (mypy, pyright).

4. Tests#

Test feature functions with small, hand-computed examples, including edge cases:

python
# tests/test_features.py
import numpy as np, pandas as pd
from aid_triage.features import dependency_ratio

def test_dependency_ratio_basic():
    assert dependency_ratio(pd.Series([5]), pd.Series([2])).iloc[0] == 1.5

def test_dependency_ratio_no_workers_is_nan():
    assert np.isnan(dependency_ratio(pd.Series([3]), pd.Series([0])).iloc[0])

5. Logging instead of print#

Use the logging module with levels and timestamps; log configuration, data versions and metrics.

6. Formatting and linting#

Automate style with tools such as ruff (linting and formatting) and pre-commit hooks — reviews can then focus on substance.

7. Command-line entry points#

Make training runnable as python -m aid_triage.train --config configs/train.yaml, so it can be scheduled, containerised and run in CI.

Code review and collaboration#

  • Use Git branches and pull requests; review code and results (metrics, plots).
  • Keep pull requests small and focused.
  • Write a README that lets a new person set up and run the project in under an hour.
  • Record decisions (why this metric, why this threshold) in the repository — future readers need the reasoning, not just the code.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

CI/CD for Machine Learning: Testing and Automating ML Systems

Continuous integration and delivery bring software-engineering discipline to ML. We cover the testing pyramid for ML — code, data and model tests — quality gates, continuous training, and a practical GitHub Actions workflow.

Intermediate⏱ 5 min#249
⚙️ MLOps & Engineering

A/B Testing and Online Evaluation of ML Models

Offline metrics do not guarantee real-world impact. Online experiments measure what a model actually changes. We design randomised A/B tests for ML, compute sample sizes, avoid common pitfalls, and discuss ethics of experimenting with people.

Intermediate⏱ 5 min#254
⚙️ MLOps & Engineering

Model Cards, Datasheets and Responsible Documentation

Documentation is how models and datasets are understood, audited and used responsibly. We cover model cards, datasheets for datasets, system cards, what to include, and how documentation supports accountability and regulation.

Beginner⏱ 5 min#256