Jupyter notebooks are wonderful for exploring data and prototyping models: immediate feedback, inline plots, narrative. But notebooks that grow into production systems become fragile: hidden state from out-of-order cell execution, copy-pasted code, no tests, hard-to-review diffs. Most of your career's ML work will be maintained by others — or by you in a year, having forgotten everything. Writing clean, well-structured ML code is a professional skill that makes your work reproducible, reviewable and reliable.
Notebooks: use them for what they are good at#
- ✅ Exploration, visualisation, quick experiments, teaching, reports.
- ❌ Core logic that must be reused, tested or scheduled.
A good workflow: explore in a notebook → move stable functions into a Python package → import them back into notebooks → run training and inference as scripts or pipelines. Restart the kernel and "run all" regularly to catch hidden-state bugs. Tools like nbstripout remove outputs before committing; jupytext pairs notebooks with plain .py files for readable diffs.
A clean project layout#
aid-triage/
├── pyproject.toml # package metadata + dependencies + tool config
├── README.md # purpose, setup, how to train/evaluate/serve
├── configs/
│ ├── train.yaml
│ └── gates.yaml
├── src/aid_triage/
│ ├── __init__.py
│ ├── data.py # loading, validation
│ ├── features.py # feature engineering (shared by train and serve)
│ ├── model.py # model definition / pipeline construction
│ ├── train.py # CLI entry point
│ ├── evaluate.py # metrics, slices, reports
│ └── serve.py # API
├── notebooks/ # exploration only; import from src
├── tests/
│ ├── test_features.py
│ └── test_model.py
└── .github/workflows/ci.ymlTemplates such as Cookiecutter Data Science provide similar starting structures.
Principles of clean ML code#
1. Small, pure functions#
Functions that take inputs and return outputs without hidden side effects are easy to test and reuse.
# src/aid_triage/features.py
from __future__ import annotations
import pandas as pd
def dependency_ratio(members: pd.Series, working_age: pd.Series) -> pd.Series:
"""Number of dependants per working-age member; 0 working-age members -> NaN."""
dependants = members - working_age
return dependants / working_age.where(working_age > 0)
def build_features(df: pd.DataFrame) -> pd.DataFrame:
out = pd.DataFrame(index=df.index)
out["dependency_ratio"] = dependency_ratio(df["members"], df["working_age_members"])
out["income_per_member"] = df["monthly_income"] / df["members"].clip(lower=1)
out["has_child_under_5"] = (df["children_under_5"] > 0).astype(int)
return out2. Configuration, not constants#
Put paths, hyperparameters, thresholds and seeds in config files (YAML/TOML) or typed config objects — not scattered through code.
from dataclasses import dataclass
import yaml
@dataclass(frozen=True)
class TrainConfig:
train_path: str
valid_path: str
target: str
learning_rate: float = 0.05
max_leaf_nodes: int = 31
seed: int = 42
cfg = TrainConfig(**yaml.safe_load(open("configs/train.yaml")))3. Type hints and docstrings#
They document intent, enable editor support and catch bugs with type checkers (mypy, pyright).
4. Tests#
Test feature functions with small, hand-computed examples, including edge cases:
# tests/test_features.py
import numpy as np, pandas as pd
from aid_triage.features import dependency_ratio
def test_dependency_ratio_basic():
assert dependency_ratio(pd.Series([5]), pd.Series([2])).iloc[0] == 1.5
def test_dependency_ratio_no_workers_is_nan():
assert np.isnan(dependency_ratio(pd.Series([3]), pd.Series([0])).iloc[0])5. Logging instead of print#
Use the logging module with levels and timestamps; log configuration, data versions and metrics.
6. Formatting and linting#
Automate style with tools such as ruff (linting and formatting) and pre-commit hooks — reviews can then focus on substance.
7. Command-line entry points#
Make training runnable as python -m aid_triage.train --config configs/train.yaml, so it can be scheduled, containerised and run in CI.
Code review and collaboration#
- Use Git branches and pull requests; review code and results (metrics, plots).
- Keep pull requests small and focused.
- Write a README that lets a new person set up and run the project in under an hour.
- Record decisions (why this metric, why this threshold) in the repository — future readers need the reasoning, not just the code.