When a model's performance suddenly drops, the cause is usually not the model code — it is the data. A column changed units, a new category appeared, a join duplicated rows, a sensor failed, an upstream team changed a definition. Managing data with the same rigour as code — versioning, validation and reproducible pipelines — is the foundation of reliable ML.
Why version data?#
- Reproducibility: re-create the exact training set behind any model.
- Debugging: compare the data of a good model and a bad one.
- Auditability: answer regulators, partners or affected people about what data a decision relied on.
- Collaboration: everyone trains on the same, identified dataset.
Git is designed for code, not multi-gigabyte files. Options:
- DVC (Data Version Control): stores small pointer files (with content hashes) in Git while the data lives in remote storage (S3, Azure Blob, Google Drive, a server). Checking out a Git commit restores the matching data version.
- Lakehouse table formats (Delta Lake, Apache Iceberg, Apache Hudi): versioned tables with time travel ("query this table as of last Tuesday").
- Dataset registries and immutable, dated snapshots in object storage with manifests.
# DVC basics
git init && dvc init
dvc remote add -d storage s3://my-bucket/dvc-store
dvc add data/registrations_2025.parquet # creates data/registrations_2025.parquet.dvc
git add data/registrations_2025.parquet.dvc data/.gitignore
git commit -m "Add registrations dataset v1"
dvc push # upload data to remote storageData validation#
Validate data before it reaches training or inference. Typical checks:
- Schema: expected columns, types, allowed categories, nullability.
- Ranges and constraints: ages between 0 and 120; dates not in the future; IDs unique.
- Distribution checks: compare feature distributions with a reference (training) dataset to detect drift or pipeline bugs.
- Volume and freshness: row counts within expected bounds; latest timestamp recent enough.
- Referential integrity: joins do not drop or duplicate records unexpectedly.
- Label quality: class balance, annotator agreement, suspicious labels.
import pandas as pd
import pandera as pa
from pandera import Column, Check
schema = pa.DataFrameSchema({
"household_id": Column(str, Check.str_matches(r"^HH-\d{6}$"), unique=True),
"members": Column(int, Check.in_range(1, 30)),
"district": Column(str, Check.isin(["Dhaka", "Chattogram", "Sylhet", "Khulna", "Rajshahi"])),
"registered_at": Column(pd.Timestamp, Check.le(pd.Timestamp.now())),
"monthly_income": Column(float, Check.ge(0), nullable=True),
})
df = pd.DataFrame({"household_id": ["HH-000001", "HH-000002"], "members": [4, 45],
"district": ["Dhaka", "Barishal"], "registered_at": pd.to_datetime(["2025-01-03", "2025-02-11"]),
"monthly_income": [12000.0, None]})
try:
schema.validate(df, lazy=True)
except pa.errors.SchemaErrors as err:
print(err.failure_cases[["column", "check", "failure_case"]])Tools include Pandera, Great Expectations, TensorFlow Data Validation and dbt tests. Decide what happens on failure: block the pipeline, quarantine bad rows, or alert a human.
Pipelines#
A pipeline is a directed acyclic graph (DAG) of steps: ingest → validate → clean → feature engineering → split → train → evaluate → register. Good pipelines are:
- Declarative and reproducible: defined in code/config, with pinned environments.
- Idempotent: re-running a step with the same inputs gives the same outputs.
- Cached: unchanged steps are skipped.
- Observable: logs, lineage and metrics for every run.
# dvc.yaml — a reproducible pipeline; `dvc repro` re-runs only what changed
stages:
validate:
cmd: python src/validate.py data/raw.parquet data/valid.parquet
deps: [src/validate.py, data/raw.parquet]
outs: [data/valid.parquet]
features:
cmd: python src/features.py data/valid.parquet data/features.parquet
deps: [src/features.py, data/valid.parquet]
outs: [data/features.parquet]
train:
cmd: python src/train.py
deps: [src/train.py, data/features.parquet, configs/train.yaml]
outs: [artifacts/model.joblib]
metrics: [artifacts/metrics.json]Orchestrators schedule and monitor pipelines in production: Apache Airflow, Prefect, Dagster, Kubeflow Pipelines, cloud-native services.
Data lineage and privacy#
Lineage records where each dataset came from and how it was transformed — vital for debugging and audits. Combine it with data-protection practices: minimise personal data, pseudonymise identifiers, restrict access, record legal bases and retention periods, and support deletion requests (which means you must know which datasets and models used a person's data).