⚙️ MLOps & Engineering · Lecture 2 of 15

Data Versioning, Validation and Pipelines

Models are only as good as their data — and data changes. We cover versioning datasets, validating schemas and distributions, building reproducible data pipelines, and orchestration tools.

When a model's performance suddenly drops, the cause is usually not the model code — it is the data. A column changed units, a new category appeared, a join duplicated rows, a sensor failed, an upstream team changed a definition. Managing data with the same rigour as code — versioning, validation and reproducible pipelines — is the foundation of reliable ML.

Why version data?#

  • Reproducibility: re-create the exact training set behind any model.
  • Debugging: compare the data of a good model and a bad one.
  • Auditability: answer regulators, partners or affected people about what data a decision relied on.
  • Collaboration: everyone trains on the same, identified dataset.

Git is designed for code, not multi-gigabyte files. Options:

  • DVC (Data Version Control): stores small pointer files (with content hashes) in Git while the data lives in remote storage (S3, Azure Blob, Google Drive, a server). Checking out a Git commit restores the matching data version.
  • Lakehouse table formats (Delta Lake, Apache Iceberg, Apache Hudi): versioned tables with time travel ("query this table as of last Tuesday").
  • Dataset registries and immutable, dated snapshots in object storage with manifests.
bash
# DVC basics
git init && dvc init
dvc remote add -d storage s3://my-bucket/dvc-store
dvc add data/registrations_2025.parquet      # creates data/registrations_2025.parquet.dvc
git add data/registrations_2025.parquet.dvc data/.gitignore
git commit -m "Add registrations dataset v1"
dvc push                                      # upload data to remote storage

Data validation#

Validate data before it reaches training or inference. Typical checks:

  1. Schema: expected columns, types, allowed categories, nullability.
  2. Ranges and constraints: ages between 0 and 120; dates not in the future; IDs unique.
  3. Distribution checks: compare feature distributions with a reference (training) dataset to detect drift or pipeline bugs.
  4. Volume and freshness: row counts within expected bounds; latest timestamp recent enough.
  5. Referential integrity: joins do not drop or duplicate records unexpectedly.
  6. Label quality: class balance, annotator agreement, suspicious labels.
python
import pandas as pd
import pandera as pa
from pandera import Column, Check

schema = pa.DataFrameSchema({
    "household_id": Column(str, Check.str_matches(r"^HH-\d{6}$"), unique=True),
    "members": Column(int, Check.in_range(1, 30)),
    "district": Column(str, Check.isin(["Dhaka", "Chattogram", "Sylhet", "Khulna", "Rajshahi"])),
    "registered_at": Column(pd.Timestamp, Check.le(pd.Timestamp.now())),
    "monthly_income": Column(float, Check.ge(0), nullable=True),
})

df = pd.DataFrame({"household_id": ["HH-000001", "HH-000002"], "members": [4, 45],
                   "district": ["Dhaka", "Barishal"], "registered_at": pd.to_datetime(["2025-01-03", "2025-02-11"]),
                   "monthly_income": [12000.0, None]})
try:
    schema.validate(df, lazy=True)
except pa.errors.SchemaErrors as err:
    print(err.failure_cases[["column", "check", "failure_case"]])

Tools include Pandera, Great Expectations, TensorFlow Data Validation and dbt tests. Decide what happens on failure: block the pipeline, quarantine bad rows, or alert a human.

Pipelines#

A pipeline is a directed acyclic graph (DAG) of steps: ingest → validate → clean → feature engineering → split → train → evaluate → register. Good pipelines are:

  • Declarative and reproducible: defined in code/config, with pinned environments.
  • Idempotent: re-running a step with the same inputs gives the same outputs.
  • Cached: unchanged steps are skipped.
  • Observable: logs, lineage and metrics for every run.
yaml
# dvc.yaml — a reproducible pipeline; `dvc repro` re-runs only what changed
stages:
  validate:
    cmd: python src/validate.py data/raw.parquet data/valid.parquet
    deps: [src/validate.py, data/raw.parquet]
    outs: [data/valid.parquet]
  features:
    cmd: python src/features.py data/valid.parquet data/features.parquet
    deps: [src/features.py, data/valid.parquet]
    outs: [data/features.parquet]
  train:
    cmd: python src/train.py
    deps: [src/train.py, data/features.parquet, configs/train.yaml]
    outs: [artifacts/model.joblib]
    metrics: [artifacts/metrics.json]

Orchestrators schedule and monitor pipelines in production: Apache Airflow, Prefect, Dagster, Kubeflow Pipelines, cloud-native services.

Data lineage and privacy#

Lineage records where each dataset came from and how it was transformed — vital for debugging and audits. Combine it with data-protection practices: minimise personal data, pseudonymise identifiers, restrict access, record legal bases and retention periods, and support deletion requests (which means you must know which datasets and models used a person's data).

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

What Is MLOps? From Notebook to Reliable Production System

Most ML projects fail not because of the model but because of everything around it. We define MLOps, walk through the ML lifecycle, describe maturity levels, and examine the hidden technical debt of machine learning systems.

Beginner⏱ 5 min#242
⚙️ MLOps & Engineering

Experiment Tracking: Never Lose a Result Again

ML development involves hundreds of runs with different data, code and hyperparameters. We cover what to track, how to use MLflow for runs, metrics and artefacts, the model registry, and good experiment hygiene.

Beginner⏱ 4 min#244
⚙️ MLOps & Engineering

Reproducibility in Machine Learning

Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.

Intermediate⏱ 4 min#245