⚙️ MLOps & Engineering · Lecture 15 of 15

Model Cards, Datasheets and Responsible Documentation

Documentation is how models and datasets are understood, audited and used responsibly. We cover model cards, datasheets for datasets, system cards, what to include, and how documentation supports accountability and regulation.

A model file tells you nothing about what it is for, how well it works, for whom it fails, or what data made it. Without documentation, models get reused in contexts they were never designed for — with predictable harm. In 2018–2019, researchers proposed simple, structured documentation practices: model cards (Mitchell et al., 2019) and datasheets for datasets (Gebru et al., 2018/2021). They have since become widely adopted norms — Hugging Face model pages use model cards — and they help meet the transparency obligations of emerging AI regulations.

Model cards#

A model card is a short document accompanying a trained model. Recommended sections (adapted from Mitchell et al.):

  1. Model details: developers, date, version, architecture, training approach, licence, contact.
  2. Intended use: primary intended uses and users; out-of-scope uses.
  3. Factors: relevant groups (demographics, languages, regions), instruments (devices, sensors) and environments that may affect performance.
  4. Metrics: which metrics, why, decision thresholds, and how uncertainty is measured.
  5. Evaluation data: datasets, why chosen, preprocessing.
  6. Training data: sources and characteristics (or reference to the datasheet).
  7. Quantitative analyses: results disaggregated by relevant factors and their intersections.
  8. Ethical considerations: sensitive data, risks, harms, mitigations.
  9. Caveats and recommendations: known limitations, conditions for safe use, monitoring advice.

The most important innovation is disaggregated evaluation: reporting performance for different groups, not just an overall number.

A model card template#

markdown
# Model Card: Household Priority Classifier v2.1

## Model details
- Developed by: Data & Analytics Unit, (organisation). Contact: data-team@example.org
- Version 2.1, trained 2025-09-10. Gradient-boosted trees (LightGBM). Licence: internal use.

## Intended use
- Supports caseworkers in **ordering** follow-up visits. Human caseworkers make all decisions.
- Out of scope: determining eligibility, reducing or denying assistance, use outside the three pilot districts.

## Factors
- District, household size, language of interview, presence of disability, head-of-household gender.

## Metrics
- Recall of "urgent" households at the operating threshold (primary), precision, calibration (ECE).
- 95% confidence intervals by bootstrap.

## Evaluation data
- 4,200 households surveyed Jan–Jun 2025, labelled by two protection officers (Cohen's kappa 0.78).

## Quantitative analyses
| Group            | n     | Recall (urgent) | Precision |
|------------------|-------|-----------------|-----------|
| All              | 4,200 | 0.86 [0.83, 0.89] | 0.61    |
| District A       | 1,900 | 0.88            | 0.63      |
| District C       |   700 | 0.79            | 0.55      |
| Disability = yes |   520 | 0.84            | 0.58      |

## Ethical considerations
- Uses sensitive personal data (processed under the data-protection policy; pseudonymised).
- Risk of under-prioritising households under-represented in training data (District C) — mitigated by manual review quotas.

## Caveats and recommendations
- Retrain or re-validate if intake forms change. Monitor recall by district monthly.
- Do not use scores as a proxy for "deservingness".

Datasheets for datasets#

Inspired by electronics datasheets, a datasheet answers questions across a dataset's lifecycle:

  • Motivation: why was the dataset created, by whom, funded by whom?
  • Composition: what do instances represent; how many; missing data; sensitive information; does it relate to people?
  • Collection process: how, when, by whom; consent; sampling strategy.
  • Preprocessing/cleaning/labelling: what was done; is raw data available; annotation guidelines and agreement.
  • Uses: what it has been used for; what it should not be used for.
  • Distribution: licence, access restrictions.
  • Maintenance: who maintains it; how errors are reported and corrected; versioning; deletion requests.

Related formats include Data Statements for NLP (describing speaker demographics, language varieties and annotator backgrounds) and Data Cards.

System cards and AI documentation at scale#

For complex systems (e.g. LLM applications combining models, retrieval, tools and filters), system cards document the whole system: components, safety evaluations, red-teaming results, mitigations and deployment safeguards. Major AI developers publish system cards for frontier models; organisations deploying AI should maintain internal equivalents.

Why documentation matters#

  • Informed use: prevents misuse outside intended conditions.
  • Accountability: records decisions, responsibilities and known risks.
  • Auditing: gives internal and external reviewers what they need.
  • Regulation: many emerging frameworks require documentation of high-risk systems — technical documentation, data governance, performance, human oversight and logging (e.g. under the EU AI Act for high-risk systems).
  • Institutional memory: survives staff turnover.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

From Notebook to Production Code: Structuring Clean ML Projects

Notebooks are great for exploration and terrible for production. We cover a clean project layout, configuration, modular code, typing and testing, logging, packaging, and a workflow for moving from exploration to maintainable software.

Beginner⏱ 5 min#255
⚙️ MLOps & Engineering

A/B Testing and Online Evaluation of ML Models

Offline metrics do not guarantee real-world impact. Online experiments measure what a model actually changes. We design randomised A/B tests for ML, compute sample sizes, avoid common pitfalls, and discuss ethics of experimenting with people.

Intermediate⏱ 5 min#254
⚙️ MLOps & Engineering

Data Labelling and Annotation: Building High-Quality Datasets

Labels are the foundation of supervised learning, yet labelling is often rushed. We cover annotation guidelines, workflows and tools, measuring agreement, handling label noise, model-assisted labelling, and fair treatment of annotators.

Beginner⏱ 5 min#253