A model file tells you nothing about what it is for, how well it works, for whom it fails, or what data made it. Without documentation, models get reused in contexts they were never designed for — with predictable harm. In 2018–2019, researchers proposed simple, structured documentation practices: model cards (Mitchell et al., 2019) and datasheets for datasets (Gebru et al., 2018/2021). They have since become widely adopted norms — Hugging Face model pages use model cards — and they help meet the transparency obligations of emerging AI regulations.
Model cards#
A model card is a short document accompanying a trained model. Recommended sections (adapted from Mitchell et al.):
- Model details: developers, date, version, architecture, training approach, licence, contact.
- Intended use: primary intended uses and users; out-of-scope uses.
- Factors: relevant groups (demographics, languages, regions), instruments (devices, sensors) and environments that may affect performance.
- Metrics: which metrics, why, decision thresholds, and how uncertainty is measured.
- Evaluation data: datasets, why chosen, preprocessing.
- Training data: sources and characteristics (or reference to the datasheet).
- Quantitative analyses: results disaggregated by relevant factors and their intersections.
- Ethical considerations: sensitive data, risks, harms, mitigations.
- Caveats and recommendations: known limitations, conditions for safe use, monitoring advice.
The most important innovation is disaggregated evaluation: reporting performance for different groups, not just an overall number.
A model card template#
# Model Card: Household Priority Classifier v2.1
## Model details
- Developed by: Data & Analytics Unit, (organisation). Contact: data-team@example.org
- Version 2.1, trained 2025-09-10. Gradient-boosted trees (LightGBM). Licence: internal use.
## Intended use
- Supports caseworkers in **ordering** follow-up visits. Human caseworkers make all decisions.
- Out of scope: determining eligibility, reducing or denying assistance, use outside the three pilot districts.
## Factors
- District, household size, language of interview, presence of disability, head-of-household gender.
## Metrics
- Recall of "urgent" households at the operating threshold (primary), precision, calibration (ECE).
- 95% confidence intervals by bootstrap.
## Evaluation data
- 4,200 households surveyed Jan–Jun 2025, labelled by two protection officers (Cohen's kappa 0.78).
## Quantitative analyses
| Group | n | Recall (urgent) | Precision |
|------------------|-------|-----------------|-----------|
| All | 4,200 | 0.86 [0.83, 0.89] | 0.61 |
| District A | 1,900 | 0.88 | 0.63 |
| District C | 700 | 0.79 | 0.55 |
| Disability = yes | 520 | 0.84 | 0.58 |
## Ethical considerations
- Uses sensitive personal data (processed under the data-protection policy; pseudonymised).
- Risk of under-prioritising households under-represented in training data (District C) — mitigated by manual review quotas.
## Caveats and recommendations
- Retrain or re-validate if intake forms change. Monitor recall by district monthly.
- Do not use scores as a proxy for "deservingness".Datasheets for datasets#
Inspired by electronics datasheets, a datasheet answers questions across a dataset's lifecycle:
- Motivation: why was the dataset created, by whom, funded by whom?
- Composition: what do instances represent; how many; missing data; sensitive information; does it relate to people?
- Collection process: how, when, by whom; consent; sampling strategy.
- Preprocessing/cleaning/labelling: what was done; is raw data available; annotation guidelines and agreement.
- Uses: what it has been used for; what it should not be used for.
- Distribution: licence, access restrictions.
- Maintenance: who maintains it; how errors are reported and corrected; versioning; deletion requests.
Related formats include Data Statements for NLP (describing speaker demographics, language varieties and annotator backgrounds) and Data Cards.
System cards and AI documentation at scale#
For complex systems (e.g. LLM applications combining models, retrieval, tools and filters), system cards document the whole system: components, safety evaluations, red-teaming results, mitigations and deployment safeguards. Major AI developers publish system cards for frontier models; organisations deploying AI should maintain internal equivalents.
Why documentation matters#
- Informed use: prevents misuse outside intended conditions.
- Accountability: records decisions, responsibilities and known risks.
- Auditing: gives internal and external reviewers what they need.
- Regulation: many emerging frameworks require documentation of high-risk systems — technical documentation, data governance, performance, human oversight and logging (e.g. under the EU AI Act for high-risk systems).
- Institutional memory: survives staff turnover.