Andrew Ng's "data-centric AI" movement argues that, for many applications, improving the data — especially labels — yields bigger gains than tweaking models. Yet labelling is often treated as a cheap afterthought. Studies have found label errors even in famous benchmark test sets (Northcutt et al., 2021, estimated an average of several percent across ten popular datasets), enough to change which model appears best. This lecture covers how to create labels you can trust.
Step 1: define the task precisely#
Write annotation guidelines before labelling begins:
- Clear definitions of every label, with positive and negative examples.
- Rules for ambiguous and edge cases ("if a message mentions both food and shelter, label both" or "label the primary need").
- What to do when unsure (an "unclear" option, or escalation).
- For spans/boxes/masks: exact boundary rules (include the title "Dr." in a PERSON entity? box around visible pixels only or the whole occluded object?).
Pilot the guidelines on a small batch, discuss disagreements, and revise. Good guidelines evolve through several iterations.
Step 2: choose annotators and tools#
- Domain experts (clinicians, protection officers, agronomists) for specialised judgements — expensive but essential for high-stakes labels.
- Trained annotators with guidelines for general tasks.
- Crowdsourcing for simple, high-volume tasks — with quality control.
- Native speakers for language tasks — critical for multilingual data.
Tools: Label Studio, CVAT (vision), doccano (text), Prodigy, Argilla, cloud labelling services. Choose tools that support your data types, quality workflows and data security requirements (self-hosting for sensitive data).
Step 3: measure agreement#
Have multiple annotators label an overlapping subset. Raw agreement overstates reliability because some agreement happens by chance. Cohen's kappa corrects for this:
where $p_o$ is observed agreement and $p_e$ is agreement expected by chance. For more than two annotators, use Fleiss' kappa or Krippendorff's alpha. Rough interpretation: above 0.8 strong, 0.6–0.8 substantial, below 0.4 weak — but acceptable levels depend on the task.
from sklearn.metrics import cohen_kappa_score, confusion_matrix
a1 = ["food", "shelter", "health", "food", "water", "health", "food", "shelter", "water", "food"]
a2 = ["food", "shelter", "health", "water", "water", "health", "food", "health", "water", "food"]
print("raw agreement:", sum(x == y for x, y in zip(a1, a2)) / len(a1))
print("Cohen's kappa:", round(cohen_kappa_score(a1, a2), 3))
labels = ["food", "water", "shelter", "health"]
print(confusion_matrix(a1, a2, labels=labels)) # where do annotators disagree?Low agreement signals unclear guidelines or a genuinely subjective task — no model can reliably exceed the consistency of its labels.
Step 4: quality control#
- Gold questions: items with known answers mixed into tasks to monitor annotator accuracy.
- Multiple annotations + aggregation: majority vote, or statistical models (Dawid–Skene) that estimate each annotator's reliability.
- Adjudication: an expert resolves disagreements.
- Regular feedback to annotators and guideline updates.
- Audit samples throughout, not only at the start.
Handling label noise#
Some noise is unavoidable. Strategies:
- Find likely errors: train a model with cross-validation and flag examples where it confidently disagrees with the label (the idea behind confident learning and the cleanlab library); send them for review.
- Robust training: label smoothing, noise-robust losses, co-teaching.
- Keep disagreement as information: for subjective tasks (e.g. offensiveness), store the distribution of annotator labels rather than forcing a single "truth" — annotators' perspectives may legitimately differ.
Model-assisted labelling and active learning#
- Pre-annotation: a model (or an LLM, or SAM for masks) proposes labels; humans correct them — often several times faster. Beware anchoring bias: annotators may accept wrong suggestions. Audit with unassisted gold items.
- Active learning: prioritise the most informative items for labelling (see the active learning lecture).
- Weak supervision: combine heuristic labelling functions statistically (e.g. Snorkel) to create noisy labels at scale.
Annotators are people#
Document your dataset#
Record who labelled the data, how, with which guidelines, agreement statistics, known limitations and intended uses — a datasheet for datasets (Gebru et al., 2021; see the documentation lecture).