⚙️ MLOps & Engineering · Lecture 12 of 15

Data Labelling and Annotation: Building High-Quality Datasets

Labels are the foundation of supervised learning, yet labelling is often rushed. We cover annotation guidelines, workflows and tools, measuring agreement, handling label noise, model-assisted labelling, and fair treatment of annotators.

Andrew Ng's "data-centric AI" movement argues that, for many applications, improving the data — especially labels — yields bigger gains than tweaking models. Yet labelling is often treated as a cheap afterthought. Studies have found label errors even in famous benchmark test sets (Northcutt et al., 2021, estimated an average of several percent across ten popular datasets), enough to change which model appears best. This lecture covers how to create labels you can trust.

Step 1: define the task precisely#

Write annotation guidelines before labelling begins:

  • Clear definitions of every label, with positive and negative examples.
  • Rules for ambiguous and edge cases ("if a message mentions both food and shelter, label both" or "label the primary need").
  • What to do when unsure (an "unclear" option, or escalation).
  • For spans/boxes/masks: exact boundary rules (include the title "Dr." in a PERSON entity? box around visible pixels only or the whole occluded object?).

Pilot the guidelines on a small batch, discuss disagreements, and revise. Good guidelines evolve through several iterations.

Step 2: choose annotators and tools#

  • Domain experts (clinicians, protection officers, agronomists) for specialised judgements — expensive but essential for high-stakes labels.
  • Trained annotators with guidelines for general tasks.
  • Crowdsourcing for simple, high-volume tasks — with quality control.
  • Native speakers for language tasks — critical for multilingual data.

Tools: Label Studio, CVAT (vision), doccano (text), Prodigy, Argilla, cloud labelling services. Choose tools that support your data types, quality workflows and data security requirements (self-hosting for sensitive data).

Step 3: measure agreement#

Have multiple annotators label an overlapping subset. Raw agreement overstates reliability because some agreement happens by chance. Cohen's kappa corrects for this:

$$ \kappa = \frac{p_o - p_e}{1 - p_e} $$

where $p_o$ is observed agreement and $p_e$ is agreement expected by chance. For more than two annotators, use Fleiss' kappa or Krippendorff's alpha. Rough interpretation: above 0.8 strong, 0.6–0.8 substantial, below 0.4 weak — but acceptable levels depend on the task.

python
from sklearn.metrics import cohen_kappa_score, confusion_matrix

a1 = ["food", "shelter", "health", "food", "water", "health", "food", "shelter", "water", "food"]
a2 = ["food", "shelter", "health", "water", "water", "health", "food", "health", "water", "food"]
print("raw agreement:", sum(x == y for x, y in zip(a1, a2)) / len(a1))
print("Cohen's kappa:", round(cohen_kappa_score(a1, a2), 3))
labels = ["food", "water", "shelter", "health"]
print(confusion_matrix(a1, a2, labels=labels))        # where do annotators disagree?

Low agreement signals unclear guidelines or a genuinely subjective task — no model can reliably exceed the consistency of its labels.

Step 4: quality control#

  • Gold questions: items with known answers mixed into tasks to monitor annotator accuracy.
  • Multiple annotations + aggregation: majority vote, or statistical models (Dawid–Skene) that estimate each annotator's reliability.
  • Adjudication: an expert resolves disagreements.
  • Regular feedback to annotators and guideline updates.
  • Audit samples throughout, not only at the start.

Handling label noise#

Some noise is unavoidable. Strategies:

  • Find likely errors: train a model with cross-validation and flag examples where it confidently disagrees with the label (the idea behind confident learning and the cleanlab library); send them for review.
  • Robust training: label smoothing, noise-robust losses, co-teaching.
  • Keep disagreement as information: for subjective tasks (e.g. offensiveness), store the distribution of annotator labels rather than forcing a single "truth" — annotators' perspectives may legitimately differ.

Model-assisted labelling and active learning#

  • Pre-annotation: a model (or an LLM, or SAM for masks) proposes labels; humans correct them — often several times faster. Beware anchoring bias: annotators may accept wrong suggestions. Audit with unassisted gold items.
  • Active learning: prioritise the most informative items for labelling (see the active learning lecture).
  • Weak supervision: combine heuristic labelling functions statistically (e.g. Snorkel) to create noisy labels at scale.

Annotators are people#

Document your dataset#

Record who labelled the data, how, with which guidelines, agreement statistics, known limitations and intended uses — a datasheet for datasets (Gebru et al., 2021; see the documentation lecture).

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

GPUs and Hardware for Machine Learning

Understanding hardware helps you train faster and cheaper. We explain why GPUs suit deep learning, the roles of memory capacity and bandwidth, precision and tensor cores, estimating requirements, and choosing between local, cloud and free resources.

Intermediate⏱ 5 min#252
⚙️ MLOps & Engineering

A/B Testing and Online Evaluation of ML Models

Offline metrics do not guarantee real-world impact. Online experiments measure what a model actually changes. We design randomised A/B tests for ML, compute sample sizes, avoid common pitfalls, and discuss ethics of experimenting with people.

Intermediate⏱ 5 min#254
⚙️ MLOps & Engineering

Edge AI and TinyML: Running Models on Phones and Microcontrollers

Running models on-device brings privacy, offline operation, low latency and low cost. We cover the edge hardware spectrum, the optimisation pipeline, TensorFlow Lite, ONNX Runtime and TinyML on microcontrollers, and field-deployment lessons.

Intermediate⏱ 5 min#251