๐Ÿ“ˆ Machine Learning ยท Lecture 45 of 47

Active Learning: Letting the Model Choose What to Label

If labelling is expensive, label the most informative examples first. We cover pool-based active learning, uncertainty and diversity sampling, query-by-committee, and the practical pitfalls of real annotation loops.

You have 200,000 unlabelled text messages and a budget to label 2,000. Which 2,000? Picking at random wastes effort on easy, redundant examples. Active learning lets the model choose the examples whose labels would teach it the most. In many applications it reaches a target accuracy with a fraction of the labels needed by random sampling.

The pool-based loop#

  1. Start with a small labelled seed set $\mathcal{L}$ and a large unlabelled pool $\mathcal{U}$.
  2. Train a model on $\mathcal{L}$.
  3. Use a query strategy to score examples in $\mathcal{U}$ by informativeness.
  4. Send the top examples (a batch) to human annotators.
  5. Move them from $\mathcal{U}$ to $\mathcal{L}$; retrain; repeat until the budget is spent or performance plateaus.

Uncertainty sampling#

Query the examples the model is least sure about. With predicted class probabilities $p(y \mid \mathbf{x})$:

  • Least confidence: $1 - \max_y p(y \mid \mathbf{x})$.
  • Margin: $p(y_1 \mid \mathbf{x}) - p(y_2 \mid \mathbf{x})$ between the top two classes (smaller = more uncertain).
  • Entropy: $H = -\sum_y p(y \mid \mathbf{x})\log p(y \mid \mathbf{x})$.

For a linear classifier, uncertainty sampling picks points near the decision boundary โ€” exactly where labels refine the boundary most.

Query-by-committee and Bayesian approaches#

Train a committee of models (bootstrap samples, different seeds or architectures) and query examples on which they disagree most (vote entropy or KL divergence from the consensus). For neural networks, BALD (Bayesian Active Learning by Disagreement) selects points maximising mutual information between the prediction and the model parameters:

$$ I(y; \boldsymbol{\theta} \mid \mathbf{x}) = H\big[\mathbb{E}_{\boldsymbol{\theta}}\,p(y \mid \mathbf{x}, \boldsymbol{\theta})\big] - \mathbb{E}_{\boldsymbol{\theta}}\,H\big[p(y \mid \mathbf{x}, \boldsymbol{\theta})\big] $$

It favours points where the model's parameters are uncertain (reducible, epistemic uncertainty) rather than points that are inherently ambiguous (irreducible, aleatoric noise) โ€” an important distinction, since labelling a truly ambiguous example teaches little.

Diversity and representativeness#

Pure uncertainty sampling in batches picks many near-duplicate examples from the same confusing region, and it may favour outliers. Better batch strategies combine uncertainty with diversity:

  • cluster the uncertain candidates and pick one per cluster;
  • core-set selection โ€” choose points that best cover the data in embedding space;
  • BADGE โ€” sample diverse points in the space of gradient embeddings, balancing uncertainty and diversity.

A simulation#

python
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True); X = X / 16
X_pool, X_test, y_pool, y_test = train_test_split(X, y, test_size=0.3, stratify=y, random_state=0)

def run(strategy, budget=300, batch=10, seed=0):
    rng = np.random.default_rng(seed)
    labelled = list(rng.choice(len(X_pool), 20, replace=False))
    curve = []
    while len(labelled) <= budget:
        m = LogisticRegression(max_iter=2000).fit(X_pool[labelled], y_pool[labelled])
        curve.append((len(labelled), m.score(X_test, y_test)))
        rest = np.setdiff1d(np.arange(len(X_pool)), labelled)
        if strategy == "random":
            pick = rng.choice(rest, batch, replace=False)
        else:                                       # margin uncertainty
            P = np.sort(m.predict_proba(X_pool[rest]), axis=1)
            pick = rest[np.argsort(P[:, -1] - P[:, -2])[:batch]]
        labelled += list(pick)
    return curve

for s in ["random", "margin"]:
    c = run(s)
    print(s, [f"{n}:{a:.3f}" for n, a in c[::7]])

Margin sampling typically reaches a given accuracy with noticeably fewer labels than random sampling.

Real-world pitfalls#

Where active learning shines#

  • Medical imaging and document review, where expert labels are costly.
  • Low-resource language tasks, where annotators are scarce.
  • Building classifiers for new categories of humanitarian feedback or incident reports, where the label set evolves.
  • Combined with pretrained embeddings or semi-supervised learning โ€” use unlabelled data for representations and active learning for label efficiency.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Semi-Supervised Learning: Learning from Few Labels and Many Unlabelled Examples

Labels are expensive; unlabelled data is cheap. We study the assumptions that make unlabelled data useful and the main techniques โ€” self-training, label propagation, consistency regularisation and FixMatch.

Intermediateโฑ 5 min#093
๐Ÿ“ˆ Machine Learning

Learning Theory: PAC Learning, VC Dimension and Generalisation Bounds

Why should a model that fits training data work on new data? Learning theory answers precisely. We develop PAC learning, finite-class bounds, the VC dimension, and discuss what these bounds do and do not explain about deep learning.

Advancedโฑ 6 min#095
๐Ÿ“ˆ Machine Learning

Time Series Forecasting: Stationarity, ARIMA and Evaluation

Forecasting demand, rainfall or arrivals needs models that respect time. We decompose series into trend and seasonality, test for stationarity, build ARIMA and SARIMA models, and evaluate with rolling-origin backtesting.

Intermediateโฑ 5 min#092