You have 200,000 unlabelled text messages and a budget to label 2,000. Which 2,000? Picking at random wastes effort on easy, redundant examples. Active learning lets the model choose the examples whose labels would teach it the most. In many applications it reaches a target accuracy with a fraction of the labels needed by random sampling.
The pool-based loop#
- Start with a small labelled seed set $\mathcal{L}$ and a large unlabelled pool $\mathcal{U}$.
- Train a model on $\mathcal{L}$.
- Use a query strategy to score examples in $\mathcal{U}$ by informativeness.
- Send the top examples (a batch) to human annotators.
- Move them from $\mathcal{U}$ to $\mathcal{L}$; retrain; repeat until the budget is spent or performance plateaus.
Uncertainty sampling#
Query the examples the model is least sure about. With predicted class probabilities $p(y \mid \mathbf{x})$:
- Least confidence: $1 - \max_y p(y \mid \mathbf{x})$.
- Margin: $p(y_1 \mid \mathbf{x}) - p(y_2 \mid \mathbf{x})$ between the top two classes (smaller = more uncertain).
- Entropy: $H = -\sum_y p(y \mid \mathbf{x})\log p(y \mid \mathbf{x})$.
For a linear classifier, uncertainty sampling picks points near the decision boundary โ exactly where labels refine the boundary most.
Query-by-committee and Bayesian approaches#
Train a committee of models (bootstrap samples, different seeds or architectures) and query examples on which they disagree most (vote entropy or KL divergence from the consensus). For neural networks, BALD (Bayesian Active Learning by Disagreement) selects points maximising mutual information between the prediction and the model parameters:
It favours points where the model's parameters are uncertain (reducible, epistemic uncertainty) rather than points that are inherently ambiguous (irreducible, aleatoric noise) โ an important distinction, since labelling a truly ambiguous example teaches little.
Diversity and representativeness#
Pure uncertainty sampling in batches picks many near-duplicate examples from the same confusing region, and it may favour outliers. Better batch strategies combine uncertainty with diversity:
- cluster the uncertain candidates and pick one per cluster;
- core-set selection โ choose points that best cover the data in embedding space;
- BADGE โ sample diverse points in the space of gradient embeddings, balancing uncertainty and diversity.
A simulation#
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True); X = X / 16
X_pool, X_test, y_pool, y_test = train_test_split(X, y, test_size=0.3, stratify=y, random_state=0)
def run(strategy, budget=300, batch=10, seed=0):
rng = np.random.default_rng(seed)
labelled = list(rng.choice(len(X_pool), 20, replace=False))
curve = []
while len(labelled) <= budget:
m = LogisticRegression(max_iter=2000).fit(X_pool[labelled], y_pool[labelled])
curve.append((len(labelled), m.score(X_test, y_test)))
rest = np.setdiff1d(np.arange(len(X_pool)), labelled)
if strategy == "random":
pick = rng.choice(rest, batch, replace=False)
else: # margin uncertainty
P = np.sort(m.predict_proba(X_pool[rest]), axis=1)
pick = rest[np.argsort(P[:, -1] - P[:, -2])[:batch]]
labelled += list(pick)
return curve
for s in ["random", "margin"]:
c = run(s)
print(s, [f"{n}:{a:.3f}" for n, a in c[::7]])Margin sampling typically reaches a given accuracy with noticeably fewer labels than random sampling.
Real-world pitfalls#
Where active learning shines#
- Medical imaging and document review, where expert labels are costly.
- Low-resource language tasks, where annotators are scarce.
- Building classifiers for new categories of humanitarian feedback or incident reports, where the label set evolves.
- Combined with pretrained embeddings or semi-supervised learning โ use unlabelled data for representations and active learning for label efficiency.