The most important question when facing a new ML problem is not "which algorithm?" but "what kind of feedback does my data provide?" The answer determines the family of methods available. Today we build a clear taxonomy.
Supervised learning#
Each training example has an input $\mathbf{x}$ and a desired output $y$ (the label). The model learns the mapping $\mathbf{x} \mapsto y$.
- Classification: $y$ is a category (spam/ham, disease A/B/C, digit 0–9).
- Regression: $y$ is a real number (house price, temperature, travel time).
- Structured prediction: $y$ is a structure (a sequence of tags, a parse tree, a segmentation mask, a translated sentence).
Typical algorithms: linear and logistic regression, decision trees, random forests, gradient boosting, SVMs, neural networks.
The bottleneck is labels — they cost human time and expertise. Labelling medical images may require specialists; labelling speech in low-resource languages may require rare language skills.
Unsupervised learning#
Only inputs $\mathbf{x}$ are available. The goal is to discover structure:
- Clustering — group similar items (customer segments, document topics).
- Dimensionality reduction — compress data while preserving structure (PCA, UMAP).
- Density estimation — model $p(\mathbf{x})$ (Gaussian mixtures, normalising flows).
- Anomaly detection — find unusual points (fraud, equipment failure).
Unsupervised results are harder to evaluate because there is no ground truth to compare against. Always validate clusters with domain experts or downstream tasks.
Self-supervised learning#
A powerful modern middle ground: create labels from the data itself by hiding part of the input and predicting it.
- Masked language modelling (BERT): hide 15% of words and predict them.
- Next-token prediction (GPT): predict each word from the previous ones.
- Contrastive learning (SimCLR, CLIP): learn that two augmented views of an image — or an image and its caption — belong together.
- Masked image modelling (MAE): reconstruct hidden image patches.
Self-supervision lets models learn from billions of unlabelled examples. The resulting representations are then adapted to downstream tasks with few labels. This recipe — pretrain, then fine-tune — is the foundation of modern foundation models.
Reinforcement learning#
An agent interacts with an environment, taking actions and receiving rewards. There are no correct labels for each action; feedback is evaluative ("that was good/bad") and often delayed (a chess move's value is revealed only at the end of the game). The agent must balance exploration and exploitation.
Applications: game playing, robotics, resource management, recommendation, and fine-tuning language models from human feedback (RLHF).
Comparison table#
| Paradigm | Feedback | Typical question | Example |
|---|---|---|---|
| Supervised | Correct answer per example | What is $y$ for this $\mathbf{x}$? | Predict loan default |
| Unsupervised | None | What structure exists? | Segment customers |
| Self-supervised | Derived from data | What is missing/next? | Pretrain a language model |
| Reinforcement | Scalar reward, often delayed | Which action maximises long-term reward? | Control a robot arm |
Hybrid settings you will meet in practice#
- Semi-supervised learning — few labels, many unlabelled examples. Techniques: pseudo-labelling, consistency regularisation.
- Weak supervision — noisy labels from heuristics, rules or crowdsourcing, combined statistically.
- Active learning — the model chooses which examples humans should label next to maximise information.
- Transfer learning — reuse a model trained on one task or domain for another.
- Few-shot and zero-shot learning — generalise from a handful of examples, or from a task description alone (as large language models do with prompts).
- Online learning — update the model continuously as data arrives.
- Multi-task learning — train one model on several related tasks to share knowledge.
Choosing a paradigm: a decision guide#
Do you have labelled outputs for your task?
├── Yes, plenty ................................ supervised learning
├── A few, plus lots of unlabelled data ........ semi-supervised / transfer / active learning
├── No, but you want structure ................. unsupervised learning
├── No, but raw data is abundant ............... self-supervised pretraining, then adapt
└── Feedback comes as rewards from actions ..... reinforcement learningfrom sklearn.datasets import make_blobs
from sklearn.cluster import KMeans
from sklearn.linear_model import LogisticRegression
X, y = make_blobs(n_samples=300, centers=3, random_state=42)
# Supervised: we use the labels y
clf = LogisticRegression(max_iter=500).fit(X, y)
print("supervised accuracy:", clf.score(X, y))
# Unsupervised: labels are ignored; we discover 3 groups
km = KMeans(n_clusters=3, n_init=10, random_state=0).fit(X)
print("cluster sizes:", [int((km.labels_ == k).sum()) for k in range(3)])