Routing support tickets, flagging urgent messages, detecting spam and hate speech, categorising news, identifying the topic of citizen feedback, triaging requests for assistance โ text classification is the most common NLP task in production. Today we build classifiers of increasing sophistication and learn to choose among them.
Task variants#
- Binary: spam vs not spam.
- Multiclass: one of $K$ categories (topic).
- Multi-label: several labels at once (a message can be both "health" and "urgent").
- Hierarchical: categories with sub-categories.
Step 1: design the labels#
Clear, mutually exclusive (or explicitly multi-label) categories with written annotation guidelines and examples. Measure inter-annotator agreement (Cohen's or Fleiss' kappa) on a sample; low agreement means the task is ill-defined, and no model can do better than the labels.
Step 2: strong classical baseline#
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score
baseline = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2), min_df=2, sublinear_tf=True),
LogisticRegression(max_iter=2000, class_weight="balanced"))
# scores = cross_val_score(baseline, texts, labels, cv=5, scoring="f1_macro")Character n-grams (analyzer="char_wb", ngram_range=(2, 5)) help with misspellings, informal text and morphologically rich languages.
Step 3: neural classifiers (historical context)#
- CNN for text (Kim, 2014): convolutions of widths 3โ5 over word embeddings, max-pooled over time โ detects key phrases regardless of position.
- BiLSTM with attention: reads the sequence in both directions and attends to the most informative words.
These improved on linear models but have been largely superseded by pretrained transformers.
Step 4: fine-tune a pretrained transformer#
Add a classification head on the [CLS] (or pooled) representation of a pretrained encoder and fine-tune the whole model:
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
TrainingArguments, Trainer, DataCollatorWithPadding)
import numpy as np, evaluate
ds = load_dataset("ag_news") # 4 news topics
ckpt = "distilbert-base-uncased"
tok = AutoTokenizer.from_pretrained(ckpt)
ds = ds.map(lambda b: tok(b["text"], truncation=True, max_length=128), batched=True)
model = AutoModelForSequenceClassification.from_pretrained(ckpt, num_labels=4)
f1 = evaluate.load("f1")
def metrics(p):
return f1.compute(predictions=np.argmax(p.predictions, -1), references=p.label_ids, average="macro")
args = TrainingArguments("out", learning_rate=2e-5, per_device_train_batch_size=32, num_train_epochs=2,
weight_decay=0.01, eval_strategy="epoch", warmup_ratio=0.06, fp16=True)
trainer = Trainer(model=model, args=args, train_dataset=ds["train"].shuffle(seed=0).select(range(20000)),
eval_dataset=ds["test"], data_collator=DataCollatorWithPadding(tok), compute_metrics=metrics)
trainer.train()Typical hyperparameters: learning rate $1\times10^{-5}$ to $5\times10^{-5}$, 2โ4 epochs, warm-up, weight decay 0.01. Fine-tuning usually gives the best accuracy when a few thousand labelled examples are available.
Other options#
- Sentence embeddings + logistic regression: embed texts with a sentence-transformer model and train a linear classifier โ fast, strong with small data, easy to update.
- Few-shot / zero-shot with LLMs: describe categories in a prompt and let a large language model classify. Excellent for bootstrapping and rare categories, but more expensive per prediction and must be evaluated carefully for consistency.
- SetFit: contrastive fine-tuning of sentence transformers with very few labels (e.g. 8 per class).
Choosing an approach#
| Labelled data | Latency/cost constraints | Recommended |
|---|---|---|
| None | Moderate | Zero-shot LLM, then collect labels |
| < 100 per class | Any | Sentence embeddings + linear model, SetFit, or few-shot LLM |
| Thousands | Low latency | Fine-tuned small transformer (DistilBERT, MiniLM) |
| Thousands | Very low resources | TF-IDF + linear model |
| Many languages | โ | Multilingual encoder (e.g. XLM-R) or multilingual embeddings |
Practical concerns#
- Imbalance: class weights, threshold tuning, macro-F1 evaluation.
- Long documents: truncate wisely (beginning + end), chunk and aggregate, or use long-context models.
- Multilingual and code-mixed text: use multilingual models and evaluate per language.
- Label drift: new topics appear (a new disease outbreak, a new policy); monitor and retrain.
- Error analysis: read misclassified examples; many will be ambiguous or mislabelled.
- Harm: classifiers used for moderation or prioritisation can encode bias against dialects or groups โ audit per subgroup.