๐Ÿ’ฌ NLP & Transformers ยท Lecture 8 of 29

Text Classification: From Linear Models to Fine-Tuned Transformers

Text classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.

Routing support tickets, flagging urgent messages, detecting spam and hate speech, categorising news, identifying the topic of citizen feedback, triaging requests for assistance โ€” text classification is the most common NLP task in production. Today we build classifiers of increasing sophistication and learn to choose among them.

Task variants#

  • Binary: spam vs not spam.
  • Multiclass: one of $K$ categories (topic).
  • Multi-label: several labels at once (a message can be both "health" and "urgent").
  • Hierarchical: categories with sub-categories.

Step 1: design the labels#

Clear, mutually exclusive (or explicitly multi-label) categories with written annotation guidelines and examples. Measure inter-annotator agreement (Cohen's or Fleiss' kappa) on a sample; low agreement means the task is ill-defined, and no model can do better than the labels.

Step 2: strong classical baseline#

python
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score

baseline = make_pipeline(
    TfidfVectorizer(ngram_range=(1, 2), min_df=2, sublinear_tf=True),
    LogisticRegression(max_iter=2000, class_weight="balanced"))
# scores = cross_val_score(baseline, texts, labels, cv=5, scoring="f1_macro")

Character n-grams (analyzer="char_wb", ngram_range=(2, 5)) help with misspellings, informal text and morphologically rich languages.

Step 3: neural classifiers (historical context)#

  • CNN for text (Kim, 2014): convolutions of widths 3โ€“5 over word embeddings, max-pooled over time โ€” detects key phrases regardless of position.
  • BiLSTM with attention: reads the sequence in both directions and attends to the most informative words.

These improved on linear models but have been largely superseded by pretrained transformers.

Step 4: fine-tune a pretrained transformer#

Add a classification head on the [CLS] (or pooled) representation of a pretrained encoder and fine-tune the whole model:

python
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
                          TrainingArguments, Trainer, DataCollatorWithPadding)
import numpy as np, evaluate

ds = load_dataset("ag_news")                                   # 4 news topics
ckpt = "distilbert-base-uncased"
tok = AutoTokenizer.from_pretrained(ckpt)
ds = ds.map(lambda b: tok(b["text"], truncation=True, max_length=128), batched=True)
model = AutoModelForSequenceClassification.from_pretrained(ckpt, num_labels=4)

f1 = evaluate.load("f1")
def metrics(p):
    return f1.compute(predictions=np.argmax(p.predictions, -1), references=p.label_ids, average="macro")

args = TrainingArguments("out", learning_rate=2e-5, per_device_train_batch_size=32, num_train_epochs=2,
                         weight_decay=0.01, eval_strategy="epoch", warmup_ratio=0.06, fp16=True)
trainer = Trainer(model=model, args=args, train_dataset=ds["train"].shuffle(seed=0).select(range(20000)),
                  eval_dataset=ds["test"], data_collator=DataCollatorWithPadding(tok), compute_metrics=metrics)
trainer.train()

Typical hyperparameters: learning rate $1\times10^{-5}$ to $5\times10^{-5}$, 2โ€“4 epochs, warm-up, weight decay 0.01. Fine-tuning usually gives the best accuracy when a few thousand labelled examples are available.

Other options#

  • Sentence embeddings + logistic regression: embed texts with a sentence-transformer model and train a linear classifier โ€” fast, strong with small data, easy to update.
  • Few-shot / zero-shot with LLMs: describe categories in a prompt and let a large language model classify. Excellent for bootstrapping and rare categories, but more expensive per prediction and must be evaluated carefully for consistency.
  • SetFit: contrastive fine-tuning of sentence transformers with very few labels (e.g. 8 per class).

Choosing an approach#

Labelled dataLatency/cost constraintsRecommended
NoneModerateZero-shot LLM, then collect labels
< 100 per classAnySentence embeddings + linear model, SetFit, or few-shot LLM
ThousandsLow latencyFine-tuned small transformer (DistilBERT, MiniLM)
ThousandsVery low resourcesTF-IDF + linear model
Many languagesโ€”Multilingual encoder (e.g. XLM-R) or multilingual embeddings

Practical concerns#

  • Imbalance: class weights, threshold tuning, macro-F1 evaluation.
  • Long documents: truncate wisely (beginning + end), chunk and aggregate, or use long-context models.
  • Multilingual and code-mixed text: use multilingual models and evaluate per language.
  • Label drift: new topics appear (a new disease outbreak, a new policy); monitor and retrain.
  • Error analysis: read misclassified examples; many will be ambiguous or mislabelled.
  • Harm: classifiers used for moderation or prioritisation can encode bias against dialects or groups โ€” audit per subgroup.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

BERT: Bidirectional Encoder Representations from Transformers

BERT showed that a bidirectional transformer pretrained on unlabelled text could be fine-tuned to beat task-specific models across NLP. We cover masked language modelling, input format, fine-tuning patterns, and successors such as RoBERTa, DeBERTa and multilingual encoders.

Intermediateโฑ 5 min#177
๐Ÿ’ฌ NLP & Transformers

Bag of Words and TF-IDF: Classical Text Representation

The simplest way to turn documents into vectors is to count words. We build bag-of-words and TF-IDF representations, derive the IDF formula, use cosine similarity for retrieval, and train strong linear text classifiers.

Beginnerโฑ 5 min#164
๐Ÿ’ฌ NLP & Transformers

Fine-Tuning Pretrained Language Models: A Practical Guide

How to adapt a pretrained language model to your task reliably โ€” choosing a model, preparing data, hyperparameters, handling small data and instability, continued pretraining, and evaluating properly.

Intermediateโฑ 5 min#180