An organisation receives 50,000 free-text survey responses, or a researcher collects a decade of news articles. What themes do they contain? Reading everything is impossible, and there are no labels. Topic models discover recurring themes — topics — automatically, and describe each document as a mixture of them. They are a staple of computational social science, digital humanities and feedback analysis.
What is a topic?#
A topic is a probability distribution over words. A "health" topic puts high probability on clinic, doctor, medicine, vaccine; a "shelter" topic on tent, roof, rain, repair. A document is a mixture of topics: a complaint may be 70% "shelter" and 30% "water".
Latent Dirichlet Allocation (LDA)#
Blei, Ng and Jordan (2003) proposed LDA as a generative probabilistic model of how documents are written:
- For each topic $k = 1, \dots, K$: draw a word distribution $\boldsymbol{\phi}_k \sim \text{Dirichlet}(\boldsymbol{\beta})$.
- For each document $d$:
- draw topic proportions $\boldsymbol{\theta}_d \sim \text{Dirichlet}(\boldsymbol{\alpha})$;
- for each word position $n$: draw a topic $z_{dn} \sim \text{Categorical}(\boldsymbol{\theta}_d)$, then a word $w_{dn} \sim \text{Categorical}(\boldsymbol{\phi}_{z_{dn}})$.
We observe only the words; topics and proportions are latent. Inference inverts the story: find the $\boldsymbol{\phi}$, $\boldsymbol{\theta}$ and $z$ that best explain the corpus, using collapsed Gibbs sampling or variational inference (the approach in the original paper and in scikit-learn and gensim).
The Dirichlet priors control sparsity. A small $\alpha$ (< 1) encourages each document to use few topics; a small $\beta$ encourages each topic to concentrate on few words.
LDA is a bag-of-words model: word order is ignored, which is why preprocessing (stop-word removal, lemmatisation, removing very rare and very common words) matters a lot for it.
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.decomposition import LatentDirichletAllocation
docs = fetch_20newsgroups(subset="train", remove=("headers", "footers", "quotes"),
categories=["sci.med", "sci.space", "rec.autos", "talk.politics.guns"]).data
vec = CountVectorizer(stop_words="english", max_df=0.5, min_df=10, token_pattern=r"(?u)\b[a-zA-Z]{3,}\b")
X = vec.fit_transform(docs)
lda = LatentDirichletAllocation(n_components=4, learning_method="online", random_state=0,
doc_topic_prior=0.1, topic_word_prior=0.01).fit(X)
vocab = vec.get_feature_names_out()
for k, comp in enumerate(lda.components_):
print(f"topic {k}:", ", ".join(vocab[comp.argsort()[-10:][::-1]]))
print("document 0 topic mixture:", lda.transform(X[:1]).round(2))Non-negative Matrix Factorisation (NMF)#
Factorise the document–term matrix (often TF-IDF) into two non-negative matrices:
Rows of $\mathbf{H}$ are topics (weights over words); rows of $\mathbf{W}$ are document–topic weights. Non-negativity yields additive, parts-based, interpretable components. NMF is fast, deterministic given initialisation, and often gives crisper topics than LDA on short texts.
Embedding-based topic models: BERTopic#
Short texts (tweets, survey answers, chat messages) contain too few words for bag-of-words models to estimate topic mixtures well. BERTopic (Grootendorst, 2022) takes a different route:
- Embed each document with a sentence-transformer (captures meaning, handles synonyms).
- Reduce dimensionality with UMAP.
- Cluster with HDBSCAN (which also flags outliers).
- Describe each cluster with a class-based TF-IDF (c-TF-IDF) — words distinctive of that cluster compared with others.
- Optionally label topics with an LLM.
It works well on short and multilingual texts (with multilingual embeddings) and supports dynamic (over-time) topics. Its trade-off: each document is typically assigned one topic rather than a mixture.
# pip install bertopic
from bertopic import BERTopic
topic_model = BERTopic(language="multilingual", min_topic_size=15)
topics, probs = topic_model.fit_transform(docs[:2000])
print(topic_model.get_topic_info().head(8))Evaluating topics#
There is no ground truth, so evaluation combines:
- Topic coherence: do a topic's top words co-occur in a reference corpus? Metrics such as NPMI or $C_V$ correlate moderately with human judgements.
- Topic diversity: fraction of unique words across topics' top lists.
- Held-out perplexity (for LDA) — notably, Chang et al. (2009) found perplexity can be negatively correlated with human interpretability.
- Human evaluation: word intrusion tests (can people spot a random word inserted into a topic's top words?) and expert review.
Choosing the number of topics#
Try several $K$, compare coherence and diversity, and — above all — inspect topics with domain experts. Too few topics merge distinct themes; too many split them into near-duplicates.