💬 NLP & Transformers · Lecture 24 of 29

Topic Modelling: LDA, NMF and Neural Topic Models

Topic models discover themes in large document collections without labels. We derive Latent Dirichlet Allocation's generative story, compare it with NMF and embedding-based BERTopic, and discuss evaluating and interpreting topics.

An organisation receives 50,000 free-text survey responses, or a researcher collects a decade of news articles. What themes do they contain? Reading everything is impossible, and there are no labels. Topic models discover recurring themes — topics — automatically, and describe each document as a mixture of them. They are a staple of computational social science, digital humanities and feedback analysis.

What is a topic?#

A topic is a probability distribution over words. A "health" topic puts high probability on clinic, doctor, medicine, vaccine; a "shelter" topic on tent, roof, rain, repair. A document is a mixture of topics: a complaint may be 70% "shelter" and 30% "water".

Latent Dirichlet Allocation (LDA)#

Blei, Ng and Jordan (2003) proposed LDA as a generative probabilistic model of how documents are written:

  1. For each topic $k = 1, \dots, K$: draw a word distribution $\boldsymbol{\phi}_k \sim \text{Dirichlet}(\boldsymbol{\beta})$.
  2. For each document $d$:
    1. draw topic proportions $\boldsymbol{\theta}_d \sim \text{Dirichlet}(\boldsymbol{\alpha})$;
    2. for each word position $n$: draw a topic $z_{dn} \sim \text{Categorical}(\boldsymbol{\theta}_d)$, then a word $w_{dn} \sim \text{Categorical}(\boldsymbol{\phi}_{z_{dn}})$.

We observe only the words; topics and proportions are latent. Inference inverts the story: find the $\boldsymbol{\phi}$, $\boldsymbol{\theta}$ and $z$ that best explain the corpus, using collapsed Gibbs sampling or variational inference (the approach in the original paper and in scikit-learn and gensim).

The Dirichlet priors control sparsity. A small $\alpha$ (< 1) encourages each document to use few topics; a small $\beta$ encourages each topic to concentrate on few words.

LDA is a bag-of-words model: word order is ignored, which is why preprocessing (stop-word removal, lemmatisation, removing very rare and very common words) matters a lot for it.

python
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.decomposition import LatentDirichletAllocation

docs = fetch_20newsgroups(subset="train", remove=("headers", "footers", "quotes"),
                          categories=["sci.med", "sci.space", "rec.autos", "talk.politics.guns"]).data
vec = CountVectorizer(stop_words="english", max_df=0.5, min_df=10, token_pattern=r"(?u)\b[a-zA-Z]{3,}\b")
X = vec.fit_transform(docs)
lda = LatentDirichletAllocation(n_components=4, learning_method="online", random_state=0,
                                doc_topic_prior=0.1, topic_word_prior=0.01).fit(X)
vocab = vec.get_feature_names_out()
for k, comp in enumerate(lda.components_):
    print(f"topic {k}:", ", ".join(vocab[comp.argsort()[-10:][::-1]]))
print("document 0 topic mixture:", lda.transform(X[:1]).round(2))

Non-negative Matrix Factorisation (NMF)#

Factorise the document–term matrix (often TF-IDF) into two non-negative matrices:

$$ \mathbf{X} \approx \mathbf{W}\mathbf{H}, \qquad \mathbf{W}, \mathbf{H} \ge 0 $$

Rows of $\mathbf{H}$ are topics (weights over words); rows of $\mathbf{W}$ are document–topic weights. Non-negativity yields additive, parts-based, interpretable components. NMF is fast, deterministic given initialisation, and often gives crisper topics than LDA on short texts.

Embedding-based topic models: BERTopic#

Short texts (tweets, survey answers, chat messages) contain too few words for bag-of-words models to estimate topic mixtures well. BERTopic (Grootendorst, 2022) takes a different route:

  1. Embed each document with a sentence-transformer (captures meaning, handles synonyms).
  2. Reduce dimensionality with UMAP.
  3. Cluster with HDBSCAN (which also flags outliers).
  4. Describe each cluster with a class-based TF-IDF (c-TF-IDF) — words distinctive of that cluster compared with others.
  5. Optionally label topics with an LLM.

It works well on short and multilingual texts (with multilingual embeddings) and supports dynamic (over-time) topics. Its trade-off: each document is typically assigned one topic rather than a mixture.

python
# pip install bertopic
from bertopic import BERTopic
topic_model = BERTopic(language="multilingual", min_topic_size=15)
topics, probs = topic_model.fit_transform(docs[:2000])
print(topic_model.get_topic_info().head(8))

Evaluating topics#

There is no ground truth, so evaluation combines:

  • Topic coherence: do a topic's top words co-occur in a reference corpus? Metrics such as NPMI or $C_V$ correlate moderately with human judgements.
  • Topic diversity: fraction of unique words across topics' top lists.
  • Held-out perplexity (for LDA) — notably, Chang et al. (2009) found perplexity can be negatively correlated with human interpretability.
  • Human evaluation: word intrusion tests (can people spot a random word inserted into a topic's top words?) and expert review.

Choosing the number of topics#

Try several $K$, compare coherence and diversity, and — above all — inspect topics with domain experts. Too few topics merge distinct themes; too many split them into near-duplicates.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Information Retrieval and Semantic Search

Search is the most used NLP application. We cover indexing, BM25, dense bi-encoder retrieval, cross-encoder re-ranking, hybrid search, approximate nearest neighbours, and evaluation with recall@k, MRR and nDCG.

Intermediate⏱ 5 min#184
💬 NLP & Transformers

Multilingual and Low-Resource NLP (with a Focus on Bangla)

Most of the world's languages have little digital data. We examine why this matters, how multilingual models enable cross-lingual transfer, the specific challenges of languages like Bangla, and practical strategies for building NLP in low-resource settings.

Intermediate⏱ 5 min#186
💬 NLP & Transformers

Evaluating NLP Systems: Perplexity, BLEU, ROUGE, BERTScore and Human Judgement

How do we know if a language system is good? We survey intrinsic and extrinsic evaluation, overlap metrics, embedding-based metrics, learned metrics, LLM judges, human evaluation, benchmarks and their pitfalls.

Intermediate⏱ 5 min#183