๐Ÿ’ฌ NLP & Transformers ยท Lecture 6 of 29

GloVe and FastText: Global Statistics and Subword Embeddings

GloVe learns embeddings from a global co-occurrence matrix with a weighted least-squares objective; FastText represents words as bags of character n-grams, handling rare and unseen words. We compare them with word2vec and discuss evaluation.

Word2vec learns from local context windows one pair at a time. Two influential follow-ups improved on it in complementary ways. GloVe (Stanford, 2014) exploits global co-occurrence statistics directly. FastText (Facebook AI Research, 2016โ€“2017) builds word vectors from character n-grams, so it can represent rare, misspelled and never-seen words โ€” invaluable for morphologically rich languages.

GloVe: Global Vectors#

Pennington, Socher and Manning started from an observation about ratios of co-occurrence probabilities. Let $P_{ij} = P(j \mid i)$ be the probability that word $j$ appears in the context of word $i$. Consider $i$ = "ice", $j$ = "steam":

Probe word $k$$P(k \mid \text{ice})/P(k \mid \text{steam})$
solidlarge (โ‰ซ 1)
gassmall (โ‰ช 1)
waterโ‰ˆ 1 (related to both)
fashionโ‰ˆ 1 (related to neither)

Ratios distinguish relevant from irrelevant words and encode meaning. GloVe designs vectors so that dot products reflect log co-occurrence counts:

$$ \mathbf{w}_i^\top\tilde{\mathbf{w}}_j + b_i + \tilde{b}_j \approx \log X_{ij} $$

where $X_{ij}$ counts how often $j$ occurs in $i$'s context. This is solved as a weighted least-squares problem over non-zero entries:

$$ J = \sum_{i,j:\,X_{ij} > 0}f(X_{ij})\left(\mathbf{w}_i^\top\tilde{\mathbf{w}}_j + b_i + \tilde{b}_j - \log X_{ij}\right)^2 $$

The weighting function

$$ f(x) = \begin{cases} (x/x_{\max})^{\alpha} & x < x_{\max} \\ 1 & \text{otherwise} \end{cases}, \qquad \alpha = 3/4,\; x_{\max} = 100 $$

prevents very frequent pairs from dominating and down-weights rare, noisy pairs. The final embedding is typically $\mathbf{w}_i + \tilde{\mathbf{w}}_i$.

GloVe combines the strengths of count-based methods (using global statistics efficiently, like LSA/SVD) and prediction-based methods (good linear structure, like word2vec). Pretrained GloVe vectors (trained on Wikipedia, web crawl and Twitter data) were widely used.

FastText: subword information#

Word2vec and GloVe treat "teach", "teacher", "teaching" and "teachers" as unrelated symbols, and have no vector for a word absent from training. Bojanowski et al. (2017) represented each word as a bag of character n-grams (typically lengths 3โ€“6), with boundary markers. For "where" with $n = 3$:

text
<wh, whe, her, ere, re>   plus the whole word <where>

A word's vector is the sum of its n-gram vectors:

$$ \mathbf{v}_w = \sum_{g \in \mathcal{G}_w}\mathbf{z}_g $$

and training uses the skip-gram negative-sampling objective. Consequences:

  • Unseen words get vectors from their n-grams ("teachable" shares n-grams with "teach" and "table").
  • Morphology is captured: words sharing roots and affixes are close.
  • Misspellings and informal spellings remain near their correct forms.
  • Especially beneficial for morphologically rich languages โ€” FastText released pretrained vectors for 157 languages, including Bangla.

N-gram vectors are stored in a fixed number of hash buckets (e.g. 2 million) to bound memory.

python
from gensim.models import FastText

sentences = [s.lower().split() for s in open("corpus.txt", encoding="utf-8")]
ft = FastText(sentences, vector_size=100, window=5, min_count=3, min_n=3, max_n=6, epochs=5)
print(ft.wv.most_similar("teacher", topn=5))
print(ft.wv["teachingly"][:5])                   # works even if never seen in training
print(ft.wv.similarity("organisation", "organization"))   # spelling variants stay close

FastText also offers a very fast text classifier: average word and n-gram embeddings, then a linear softmax โ€” trains on millions of documents in minutes on a CPU and is a strong baseline for tasks such as language identification.

Comparing static embeddings#

Word2vecGloVeFastText
Training signalLocal windows (prediction)Global co-occurrence matrix (regression)Local windows over subwords
OOV wordsNo vectorNo vectorYes, from n-grams
MorphologyNot capturedNot capturedCaptured
Rare wordsWeakWeakBetter
MemoryVocabulary ร— dimVocabulary ร— dim+ n-gram buckets

Evaluating embeddings#

  • Intrinsic: word similarity datasets (correlate cosine similarity with human ratings, e.g. SimLex-999, WordSim-353) and analogy tasks.
  • Extrinsic: performance of downstream tasks (NER, classification) using the embeddings as features โ€” what ultimately matters.

Intrinsic and extrinsic results do not always agree.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Word2Vec: Learning Word Embeddings from Context

Word2Vec learns dense word vectors by predicting context words. We derive the skip-gram and CBOW objectives, negative sampling, and explore analogies, similarity and the limitations of static embeddings.

Intermediateโฑ 5 min#166
๐Ÿ’ฌ NLP & Transformers

Subword Tokenisation: BPE, WordPiece, Unigram and SentencePiece

Modern language models split text into subword units. We derive byte-pair encoding step by step, compare WordPiece and Unigram LM tokenisation, discuss byte-level BPE, and examine how tokenisation affects multilingual fairness and cost.

Intermediateโฑ 5 min#168
๐Ÿ’ฌ NLP & Transformers

N-gram Language Models, Smoothing and Perplexity

A language model assigns probabilities to sequences of words. We derive n-gram models from the chain rule and Markov assumption, fix zero probabilities with smoothing, generate text, and evaluate with perplexity.

Intermediateโฑ 5 min#165