Word2vec learns from local context windows one pair at a time. Two influential follow-ups improved on it in complementary ways. GloVe (Stanford, 2014) exploits global co-occurrence statistics directly. FastText (Facebook AI Research, 2016โ2017) builds word vectors from character n-grams, so it can represent rare, misspelled and never-seen words โ invaluable for morphologically rich languages.
GloVe: Global Vectors#
Pennington, Socher and Manning started from an observation about ratios of co-occurrence probabilities. Let $P_{ij} = P(j \mid i)$ be the probability that word $j$ appears in the context of word $i$. Consider $i$ = "ice", $j$ = "steam":
| Probe word $k$ | $P(k \mid \text{ice})/P(k \mid \text{steam})$ |
|---|---|
| solid | large (โซ 1) |
| gas | small (โช 1) |
| water | โ 1 (related to both) |
| fashion | โ 1 (related to neither) |
Ratios distinguish relevant from irrelevant words and encode meaning. GloVe designs vectors so that dot products reflect log co-occurrence counts:
where $X_{ij}$ counts how often $j$ occurs in $i$'s context. This is solved as a weighted least-squares problem over non-zero entries:
The weighting function
prevents very frequent pairs from dominating and down-weights rare, noisy pairs. The final embedding is typically $\mathbf{w}_i + \tilde{\mathbf{w}}_i$.
GloVe combines the strengths of count-based methods (using global statistics efficiently, like LSA/SVD) and prediction-based methods (good linear structure, like word2vec). Pretrained GloVe vectors (trained on Wikipedia, web crawl and Twitter data) were widely used.
FastText: subword information#
Word2vec and GloVe treat "teach", "teacher", "teaching" and "teachers" as unrelated symbols, and have no vector for a word absent from training. Bojanowski et al. (2017) represented each word as a bag of character n-grams (typically lengths 3โ6), with boundary markers. For "where" with $n = 3$:
<wh, whe, her, ere, re> plus the whole word <where>A word's vector is the sum of its n-gram vectors:
and training uses the skip-gram negative-sampling objective. Consequences:
- Unseen words get vectors from their n-grams ("teachable" shares n-grams with "teach" and "table").
- Morphology is captured: words sharing roots and affixes are close.
- Misspellings and informal spellings remain near their correct forms.
- Especially beneficial for morphologically rich languages โ FastText released pretrained vectors for 157 languages, including Bangla.
N-gram vectors are stored in a fixed number of hash buckets (e.g. 2 million) to bound memory.
from gensim.models import FastText
sentences = [s.lower().split() for s in open("corpus.txt", encoding="utf-8")]
ft = FastText(sentences, vector_size=100, window=5, min_count=3, min_n=3, max_n=6, epochs=5)
print(ft.wv.most_similar("teacher", topn=5))
print(ft.wv["teachingly"][:5]) # works even if never seen in training
print(ft.wv.similarity("organisation", "organization")) # spelling variants stay closeFastText also offers a very fast text classifier: average word and n-gram embeddings, then a linear softmax โ trains on millions of documents in minutes on a CPU and is a strong baseline for tasks such as language identification.
Comparing static embeddings#
| Word2vec | GloVe | FastText | |
|---|---|---|---|
| Training signal | Local windows (prediction) | Global co-occurrence matrix (regression) | Local windows over subwords |
| OOV words | No vector | No vector | Yes, from n-grams |
| Morphology | Not captured | Not captured | Captured |
| Rare words | Weak | Weak | Better |
| Memory | Vocabulary ร dim | Vocabulary ร dim | + n-gram buckets |
Evaluating embeddings#
- Intrinsic: word similarity datasets (correlate cosine similarity with human ratings, e.g. SimLex-999, WordSim-353) and analogy tasks.
- Extrinsic: performance of downstream tasks (NER, classification) using the embeddings as features โ what ultimately matters.
Intrinsic and extrinsic results do not always agree.