In 2013 Tomas Mikolov and colleagues at Google released word2vec, a simple and fast method for learning word vectors from raw text. Its vectors captured semantic and syntactic relationships so strikingly โ the famous "king โ man + woman โ queen" โ that word embeddings became the standard input representation for neural NLP for the next five years. The ideas behind it, especially contrastive learning with negative sampling, remain central today.
The distributional hypothesis#
"You shall know a word by the company it keeps" (J.R. Firth, 1957). Words appearing in similar contexts tend to have similar meanings: "doctor" and "nurse" both appear near "hospital", "patient", "treatment". Word2vec turns this idea into a prediction task.
Two architectures#
For a corpus of words, with a context window of size $c$ (e.g. 5 words either side):
- Skip-gram: given a centre word, predict each surrounding context word.
- CBOW (Continuous Bag of Words): given the (averaged) context words, predict the centre word.
Skip-gram works better for rare words and smaller corpora; CBOW trains faster.
The skip-gram objective#
Each word $w$ has two vectors: $\mathbf{v}_w$ (as centre word) and $\mathbf{u}_w$ (as context word). The probability of context word $o$ given centre word $c$ is a softmax:
and we maximise the average log-probability of observed (centre, context) pairs. The problem: the denominator sums over the whole vocabulary (hundreds of thousands of words) for every training pair โ far too expensive.
Negative sampling#
Instead of a full softmax, train a binary classifier: is this (centre, context) pair real, or a random fake? For a real pair $(c, o)$ and $k$ negative words $n_1, \dots, n_k$ sampled from a noise distribution:
Only $k + 1$ dot products per pair ($k$ = 5โ20 for small data, 2โ5 for large data). Negatives are drawn from the unigram distribution raised to the $3/4$ power, $P_n(w) \propto f(w)^{3/4}$, which boosts rare words relative to raw frequency. This is an early instance of the contrastive learning idea used by CLIP and SimCLR.
Subsampling of very frequent words (discarding "the" with high probability) speeds training and improves rare-word vectors.
Implementation sketch#
import torch
import torch.nn as nn
import torch.nn.functional as F
class SkipGramNS(nn.Module):
def __init__(self, vocab_size, dim=100):
super().__init__()
self.center = nn.Embedding(vocab_size, dim)
self.context = nn.Embedding(vocab_size, dim)
nn.init.uniform_(self.center.weight, -0.5 / dim, 0.5 / dim)
nn.init.zeros_(self.context.weight)
def forward(self, c, o, neg): # c, o: (B,), neg: (B, k)
vc = self.center(c) # (B, d)
pos = F.logsigmoid((self.context(o) * vc).sum(-1)) # (B,)
negs = F.logsigmoid(-(self.context(neg) @ vc.unsqueeze(-1)).squeeze(-1)).sum(-1)
return -(pos + negs).mean()
model = SkipGramNS(vocab_size=10_000)
c, o = torch.randint(0, 10_000, (32,)), torch.randint(0, 10_000, (32,))
neg = torch.randint(0, 10_000, (32, 5))
print(model(c, o, neg))In practice, use the optimised gensim implementation:
from gensim.models import Word2Vec
sentences = [line.lower().split() for line in open("corpus.txt", encoding="utf-8")]
w2v = Word2Vec(sentences, vector_size=100, window=5, min_count=5, sg=1, negative=10, epochs=5)
print(w2v.wv.most_similar("doctor", topn=5))
print(w2v.wv.most_similar(positive=["king", "woman"], negative=["man"], topn=3))What the vectors capture#
- Similarity: nearest neighbours are semantically related words.
- Analogies: vector offsets encode relations โ gender (man โ woman), countryโcapital (France โ Paris), verb tense (walk โ walked), comparatives (good โ better). Solve "a is to b as c is to ?" by finding the word nearest $\mathbf{v}_b - \mathbf{v}_a + \mathbf{v}_c$.
- Clusters of topics, and some interpretable directions.
Limitations#
- One vector per word type: "bank" gets a single vector blending riverbank and financial meanings. Contextual embeddings (ELMo, BERT) solve this.
- Out-of-vocabulary words get no vector (FastText addresses this with subwords).
- Bias: embeddings reproduce stereotypes present in the training text; analogy tests have revealed associations such as occupations with genders.
- Analogy results are overstated: many famous analogies only work when the input words are excluded from the candidates, and performance varies widely across relation types.