๐Ÿ’ฌ NLP & Transformers ยท Lecture 5 of 29

Word2Vec: Learning Word Embeddings from Context

Word2Vec learns dense word vectors by predicting context words. We derive the skip-gram and CBOW objectives, negative sampling, and explore analogies, similarity and the limitations of static embeddings.

In 2013 Tomas Mikolov and colleagues at Google released word2vec, a simple and fast method for learning word vectors from raw text. Its vectors captured semantic and syntactic relationships so strikingly โ€” the famous "king โˆ’ man + woman โ‰ˆ queen" โ€” that word embeddings became the standard input representation for neural NLP for the next five years. The ideas behind it, especially contrastive learning with negative sampling, remain central today.

The distributional hypothesis#

"You shall know a word by the company it keeps" (J.R. Firth, 1957). Words appearing in similar contexts tend to have similar meanings: "doctor" and "nurse" both appear near "hospital", "patient", "treatment". Word2vec turns this idea into a prediction task.

Two architectures#

For a corpus of words, with a context window of size $c$ (e.g. 5 words either side):

  • Skip-gram: given a centre word, predict each surrounding context word.
  • CBOW (Continuous Bag of Words): given the (averaged) context words, predict the centre word.

Skip-gram works better for rare words and smaller corpora; CBOW trains faster.

The skip-gram objective#

Each word $w$ has two vectors: $\mathbf{v}_w$ (as centre word) and $\mathbf{u}_w$ (as context word). The probability of context word $o$ given centre word $c$ is a softmax:

$$ P(o \mid c) = \frac{\exp(\mathbf{u}_o^\top\mathbf{v}_c)}{\sum_{w \in V}\exp(\mathbf{u}_w^\top\mathbf{v}_c)} $$

and we maximise the average log-probability of observed (centre, context) pairs. The problem: the denominator sums over the whole vocabulary (hundreds of thousands of words) for every training pair โ€” far too expensive.

Negative sampling#

Instead of a full softmax, train a binary classifier: is this (centre, context) pair real, or a random fake? For a real pair $(c, o)$ and $k$ negative words $n_1, \dots, n_k$ sampled from a noise distribution:

$$ \mathcal{L} = -\log\sigma(\mathbf{u}_o^\top\mathbf{v}_c) - \sum_{i=1}^{k}\log\sigma(-\mathbf{u}_{n_i}^\top\mathbf{v}_c) $$

Only $k + 1$ dot products per pair ($k$ = 5โ€“20 for small data, 2โ€“5 for large data). Negatives are drawn from the unigram distribution raised to the $3/4$ power, $P_n(w) \propto f(w)^{3/4}$, which boosts rare words relative to raw frequency. This is an early instance of the contrastive learning idea used by CLIP and SimCLR.

Subsampling of very frequent words (discarding "the" with high probability) speeds training and improves rare-word vectors.

Implementation sketch#

python
import torch
import torch.nn as nn
import torch.nn.functional as F

class SkipGramNS(nn.Module):
    def __init__(self, vocab_size, dim=100):
        super().__init__()
        self.center = nn.Embedding(vocab_size, dim)
        self.context = nn.Embedding(vocab_size, dim)
        nn.init.uniform_(self.center.weight, -0.5 / dim, 0.5 / dim)
        nn.init.zeros_(self.context.weight)
    def forward(self, c, o, neg):               # c, o: (B,), neg: (B, k)
        vc = self.center(c)                                    # (B, d)
        pos = F.logsigmoid((self.context(o) * vc).sum(-1))     # (B,)
        negs = F.logsigmoid(-(self.context(neg) @ vc.unsqueeze(-1)).squeeze(-1)).sum(-1)
        return -(pos + negs).mean()

model = SkipGramNS(vocab_size=10_000)
c, o = torch.randint(0, 10_000, (32,)), torch.randint(0, 10_000, (32,))
neg = torch.randint(0, 10_000, (32, 5))
print(model(c, o, neg))

In practice, use the optimised gensim implementation:

python
from gensim.models import Word2Vec
sentences = [line.lower().split() for line in open("corpus.txt", encoding="utf-8")]
w2v = Word2Vec(sentences, vector_size=100, window=5, min_count=5, sg=1, negative=10, epochs=5)
print(w2v.wv.most_similar("doctor", topn=5))
print(w2v.wv.most_similar(positive=["king", "woman"], negative=["man"], topn=3))

What the vectors capture#

  • Similarity: nearest neighbours are semantically related words.
  • Analogies: vector offsets encode relations โ€” gender (man โ†’ woman), countryโ€“capital (France โ†’ Paris), verb tense (walk โ†’ walked), comparatives (good โ†’ better). Solve "a is to b as c is to ?" by finding the word nearest $\mathbf{v}_b - \mathbf{v}_a + \mathbf{v}_c$.
  • Clusters of topics, and some interpretable directions.

Limitations#

  • One vector per word type: "bank" gets a single vector blending riverbank and financial meanings. Contextual embeddings (ELMo, BERT) solve this.
  • Out-of-vocabulary words get no vector (FastText addresses this with subwords).
  • Bias: embeddings reproduce stereotypes present in the training text; analogy tests have revealed associations such as occupations with genders.
  • Analogy results are overstated: many famous analogies only work when the input words are excluded from the candidates, and performance varies widely across relation types.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

GloVe and FastText: Global Statistics and Subword Embeddings

GloVe learns embeddings from a global co-occurrence matrix with a weighted least-squares objective; FastText represents words as bags of character n-grams, handling rare and unseen words. We compare them with word2vec and discuss evaluation.

Intermediateโฑ 5 min#167
๐Ÿ’ฌ NLP & Transformers

N-gram Language Models, Smoothing and Perplexity

A language model assigns probabilities to sequences of words. We derive n-gram models from the chain rule and Markov assumption, fix zero probabilities with smoothing, generate text, and evaluate with perplexity.

Intermediateโฑ 5 min#165
๐Ÿ’ฌ NLP & Transformers

Bag of Words and TF-IDF: Classical Text Representation

The simplest way to turn documents into vectors is to count words. We build bag-of-words and TF-IDF representations, derive the IDF formula, use cosine similarity for retrieval, and train strong linear text classifiers.

Beginnerโฑ 5 min#164