Neural networks work with numbers, but much of the world is discrete: words, product IDs, user IDs, diagnoses, postcodes. One-hot encoding a vocabulary of 50,000 words produces 50,000-dimensional vectors in which every pair of words is equally distant โ "cat" is as far from "kitten" as from "carburettor". Embeddings fix this by mapping each discrete item to a dense, low-dimensional vector learned so that similar items end up close together. Embeddings are one of the most important ideas in modern AI.
The embedding layer#
An embedding layer is simply a lookup table โ a matrix $\mathbf{E} \in \mathbb{R}^{V \times d}$ with one row per item:
where $\mathbf{e}_i$ is the one-hot vector for item $i$. Mathematically it is a linear layer applied to a one-hot input; computationally it is an efficient row lookup. The rows are ordinary parameters, learned by backpropagation like any weights. Typical sizes: $d$ from 16 (small categorical features) to several thousand (large language models).
import torch
import torch.nn as nn
emb = nn.Embedding(num_embeddings=10_000, embedding_dim=64)
token_ids = torch.tensor([[12, 845, 3], [7, 7, 9999]]) # (batch, sequence)
vectors = emb(token_ids)
print(vectors.shape) # (2, 3, 64)How embeddings acquire meaning#
An embedding has no meaning on its own; it acquires meaning from the task it is trained on.
- In a sentiment classifier, word embeddings organise along positive/negative lines.
- In word2vec, trained to predict neighbouring words, words appearing in similar contexts get similar vectors โ the distributional hypothesis: "you shall know a word by the company it keeps" (Firth, 1957).
- In a recommender, user and item embeddings are trained so their dot product predicts interactions โ similar users and similar items cluster.
- In contrastive models like CLIP, images and captions are embedded in a shared space where matching pairs are close.
- In language models, token embeddings are trained jointly with the whole network; deeper layers produce contextual embeddings in which "bank" in "river bank" differs from "bank" in "bank account".
Measuring similarity#
The standard measure is cosine similarity:
which compares direction and ignores length. Many systems L2-normalise embeddings so that cosine similarity equals the dot product and nearest-neighbour search is simple.
Geometry of meaning#
Well-trained embeddings exhibit remarkable structure:
- Clusters of related items (countries, verbs, sports).
- Linear relationships: $\mathbf{v}_{\text{Paris}} - \mathbf{v}_{\text{France}} + \mathbf{v}_{\text{Japan}} \approx \mathbf{v}_{\text{Tokyo}}$ โ the famous analogies of word2vec (these hold approximately and less reliably than early demonstrations suggested).
- Directions corresponding to attributes such as gender, tense or sentiment.
Embeddings as a universal interface#
Because embeddings turn anything into vectors with meaningful distances, they power a wide range of systems:
- Semantic search โ embed documents and queries; retrieve nearest neighbours. This finds relevant results even without shared keywords ("How do I register a newborn?" matches "birth registration procedure").
- Retrieval-augmented generation โ retrieve relevant passages by embedding similarity and give them to a language model.
- Recommendation โ nearest item embeddings to a user embedding.
- Clustering and deduplication โ group similar support tickets or detect near-duplicate records.
- Classification with few labels โ train a simple classifier on top of pretrained embeddings.
- Categorical features in tabular models โ entity embeddings for high-cardinality columns.
import numpy as np
from sentence_transformers import SentenceTransformer # pip install sentence-transformers
model = SentenceTransformer("all-MiniLM-L6-v2")
docs = ["How to register the birth of a child",
"Where can I get vaccinations for my baby?",
"Steps to renew an expired passport",
"Cash assistance eligibility criteria"]
query = "My newborn needs a birth certificate"
D = model.encode(docs, normalize_embeddings=True)
q = model.encode([query], normalize_embeddings=True)[0]
for i in np.argsort(-(D @ q)):
print(f"{D[i] @ q:.3f} {docs[i]}")The top result shares almost no keywords with the query, yet is semantically closest.
Practical tips#
- Pretrained embeddings (from sentence encoders, CLIP, language models) are often better than training from scratch when data is limited.
- Use multilingual embedding models for multilingual content โ essential in contexts where users write in Bangla, Arabic, English and other languages.
- Evaluate embeddings on your task: retrieval metrics (recall@k) with real queries matter more than generic benchmarks.
- For large collections, index embeddings with an approximate nearest-neighbour library or vector database.
- Store the model version with every embedding; vectors from different models are not comparable.