๐Ÿ”— Deep Learning ยท Lecture 26 of 38

Embeddings: Turning Discrete Things into Meaningful Vectors

Words, users, products and categories become dense vectors whose geometry encodes meaning. We explain embedding layers, how embeddings are learned, how to measure similarity, and how they power search and recommendation.

Neural networks work with numbers, but much of the world is discrete: words, product IDs, user IDs, diagnoses, postcodes. One-hot encoding a vocabulary of 50,000 words produces 50,000-dimensional vectors in which every pair of words is equally distant โ€” "cat" is as far from "kitten" as from "carburettor". Embeddings fix this by mapping each discrete item to a dense, low-dimensional vector learned so that similar items end up close together. Embeddings are one of the most important ideas in modern AI.

The embedding layer#

An embedding layer is simply a lookup table โ€” a matrix $\mathbf{E} \in \mathbb{R}^{V \times d}$ with one row per item:

$$ \text{embed}(i) = \mathbf{E}[i, :] = \mathbf{e}_i^\top\mathbf{E} $$

where $\mathbf{e}_i$ is the one-hot vector for item $i$. Mathematically it is a linear layer applied to a one-hot input; computationally it is an efficient row lookup. The rows are ordinary parameters, learned by backpropagation like any weights. Typical sizes: $d$ from 16 (small categorical features) to several thousand (large language models).

python
import torch
import torch.nn as nn

emb = nn.Embedding(num_embeddings=10_000, embedding_dim=64)
token_ids = torch.tensor([[12, 845, 3], [7, 7, 9999]])     # (batch, sequence)
vectors = emb(token_ids)
print(vectors.shape)                                        # (2, 3, 64)

How embeddings acquire meaning#

An embedding has no meaning on its own; it acquires meaning from the task it is trained on.

  • In a sentiment classifier, word embeddings organise along positive/negative lines.
  • In word2vec, trained to predict neighbouring words, words appearing in similar contexts get similar vectors โ€” the distributional hypothesis: "you shall know a word by the company it keeps" (Firth, 1957).
  • In a recommender, user and item embeddings are trained so their dot product predicts interactions โ€” similar users and similar items cluster.
  • In contrastive models like CLIP, images and captions are embedded in a shared space where matching pairs are close.
  • In language models, token embeddings are trained jointly with the whole network; deeper layers produce contextual embeddings in which "bank" in "river bank" differs from "bank" in "bank account".

Measuring similarity#

The standard measure is cosine similarity:

$$ \cos(\mathbf{u}, \mathbf{v}) = \frac{\mathbf{u}^\top\mathbf{v}}{\|\mathbf{u}\|\,\|\mathbf{v}\|} $$

which compares direction and ignores length. Many systems L2-normalise embeddings so that cosine similarity equals the dot product and nearest-neighbour search is simple.

Geometry of meaning#

Well-trained embeddings exhibit remarkable structure:

  • Clusters of related items (countries, verbs, sports).
  • Linear relationships: $\mathbf{v}_{\text{Paris}} - \mathbf{v}_{\text{France}} + \mathbf{v}_{\text{Japan}} \approx \mathbf{v}_{\text{Tokyo}}$ โ€” the famous analogies of word2vec (these hold approximately and less reliably than early demonstrations suggested).
  • Directions corresponding to attributes such as gender, tense or sentiment.

Embeddings as a universal interface#

Because embeddings turn anything into vectors with meaningful distances, they power a wide range of systems:

  1. Semantic search โ€” embed documents and queries; retrieve nearest neighbours. This finds relevant results even without shared keywords ("How do I register a newborn?" matches "birth registration procedure").
  2. Retrieval-augmented generation โ€” retrieve relevant passages by embedding similarity and give them to a language model.
  3. Recommendation โ€” nearest item embeddings to a user embedding.
  4. Clustering and deduplication โ€” group similar support tickets or detect near-duplicate records.
  5. Classification with few labels โ€” train a simple classifier on top of pretrained embeddings.
  6. Categorical features in tabular models โ€” entity embeddings for high-cardinality columns.
python
import numpy as np
from sentence_transformers import SentenceTransformer   # pip install sentence-transformers

model = SentenceTransformer("all-MiniLM-L6-v2")
docs = ["How to register the birth of a child",
        "Where can I get vaccinations for my baby?",
        "Steps to renew an expired passport",
        "Cash assistance eligibility criteria"]
query = "My newborn needs a birth certificate"
D = model.encode(docs, normalize_embeddings=True)
q = model.encode([query], normalize_embeddings=True)[0]
for i in np.argsort(-(D @ q)):
    print(f"{D[i] @ q:.3f}  {docs[i]}")

The top result shares almost no keywords with the query, yet is semantically closest.

Practical tips#

  • Pretrained embeddings (from sentence encoders, CLIP, language models) are often better than training from scratch when data is limited.
  • Use multilingual embedding models for multilingual content โ€” essential in contexts where users write in Bangla, Arabic, English and other languages.
  • Evaluate embeddings on your task: retrieval metrics (recall@k) with real queries matter more than generic benchmarks.
  • For large collections, index embeddings with an approximate nearest-neighbour library or vector database.
  • Store the model version with every embedding; vectors from different models are not comparable.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Autoencoders: Compression, Denoising and Representation Learning

An autoencoder learns to reconstruct its input through a bottleneck, discovering compact representations without labels. We cover undercomplete, denoising, sparse and convolutional autoencoders and their uses.

Intermediateโฑ 5 min#121
๐Ÿ”— Deep Learning

From Biological Neurons to Artificial Neural Networks

We open the Deep Learning track by tracing the path from biological neurons to artificial ones, defining a neural network precisely, and explaining why depth and learned representations changed AI.

Beginnerโฑ 4 min#097
๐Ÿ”— Deep Learning

Transfer Learning and Fine-Tuning

Pretrained models let you achieve strong results with small datasets. We compare feature extraction and fine-tuning, explain discriminative learning rates and layer freezing, and discuss when transfer helps or hurts.

Intermediateโฑ 5 min#123