Embeddings Explained
Learning Outcomes
- Explain what an embedding is and why dense vector representations capture meaning
- Describe how encoder models produce embeddings from raw text
- Calculate and compare cosine similarity, dot product, and Euclidean distance between vectors
- Evaluate embedding models by dimension, max-token limit, and MTEB benchmark score
- Produce sentence embeddings with sentence-transformers and the OpenAI Embeddings API
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why embeddings are the heart of RAG |
| Explain | 7 min | What an embedding is — vectors, dimensions, meaning |
| Explain | 7 min | How encoder models produce embeddings |
| Demo | 8 min | Hands-on: embed sentences, inspect the output |
| Explain | 8 min | Distance metrics — cosine, dot product, Euclidean |
| Demo | 8 min | Choosing an embedding model — open-source, OpenAI, Cohere |
| Wrap-up | 4 min | Key takeaways, preview of vector databases |
Before You Begin
Pre-work:
- Complete Lesson 1: What Is RAG? — this lesson builds directly on the RAG architecture overview
- Have Python 3.10+ and pip available in your environment
- An OpenAI API key is optional; all core examples use a free, local model
Shopping List:
- Python 3.10 or later
pip install -U sentence-transformers numpy(open-source, CPU-only, no API key required)- Optionally
pip install openaifor the API-based examples - A terminal and a code editor — no GPU is required for the sentence-transformers examples in this lesson
An embedding is a list of floating-point numbers — a vector — that represents the meaning of a piece of text. Embedding models are trained so that texts with similar meaning produce vectors that are close together in the high-dimensional space those numbers define.
The key insight: geometry encodes semantics. "Dog" and "puppy" end up nearby. "Paris" and "London" are near each other, and both are far from "calculus." Sentences about subscription cancellation cluster together regardless of whether the word "cancel," "terminate," or "unsubscribe" is used — which is exactly what you need for a search engine that understands intent rather than keywords.
Here is a toy three-dimensional illustration:
# Three-dimensional toy embeddings (not real values — for illustration only)
# In practice embeddings have hundreds or thousands of dimensions.
dog = [0.82, 0.10, 0.05]
puppy = [0.79, 0.12, 0.08] # very close to dog
cat = [0.75, 0.08, 0.20] # same animal neighbourhood
snake = [0.05, 0.90, 0.30] # very different direction
# Real embeddings: text-embedding-3-small produces 1536-element vectors,
# bge-m3 produces 1024-element vectors.
A real embedding for "The server returned a 404 error" might be a 1 536-element list — one float per dimension — but the principle is the same: the model learns to place meaning in geometry.
What embeddings are not:
- They are not token IDs (discrete integer look-ups)
- They are not one-hot vectors (sparse with no geometry)
- They do not "contain" the original text — they are a lossy compression of meaning
Dimensions in the wild. Higher-dimensional embeddings can capture more nuance, but cost more memory and compute:
| Dimensions | Typical Example |
|---|---|
| 384 | all-MiniLM-L6-v2 (fast, lightweight) |
| 768 | BERT-base, bge-small |
| 1 024 | BGE-M3, bge-large |
| 1 536 | OpenAI text-embedding-3-small |
| 3 072 | OpenAI text-embedding-3-large |
Embedding models are built on transformer encoders. Unlike decoder-only models (GPT, Claude, Llama) which generate text left-to-right, an encoder reads the entire input bidirectionally — every token attends to every other token simultaneously. That full-context attention is what lets the encoder build a representation capturing the overall meaning of a phrase rather than predicting the next word.
The most influential encoder architecture is BERT (Bidirectional Encoder Representations from Transformers). Modern embedding models like BGE, E5, and many sentence-transformers models all descend from this lineage.
Processing pipeline (simplified):
Raw text input
|
Tokenise --> ["The", "quick", "brown", "fox"] --> [token IDs]
|
Token embeddings (look-up table: token ID -> d-dim vector)
+
Positional embeddings (encode position in sequence)
|
N x Transformer encoder layers (bidirectional self-attention + FFN)
|
Per-token hidden states [seq_len x d_model]
|
Pooling --> single vector [1 x d_model]
|
(optional) Linear projection + L2 normalisation
|
Final embedding vector [1 x embedding_dim]
Pooling is how you collapse per-token outputs into a single vector for the whole sentence. Common strategies:
| Strategy | Description | Notes |
|---|---|---|
[CLS] token |
Use the special classification token output | Original BERT approach |
| Mean pooling | Average all token outputs | Default for most sentence-transformers |
| Max pooling | Element-wise max across token outputs | Captures peak activations |
| Weighted mean | Weight by attention scores | Some newer models |
sentence-transformers defaults to mean pooling because it is empirically more stable than using the [CLS] token alone.
Training objective. A raw pretrained encoder is not yet a good sentence encoder. Embedding models are fine-tuned with contrastive learning: the model is shown pairs of semantically similar sentences (positive pairs) and trained to maximise their similarity while minimising similarity with unrelated sentences (hard negatives). This contrastive training is what gives the vector geometry its meaning — without it, the vectors cluster arbitrarily.
sentence-transformers is the de-facto Python library for running embedding models locally. It wraps Hugging Face models behind a clean API and handles tokenisation, batching, and pooling automatically.
Install the library:
pip install -U sentence-transformers numpy
Embed a small set of sentences and inspect the output:
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
sentences = [
"How do I cancel my subscription?",
"What is the process to terminate my account?",
"Where can I find the pricing page?",
"The weather in Paris is lovely in spring.",
]
embeddings = model.encode(sentences)
print(f"Shape: {embeddings.shape}")
# Shape: (4, 384) -- 4 sentences, 384 dimensions each
print(f"dtype: {embeddings.dtype}")
# dtype: float32
print(f"First 5 dims of sentence 0: {embeddings[0][:5]}")
# e.g. [ 0.032 -0.118 0.204 0.071 -0.056 ]
Now measure how similar the embeddings are to each other:
def cosine_similarity(a, b):
"""Cosine similarity between two 1-D numpy arrays."""
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
sim_cancel_terminate = cosine_similarity(embeddings[0], embeddings[1])
sim_cancel_pricing = cosine_similarity(embeddings[0], embeddings[2])
sim_cancel_weather = cosine_similarity(embeddings[0], embeddings[3])
print(f"'cancel subscription' vs 'terminate account': {sim_cancel_terminate:.4f}")
print(f"'cancel subscription' vs 'pricing page': {sim_cancel_pricing:.4f}")
print(f"'cancel subscription' vs 'Paris weather': {sim_cancel_weather:.4f}")
# Typical output:
# 'cancel subscription' vs 'terminate account': 0.8312
# 'cancel subscription' vs 'pricing page': 0.3150
# 'cancel subscription' vs 'Paris weather': 0.0421
The model was never explicitly told "cancel" and "terminate account" mean the same thing. It inferred this geometry from contrastive training data. This mechanism is the engine of semantic search.
model.encode(sentences) call batches inputs automatically. For large corpora pass batch_size=64 and show_progress_bar=True. On a modern CPU, MiniLM-L6-v2 can embed tens of thousands of short sentences per minute — no GPU required for experimentation or moderate-scale workloads.About all-MiniLM-L6-v2: it is a 22 M-parameter, 384-dimension model optimised for speed and is the default recommendation for local experimentation. When you move to production you will choose a larger model based on your quality requirements — covered in Step 6.
When you query a vector database, it compares your query embedding against every stored embedding using a distance metric. Choosing the right metric matters because different metrics make different assumptions about your vectors.
The three metrics you will encounter:
1. Cosine similarity — measures the angle between two vectors, ignoring their magnitude. If two vectors point in exactly the same direction, cosine similarity is 1.0. Opposite directions give -1.0. Perpendicular gives 0.
import numpy as np
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
# cosine distance = 1 - cosine_similarity (some libraries use distance)
Cosine similarity is the most common choice for text embeddings because it is invariant to vector magnitude — the length of a vector depends on pooling and normalisation decisions, not meaning.
2. Dot product — multiplies corresponding elements and sums them:
def dot_product(a, b):
return float(np.dot(a, b))
For unit-normalised vectors (L2 norm = 1), dot product and cosine similarity are identical. The difference only matters when vectors are not normalised. Most production embedding models output L2-normalised vectors by default, making dot product the preferred choice in vector databases — it is faster because it avoids the norm division in the hot path.
3. Euclidean distance (L2) — the straight-line distance between two points in vector space:
def euclidean_distance(a, b):
return float(np.linalg.norm(a - b))
Euclidean distance is sensitive to vector magnitude, which makes it appropriate for embeddings that encode quantity or magnitude (image pixel intensities, count features). For text embeddings it is generally inferior to cosine similarity unless the vectors are normalised.
A worked comparison:
a = np.array([0.8, 0.3, 0.5])
b = np.array([0.7, 0.4, 0.6])
c = np.array([0.1, 0.9, 0.2])
print(f"cosine(a, b) = {cosine_similarity(a, b):.4f}") # high: similar direction
print(f"cosine(a, c) = {cosine_similarity(a, c):.4f}") # low: different direction
print(f"dot(a, b) = {dot_product(a, b):.4f}")
print(f"euclid(a, b) = {euclidean_distance(a, b):.4f}") # small: close in space
print(f"euclid(a, c) = {euclidean_distance(a, c):.4f}") # larger: farther apart
Decision table:
| Metric | Use When |
|---|---|
| Cosine similarity | Vectors may not be normalised; focus is on directional similarity |
| Dot product | Vectors are L2-normalised (most embedding models); fastest at query time |
| Euclidean (L2) | Embeddings encode magnitude as well as direction (rare for text) |
model.encode(sentences, normalize_embeddings=True) in sentence-transformers, the output vectors are L2-normalised to unit length. That makes cosine similarity and dot product equivalent, and Euclidean distance proportional to both. Many models normalise by default — check the model card.Before committing to a model for production, it is worth testing it on your actual domain vocabulary. Two experiments that reveal a model's quality quickly are analogy probing and cluster inspection.
Analogy probing. Good embeddings capture analogical relationships. The classic example from the original Word2Vec paper: "king" - "man" + "woman" lands near "queen." Modern sentence encoders support richer analogies:
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
# Test semantic clustering on your domain vocabulary
domain_terms = [
# Customer support
"subscription cancellation",
"account termination",
"refund request",
# Technical
"database connection error",
"server timeout",
"memory leak",
# Unrelated
"chocolate cake recipe",
"mountain hiking trail",
]
vecs = model.encode(domain_terms, normalize_embeddings=True)
# Build a similarity matrix
sim_matrix = vecs @ vecs.T # dot product == cosine for normalised vectors
# Print the top-2 nearest neighbours for each term
for i, term in enumerate(domain_terms):
row = sim_matrix[i].copy()
row[i] = -1 # exclude self
top2 = np.argsort(row)[::-1][:2]
print(f"{term!r}")
for j in top2:
print(f" -> {domain_terms[j]!r} ({sim_matrix[i, j]:.3f})")
print()
You should see the customer-support terms cluster together, the technical terms cluster together, and both groups stay far from the recipes and hiking content.
What to look for in the output:
- Terms you know are synonymous should have similarity above ~0.75
- Terms from different domains should be below ~0.3
- If near-synonyms score below 0.5, the model may be too small or not domain-adapted
The MTEB (Massive Text Embedding Benchmark) leaderboard is the standard reference for comparing embedding model quality across retrieval, classification, clustering, and other tasks. The benchmark covers dozens of tasks and multiple languages.
Key models to know in 2026:
Open-source / self-hosted:
| Model | Dims | Max Tokens | MTEB Class | Notes |
|---|---|---|---|---|
| all-MiniLM-L6-v2 | 384 | 256 | Good | Best for fast experimentation; very small |
| BGE-M3 (BAAI) | 1 024 | 8 192 | Strong | Dense + sparse + multi-vector; 100+ languages |
| bge-large-en-v1.5 | 1 024 | 512 | Strong | English-only, well-balanced quality/size |
| Qwen3-Embedding-8B | 4 096 | 32 768 | Near SOTA | Apache 2.0; decoder-based; GPU required |
API-based (managed):
| Model | Dims | Max Tokens | Normalised | Notes |
|---|---|---|---|---|
| OpenAI text-embedding-3-small | 1 536 | 8 192 | Yes | Most widely used API model; cost-effective |
| OpenAI text-embedding-3-large | 3 072 | 8 192 | Yes | Higher MTEB score; ~6x more expensive |
| Cohere embed-v4 | 256–1 536 | 128 000 | Yes | Multimodal; long-context; Matryoshka dims |
Code example — OpenAI Embeddings API:
from openai import OpenAI
import numpy as np
client = OpenAI() # reads OPENAI_API_KEY from environment
response = client.embeddings.create(
model="text-embedding-3-small",
input=["How do I cancel my subscription?",
"What is the process to terminate my account?"],
dimensions=512, # Matryoshka: request fewer dims to save cost
)
vecs = np.array([item.embedding for item in response.data])
print(vecs.shape) # (2, 512)
Code example — local BGE-M3 via sentence-transformers:
from sentence_transformers import SentenceTransformer
# BGE-M3 is ~2 GB download; cached after first use
model = SentenceTransformer("BAAI/bge-m3")
docs = [
"Subscription cancellation policy",
"How to request a refund",
"Server configuration guide",
]
# Dense embeddings; normalised by default for BGE models
embeddings = model.encode(docs, normalize_embeddings=True)
print(embeddings.shape) # (3, 1024)
Decision framework:
Is your corpus multilingual or longer than 512 tokens?
YES --> BGE-M3 (local) or Cohere embed-v4 (API)
NO --> continue
Do you need zero-infra, managed, always-up-to-date?
YES --> OpenAI text-embedding-3-small (good default)
NO --> continue
Do you need the lowest possible latency at high throughput?
YES --> all-MiniLM-L6-v2 or bge-small-en-v1.5 (tiny, fast)
NO --> bge-large-en-v1.5 or text-embedding-3-small
Everything in this lesson comes together in a minimal semantic search function. This is the retrieval core of every RAG pipeline — the query goes in, the most relevant passages come out.
from sentence_transformers import SentenceTransformer
import numpy as np
# ---- Index time (done once per corpus) ----
model = SentenceTransformer(
"sentence-transformers/all-MiniLM-L6-v2"
)
# Your knowledge base: list of text chunks
corpus = [
"To cancel your subscription, go to Account Settings and click Cancel Plan.",
"Refunds are processed within 5-7 business days to the original payment method.",
"You can upgrade your plan at any time from the Billing section.",
"Contact support at [email protected] for billing issues.",
"The API rate limit is 1000 requests per minute on the Pro plan.",
"Two-factor authentication can be enabled under Security Settings.",
]
# Embed the corpus (do this once; store the result in a vector DB)
corpus_embeddings = model.encode(corpus, normalize_embeddings=True)
# shape: (6, 384)
# ---- Query time (done for every user query) ----
def semantic_search(query: str, top_k: int = 3) -> list[dict]:
query_embedding = model.encode(
[query], normalize_embeddings=True
)[0]
# Cosine similarity == dot product for normalised vectors
scores = corpus_embeddings @ query_embedding
top_indices = np.argsort(scores)[::-1][:top_k]
return [
{"text": corpus[i], "score": float(scores[i])}
for i in top_indices
]
results = semantic_search("How do I get my money back?")
for r in results:
print(f"[{r['score']:.3f}] {r['text']}")
# Example output:
# [0.791] Refunds are processed within 5-7 business days ...
# [0.423] To cancel your subscription, go to Account Settings ...
# [0.311] Contact support at [email protected] for billing issues.
The query "How do I get my money back?" contains none of the words in the top result — "refund," "processed," "payment method" — but the embeddings place them in the same semantic neighbourhood. That is the payoff of the whole lesson.
What you would add in a real system:
- Replace the in-memory NumPy matrix with a vector database (covered in Lesson 3)
- Replace the simple text list with properly chunked documents (covered in Lesson 4)
- Pass the top-k results to an LLM as context (covered in Lesson 5)
pip install faiss-cpu. For larger corpora or when you need metadata filtering, a dedicated vector database is the right tool.Questions & Answers
Key Takeaways
- Embeddings are meaning as geometry — dense floating-point vectors where similar meanings produce geometrically close vectors, enabling semantic search that matches intent rather than keywords.
- Encoder models read bidirectionally — unlike generative LLMs, encoder architectures attend to the full input context simultaneously and are fine-tuned with contrastive learning to give the vector space its semantic structure.
- Use dot product for normalised vectors — for L2-normalised embeddings (the common case), dot product equals cosine similarity and is faster; only fall back to Euclidean when your embeddings encode magnitude information.
- The MTEB leaderboard is your starting point — use it to shortlist models, then test on your own domain vocabulary before committing; lock your model choice early because changing it requires re-embedding the entire corpus.
- sentence-transformers gets you running in three lines — open-source, CPU-capable, and wraps hundreds of MTEB-ranked models behind a consistent API; BGE-M3 is a strong open-source default for multilingual or long-document corpora.
- Model selection is a production dependency — once vectors are stored, the embedding model is pinned; mixing vectors from different models silently breaks similarity scores, so treat the model choice with the same care as a database schema migration.
Next Steps: Lesson 3: Vector Databases