Embeddings Explained

45 min beginner Lesson 2

Learning Outcomes

  • Explain what an embedding is and why dense vector representations capture meaning
  • Describe how encoder models produce embeddings from raw text
  • Calculate and compare cosine similarity, dot product, and Euclidean distance between vectors
  • Evaluate embedding models by dimension, max-token limit, and MTEB benchmark score
  • Produce sentence embeddings with sentence-transformers and the OpenAI Embeddings API

Lesson Plan

Segment Duration Topic
Intro 3 min Why embeddings are the heart of RAG
Explain 7 min What an embedding is — vectors, dimensions, meaning
Explain 7 min How encoder models produce embeddings
Demo 8 min Hands-on: embed sentences, inspect the output
Explain 8 min Distance metrics — cosine, dot product, Euclidean
Demo 8 min Choosing an embedding model — open-source, OpenAI, Cohere
Wrap-up 4 min Key takeaways, preview of vector databases

Before You Begin

Pre-work:

  • Complete Lesson 1: What Is RAG? — this lesson builds directly on the RAG architecture overview
  • Have Python 3.10+ and pip available in your environment
  • An OpenAI API key is optional; all core examples use a free, local model

Shopping List:

  • Python 3.10 or later
  • pip install -U sentence-transformers numpy (open-source, CPU-only, no API key required)
  • Optionally pip install openai for the API-based examples
  • A terminal and a code editor — no GPU is required for the sentence-transformers examples in this lesson

1 What Is an Embedding?

An embedding is a list of floating-point numbers — a vector — that represents the meaning of a piece of text. Embedding models are trained so that texts with similar meaning produce vectors that are close together in the high-dimensional space those numbers define.

The key insight: geometry encodes semantics. "Dog" and "puppy" end up nearby. "Paris" and "London" are near each other, and both are far from "calculus." Sentences about subscription cancellation cluster together regardless of whether the word "cancel," "terminate," or "unsubscribe" is used — which is exactly what you need for a search engine that understands intent rather than keywords.

Here is a toy three-dimensional illustration:

# Three-dimensional toy embeddings (not real values — for illustration only)
# In practice embeddings have hundreds or thousands of dimensions.

dog    = [0.82, 0.10, 0.05]
puppy  = [0.79, 0.12, 0.08]   # very close to dog
cat    = [0.75, 0.08, 0.20]   # same animal neighbourhood
snake  = [0.05, 0.90, 0.30]   # very different direction

# Real embeddings: text-embedding-3-small produces 1536-element vectors,
# bge-m3 produces 1024-element vectors.

A real embedding for "The server returned a 404 error" might be a 1 536-element list — one float per dimension — but the principle is the same: the model learns to place meaning in geometry.

What embeddings are not:

  • They are not token IDs (discrete integer look-ups)
  • They are not one-hot vectors (sparse with no geometry)
  • They do not "contain" the original text — they are a lossy compression of meaning

Dimensions in the wild. Higher-dimensional embeddings can capture more nuance, but cost more memory and compute:

Dimensions Typical Example
384 all-MiniLM-L6-v2 (fast, lightweight)
768 BERT-base, bge-small
1 024 BGE-M3, bge-large
1 536 OpenAI text-embedding-3-small
3 072 OpenAI text-embedding-3-large
NOTE
Why Not Just Use Keywords?
Keyword search matches tokens. A query for 'price list' misses a document that says 'cost schedule'. Embeddings match meaning, so both phrases land in the same region of vector space. That gap is the entire reason RAG systems use embeddings rather than full-text search alone — although, as you will see in Lesson 6, combining both gives the best results.
TIP
Rule of Thumb on Dimensions
For most RAG workloads, 768–1 536 dimensions hits a good quality-cost-latency balance. You rarely need more than that unless you are doing fine-grained scientific, legal, or multilingual retrieval where the extra capacity pays off measurably.

2 How Encoder Models Produce Embeddings

Embedding models are built on transformer encoders. Unlike decoder-only models (GPT, Claude, Llama) which generate text left-to-right, an encoder reads the entire input bidirectionally — every token attends to every other token simultaneously. That full-context attention is what lets the encoder build a representation capturing the overall meaning of a phrase rather than predicting the next word.

The most influential encoder architecture is BERT (Bidirectional Encoder Representations from Transformers). Modern embedding models like BGE, E5, and many sentence-transformers models all descend from this lineage.

Processing pipeline (simplified):

Raw text input
      |
  Tokenise  -->  ["The", "quick", "brown", "fox"]  -->  [token IDs]
      |
  Token embeddings  (look-up table: token ID -> d-dim vector)
      +
  Positional embeddings  (encode position in sequence)
      |
  N x Transformer encoder layers  (bidirectional self-attention + FFN)
      |
  Per-token hidden states  [seq_len x d_model]
      |
  Pooling  -->  single vector  [1 x d_model]
      |
  (optional) Linear projection + L2 normalisation
      |
  Final embedding vector  [1 x embedding_dim]

Pooling is how you collapse per-token outputs into a single vector for the whole sentence. Common strategies:

Strategy Description Notes
[CLS] token Use the special classification token output Original BERT approach
Mean pooling Average all token outputs Default for most sentence-transformers
Max pooling Element-wise max across token outputs Captures peak activations
Weighted mean Weight by attention scores Some newer models

sentence-transformers defaults to mean pooling because it is empirically more stable than using the [CLS] token alone.

Training objective. A raw pretrained encoder is not yet a good sentence encoder. Embedding models are fine-tuned with contrastive learning: the model is shown pairs of semantically similar sentences (positive pairs) and trained to maximise their similarity while minimising similarity with unrelated sentences (hard negatives). This contrastive training is what gives the vector geometry its meaning — without it, the vectors cluster arbitrarily.

NOTE
Decoder LLMs as Embedding Models
Decoder-only LLMs (Mistral, Qwen, Llama) can also produce embeddings by pooling their hidden states. Some large models fine-tuned this way — such as Alibaba's Qwen3-Embedding-8B — score near the top of MTEB benchmarks. The trade-off: decoder-based embedding models are much larger and slower than dedicated encoder models. For a production RAG ingestion pipeline that must embed millions of documents, a purpose-built encoder model is usually the right choice for the hot path.

3 Your First Embeddings — Hands-On with sentence-transformers

sentence-transformers is the de-facto Python library for running embedding models locally. It wraps Hugging Face models behind a clean API and handles tokenisation, batching, and pooling automatically.

Install the library:

pip install -U sentence-transformers numpy

Embed a small set of sentences and inspect the output:

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")

sentences = [
    "How do I cancel my subscription?",
    "What is the process to terminate my account?",
    "Where can I find the pricing page?",
    "The weather in Paris is lovely in spring.",
]

embeddings = model.encode(sentences)

print(f"Shape: {embeddings.shape}")
# Shape: (4, 384)   -- 4 sentences, 384 dimensions each

print(f"dtype:  {embeddings.dtype}")
# dtype: float32

print(f"First 5 dims of sentence 0: {embeddings[0][:5]}")
# e.g. [ 0.032  -0.118   0.204   0.071  -0.056 ]

Now measure how similar the embeddings are to each other:

def cosine_similarity(a, b):
    """Cosine similarity between two 1-D numpy arrays."""
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

sim_cancel_terminate = cosine_similarity(embeddings[0], embeddings[1])
sim_cancel_pricing   = cosine_similarity(embeddings[0], embeddings[2])
sim_cancel_weather   = cosine_similarity(embeddings[0], embeddings[3])

print(f"'cancel subscription' vs 'terminate account':  {sim_cancel_terminate:.4f}")
print(f"'cancel subscription' vs 'pricing page':       {sim_cancel_pricing:.4f}")
print(f"'cancel subscription' vs 'Paris weather':      {sim_cancel_weather:.4f}")

# Typical output:
# 'cancel subscription' vs 'terminate account':  0.8312
# 'cancel subscription' vs 'pricing page':       0.3150
# 'cancel subscription' vs 'Paris weather':      0.0421

The model was never explicitly told "cancel" and "terminate account" mean the same thing. It inferred this geometry from contrastive training data. This mechanism is the engine of semantic search.

TIP
Batch Encoding at Scale
The model.encode(sentences) call batches inputs automatically. For large corpora pass batch_size=64 and show_progress_bar=True. On a modern CPU, MiniLM-L6-v2 can embed tens of thousands of short sentences per minute — no GPU required for experimentation or moderate-scale workloads.

About all-MiniLM-L6-v2: it is a 22 M-parameter, 384-dimension model optimised for speed and is the default recommendation for local experimentation. When you move to production you will choose a larger model based on your quality requirements — covered in Step 6.


4 Distance Metrics — Cosine, Dot Product, and Euclidean

When you query a vector database, it compares your query embedding against every stored embedding using a distance metric. Choosing the right metric matters because different metrics make different assumptions about your vectors.

The three metrics you will encounter:

1. Cosine similarity — measures the angle between two vectors, ignoring their magnitude. If two vectors point in exactly the same direction, cosine similarity is 1.0. Opposite directions give -1.0. Perpendicular gives 0.

import numpy as np

def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

# cosine distance = 1 - cosine_similarity  (some libraries use distance)

Cosine similarity is the most common choice for text embeddings because it is invariant to vector magnitude — the length of a vector depends on pooling and normalisation decisions, not meaning.

2. Dot product — multiplies corresponding elements and sums them:

def dot_product(a, b):
    return float(np.dot(a, b))

For unit-normalised vectors (L2 norm = 1), dot product and cosine similarity are identical. The difference only matters when vectors are not normalised. Most production embedding models output L2-normalised vectors by default, making dot product the preferred choice in vector databases — it is faster because it avoids the norm division in the hot path.

3. Euclidean distance (L2) — the straight-line distance between two points in vector space:

def euclidean_distance(a, b):
    return float(np.linalg.norm(a - b))

Euclidean distance is sensitive to vector magnitude, which makes it appropriate for embeddings that encode quantity or magnitude (image pixel intensities, count features). For text embeddings it is generally inferior to cosine similarity unless the vectors are normalised.

A worked comparison:

a = np.array([0.8, 0.3, 0.5])
b = np.array([0.7, 0.4, 0.6])
c = np.array([0.1, 0.9, 0.2])

print(f"cosine(a, b) = {cosine_similarity(a, b):.4f}")   # high: similar direction
print(f"cosine(a, c) = {cosine_similarity(a, c):.4f}")   # low: different direction
print(f"dot(a, b)    = {dot_product(a, b):.4f}")
print(f"euclid(a, b) = {euclidean_distance(a, b):.4f}")  # small: close in space
print(f"euclid(a, c) = {euclidean_distance(a, c):.4f}")  # larger: farther apart

Decision table:

Metric Use When
Cosine similarity Vectors may not be normalised; focus is on directional similarity
Dot product Vectors are L2-normalised (most embedding models); fastest at query time
Euclidean (L2) Embeddings encode magnitude as well as direction (rare for text)
WARNING
Check Your Vector Database's Default
Different vector databases default to different metrics. Pinecone uses cosine by default; pgvector uses L2 by default; Qdrant lets you specify at collection creation time. Using the wrong metric with non-normalised vectors gives meaningless similarity scores. Always check the docs and verify your embedding model's output normalisation before choosing a metric.
NOTE
Normalised vs Non-Normalised
If you call model.encode(sentences, normalize_embeddings=True) in sentence-transformers, the output vectors are L2-normalised to unit length. That makes cosine similarity and dot product equivalent, and Euclidean distance proportional to both. Many models normalise by default — check the model card.

5 Visualising What Embeddings Capture

Before committing to a model for production, it is worth testing it on your actual domain vocabulary. Two experiments that reveal a model's quality quickly are analogy probing and cluster inspection.

Analogy probing. Good embeddings capture analogical relationships. The classic example from the original Word2Vec paper: "king" - "man" + "woman" lands near "queen." Modern sentence encoders support richer analogies:

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")

# Test semantic clustering on your domain vocabulary
domain_terms = [
    # Customer support
    "subscription cancellation",
    "account termination",
    "refund request",
    # Technical
    "database connection error",
    "server timeout",
    "memory leak",
    # Unrelated
    "chocolate cake recipe",
    "mountain hiking trail",
]

vecs = model.encode(domain_terms, normalize_embeddings=True)

# Build a similarity matrix
sim_matrix = vecs @ vecs.T   # dot product == cosine for normalised vectors

# Print the top-2 nearest neighbours for each term
for i, term in enumerate(domain_terms):
    row = sim_matrix[i].copy()
    row[i] = -1             # exclude self
    top2 = np.argsort(row)[::-1][:2]
    print(f"{term!r}")
    for j in top2:
        print(f"  -> {domain_terms[j]!r}  ({sim_matrix[i, j]:.3f})")
    print()

You should see the customer-support terms cluster together, the technical terms cluster together, and both groups stay far from the recipes and hiking content.

What to look for in the output:

  • Terms you know are synonymous should have similarity above ~0.75
  • Terms from different domains should be below ~0.3
  • If near-synonyms score below 0.5, the model may be too small or not domain-adapted
TIP
Test on Your Own Data First
Before spending time benchmarking models on public datasets, test them on 20–30 sentence pairs from your actual corpus — a few known-similar pairs and a few known-unrelated pairs. A five-minute spot check often tells you more than an MTEB score about whether a model will work for your specific domain.
NOTE
Matryoshka Embeddings
Several modern models (including OpenAI's text-embedding-3 series and Cohere embed-v4) support Matryoshka Representation Learning — the model is trained so that the first N dimensions of a larger embedding are themselves a high-quality N-dimensional embedding. This means you can truncate a 1 536-dim vector to 256 dims with only a small quality loss, enabling a cost-quality slider without re-embedding your corpus.

6 Choosing an Embedding Model

The MTEB (Massive Text Embedding Benchmark) leaderboard is the standard reference for comparing embedding model quality across retrieval, classification, clustering, and other tasks. The benchmark covers dozens of tasks and multiple languages.

Key models to know in 2026:

Open-source / self-hosted:

Model Dims Max Tokens MTEB Class Notes
all-MiniLM-L6-v2 384 256 Good Best for fast experimentation; very small
BGE-M3 (BAAI) 1 024 8 192 Strong Dense + sparse + multi-vector; 100+ languages
bge-large-en-v1.5 1 024 512 Strong English-only, well-balanced quality/size
Qwen3-Embedding-8B 4 096 32 768 Near SOTA Apache 2.0; decoder-based; GPU required

API-based (managed):

Model Dims Max Tokens Normalised Notes
OpenAI text-embedding-3-small 1 536 8 192 Yes Most widely used API model; cost-effective
OpenAI text-embedding-3-large 3 072 8 192 Yes Higher MTEB score; ~6x more expensive
Cohere embed-v4 256–1 536 128 000 Yes Multimodal; long-context; Matryoshka dims

Code example — OpenAI Embeddings API:

from openai import OpenAI
import numpy as np

client = OpenAI()   # reads OPENAI_API_KEY from environment

response = client.embeddings.create(
    model="text-embedding-3-small",
    input=["How do I cancel my subscription?",
           "What is the process to terminate my account?"],
    dimensions=512,    # Matryoshka: request fewer dims to save cost
)

vecs = np.array([item.embedding for item in response.data])
print(vecs.shape)   # (2, 512)

Code example — local BGE-M3 via sentence-transformers:

from sentence_transformers import SentenceTransformer

# BGE-M3 is ~2 GB download; cached after first use
model = SentenceTransformer("BAAI/bge-m3")

docs = [
    "Subscription cancellation policy",
    "How to request a refund",
    "Server configuration guide",
]

# Dense embeddings; normalised by default for BGE models
embeddings = model.encode(docs, normalize_embeddings=True)
print(embeddings.shape)   # (3, 1024)

Decision framework:

Is your corpus multilingual or longer than 512 tokens?
  YES  -->  BGE-M3 (local) or Cohere embed-v4 (API)
  NO   -->  continue

Do you need zero-infra, managed, always-up-to-date?
  YES  -->  OpenAI text-embedding-3-small  (good default)
  NO   -->  continue

Do you need the lowest possible latency at high throughput?
  YES  -->  all-MiniLM-L6-v2 or bge-small-en-v1.5 (tiny, fast)
  NO   -->  bge-large-en-v1.5 or text-embedding-3-small
WARNING
Lock Your Embedding Model
Once you have embedded your corpus and stored the vectors, you CANNOT change the embedding model without re-embedding everything. Vectors from different models live in incompatible geometric spaces — mixing them silently corrupts your search results. Choose your model deliberately and treat the choice as a production dependency to pin and version.

7 Putting It Together — Semantic Search in 30 Lines

Everything in this lesson comes together in a minimal semantic search function. This is the retrieval core of every RAG pipeline — the query goes in, the most relevant passages come out.

from sentence_transformers import SentenceTransformer
import numpy as np

# ---- Index time (done once per corpus) ----

model = SentenceTransformer(
    "sentence-transformers/all-MiniLM-L6-v2"
)

# Your knowledge base: list of text chunks
corpus = [
    "To cancel your subscription, go to Account Settings and click Cancel Plan.",
    "Refunds are processed within 5-7 business days to the original payment method.",
    "You can upgrade your plan at any time from the Billing section.",
    "Contact support at [email protected] for billing issues.",
    "The API rate limit is 1000 requests per minute on the Pro plan.",
    "Two-factor authentication can be enabled under Security Settings.",
]

# Embed the corpus (do this once; store the result in a vector DB)
corpus_embeddings = model.encode(corpus, normalize_embeddings=True)
# shape: (6, 384)

# ---- Query time (done for every user query) ----

def semantic_search(query: str, top_k: int = 3) -> list[dict]:
    query_embedding = model.encode(
        [query], normalize_embeddings=True
    )[0]

    # Cosine similarity == dot product for normalised vectors
    scores = corpus_embeddings @ query_embedding

    top_indices = np.argsort(scores)[::-1][:top_k]

    return [
        {"text": corpus[i], "score": float(scores[i])}
        for i in top_indices
    ]

results = semantic_search("How do I get my money back?")
for r in results:
    print(f"[{r['score']:.3f}]  {r['text']}")

# Example output:
# [0.791]  Refunds are processed within 5-7 business days ...
# [0.423]  To cancel your subscription, go to Account Settings ...
# [0.311]  Contact support at [email protected] for billing issues.

The query "How do I get my money back?" contains none of the words in the top result — "refund," "processed," "payment method" — but the embeddings place them in the same semantic neighbourhood. That is the payoff of the whole lesson.

What you would add in a real system:

  • Replace the in-memory NumPy matrix with a vector database (covered in Lesson 3)
  • Replace the simple text list with properly chunked documents (covered in Lesson 4)
  • Pass the top-k results to an LLM as context (covered in Lesson 5)
NOTE
FAISS for In-Process Search
For corpora up to a few hundred thousand chunks, Facebook's FAISS library provides fast approximate nearest-neighbour search entirely in-process — no network round-trip. It is a useful middle ground between a NumPy dot product and a full vector database. Install with pip install faiss-cpu. For larger corpora or when you need metadata filtering, a dedicated vector database is the right tool.

Questions & Answers

Q: My embedding model has a 512-token limit, but my documents are much longer. Do I need to truncate them?
Yes — if you pass more tokens than the model's context window, most libraries silently truncate the input. The embedding then only captures the first 512 tokens, which may miss the key content. The solution is chunking: split documents into segments that fit within the model's context before embedding. This is covered in depth in Lesson 4: Chunking Strategies. For corpora with many long documents, consider a model with a longer context window — BGE-M3 supports up to 8 192 tokens, and Cohere embed-v4 extends to 128 000 tokens.
Q: Should I fine-tune an embedding model on my domain data, or just use a general-purpose model?
Start with a general-purpose model. For most production workloads, the gap between a well-chosen off-the-shelf model and a fine-tuned one is smaller than the improvements you will get from better chunking, hybrid search, and re-ranking. Fine-tuning pays off when: (a) your domain has highly specialised vocabulary that general models have rarely seen (biomedical, legal, code in niche languages), and (b) you have labelled query-passage pairs to train on. If you do fine-tune, start from a strong checkpoint (bge-large, E5-large) rather than from scratch — the contrastive training data already in those models is expensive to replicate.
Q: Cosine similarity scores look low (0.5–0.6) even for documents I know are relevant. Is the model broken?
Not necessarily — cosine similarity is not a probability. Raw scores of 0.5–0.7 can still represent excellent retrieval, especially for long documents or domain-specific content. What matters is the ranking: relevant documents should score consistently higher than irrelevant ones. If the relevant document ranks third instead of first, that is a retrieval problem. Absolute score thresholds are unreliable across different models; relative ranking is what your pipeline depends on. When setting a minimum-score filter, calibrate it empirically on your test set rather than picking a round number.
Q: I need to embed 10 million documents. How long will that take and what will it cost?
It depends heavily on your model and infrastructure. With a local model on a single GPU (e.g. NVIDIA A10), you can expect roughly 5 000–20 000 short sentences per second — so 10 million chunks might take minutes to hours. On CPU with MiniLM, expect tens of thousands per minute, meaning several hours. With the OpenAI API, at 10 million chunks of ~200 tokens each, you are sending roughly 2 billion tokens — check current API pricing for that token count. For large-scale indexing jobs, batch the requests, use a GPU if available, and cache the output immediately. Re-embedding 10 M documents because you forgot to save the vectors is painful. Store embeddings in the vector database as you go, do not buffer them all in memory.
Q: Can I use the same embedding model for both documents and queries, or do I need different models?
Most embedding models use the same model for both — you encode documents at index time and encode the query at retrieval time, then compare them in the same vector space. That is the standard approach and works well for symmetric tasks (document-to-document similarity, paraphrase detection). Some models are designed for asymmetric retrieval, where short queries are matched against long passages. BGE and E5 models expose this via instruction prefixes: you prepend "Represent this sentence for searching relevant passages:" to the query and a different prefix to the document. Check your model's card for the recommended asymmetric usage pattern, especially for question-answering workloads.

Key Takeaways

  1. Embeddings are meaning as geometry — dense floating-point vectors where similar meanings produce geometrically close vectors, enabling semantic search that matches intent rather than keywords.
  2. Encoder models read bidirectionally — unlike generative LLMs, encoder architectures attend to the full input context simultaneously and are fine-tuned with contrastive learning to give the vector space its semantic structure.
  3. Use dot product for normalised vectors — for L2-normalised embeddings (the common case), dot product equals cosine similarity and is faster; only fall back to Euclidean when your embeddings encode magnitude information.
  4. The MTEB leaderboard is your starting point — use it to shortlist models, then test on your own domain vocabulary before committing; lock your model choice early because changing it requires re-embedding the entire corpus.
  5. sentence-transformers gets you running in three lines — open-source, CPU-capable, and wraps hundreds of MTEB-ranked models behind a consistent API; BGE-M3 is a strong open-source default for multilingual or long-document corpora.
  6. Model selection is a production dependency — once vectors are stored, the embedding model is pinned; mixing vectors from different models silently breaks similarity scores, so treat the model choice with the same care as a database schema migration.

Next Steps: Lesson 3: Vector Databases