Vector Databases
Learning Outcomes
- Explain why general-purpose databases are unsuitable for embedding search and what vector databases provide instead
- Describe how HNSW and IVF indexing algorithms trade recall for query speed and when to choose each
- Set up Chroma locally and pgvector on PostgreSQL, load embeddings, and run similarity queries in Python
- Apply metadata filters to narrow search results without sacrificing index performance
- Select the right vector database for a given project using a structured decision framework
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why ordinary databases fail at vector search |
| Explain | 7 min | Indexing algorithms — HNSW and IVF in depth |
| Demo | 8 min | Chroma — local development setup and queries |
| Demo | 8 min | pgvector — adding vector search to PostgreSQL |
| Explain | 6 min | Managed options — Pinecone and Weaviate overview |
| Demo | 8 min | Metadata filtering — patterns and pitfalls |
| Explain | 6 min | Selection framework and capacity planning |
| Wrap-up | 4 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- Complete Lesson 1: What Is RAG? — you need the conceptual picture of how retrieval fits into a RAG pipeline
- Complete Lesson 2: Embeddings Explained — this lesson assumes you can produce embeddings and understand cosine similarity
- Have Python 3.9 or later installed and a working virtual environment
Shopping List:
- Python with
pip install chromadb sentence-transformers pgvector psycopg[binary] - PostgreSQL 13 or later (local or a free tier on Supabase or Neon) for the pgvector steps
- Optional: a Pinecone free-tier account for the managed service section
- Optional: Docker, if you prefer to run PostgreSQL in a container
If you have worked through Lesson 2, you know that an embedding is a dense vector of floating-point numbers — typically 384 to 3072 dimensions depending on the model. Finding the documents most similar to a query means finding the vectors in your corpus that are closest to the query vector, usually measured by cosine similarity or L2 (Euclidean) distance.
The naive approach — loading all vectors and computing pairwise distances — is called exact nearest-neighbor search (also written k-NN). On a million vectors of 768 dimensions each, exact search requires roughly 768 million floating-point operations per query. At hundreds of millions of FLOP/s per CPU core, that means hundreds of milliseconds per query before you touch the network or the LLM. At ten million vectors it becomes completely impractical.
A conventional relational database has no data structure optimized for this kind of geometric comparison. Its B-tree indexes are designed for equality lookups and range scans on scalar values, not for navigating a high-dimensional vector space.
What a vector database adds:
| Capability | How it is delivered |
|---|---|
| Sub-millisecond ANN queries at scale | Graph or cluster-based indexes (HNSW, IVF) |
| Per-document metadata | Columnar or JSON side-storage, filterable at query time |
| Distance metric choice | L2, cosine, dot-product operators wired into the index |
| Incremental inserts | Index structures that support live upserts |
| Recall/speed tuning | Parameters that let you trade accuracy for latency |
ANN stands for approximate nearest-neighbor search. Rather than guaranteeing you find the absolute closest vectors, ANN indexes find vectors that are very likely to be among the closest — in practice achieving 95–99% recall at a fraction of the cost of exact search. For RAG, 95% recall is almost always sufficient: missing one relevant chunk occasionally is far less damaging than adding 200 ms to every query.
Every vector database lets you choose or configure an indexing algorithm. You will encounter two algorithms in almost every system: HNSW and IVF. Understanding them at an intuitive level lets you tune the parameters correctly when recall starts drifting.
HNSW — Hierarchical Navigable Small World
HNSW builds a multi-layer graph. Each vector becomes a node. At the bottom layer (layer 0) every node is connected to a small number of its nearest neighbors. Higher layers contain progressively fewer nodes and act as "express lanes" that allow the search to jump across the space in large strides before descending for fine-grained navigation.
At query time the algorithm enters at the top layer, greedily follows edges toward the query vector, descends to the next layer, and repeats until it reaches layer 0, returning the best candidates found.
The two key build-time parameters are:
| Parameter | Meaning | Default | Effect of increasing |
|---|---|---|---|
m |
Max connections per node per layer | 16 | Higher recall, more memory, slower build |
ef_construction |
Candidate list size during build | 64 | Higher recall, slower build |
The query-time parameter ef_search controls how many candidates are tracked during the search — a larger value gives higher recall at the cost of slightly more latency. HNSW supports incremental inserts without rebuilding the index, making it the default choice for most production deployments.
-- pgvector: HNSW index with explicit parameters
CREATE INDEX ON documents
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
-- Tune recall at query time (session-scoped):
SET hnsw.ef_search = 100;
IVF — Inverted File Index
IVF first runs k-means clustering on your vectors, partitioning them into nlist clusters (cells). Each vector is assigned to its nearest cluster centroid. At query time the algorithm identifies the nprobe clusters closest to the query and searches only within those clusters.
| Parameter | Meaning | Build-time or query-time |
|---|---|---|
nlist |
Number of clusters | Build-time — requires rebuild to change |
nprobe |
How many clusters to search per query | Query-time — safe to tune live |
IVF indexes build faster and use less memory than HNSW, but query performance is generally lower for a given recall target. IVF also requires you to have enough vectors before the k-means clustering is meaningful — a common guideline is at least 100 rows per list value.
-- pgvector: IVFFlat index
CREATE INDEX ON documents
USING ivfflat (embedding vector_l2_ops)
WITH (lists = 200);
-- Tune probe count at query time:
SET ivfflat.probes = 20;
Choosing between them
| Situation | Recommended algorithm |
|---|---|
| Corpus grows continuously with live inserts | HNSW |
| Very large static corpus, memory is tight | IVF |
| You need the fastest possible query at a given recall | HNSW |
| Corpus is under ~50 000 vectors | No index needed — exact search is fast enough |
Chroma is an open-source vector database designed for fast developer iteration. The 1.5.x release (May 2026) ships as a single Python package with no external services required for local use.
Install
pip install chromadb sentence-transformers
Three client modes
Chroma offers three client types that match different stages of the development lifecycle:
import chromadb
# 1. In-memory: data lost when process exits. Use for unit tests.
client = chromadb.EphemeralClient()
# 2. Persistent: saves to disk. Use for local prototyping.
client = chromadb.PersistentClient(path="./chroma_db")
# 3. Client-server: connects to a running Chroma server.
# Launch server first: chroma run --path ./chroma_db
client = chromadb.HttpClient(host="localhost", port=8000)
Create a collection and load data
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
collection = client.get_or_create_collection(
name="docs",
metadata={"hnsw:space": "cosine"}, # distance metric for the index
)
documents = [
"The RAG pipeline retrieves context before generating answers.",
"Embeddings map text to dense vectors that capture semantic meaning.",
"Vector databases store and index high-dimensional embeddings.",
"Chunking splits documents into fragments suitable for retrieval.",
]
ids = [f"doc_{i}" for i in range(len(documents))]
embeddings = model.encode(documents).tolist()
collection.add(
ids=ids,
embeddings=embeddings,
documents=documents,
metadatas=[
{"source": "lesson-01", "topic": "rag"},
{"source": "lesson-02", "topic": "embeddings"},
{"source": "lesson-03", "topic": "vector-db"},
{"source": "lesson-04", "topic": "chunking"},
],
)
Query
query = "How do I search embeddings efficiently?"
query_embedding = model.encode(query).tolist()
results = collection.query(
query_embeddings=[query_embedding],
n_results=2,
include=["documents", "distances", "metadatas"],
)
for doc, dist, meta in zip(
results["documents"][0],
results["distances"][0],
results["metadatas"][0],
):
print(f"[distance={dist:.4f}] {doc} (source: {meta['source']})")
Chroma returns distances where a lower value means more similar (for cosine space, 0 = identical). You will typically see the vector-databases document returned at the top for this query.
metadata={'hnsw:space': 'cosine'} when creating a collection to use cosine similarity. The default space is L2. You cannot change the metric after data is loaded — you must recreate the collection.If your application already runs on PostgreSQL, pgvector lets you add vector similarity search without adopting a second database. You store embeddings in a standard Postgres column, create an index, and query with SQL — all within the same transaction model, connection pool, and backup strategy you already operate.
Enable the extension
-- Run once per database (requires superuser or the extension to be trusted)
CREATE EXTENSION IF NOT EXISTS vector;
Schema
CREATE TABLE documents (
id SERIAL PRIMARY KEY,
content TEXT NOT NULL,
source TEXT,
topic TEXT,
embedding VECTOR(384) -- dimension must match your embedding model
);
Install the Python helper library
pip install pgvector "psycopg[binary]"
Load data with psycopg (v3)
import psycopg
from pgvector.psycopg import register_vector
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
rows = [
("lesson-01", "rag", "The RAG pipeline retrieves context before generating answers."),
("lesson-02", "embeddings", "Embeddings map text to dense vectors that capture semantic meaning."),
("lesson-03", "vector-db", "Vector databases store and index high-dimensional embeddings."),
]
with psycopg.connect("postgresql://user:pass@localhost/mydb") as conn:
register_vector(conn) # teaches psycopg how to serialise VECTOR columns
with conn.cursor() as cur:
for source, topic, content in rows:
embedding = model.encode(content) # numpy array, shape (384,)
cur.execute(
"INSERT INTO documents (content, source, topic, embedding)"
" VALUES (%s, %s, %s, %s)",
(content, source, topic, embedding),
)
conn.commit()
Create an HNSW index
CREATE INDEX ON documents
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
Query — nearest neighbors with SQL
query_vec = model.encode("How does semantic search work?")
with psycopg.connect("postgresql://user:pass@localhost/mydb") as conn:
register_vector(conn)
with conn.cursor() as cur:
cur.execute(
"""
SELECT content, source,
1 - (embedding <=> %s) AS cosine_similarity
FROM documents
ORDER BY embedding <=> %s
LIMIT 3
""",
(query_vec, query_vec),
)
for content, source, sim in cur.fetchall():
print(f"[sim={sim:.4f}] ({source}) {content}")
The <=> operator returns cosine distance (0 = identical, 2 = opposite). Subtracting from 1 converts it to cosine similarity between -1 and 1, matching the convention introduced in Lesson 2.
| pgvector distance operator | Metric | Index ops class |
|---|---|---|
<-> |
L2 (Euclidean) | vector_l2_ops |
<=> |
Cosine distance | vector_cosine_ops |
<#> |
Negative inner product | vector_ip_ops |
VECTOR(384) must match the output size of your embedding model exactly. If you switch to a different model later, you must recreate the column and re-embed every document. Choose your embedding model before designing your schema.For teams that want to skip infrastructure management entirely, two managed services dominate the market.
Pinecone
Pinecone is a fully managed vector database that exposes a simple REST/gRPC API. You create an index, upsert vectors with metadata, and query — no servers to provision, no index parameters to tune by hand.
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key="YOUR_API_KEY")
# Create a serverless index
pc.create_index(
name="my-rag-index",
dimension=384,
metric="cosine",
spec=ServerlessSpec(cloud="aws", region="us-east-1"),
)
index = pc.Index("my-rag-index")
# Upsert vectors with metadata
index.upsert(
vectors=[
{
"id": "doc-0",
"values": [0.1, 0.2, ...], # 384-dim list
"metadata": {"source": "lesson-01", "topic": "rag"},
},
{
"id": "doc-1",
"values": [0.3, 0.1, ...],
"metadata": {"source": "lesson-02", "topic": "embeddings"},
},
],
namespace="production",
)
# Query
query_result = index.query(
vector=query_embedding,
top_k=5,
include_metadata=True,
namespace="production",
filter={"topic": {"$eq": "rag"}},
)
for match in query_result["matches"]:
print(f"[score={match['score']:.4f}] {match['id']} — {match['metadata']}")
Pinecone returns similarity scores (higher = more similar) rather than distances, which is the reverse of Chroma's convention — worth noting when you normalise results downstream.
Namespaces partition data within a single index, useful for multi-tenant applications where you want customer A's data isolated from customer B's without maintaining separate indexes.
Weaviate
Weaviate is an open-source vector database available as a managed cloud service or self-hosted. Its headline feature is first-class hybrid search — combining dense vector search with BM25 keyword search in a single query. This matters for RAG: product names, proper nouns, and technical identifiers (like a ticket number or a function name) are retrieved better by keyword matching than by semantics alone.
import weaviate
from weaviate.classes.query import Filter, HybridFusion
client = weaviate.connect_to_weaviate_cloud(
cluster_url="https://YOUR_CLUSTER.weaviate.network",
auth_credentials=weaviate.auth.AuthApiKey("YOUR_API_KEY"),
)
collection = client.collections.use("Documents")
# Hybrid search: alpha=0.5 is an even blend of vector and BM25
response = collection.query.hybrid(
query="semantic search embeddings",
alpha=0.5, # 0 = BM25 only, 1 = vector only
fusion_type=HybridFusion.RELATIVE_SCORE,
filters=Filter.by_property("topic").equal("vector-db"),
limit=5,
)
for obj in response.objects:
print(obj.properties)
client.close()
The alpha parameter is the primary lever for tuning hybrid search. Start at 0.5 and adjust based on your evaluation results from Lesson 7.
Qdrant — an honourable mention
Qdrant is an open-source vector database written in Rust, available as a managed cloud service or self-hosted. Its Python client API is clean and idiomatic. Qdrant calls metadata "payload" and offers sophisticated payload filtering with a built-in payload index — filtering is genuinely fast even on large collections because payload conditions can prune the HNSW graph during traversal rather than being applied after the fact.
from qdrant_client import QdrantClient
from qdrant_client.http.models import (
Distance, VectorParams, Filter, FieldCondition, MatchValue
)
client = QdrantClient(url="http://localhost:6333")
client.create_collection(
collection_name="docs",
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)
# Filtered nearest-neighbor search
results = client.search(
collection_name="docs",
query_vector=query_embedding,
query_filter=Filter(
must=[FieldCondition(key="topic", match=MatchValue(value="rag"))]
),
limit=5,
)
Metadata filtering narrows vector search to a subset of your corpus before (pre-filter) or after (post-filter) the ANN search. For RAG pipelines it is indispensable: you typically want to search only the documents relevant to a tenant, a date range, a document type, or an access-control group.
Pre-filter vs post-filter
Pre-filtering restricts the ANN search to only the matching rows. It guarantees you get exactly k results (if at least k match), but disrupts HNSW graph traversal because the algorithm can only follow edges that satisfy the filter. For very selective filters this is fine; for broad filters it becomes expensive.
Post-filtering performs the full ANN search and then discards rows that do not match the filter. It is fast (the index works unimpeded) but can return fewer than k results if many top-k candidates are filtered out.
Most modern vector databases use a hybrid approach: they attempt pre-filtering when the filter is highly selective (matching a small fraction of vectors) and fall back to post-filtering otherwise. As an application developer you typically declare the filter and trust the database to choose, but it helps to know the difference when debugging low recall.
Chroma: filtering with the where clause
# Match a single metadata field
results = collection.query(
query_embeddings=[query_embedding],
n_results=3,
where={"topic": "rag"}, # simple equality
)
# Combine conditions with $and
results = collection.query(
query_embeddings=[query_embedding],
n_results=3,
where={
"$and": [
{"topic": {"$in": ["rag", "vector-db"]}},
{"source": {"$ne": "lesson-04"}},
]
},
)
pgvector: filtering is just SQL
With pgvector, metadata filtering is a standard SQL WHERE clause. Because PostgreSQL's query planner can combine the vector index scan with regular B-tree index scans on your metadata columns, the filtering story is more expressive than in purpose-built databases.
-- Add a B-tree index on the metadata column you filter most often
CREATE INDEX ON documents (topic);
-- Combined vector search + metadata filter
SELECT content, source,
1 - (embedding <=> $1) AS cosine_similarity
FROM documents
WHERE topic = 'rag' -- metadata filter
ORDER BY embedding <=> $1 -- vector distance
LIMIT 5;
Pinecone: filter in the query call
Pinecone metadata filters use a MongoDB-style JSON syntax in the filter argument to index.query():
results = index.query(
vector=query_embedding,
top_k=5,
include_metadata=True,
filter={
"$and": [
{"topic": {"$in": ["rag", "embeddings"]}},
{"published_year": {"$gte": 2024}},
]
},
)
Common mistakes
| Mistake | Symptom | Fix |
|---|---|---|
| Filtering on a high-cardinality field with no index | Slow queries as corpus grows | Add a payload/metadata index on that field |
| Overly selective filter with small corpus | Fewer than k results returned | Increase n_results / top_k and trim after retrieval |
| Storing user/tenant ID only in metadata, not in the vector space | Cross-tenant leaks if filter is accidentally omitted | Always include tenant_id in query filter; enforce at the middleware layer |
Mixing types for the same key (e.g., year as "2024" in one doc and 2024 int in another) |
Filter silently misses records | Enforce a schema in the ingestion pipeline |
With four credible options covered, here is a decision framework you can apply to real projects.
Decision tree
Does your stack already run PostgreSQL?
├─ YES → Use pgvector. You get vector search with no new
│ infrastructure and full SQL expressiveness.
│
└─ NO → Is hybrid search (vector + keyword) important?
├─ YES → Weaviate (self-hosted or cloud). First-class
│ BM25 + vector fusion with the alpha parameter.
│
└─ NO → Is this local/dev-only or early-stage?
├─ YES → Chroma with PersistentClient. Zero-config,
│ pure Python, swap to any other store later.
│
└─ NO → Do you want fully managed with no ops overhead?
├─ YES → Pinecone. Serverless, auto-scales,
│ minimal configuration required.
│
└─ NO → Qdrant (self-hosted). Open source,
Rust-speed, fine-grained HNSW tuning
and efficient filtered search.
Comparison table
| Chroma | pgvector | Pinecone | Weaviate | Qdrant | |
|---|---|---|---|---|---|
| Hosting | Local / Cloud | Anywhere Postgres runs | Fully managed | Cloud or self-hosted | Cloud or self-hosted |
| Hybrid search | No | No (vector only) | No | Yes (BM25 + vector) | Yes (sparse + dense) |
| Index algorithm | HNSW | HNSW or IVFFlat | Managed | HNSW | HNSW |
| Metadata filter | Yes (where) |
Yes (SQL) | Yes (JSON filter) | Yes (GraphQL/Python) | Yes (payload filter) |
| Best for | Dev / prototyping | Existing Postgres teams | No-ops production | Hybrid retrieval | High-throughput self-hosted |
Capacity planning rules of thumb
The memory footprint of an HNSW index depends on the number of vectors, their dimension, and the m parameter. A rough working formula:
memory_bytes ≈ num_vectors × dimension × 4 (float32) × 1.5 (index overhead)
| Corpus size | Dimension 384 (MiniLM) | Dimension 1536 (OpenAI) |
|---|---|---|
| 100 000 vectors | ~230 MB | ~920 MB |
| 1 000 000 vectors | ~2.3 GB | ~9.2 GB |
| 10 000 000 vectors | ~23 GB | ~92 GB |
These are rough estimates; actual overhead varies by database and m value. Use them to size your instance class before you load data, not after.
Query latency targets: aim for under 50 ms for the vector search step on your target corpus. If you are seeing higher latency, the first levers are: reducing ef_search (HNSW) or nprobe (IVF) and checking whether metadata filtering is triggering a full scan.
Questions & Answers
dense_vector field and query with the knn clause. The advantage is zero new infrastructure and native integration with your existing BM25 full-text search (which gives you hybrid search out of the box). The tradeoff is that Elasticsearch's kNN implementation uses HNSW but with fewer tuning knobs than dedicated vector databases, and at very large scale it consumes significantly more memory than purpose-built solutions. If your team already operates Elasticsearch competently, it is a legitimate starting point and an easy migration to a dedicated store later if needed.m=16, ef_construction=64, ef_search=64) you should see 95–99% recall compared to exact search, meaning roughly 1 in 50 results may differ. If you are seeing larger differences, increase ef_construction at build time and ef_search at query time. Also verify that the distance metric in the index (e.g. vector_cosine_ops) matches the metric you are comparing against in brute-force code.Key Takeaways
-
General-purpose databases cannot search embeddings at scale — B-tree indexes are designed for scalar comparisons, not high-dimensional geometry. Vector databases provide ANN indexes (HNSW, IVF) that deliver sub-millisecond queries at millions of vectors.
-
HNSW is the right default — graph-based, supports incremental inserts, high recall at fast query times. IVF is an alternative when memory is tight and the corpus is mostly static. Neither index is needed below around 50 000 vectors.
-
Chroma is the fastest path to a working local setup — three-line install, three client modes, no external services. Use it for development and swap to a production store later without changing the rest of the pipeline.
-
pgvector earns its place if you already run PostgreSQL — the extension adds vector columns, HNSW and IVFFlat indexes, and three distance operators to standard SQL. Metadata filtering is just a WHERE clause, and every hosted Postgres service supports it.
-
Metadata filtering is mandatory in production but has traps — always index the fields you filter on, do not rely on filters as a security boundary, and be aware of the pre-filter vs post-filter tradeoff when recall is unexpectedly low.
-
Choose by operational fit, not benchmark scores — pgvector for Postgres shops, Chroma for fast prototypes, Weaviate when hybrid search matters, Pinecone when you want no infrastructure to manage. All four are legitimate production choices used by large-scale systems.
Next Steps: Lesson 4: Chunking Strategies