Vector Databases

50 min beginner Lesson 3

Learning Outcomes

  • Explain why general-purpose databases are unsuitable for embedding search and what vector databases provide instead
  • Describe how HNSW and IVF indexing algorithms trade recall for query speed and when to choose each
  • Set up Chroma locally and pgvector on PostgreSQL, load embeddings, and run similarity queries in Python
  • Apply metadata filters to narrow search results without sacrificing index performance
  • Select the right vector database for a given project using a structured decision framework

Lesson Plan

Segment Duration Topic
Intro 3 min Why ordinary databases fail at vector search
Explain 7 min Indexing algorithms — HNSW and IVF in depth
Demo 8 min Chroma — local development setup and queries
Demo 8 min pgvector — adding vector search to PostgreSQL
Explain 6 min Managed options — Pinecone and Weaviate overview
Demo 8 min Metadata filtering — patterns and pitfalls
Explain 6 min Selection framework and capacity planning
Wrap-up 4 min Key takeaways and next steps

Before You Begin

Pre-work:

  • Complete Lesson 1: What Is RAG? — you need the conceptual picture of how retrieval fits into a RAG pipeline
  • Complete Lesson 2: Embeddings Explained — this lesson assumes you can produce embeddings and understand cosine similarity
  • Have Python 3.9 or later installed and a working virtual environment

Shopping List:

  • Python with pip install chromadb sentence-transformers pgvector psycopg[binary]
  • PostgreSQL 13 or later (local or a free tier on Supabase or Neon) for the pgvector steps
  • Optional: a Pinecone free-tier account for the managed service section
  • Optional: Docker, if you prefer to run PostgreSQL in a container

1 Why Ordinary Databases Cannot Search Embeddings Efficiently

If you have worked through Lesson 2, you know that an embedding is a dense vector of floating-point numbers — typically 384 to 3072 dimensions depending on the model. Finding the documents most similar to a query means finding the vectors in your corpus that are closest to the query vector, usually measured by cosine similarity or L2 (Euclidean) distance.

The naive approach — loading all vectors and computing pairwise distances — is called exact nearest-neighbor search (also written k-NN). On a million vectors of 768 dimensions each, exact search requires roughly 768 million floating-point operations per query. At hundreds of millions of FLOP/s per CPU core, that means hundreds of milliseconds per query before you touch the network or the LLM. At ten million vectors it becomes completely impractical.

A conventional relational database has no data structure optimized for this kind of geometric comparison. Its B-tree indexes are designed for equality lookups and range scans on scalar values, not for navigating a high-dimensional vector space.

What a vector database adds:

Capability How it is delivered
Sub-millisecond ANN queries at scale Graph or cluster-based indexes (HNSW, IVF)
Per-document metadata Columnar or JSON side-storage, filterable at query time
Distance metric choice L2, cosine, dot-product operators wired into the index
Incremental inserts Index structures that support live upserts
Recall/speed tuning Parameters that let you trade accuracy for latency

ANN stands for approximate nearest-neighbor search. Rather than guaranteeing you find the absolute closest vectors, ANN indexes find vectors that are very likely to be among the closest — in practice achieving 95–99% recall at a fraction of the cost of exact search. For RAG, 95% recall is almost always sufficient: missing one relevant chunk occasionally is far less damaging than adding 200 ms to every query.

NOTE
Exact vs. Approximate Search
Chroma, pgvector, and every production vector database default to approximate search once an index is built. Without an index they fall back to exact (sequential) search — which is fine during development on a few thousand vectors, but becomes painfully slow as your corpus grows into the hundreds of thousands.

2 Indexing Algorithms — HNSW and IVF

Every vector database lets you choose or configure an indexing algorithm. You will encounter two algorithms in almost every system: HNSW and IVF. Understanding them at an intuitive level lets you tune the parameters correctly when recall starts drifting.

HNSW — Hierarchical Navigable Small World

HNSW builds a multi-layer graph. Each vector becomes a node. At the bottom layer (layer 0) every node is connected to a small number of its nearest neighbors. Higher layers contain progressively fewer nodes and act as "express lanes" that allow the search to jump across the space in large strides before descending for fine-grained navigation.

At query time the algorithm enters at the top layer, greedily follows edges toward the query vector, descends to the next layer, and repeats until it reaches layer 0, returning the best candidates found.

The two key build-time parameters are:

Parameter Meaning Default Effect of increasing
m Max connections per node per layer 16 Higher recall, more memory, slower build
ef_construction Candidate list size during build 64 Higher recall, slower build

The query-time parameter ef_search controls how many candidates are tracked during the search — a larger value gives higher recall at the cost of slightly more latency. HNSW supports incremental inserts without rebuilding the index, making it the default choice for most production deployments.

-- pgvector: HNSW index with explicit parameters
CREATE INDEX ON documents
  USING hnsw (embedding vector_cosine_ops)
  WITH (m = 16, ef_construction = 64);

-- Tune recall at query time (session-scoped):
SET hnsw.ef_search = 100;

IVF — Inverted File Index

IVF first runs k-means clustering on your vectors, partitioning them into nlist clusters (cells). Each vector is assigned to its nearest cluster centroid. At query time the algorithm identifies the nprobe clusters closest to the query and searches only within those clusters.

Parameter Meaning Build-time or query-time
nlist Number of clusters Build-time — requires rebuild to change
nprobe How many clusters to search per query Query-time — safe to tune live

IVF indexes build faster and use less memory than HNSW, but query performance is generally lower for a given recall target. IVF also requires you to have enough vectors before the k-means clustering is meaningful — a common guideline is at least 100 rows per list value.

-- pgvector: IVFFlat index
CREATE INDEX ON documents
  USING ivfflat (embedding vector_l2_ops)
  WITH (lists = 200);

-- Tune probe count at query time:
SET ivfflat.probes = 20;

Choosing between them

Situation Recommended algorithm
Corpus grows continuously with live inserts HNSW
Very large static corpus, memory is tight IVF
You need the fastest possible query at a given recall HNSW
Corpus is under ~50 000 vectors No index needed — exact search is fast enough
WARNING
IVF needs data before indexing
Build an IVFFlat index only after loading a substantial number of vectors. If you create the index on an empty table and then insert data, the cluster centroids will be meaningless and recall will be poor. The pgvector docs recommend at least 100 rows per list value.
TIP
Start with HNSW unless you have a reason not to
HNSW is the default in Chroma, Qdrant, and pgvector for good reason. Its incremental insert support and consistently high recall make it the lowest-risk choice for new projects. Switch to IVF only if memory constraints force your hand.

3 Chroma — Local Development in Minutes

Chroma is an open-source vector database designed for fast developer iteration. The 1.5.x release (May 2026) ships as a single Python package with no external services required for local use.

Install

pip install chromadb sentence-transformers

Three client modes

Chroma offers three client types that match different stages of the development lifecycle:

import chromadb

# 1. In-memory: data lost when process exits. Use for unit tests.
client = chromadb.EphemeralClient()

# 2. Persistent: saves to disk. Use for local prototyping.
client = chromadb.PersistentClient(path="./chroma_db")

# 3. Client-server: connects to a running Chroma server.
#    Launch server first:  chroma run --path ./chroma_db
client = chromadb.HttpClient(host="localhost", port=8000)

Create a collection and load data

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")

collection = client.get_or_create_collection(
    name="docs",
    metadata={"hnsw:space": "cosine"},  # distance metric for the index
)

documents = [
    "The RAG pipeline retrieves context before generating answers.",
    "Embeddings map text to dense vectors that capture semantic meaning.",
    "Vector databases store and index high-dimensional embeddings.",
    "Chunking splits documents into fragments suitable for retrieval.",
]
ids = [f"doc_{i}" for i in range(len(documents))]
embeddings = model.encode(documents).tolist()

collection.add(
    ids=ids,
    embeddings=embeddings,
    documents=documents,
    metadatas=[
        {"source": "lesson-01", "topic": "rag"},
        {"source": "lesson-02", "topic": "embeddings"},
        {"source": "lesson-03", "topic": "vector-db"},
        {"source": "lesson-04", "topic": "chunking"},
    ],
)

Query

query = "How do I search embeddings efficiently?"
query_embedding = model.encode(query).tolist()

results = collection.query(
    query_embeddings=[query_embedding],
    n_results=2,
    include=["documents", "distances", "metadatas"],
)

for doc, dist, meta in zip(
    results["documents"][0],
    results["distances"][0],
    results["metadatas"][0],
):
    print(f"[distance={dist:.4f}] {doc}  (source: {meta['source']})")

Chroma returns distances where a lower value means more similar (for cosine space, 0 = identical). You will typically see the vector-databases document returned at the top for this query.

TIP
Collection metadata controls the distance metric
Set metadata={'hnsw:space': 'cosine'} when creating a collection to use cosine similarity. The default space is L2. You cannot change the metric after data is loaded — you must recreate the collection.
NOTE
Chroma Cloud for production
Chroma also offers a managed serverless service when you outgrow local persistence. The Python API is identical — you swap only the client initialisation line.

4 pgvector — Vector Search Inside PostgreSQL

If your application already runs on PostgreSQL, pgvector lets you add vector similarity search without adopting a second database. You store embeddings in a standard Postgres column, create an index, and query with SQL — all within the same transaction model, connection pool, and backup strategy you already operate.

Enable the extension

-- Run once per database (requires superuser or the extension to be trusted)
CREATE EXTENSION IF NOT EXISTS vector;

Schema

CREATE TABLE documents (
    id          SERIAL PRIMARY KEY,
    content     TEXT    NOT NULL,
    source      TEXT,
    topic       TEXT,
    embedding   VECTOR(384)   -- dimension must match your embedding model
);

Install the Python helper library

pip install pgvector "psycopg[binary]"

Load data with psycopg (v3)

import psycopg
from pgvector.psycopg import register_vector
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")

rows = [
    ("lesson-01", "rag",        "The RAG pipeline retrieves context before generating answers."),
    ("lesson-02", "embeddings", "Embeddings map text to dense vectors that capture semantic meaning."),
    ("lesson-03", "vector-db",  "Vector databases store and index high-dimensional embeddings."),
]

with psycopg.connect("postgresql://user:pass@localhost/mydb") as conn:
    register_vector(conn)          # teaches psycopg how to serialise VECTOR columns
    with conn.cursor() as cur:
        for source, topic, content in rows:
            embedding = model.encode(content)  # numpy array, shape (384,)
            cur.execute(
                "INSERT INTO documents (content, source, topic, embedding)"
                " VALUES (%s, %s, %s, %s)",
                (content, source, topic, embedding),
            )
    conn.commit()

Create an HNSW index

CREATE INDEX ON documents
  USING hnsw (embedding vector_cosine_ops)
  WITH (m = 16, ef_construction = 64);

Query — nearest neighbors with SQL

query_vec = model.encode("How does semantic search work?")

with psycopg.connect("postgresql://user:pass@localhost/mydb") as conn:
    register_vector(conn)
    with conn.cursor() as cur:
        cur.execute(
            """
            SELECT content, source,
                   1 - (embedding <=> %s) AS cosine_similarity
            FROM documents
            ORDER BY embedding <=> %s
            LIMIT 3
            """,
            (query_vec, query_vec),
        )
        for content, source, sim in cur.fetchall():
            print(f"[sim={sim:.4f}] ({source}) {content}")

The <=> operator returns cosine distance (0 = identical, 2 = opposite). Subtracting from 1 converts it to cosine similarity between -1 and 1, matching the convention introduced in Lesson 2.

pgvector distance operator Metric Index ops class
<-> L2 (Euclidean) vector_l2_ops
<=> Cosine distance vector_cosine_ops
<#> Negative inner product vector_ip_ops
WARNING
Dimension must be fixed at schema time
The number inside VECTOR(384) must match the output size of your embedding model exactly. If you switch to a different model later, you must recreate the column and re-embed every document. Choose your embedding model before designing your schema.
TIP
pgvector is already in every major hosted PostgreSQL
Google Cloud SQL, AWS RDS, Azure Database for PostgreSQL, Supabase, and Neon all support pgvector. If your stack is already PostgreSQL, enabling vector search may be as simple as running CREATE EXTENSION — no new infrastructure required.

5 Managed Vector Databases — Pinecone and Weaviate

For teams that want to skip infrastructure management entirely, two managed services dominate the market.

Pinecone

Pinecone is a fully managed vector database that exposes a simple REST/gRPC API. You create an index, upsert vectors with metadata, and query — no servers to provision, no index parameters to tune by hand.

from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key="YOUR_API_KEY")

# Create a serverless index
pc.create_index(
    name="my-rag-index",
    dimension=384,
    metric="cosine",
    spec=ServerlessSpec(cloud="aws", region="us-east-1"),
)

index = pc.Index("my-rag-index")

# Upsert vectors with metadata
index.upsert(
    vectors=[
        {
            "id": "doc-0",
            "values": [0.1, 0.2, ...],    # 384-dim list
            "metadata": {"source": "lesson-01", "topic": "rag"},
        },
        {
            "id": "doc-1",
            "values": [0.3, 0.1, ...],
            "metadata": {"source": "lesson-02", "topic": "embeddings"},
        },
    ],
    namespace="production",
)

# Query
query_result = index.query(
    vector=query_embedding,
    top_k=5,
    include_metadata=True,
    namespace="production",
    filter={"topic": {"$eq": "rag"}},
)

for match in query_result["matches"]:
    print(f"[score={match['score']:.4f}] {match['id']} — {match['metadata']}")

Pinecone returns similarity scores (higher = more similar) rather than distances, which is the reverse of Chroma's convention — worth noting when you normalise results downstream.

Namespaces partition data within a single index, useful for multi-tenant applications where you want customer A's data isolated from customer B's without maintaining separate indexes.

Weaviate

Weaviate is an open-source vector database available as a managed cloud service or self-hosted. Its headline feature is first-class hybrid search — combining dense vector search with BM25 keyword search in a single query. This matters for RAG: product names, proper nouns, and technical identifiers (like a ticket number or a function name) are retrieved better by keyword matching than by semantics alone.

import weaviate
from weaviate.classes.query import Filter, HybridFusion

client = weaviate.connect_to_weaviate_cloud(
    cluster_url="https://YOUR_CLUSTER.weaviate.network",
    auth_credentials=weaviate.auth.AuthApiKey("YOUR_API_KEY"),
)

collection = client.collections.use("Documents")

# Hybrid search: alpha=0.5 is an even blend of vector and BM25
response = collection.query.hybrid(
    query="semantic search embeddings",
    alpha=0.5,                            # 0 = BM25 only, 1 = vector only
    fusion_type=HybridFusion.RELATIVE_SCORE,
    filters=Filter.by_property("topic").equal("vector-db"),
    limit=5,
)

for obj in response.objects:
    print(obj.properties)

client.close()

The alpha parameter is the primary lever for tuning hybrid search. Start at 0.5 and adjust based on your evaluation results from Lesson 7.

NOTE
Weaviate hybrid search is covered in depth in Lesson 6
Advanced Retrieval (Lesson 6) walks through adding hybrid search and BM25 scoring to a complete RAG pipeline, including how to measure the improvement over pure vector search.

Qdrant — an honourable mention

Qdrant is an open-source vector database written in Rust, available as a managed cloud service or self-hosted. Its Python client API is clean and idiomatic. Qdrant calls metadata "payload" and offers sophisticated payload filtering with a built-in payload index — filtering is genuinely fast even on large collections because payload conditions can prune the HNSW graph during traversal rather than being applied after the fact.

from qdrant_client import QdrantClient
from qdrant_client.http.models import (
    Distance, VectorParams, Filter, FieldCondition, MatchValue
)

client = QdrantClient(url="http://localhost:6333")

client.create_collection(
    collection_name="docs",
    vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)

# Filtered nearest-neighbor search
results = client.search(
    collection_name="docs",
    query_vector=query_embedding,
    query_filter=Filter(
        must=[FieldCondition(key="topic", match=MatchValue(value="rag"))]
    ),
    limit=5,
)

6 Metadata Filtering — Patterns and Pitfalls

Metadata filtering narrows vector search to a subset of your corpus before (pre-filter) or after (post-filter) the ANN search. For RAG pipelines it is indispensable: you typically want to search only the documents relevant to a tenant, a date range, a document type, or an access-control group.

Pre-filter vs post-filter

Pre-filtering restricts the ANN search to only the matching rows. It guarantees you get exactly k results (if at least k match), but disrupts HNSW graph traversal because the algorithm can only follow edges that satisfy the filter. For very selective filters this is fine; for broad filters it becomes expensive.

Post-filtering performs the full ANN search and then discards rows that do not match the filter. It is fast (the index works unimpeded) but can return fewer than k results if many top-k candidates are filtered out.

Most modern vector databases use a hybrid approach: they attempt pre-filtering when the filter is highly selective (matching a small fraction of vectors) and fall back to post-filtering otherwise. As an application developer you typically declare the filter and trust the database to choose, but it helps to know the difference when debugging low recall.

Chroma: filtering with the where clause

# Match a single metadata field
results = collection.query(
    query_embeddings=[query_embedding],
    n_results=3,
    where={"topic": "rag"},                     # simple equality
)

# Combine conditions with $and
results = collection.query(
    query_embeddings=[query_embedding],
    n_results=3,
    where={
        "$and": [
            {"topic": {"$in": ["rag", "vector-db"]}},
            {"source": {"$ne": "lesson-04"}},
        ]
    },
)

pgvector: filtering is just SQL

With pgvector, metadata filtering is a standard SQL WHERE clause. Because PostgreSQL's query planner can combine the vector index scan with regular B-tree index scans on your metadata columns, the filtering story is more expressive than in purpose-built databases.

-- Add a B-tree index on the metadata column you filter most often
CREATE INDEX ON documents (topic);

-- Combined vector search + metadata filter
SELECT content, source,
       1 - (embedding <=> $1) AS cosine_similarity
FROM documents
WHERE topic = 'rag'                      -- metadata filter
ORDER BY embedding <=> $1               -- vector distance
LIMIT 5;

Pinecone: filter in the query call

Pinecone metadata filters use a MongoDB-style JSON syntax in the filter argument to index.query():

results = index.query(
    vector=query_embedding,
    top_k=5,
    include_metadata=True,
    filter={
        "$and": [
            {"topic": {"$in": ["rag", "embeddings"]}},
            {"published_year": {"$gte": 2024}},
        ]
    },
)

Common mistakes

Mistake Symptom Fix
Filtering on a high-cardinality field with no index Slow queries as corpus grows Add a payload/metadata index on that field
Overly selective filter with small corpus Fewer than k results returned Increase n_results / top_k and trim after retrieval
Storing user/tenant ID only in metadata, not in the vector space Cross-tenant leaks if filter is accidentally omitted Always include tenant_id in query filter; enforce at the middleware layer
Mixing types for the same key (e.g., year as "2024" in one doc and 2024 int in another) Filter silently misses records Enforce a schema in the ingestion pipeline
WARNING
Metadata is not a security boundary on its own
Metadata filters reduce what the model sees, but they are not an access control mechanism. For multi-tenant RAG, enforce tenant isolation at the API layer and use separate namespaces (Pinecone) or separate collections (Chroma/Qdrant) rather than relying solely on a metadata filter.
TIP
Index the fields you filter on
Most vector databases require you to explicitly create a payload or metadata index on fields you intend to filter. Without it, filtering may trigger a full scan of all metadata records, negating the benefit of the vector index entirely.

7 Selection Framework and Capacity Planning

With four credible options covered, here is a decision framework you can apply to real projects.

Decision tree

Does your stack already run PostgreSQL?
├─ YES → Use pgvector. You get vector search with no new
│        infrastructure and full SQL expressiveness.
│
└─ NO → Is hybrid search (vector + keyword) important?
    ├─ YES → Weaviate (self-hosted or cloud). First-class
    │        BM25 + vector fusion with the alpha parameter.
    │
    └─ NO → Is this local/dev-only or early-stage?
        ├─ YES → Chroma with PersistentClient. Zero-config,
        │        pure Python, swap to any other store later.
        │
        └─ NO → Do you want fully managed with no ops overhead?
            ├─ YES → Pinecone. Serverless, auto-scales,
            │        minimal configuration required.
            │
            └─ NO → Qdrant (self-hosted). Open source,
                     Rust-speed, fine-grained HNSW tuning
                     and efficient filtered search.

Comparison table

Chroma pgvector Pinecone Weaviate Qdrant
Hosting Local / Cloud Anywhere Postgres runs Fully managed Cloud or self-hosted Cloud or self-hosted
Hybrid search No No (vector only) No Yes (BM25 + vector) Yes (sparse + dense)
Index algorithm HNSW HNSW or IVFFlat Managed HNSW HNSW
Metadata filter Yes (where) Yes (SQL) Yes (JSON filter) Yes (GraphQL/Python) Yes (payload filter)
Best for Dev / prototyping Existing Postgres teams No-ops production Hybrid retrieval High-throughput self-hosted

Capacity planning rules of thumb

The memory footprint of an HNSW index depends on the number of vectors, their dimension, and the m parameter. A rough working formula:

memory_bytes ≈ num_vectors × dimension × 4 (float32) × 1.5 (index overhead)
Corpus size Dimension 384 (MiniLM) Dimension 1536 (OpenAI)
100 000 vectors ~230 MB ~920 MB
1 000 000 vectors ~2.3 GB ~9.2 GB
10 000 000 vectors ~23 GB ~92 GB

These are rough estimates; actual overhead varies by database and m value. Use them to size your instance class before you load data, not after.

Query latency targets: aim for under 50 ms for the vector search step on your target corpus. If you are seeing higher latency, the first levers are: reducing ef_search (HNSW) or nprobe (IVF) and checking whether metadata filtering is triggering a full scan.

NOTE
You can swap vector stores without touching the rest of the pipeline
In LlamaIndex and LangChain, vector stores are pluggable components with a common interface. If you start with Chroma locally and later move to Pinecone in production, you change a single constructor call — the embedding model, chunking, and LLM layers are unaffected. Design your pipeline around that abstraction from day one.

Questions & Answers

Q: We already have Elasticsearch in production. Can we just use that instead of a vector database?
Elasticsearch added dense vector support and approximate kNN search in the 8.x series, so yes — if you are already running it, you can store embeddings in a dense_vector field and query with the knn clause. The advantage is zero new infrastructure and native integration with your existing BM25 full-text search (which gives you hybrid search out of the box). The tradeoff is that Elasticsearch's kNN implementation uses HNSW but with fewer tuning knobs than dedicated vector databases, and at very large scale it consumes significantly more memory than purpose-built solutions. If your team already operates Elasticsearch competently, it is a legitimate starting point and an easy migration to a dedicated store later if needed.
Q: How do I handle embedding model upgrades? All my stored vectors become invalid if I switch models.
This is one of the most underestimated operational risks in RAG systems. When you switch embedding models (or when a provider updates a model's weights), your stored vectors become incompatible with new query vectors — cosine similarity between vectors from different models is meaningless. The standard approach is to store the model identifier as metadata alongside each vector, maintain a versioned index (e.g. separate Pinecone namespace or Chroma collection per model version), and run a background re-indexing job when you upgrade. During the migration window, route queries to both the old and new index and merge results. Never hot-swap the embedding model in production without a re-indexing plan. This is covered in more depth in Lesson 9: Production RAG.
Q: My recall metrics look good on the test set but users are complaining that the chatbot gives wrong answers. What is going wrong?
Retrieval recall and generation faithfulness are separate concerns. High recall means you are returning the right documents in the top k — but if the LLM ignores them, summarises them incorrectly, or the retrieved chunks do not contain the specific passage that answers the question, the answer is still wrong. Common culprits: chunks are too large (the relevant sentence is buried in noise), the retrieval is finding topically similar but not answer-specific documents, or the generation prompt does not strongly instruct the model to ground its answer in the provided context. Lesson 7 covers evaluation frameworks that measure faithfulness (does the answer match the source text?) separately from retrieval precision.
Q: Is it safe to put all my tenants' data in one Chroma collection separated only by a metadata filter?
Not for strict data isolation requirements. Metadata filters are a query-time convenience, not an access control boundary — a bug in your filter logic or a missing filter on one code path could expose one tenant's data to another. For production multi-tenant RAG where tenants are distinct organisations, use separate collections (Chroma, Qdrant) or separate namespaces (Pinecone) per tenant. The extra operational overhead is small — collections are lightweight — and the isolation is structural rather than depending on filter correctness. Reserve shared collections with metadata filters for scenarios where the tenants are internal teams or risk tolerance is high.
Q: The Chroma or pgvector index gives different top results than exact brute-force search on the same data. Is my index broken?
No — this is the expected behaviour of approximate nearest-neighbor search. ANN indexes trade a small amount of recall for large speed gains. The degree of difference depends on your index parameters: with HNSW at default settings (m=16, ef_construction=64, ef_search=64) you should see 95–99% recall compared to exact search, meaning roughly 1 in 50 results may differ. If you are seeing larger differences, increase ef_construction at build time and ef_search at query time. Also verify that the distance metric in the index (e.g. vector_cosine_ops) matches the metric you are comparing against in brute-force code.

Key Takeaways

  1. General-purpose databases cannot search embeddings at scale — B-tree indexes are designed for scalar comparisons, not high-dimensional geometry. Vector databases provide ANN indexes (HNSW, IVF) that deliver sub-millisecond queries at millions of vectors.

  2. HNSW is the right default — graph-based, supports incremental inserts, high recall at fast query times. IVF is an alternative when memory is tight and the corpus is mostly static. Neither index is needed below around 50 000 vectors.

  3. Chroma is the fastest path to a working local setup — three-line install, three client modes, no external services. Use it for development and swap to a production store later without changing the rest of the pipeline.

  4. pgvector earns its place if you already run PostgreSQL — the extension adds vector columns, HNSW and IVFFlat indexes, and three distance operators to standard SQL. Metadata filtering is just a WHERE clause, and every hosted Postgres service supports it.

  5. Metadata filtering is mandatory in production but has traps — always index the fields you filter on, do not rely on filters as a security boundary, and be aware of the pre-filter vs post-filter tradeoff when recall is unexpectedly low.

  6. Choose by operational fit, not benchmark scores — pgvector for Postgres shops, Chroma for fast prototypes, Weaviate when hybrid search matters, Pinecone when you want no infrastructure to manage. All four are legitimate production choices used by large-scale systems.

Next Steps: Lesson 4: Chunking Strategies