Vector DB Reference
Reference
intermediate
Vector Store Comparison
| Store | Hosting | Hybrid search | Index | Best fit |
|---|---|---|---|---|
| Pinecone | Fully managed SaaS (serverless + pod) | Dense + sparse + BM25 full-text on string fields | Proprietary graph ANN | Zero-ops; prototype-to-production |
| Weaviate | Docker/K8s, managed cloud, embedded | BM25 + vector; RelativeScoreFusion or RankedFusion | HNSW, Flat, Dynamic (auto flat→HNSW at 10 k objects) | Hybrid-search workloads; multi-tenant SaaS |
| Chroma | Local in-process / persistent; Chroma Cloud | Dense, sparse, hybrid, keyword + regex | HNSW (dense); sparse inverted index | Local dev, prototyping; Apache 2.0 |
| Qdrant | OSS, managed cloud, hybrid cloud, private cloud | Dense + sparse via Query API; RRF built-in | Filterable HNSW + sparse index | Self-hosted price-performance; strict filtered search |
| pgvector | Any Postgres host (RDS, Cloud SQL, Neon, Supabase…) | Manual — <=> + tsvector; fusion DIY |
HNSW (recommended) or IVFFlat | Postgres-first; need ACID / JOINs |
Caveats
- Pinecone serverless hybrid differs from pod-based; use separate dense/sparse indexes for predictable accuracy.
- Weaviate Dynamic requires
ASYNC_INDEXING=true(self-hosted); one-way flat→HNSW only. - pgvector hybrid is DIY: combine
<=>withtsvector/tsquery. - Chroma suits single-machine workloads; purpose-built stores scale further.
Index-Type Reference: HNSW vs IVFFlat
Both solve ANN search — finding close vectors without a full scan.
| Property | HNSW | IVFFlat |
|---|---|---|
| Structure | Multi-layer proximity graph (sparse long-range upper layers, dense short-range bottom layer) | Inverted file: vectors partitioned into lists centroid clusters; query scans probes clusters |
| Build time | Slow — incremental graph insertion | Fast — one-shot k-means training |
| Needs training data | No — builds on empty table | Yes — needs representative data upfront |
| Memory | High — ~2–5× IVFFlat (graph edges per layer) | Low — centroid metadata + raw vectors |
| Default recall | High out-of-the-box | Poor at probes=1; raise probes before benchmarking |
| Recall tuning | Raise ef_search at query time, no rebuild |
Raise probes at query time, no rebuild |
| Incremental inserts | Excellent | Degrades; needs periodic reindex |
| Filtered search | Graceful (Qdrant filterable HNSW weaves filter edges into graph) | Can degrade — post-filter may miss target subset |
| Sweet spot | General RAG, most production workloads | Memory-constrained; large static datasets; batch indexing |
HNSW parameters
| Parameter | pgvector | Qdrant | Weaviate | Effect |
|---|---|---|---|---|
| Max connections | m (16) |
m |
maxConnections |
Higher = better recall, larger index |
| Build candidate list | ef_construction (64) |
ef_construct |
efConstruction |
Higher = better build recall, slower build |
| Search candidate list | ef_search (40) |
ef |
ef |
Higher = better query recall, higher latency |
Tune ef_search/ef first (no rebuild). Raise m/ef_construction only if willing to rebuild.
IVFFlat (pgvector)
| Parameter | Default | Rule of thumb |
|---|---|---|
lists |
100 | rows/1000 for ≤1 M rows; sqrt(rows) above |
probes |
1 | Start at lists/10; raise until recall target met |
Common mistake: judging IVFFlat at probes=1. Set probes to at least lists/10 before comparing to HNSW.
Distance Metrics
| Metric | pgvector op | Notes |
|---|---|---|
| Cosine | <=> |
Default for text — magnitude-insensitive; most models (sentence-transformers, OpenAI, Cohere) |
| Dot product | <#> |
Equivalent to cosine on unit-norm vectors; use when model was trained with dot-product loss |
| L2 / Euclidean | <-> |
Image embeddings; magnitude-sensitive |
| L1 / Manhattan | <+> |
Rarely used for text |
Code Patterns
pgvector — HNSW index and cosine search
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (id BIGSERIAL PRIMARY KEY,
content TEXT, embedding vector(1536));
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
cur.execute("SET hnsw.ef_search = 80") # raise for better recall
cur.execute("SELECT id, content, 1-(embedding<=>%s) AS score "
"FROM documents ORDER BY embedding<=>%s LIMIT 10",
(query_vec, query_vec))
Qdrant — hybrid query (RRF)
from qdrant_client import QdrantClient
from qdrant_client.models import FusionQuery, Fusion
client = QdrantClient(url="http://localhost:6333")
results = client.query_points(
collection_name="docs",
query=FusionQuery(fusion=Fusion.RRF),
prefetch=[
{"query": {"name": "dense", "vector": dense_vec}, "limit": 20},
{"query": {"name": "sparse", "vector": sparse_vec}, "limit": 20},
],
limit=10,
)
Weaviate — BM25 + vector hybrid
import weaviate
client = weaviate.connect_to_local()
client.collections.get("Document").query.hybrid(
query="retrieval augmented generation",
alpha=0.6, # 1.0=pure vector, 0.0=pure BM25
limit=10,
)
client.close()
Chroma — local collection
import chromadb
col = chromadb.PersistentClient("/tmp/chroma").get_or_create_collection("docs")
col.add(ids=["d1"], documents=["RAG grounds LLMs in external data."])
results = col.query(query_texts=["how does retrieval work"], n_results=5)
Choosing a Vector Store
Already on Postgres? YES -> pgvector (add a dedicated store only if you hit
scaling/recall limits you cannot tune away).
Need zero-ops SaaS? YES -> Pinecone serverless.
Hybrid search first? YES -> Weaviate (BM25+vector fusion; OSS + cloud).
Prototyping locally? YES -> Chroma (in-process, no server, Apache 2.0).
Otherwise --> Qdrant (self-hosted; filterable HNSW; named
multi-vectors; best price-performance).