Vector DB Reference

Reference intermediate

Vector Store Comparison

Store Hosting Hybrid search Index Best fit
Pinecone Fully managed SaaS (serverless + pod) Dense + sparse + BM25 full-text on string fields Proprietary graph ANN Zero-ops; prototype-to-production
Weaviate Docker/K8s, managed cloud, embedded BM25 + vector; RelativeScoreFusion or RankedFusion HNSW, Flat, Dynamic (auto flat→HNSW at 10 k objects) Hybrid-search workloads; multi-tenant SaaS
Chroma Local in-process / persistent; Chroma Cloud Dense, sparse, hybrid, keyword + regex HNSW (dense); sparse inverted index Local dev, prototyping; Apache 2.0
Qdrant OSS, managed cloud, hybrid cloud, private cloud Dense + sparse via Query API; RRF built-in Filterable HNSW + sparse index Self-hosted price-performance; strict filtered search
pgvector Any Postgres host (RDS, Cloud SQL, Neon, Supabase…) Manual — <=> + tsvector; fusion DIY HNSW (recommended) or IVFFlat Postgres-first; need ACID / JOINs

Caveats

  • Pinecone serverless hybrid differs from pod-based; use separate dense/sparse indexes for predictable accuracy.
  • Weaviate Dynamic requires ASYNC_INDEXING=true (self-hosted); one-way flat→HNSW only.
  • pgvector hybrid is DIY: combine <=> with tsvector/tsquery.
  • Chroma suits single-machine workloads; purpose-built stores scale further.

Index-Type Reference: HNSW vs IVFFlat

Both solve ANN search — finding close vectors without a full scan.

Property HNSW IVFFlat
Structure Multi-layer proximity graph (sparse long-range upper layers, dense short-range bottom layer) Inverted file: vectors partitioned into lists centroid clusters; query scans probes clusters
Build time Slow — incremental graph insertion Fast — one-shot k-means training
Needs training data No — builds on empty table Yes — needs representative data upfront
Memory High — ~2–5× IVFFlat (graph edges per layer) Low — centroid metadata + raw vectors
Default recall High out-of-the-box Poor at probes=1; raise probes before benchmarking
Recall tuning Raise ef_search at query time, no rebuild Raise probes at query time, no rebuild
Incremental inserts Excellent Degrades; needs periodic reindex
Filtered search Graceful (Qdrant filterable HNSW weaves filter edges into graph) Can degrade — post-filter may miss target subset
Sweet spot General RAG, most production workloads Memory-constrained; large static datasets; batch indexing

HNSW parameters

Parameter pgvector Qdrant Weaviate Effect
Max connections m (16) m maxConnections Higher = better recall, larger index
Build candidate list ef_construction (64) ef_construct efConstruction Higher = better build recall, slower build
Search candidate list ef_search (40) ef ef Higher = better query recall, higher latency

Tune ef_search/ef first (no rebuild). Raise m/ef_construction only if willing to rebuild.

IVFFlat (pgvector)

Parameter Default Rule of thumb
lists 100 rows/1000 for ≤1 M rows; sqrt(rows) above
probes 1 Start at lists/10; raise until recall target met

Common mistake: judging IVFFlat at probes=1. Set probes to at least lists/10 before comparing to HNSW.


Distance Metrics

Metric pgvector op Notes
Cosine <=> Default for text — magnitude-insensitive; most models (sentence-transformers, OpenAI, Cohere)
Dot product <#> Equivalent to cosine on unit-norm vectors; use when model was trained with dot-product loss
L2 / Euclidean <-> Image embeddings; magnitude-sensitive
L1 / Manhattan <+> Rarely used for text

Code Patterns

pgvector — HNSW index and cosine search

CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (id BIGSERIAL PRIMARY KEY,
    content TEXT, embedding vector(1536));
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops)
    WITH (m = 16, ef_construction = 64);
cur.execute("SET hnsw.ef_search = 80")  # raise for better recall
cur.execute("SELECT id, content, 1-(embedding<=>%s) AS score "
    "FROM documents ORDER BY embedding<=>%s LIMIT 10",
    (query_vec, query_vec))

Qdrant — hybrid query (RRF)

from qdrant_client import QdrantClient
from qdrant_client.models import FusionQuery, Fusion

client = QdrantClient(url="http://localhost:6333")
results = client.query_points(
    collection_name="docs",
    query=FusionQuery(fusion=Fusion.RRF),
    prefetch=[
        {"query": {"name": "dense", "vector": dense_vec}, "limit": 20},
        {"query": {"name": "sparse", "vector": sparse_vec}, "limit": 20},
    ],
    limit=10,
)

Weaviate — BM25 + vector hybrid

import weaviate
client = weaviate.connect_to_local()
client.collections.get("Document").query.hybrid(
    query="retrieval augmented generation",
    alpha=0.6,   # 1.0=pure vector, 0.0=pure BM25
    limit=10,
)
client.close()

Chroma — local collection

import chromadb
col = chromadb.PersistentClient("/tmp/chroma").get_or_create_collection("docs")
col.add(ids=["d1"], documents=["RAG grounds LLMs in external data."])
results = col.query(query_texts=["how does retrieval work"], n_results=5)

Choosing a Vector Store

Already on Postgres?   YES -> pgvector (add a dedicated store only if you hit
                              scaling/recall limits you cannot tune away).
Need zero-ops SaaS?    YES -> Pinecone serverless.
Hybrid search first?   YES -> Weaviate (BM25+vector fusion; OSS + cloud).
Prototyping locally?   YES -> Chroma (in-process, no server, Apache 2.0).
Otherwise             -->   Qdrant (self-hosted; filterable HNSW; named
                              multi-vectors; best price-performance).

Related Pages