Advanced Retrieval

55 min intermediate Lesson 6

Learning Outcomes

  • Implement hybrid search by combining dense vector retrieval with BM25 sparse retrieval using Reciprocal Rank Fusion
  • Apply cross-encoder re-ranking to improve precision on a shortlist of candidates retrieved by a bi-encoder
  • Generate multiple query variants and use HyDE (Hypothetical Document Embeddings) to bridge the semantic gap between short queries and long documents
  • Build a parent-child retrieval pipeline that indexes fine-grained chunks but returns full parent sections to the LLM
  • Measure retrieval quality with precision@k, recall@k, and MRR after adding each technique to the pipeline

Lesson Plan

Segment Duration Topic
Intro 3 min Why basic similarity search fails and what we'll fix
Step 1–2 12 min Hybrid search: BM25 + dense + RRF fusion
Step 3 10 min Cross-encoder re-ranking
Step 4 8 min Query expansion and multi-query retrieval
Step 5 8 min HyDE — hypothetical document embeddings
Step 6–7 10 min Parent-child retrieval and pipeline assembly
Wrap-up 4 min Measuring quality at each stage, key takeaways

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ virtual environment
  • pip install rank-bm25 sentence-transformers langchain langchain-community chromadb openai
  • An OpenAI API key (or any LLM API) for HyDE query generation
  • The document corpus and Chroma collection from Lesson 5

1 Why Pure Similarity Search Fails — and the Hybrid Fix

Dense vector retrieval is excellent at semantic paraphrase — "car" and "automobile" land close together in embedding space. But it struggles with rare tokens, product codes, and exact-match queries where the text must contain a specific string. BM25 is the opposite: it excels at exact matches but knows nothing about meaning.

The vocabulary mismatch problem. A user searches for LDAP CVE-2023-1234. Your documentation uses the exact phrase. A dense retriever may confidently return paragraphs about "directory authentication" that never mention the CVE number. BM25 finds the document immediately.

Hybrid search runs both retrievers in parallel and merges results with Reciprocal Rank Fusion (RRF).

RRF assigns each document a score based on its rank position in each list, not its raw similarity score. This solves score normalization — BM25 and cosine similarity live on incompatible scales, but ranks are universal.

The formula for a single list:

RRF_score(doc) = sum over lists: 1 / (rank_in_list + k)

where k is a constant (typically 60) that dampens the outsized influence of top ranks. Documents ranked first in every list win; documents absent from a list contribute zero.

NOTE
Why Hybrid is the Minimum Viable Baseline
Leading vector databases (Weaviate, Qdrant, Pinecone) and frameworks (LangChain, LlamaIndex) all ship hybrid retrieval as a built-in feature. There is no good reason to deploy production RAG with pure dense search when hybrid adds negligible latency and consistently improves recall.
Query type BM25 recall Dense recall Hybrid recall
Exact product codes / error numbers High Low High
Conceptual / paraphrase queries Low High High
Mixed (code + description) Medium Medium Higher than either

2 Implementing Hybrid Search with BM25 and LangChain's EnsembleRetriever

We will use the rank-bm25 library for sparse retrieval and LangChain's EnsembleRetriever to fuse the two ranked lists with RRF.

Install dependencies:

pip install rank-bm25 langchain langchain-community chromadb

Build the hybrid retriever:

import re
from rank_bm25 import BM25Okapi
from langchain.retrievers import BM25Retriever, EnsembleRetriever
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings

# --- 1. Load your existing Chroma vector store from Lesson 5 ---
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma(
    persist_directory="./chroma_db",
    embedding_function=embeddings,
)

# --- 2. Load the same raw docs (needed for BM25) ---
# Replace with your actual document list from Lesson 5
from langchain_community.document_loaders import DirectoryLoader, TextLoader
loader = DirectoryLoader("./docs", glob="**/*.md", loader_cls=TextLoader)
documents = loader.load()

# --- 3. Build BM25 retriever from the same chunks ---
# BM25Retriever expects LangChain Document objects
bm25_retriever = BM25Retriever.from_documents(documents)
bm25_retriever.k = 10  # fetch top-10 candidates from BM25

# --- 4. Build dense retriever from Chroma ---
dense_retriever = vectorstore.as_retriever(search_kwargs={"k": 10})

# --- 5. Combine with EnsembleRetriever (uses RRF under the hood) ---
# weights=[0.5, 0.5] gives equal weight; tune based on your query mix
hybrid_retriever = EnsembleRetriever(
    retrievers=[bm25_retriever, dense_retriever],
    weights=[0.5, 0.5],
)

# --- 6. Test it ---
results = hybrid_retriever.invoke("LDAP authentication CVE-2023-1234")
for doc in results[:3]:
    print(doc.page_content[:200])
    print("---")
TIP
Tuning the Weights
Start with weights=[0.5, 0.5]. If your queries are mostly conceptual, shift toward [0.3, 0.7] (more dense). If you have lots of exact-match queries (codes, names, error IDs), shift toward [0.7, 0.3]. Measure with your golden test set from Lesson 7 before locking in a ratio.
WARNING
Keep BM25 and Vector Chunks in Sync
BM25Retriever and your vector store must be built from the same chunks at the same granularity. If you re-chunk your docs, rebuild both. Mismatched chunk sets will cause duplicates and missed results after fusion.

Measuring improvement. For a quick sanity check before you have a full evaluation harness, run your 10 hardest queries (the ones where your Lesson 5 pipeline returned wrong results) and count how many are now correct. A formal eval setup is covered in Lesson 7: Evaluation & Quality.


3 Cross-Encoder Re-Ranking — Trading Latency for Precision

After hybrid search, you have a candidate set of perhaps 15–20 documents. Many are plausible but not actually relevant. A cross-encoder re-ranker takes each (query, document) pair, passes both through a single model with full cross-attention, and outputs a calibrated relevance score.

Why cross-encoders are more accurate than bi-encoders. Bi-encoders embed the query and document independently, then measure distance. Cross-encoders see query and document together, so they can attend to exact query terms anywhere in the document. This makes them far more accurate — but too slow to run over millions of documents. The standard pattern is:

query → bi-encoder retrieval (top 50) → cross-encoder rerank (top 10) → LLM

Install sentence-transformers:

pip install sentence-transformers

Re-ranking pipeline:

import torch
from sentence_transformers import CrossEncoder

# cross-encoder/ms-marco-MiniLM-L6-v2: 74.30 NDCG@10, ~1800 docs/sec
# cross-encoder/ms-marco-MiniLM-L12-v2: 74.31 NDCG@10, ~960 docs/sec (marginally better, 2x slower)
# For production latency constraints, L6-v2 is the default choice
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")

def rerank(query: str, docs: list, top_k: int = 5) -> list:
    """
    Re-score and sort candidate documents by cross-encoder relevance.

    Args:
        query: The user's original query string.
        docs: LangChain Document objects (or any objects with .page_content).
        top_k: Number of top documents to return.

    Returns:
        List of (score, doc) tuples, sorted descending by score.
    """
    pairs = [(query, doc.page_content) for doc in docs]
    scores = model.predict(pairs)  # returns list of floats
    ranked = sorted(zip(scores, docs), key=lambda x: x[0], reverse=True)
    return ranked[:top_k]

# In your pipeline:
query = "How does the LDAP connector handle certificate validation?"
candidates = hybrid_retriever.invoke(query)  # ~20 docs from Step 2
top_docs = rerank(query, candidates, top_k=5)

context = "\n\n".join(doc.page_content for _, doc in top_docs)
NOTE
Latency Budget
A cross-encoder scoring 20 candidates on a laptop GPU (or fast CPU) takes roughly 100–300 ms. In a web application this is acceptable. If you need sub-50 ms end-to-end, consider running the reranker asynchronously while streaming the LLM response, or reducing the candidate pool to 10.

MS-MARCO model performance reference:

Model NDCG@10 Throughput
cross-encoder/ms-marco-TinyBERT-L2-v2 69.84 ~9000 docs/sec
cross-encoder/ms-marco-MiniLM-L6-v2 74.30 ~1800 docs/sec
cross-encoder/ms-marco-MiniLM-L12-v2 74.31 ~960 docs/sec

The L6 and L12 models are nearly identical in quality. Choose L6 unless you have a compelling reason to take the 2× latency hit.

WARNING
Domain Shift
MS-MARCO models are trained on web search pairs. They work well for general text but may under-perform on highly technical or domain-specific corpora (e.g., legal documents, medical literature, internal codebase docs). If you see poor re-ranking on your domain, fine-tune a cross-encoder on labeled pairs from your corpus.

4 Query Expansion — Covering What the User Didn't Say

Users write short, ambiguous queries. "connection timeout" could mean a database driver, a load balancer, a Kubernetes readiness probe, or a browser idle session. A single query embedding misses relevant documents that use different terminology.

Query expansion generates several alternative phrasings of the query, retrieves candidates for each, deduplicates, and merges before passing to the re-ranker.

LangChain's MultiQueryRetriever implements this pattern: it prompts an LLM to produce N reformulations (default 3), runs each against your retriever, and merges the unique results.

from langchain.retrievers.multi_query import MultiQueryRetriever
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)

# Wrap our hybrid_retriever with multi-query expansion
multi_query_retriever = MultiQueryRetriever.from_llm(
    retriever=hybrid_retriever,
    llm=llm,
)

# Optional: enable logging to inspect generated queries
import logging
logging.getLogger("langchain.retrievers.multi_query").setLevel(logging.INFO)

results = multi_query_retriever.invoke(
    "connection timeout handling in the data pipeline"
)
print(f"Retrieved {len(results)} unique documents across all query variants")

When you run this, the logger will show the three auto-generated alternatives, for example:

Generated queries:
 - "data pipeline socket timeout configuration"
 - "how to handle connection errors in ETL workflows"
 - "retry logic for database connection failures"

Each variant retrieves candidates independently, and the union is deduplicated by document ID before passing downstream.

TIP
Custom Prompt for Domain Terminology
The default MultiQueryRetriever prompt is domain-agnostic. For technical corpora, override the prompt to instruct the LLM to use domain-specific synonyms: PROMPT = PromptTemplate(input_variables=['question'], template='You are an expert in cloud infrastructure. Generate 3 alternative phrasings of: {question}'). Pass it as prompt=PROMPT to from_llm.
WARNING
Latency Multiplication
Each query variant adds a full retrieval round-trip. With 3 variants, expect 3× the latency of a single retrieval pass. Cache query expansions keyed on the original query string if your query distribution repeats (which it usually does in production).

5 HyDE — Searching with the Answer Instead of the Question

HyDE (Hypothetical Document Embeddings) is a query transformation technique introduced in a 2022 paper by Gao et al. The insight: a short question and a long answer live in different regions of embedding space. A question like "What is RAG?" embeds near other questions; the relevant documentation embeds near other documentation. HyDE closes this gap by:

  1. Prompting an LLM to write a short hypothetical answer to the query (without access to the knowledge base)
  2. Embedding the hypothetical answer (not the query)
  3. Searching for documents similar to that hypothetical answer embedding

The hypothetical answer is never shown to the user — it is only used as a search vector. The LLM generates it based on its parametric knowledge, which may be incomplete or wrong, but that does not matter: you need the shape of an answer, not a correct answer.

from openai import OpenAI
from langchain_openai import OpenAIEmbeddings

openai_client = OpenAI()
embeddings_model = OpenAIEmbeddings(model="text-embedding-3-small")

def hyde_retrieve(query: str, vectorstore, top_k: int = 10) -> list:
    """
    Retrieve documents using Hypothetical Document Embeddings (HyDE).

    1. Generate a hypothetical answer from the LLM.
    2. Embed the hypothetical answer.
    3. Search the vector store with that embedding.
    """
    # Step 1: Generate a plausible hypothetical answer
    response = openai_client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {
                "role": "system",
                "content": (
                    "Write a concise, factual paragraph that directly answers "
                    "the question below. This is a hypothetical passage from "
                    "technical documentation. Be specific and use domain terminology."
                ),
            },
            {"role": "user", "content": query},
        ],
        temperature=0.0,
        max_tokens=200,
    )
    hypothetical_doc = response.choices[0].message.content

    # Step 2: Embed the hypothetical document
    hypo_embedding = embeddings_model.embed_query(hypothetical_doc)

    # Step 3: Search using the hypothetical embedding directly
    # Chroma supports passing a pre-computed embedding vector
    results = vectorstore._collection.query(
        query_embeddings=[hypo_embedding],
        n_results=top_k,
    )
    return results

# Example
query = "How does HNSW indexing affect recall at high load?"
docs = hyde_retrieve(query, vectorstore, top_k=10)

When HyDE helps most. HyDE is particularly effective when:

  • Queries are short but documents are long and detailed
  • The domain has precise terminology that users rarely use in their queries
  • You are doing zero-shot retrieval on a new domain

When HyDE is less useful. If your queries are already long and specific (e.g., users paste error messages), HyDE adds LLM latency for no gain. Use query expansion instead.

NOTE
Averaging Multiple Hypothetical Documents
For higher accuracy, generate 3-5 hypothetical documents and average their embeddings before searching. This reduces variance from any single hallucinated answer. The Haystack framework includes a HypotheticalDocumentEmbedder component that handles this automatically.

6 Parent-Child Retrieval — Small Chunks for Search, Large Chunks for Context

Fine-grained chunks (128–256 tokens) are excellent for embedding precision — short passages have focused meaning, and their embeddings are accurate. But returning 128-token snippets to the LLM robs it of context: the answer might require the surrounding section.

Parent-child retrieval stores two levels of chunks:

  • Child nodes (small, e.g. 128 tokens) — indexed in the vector store, used for search
  • Parent nodes (large, e.g. 512 tokens or full section) — stored in a document store, returned to the LLM

At query time: embed the query, find the most similar child chunks, look up their parent IDs, return the parent chunks to the LLM. The LLM gets richer context; the retrieval stays precise.

LlamaIndex's AutoMergingRetriever implements this with automatic promotion: if enough child nodes from the same parent are retrieved, the system replaces them with the parent.

from llama_index.core import VectorStoreIndex, StorageContext
from llama_index.core.node_parser import HierarchicalNodeParser, get_leaf_nodes
from llama_index.core.storage.docstore import SimpleDocumentStore
from llama_index.core.retrievers import AutoMergingRetriever
from llama_index.core.query_engine import RetrieverQueryEngine
from llama_index.core import SimpleDirectoryReader

# --- 1. Load documents ---
docs = SimpleDirectoryReader("./docs").load_data()

# --- 2. Build hierarchical chunks ---
# Default: chunk_sizes=[2048, 512, 128]
# 2048 = grandparent, 512 = parent, 128 = child (leaf)
node_parser = HierarchicalNodeParser.from_defaults(
    chunk_sizes=[512, 128]
)
nodes = node_parser.get_nodes_from_documents(docs)
leaf_nodes = get_leaf_nodes(nodes)

# --- 3. Store ALL nodes in docstore (parent nodes must be accessible) ---
docstore = SimpleDocumentStore()
docstore.add_documents(nodes)
storage_context = StorageContext.from_defaults(docstore=docstore)

# --- 4. Index ONLY leaf nodes in vector store ---
# Only fine-grained chunks are searchable; parents are in docstore only
base_index = VectorStoreIndex(
    leaf_nodes,
    storage_context=storage_context,
)
base_retriever = base_index.as_retriever(similarity_top_k=12)

# --- 5. Wrap with AutoMergingRetriever ---
# If >= 50% of a parent's children are retrieved, it promotes to the parent
retriever = AutoMergingRetriever(
    base_retriever,
    storage_context,
    verbose=True,
)

# --- 6. Build a query engine ---
query_engine = RetrieverQueryEngine.from_args(retriever)
response = query_engine.query(
    "Explain the certificate pinning configuration options"
)
print(response)
TIP
Tuning Chunk Sizes
The default [2048, 512, 128] hierarchy works well for long-form technical docs (API references, runbooks). For shorter content like FAQ entries or news articles, try [512, 128]. The rule of thumb: child chunks should be short enough that their embedding is unambiguous; parent chunks should be long enough to fully answer a question.
WARNING
LangChain's ParentDocumentRetriever is a Simpler Alternative
LangChain ships a ParentDocumentRetriever that implements the same idea with a flat two-level hierarchy. It is easier to configure if you are already using LangChain's abstractions, but does not support the auto-merging promotion that LlamaIndex provides.

Comparing approaches side by side:

Technique Retrieves Returns to LLM Key benefit
Standard chunking 512-token chunks Same chunks Simplicity
Parent-child 128-token leaf chunks 512-token parent chunks Better context for generation
Auto-merging 128-token leaf chunks Parent if majority matched Adaptive context window

7 Assembling the Full Advanced Retrieval Pipeline

Now that you have five techniques individually, the question is: which ones to stack, and in what order? The answer depends on your latency budget and the failure modes you observed in your baseline pipeline.

Recommended order of adoption (by impact-to-effort ratio):

  1. Hybrid search — always on; adds ~5 ms
  2. Cross-encoder re-ranking — always on for precision-critical use cases; adds 100–300 ms
  3. Query expansion — add when queries are ambiguous or short; adds 1 full LLM call
  4. Parent-child retrieval — add when context truncation is a problem (LLM gives partial answers)
  5. HyDE — add when there is a large semantic gap between query style and document style

Here is a complete pipeline that stacks hybrid search, query expansion, and cross-encoder re-ranking:

from langchain.retrievers import BM25Retriever, EnsembleRetriever
from langchain.retrievers.multi_query import MultiQueryRetriever
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_community.vectorstores import Chroma
from sentence_transformers import CrossEncoder
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough

# ---------------------------------------------------------------
# Component setup (replace with your actual values from Lesson 5)
# ---------------------------------------------------------------
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma(persist_directory="./chroma_db", embedding_function=embeddings)
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")

# ---------------------------------------------------------------
# Layer 1: Hybrid retriever (BM25 + dense vector)
# ---------------------------------------------------------------
bm25_retriever = BM25Retriever.from_documents(raw_documents)
bm25_retriever.k = 15
dense_retriever = vectorstore.as_retriever(search_kwargs={"k": 15})
hybrid_retriever = EnsembleRetriever(
    retrievers=[bm25_retriever, dense_retriever],
    weights=[0.5, 0.5],
)

# ---------------------------------------------------------------
# Layer 2: Query expansion wrapping the hybrid retriever
# ---------------------------------------------------------------
multi_query_retriever = MultiQueryRetriever.from_llm(
    retriever=hybrid_retriever,
    llm=llm,
)

# ---------------------------------------------------------------
# Layer 3: Cross-encoder re-ranking function
# ---------------------------------------------------------------
def retrieve_and_rerank(query: str, top_k: int = 5) -> list:
    candidates = multi_query_retriever.invoke(query)
    if not candidates:
        return []
    pairs = [(query, doc.page_content) for doc in candidates]
    scores = reranker.predict(pairs)
    ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
    return [doc for _, doc in ranked[:top_k]]

# ---------------------------------------------------------------
# Layer 4: Generation with retrieved context
# ---------------------------------------------------------------
prompt = ChatPromptTemplate.from_messages([
    (
        "system",
        "You are a technical assistant. Answer the question using ONLY "
        "the provided context. If the context does not contain enough "
        "information, say so explicitly.\n\nContext:\n{context}",
    ),
    ("human", "{question}"),
])

def format_docs(docs):
    return "\n\n---\n\n".join(doc.page_content for doc in docs)

def rag_chain(query: str) -> str:
    docs = retrieve_and_rerank(query)
    context = format_docs(docs)
    messages = prompt.format_messages(context=context, question=query)
    return llm.invoke(messages).content

# Usage
answer = rag_chain("How does the LDAP connector handle certificate validation errors?")
print(answer)

Measuring improvement at each step. Before you move to full evaluation (Lesson 7), a quick precision measurement on a small golden set tells you whether each layer is helping:

# Minimal precision@5 measurement
golden_set = [
    {
        "query": "LDAP certificate validation errors",
        "relevant_doc_ids": ["doc_42", "doc_107"],
    },
    # ... add 10-20 queries
]

def precision_at_k(retrieved_ids, relevant_ids, k=5):
    top_k = set(retrieved_ids[:k])
    relevant = set(relevant_ids)
    return len(top_k & relevant) / k

# Evaluate baseline vs advanced pipeline
for item in golden_set:
    baseline_ids = [d.metadata.get("doc_id") for d in dense_retriever.invoke(item["query"])[:5]]
    advanced_ids = [d.metadata.get("doc_id") for d in retrieve_and_rerank(item["query"])]
    baseline_p = precision_at_k(baseline_ids, item["relevant_doc_ids"])
    advanced_p = precision_at_k(advanced_ids, item["relevant_doc_ids"])
    print(f"Query: {item['query'][:50]}")
    print(f"  Baseline P@5: {baseline_p:.2f} | Advanced P@5: {advanced_p:.2f}")
NOTE
Incremental Adoption Pays Off
Stack one technique at a time and measure after each addition. In practice, hybrid search alone often gives you 15–20% precision improvement. Re-ranking adds another 10–15%. Query expansion helps most on ambiguous short queries. Adding all three without measuring means you cannot diagnose regressions.

Questions & Answers

Q: My BM25 retriever keeps surfacing documents that contain the query words but are totally off-topic. Is hybrid search making things worse?
BM25 is doing its job — it is finding exact term matches. The issue is that these off-topic matches are dragging up through RRF because they score highly in the BM25 list. Two fixes: (1) apply metadata filtering before the BM25 pass to restrict to relevant document categories; (2) add a cross-encoder re-ranking step (Step 3) which will demote irrelevant results even if they matched on surface keywords. Also check whether your BM25 tokenization is too permissive — stop-word removal and stemming help.
Q: The cross-encoder re-ranker is adding 400 ms to every query. What are my options to reduce latency?
Several levers: (1) Reduce the candidate pool from 20 to 10 before re-ranking — this roughly halves the latency. (2) Switch from ms-marco-MiniLM-L12-v2 to ms-marco-MiniLM-L6-v2 — nearly identical quality, 2× faster. (3) Run the re-ranker on GPU (even a modest one). (4) Cache re-ranking results keyed on (query hash, doc ID set hash) — repeated queries are free. (5) Run the re-ranker asynchronously while streaming the LLM response so the user perceives less latency.
Q: HyDE is generating plausible-sounding but factually wrong hypothetical answers. Does that corrupt my retrieval?
No — and this is the key insight behind HyDE. Factual correctness of the hypothetical answer does not matter. What matters is that its embedding lands in the same region of vector space as real documents that answer the question. A hallucinated but stylistically correct answer ("The HNSW index uses a greedy beam search with ef_search=200...") will embed near the real documentation even if the numbers are wrong. Where HyDE fails is when the LLM generates a hypothetical answer in the wrong domain entirely — guard against this with a well-crafted system prompt that anchors the LLM to your domain.
Q: Should I use LlamaIndex's AutoMergingRetriever or LangChain's ParentDocumentRetriever? They seem to do the same thing.
They are architecturally similar but differ in control. LangChain's ParentDocumentRetriever always returns the parent chunk — it is a flat two-level hierarchy. LlamaIndex's AutoMergingRetriever supports deeper hierarchies (three levels) and has the auto-promotion logic: if a majority of a parent's children are retrieved, it replaces them with the parent. For simple cases, LangChain's version is easier to set up. For complex, long-form documentation where you want adaptive context depth, LlamaIndex's version is more capable. If you covered tool use in the Agentic AI course, note that both integrations work well as tool-enabled retrievers in agent pipelines — see the Agentic AI course for that pattern.
Q: I have a fully managed vector database (Pinecone, Qdrant Cloud) that does hybrid search natively. Do I still need rank-bm25 and EnsembleRetriever?
No. If your vector database handles hybrid search at the index level — which Weaviate, Qdrant, and Pinecone all support natively — use their built-in hybrid query API. It is faster (no Python-side fusion), handles sparse index maintenance automatically, and often uses more sophisticated sparse encodings than plain BM25 (e.g., SPLADE). The rank-bm25 + EnsembleRetriever approach is the right choice when you are using Chroma, pgvector, or any store that does not support hybrid queries natively, or when you want full control over the fusion logic in Python.

Key Takeaways

  1. Hybrid search is the minimum viable retrieval baseline — combining BM25 sparse retrieval with dense vectors via RRF costs almost nothing in latency and reliably improves recall for exact-match and semantic queries alike.

  2. Cross-encoder re-ranking raises precision at the cost of latency — scoring 20 candidates with cross-encoder/ms-marco-MiniLM-L6-v2 takes ~100–300 ms and consistently removes the off-topic documents that dense retrieval includes.

  3. Query expansion catches what the user did not say — LangChain's MultiQueryRetriever generates alternative phrasings automatically; apply it when users write short, ambiguous queries and you observe low recall on diverse terminology.

  4. HyDE closes the query-to-document semantic gap — embedding a hypothetical answer rather than the raw query improves retrieval when the linguistic style of queries differs significantly from the style of your documents; factual accuracy of the hypothetical is irrelevant.

  5. Parent-child retrieval gives the LLM context without sacrificing index precision — index small leaf chunks (128 tokens) for accurate retrieval, return parent sections (512 tokens) for rich generation context; LlamaIndex's AutoMergingRetriever automates the promotion logic.

  6. Add techniques one at a time and measure after each — measure precision@k, recall@k, or MRR on a small golden test set before stacking the next technique; this is the only reliable way to isolate what is helping and diagnose regressions.

Next Steps: Lesson 7: Evaluation & Quality