Advanced Retrieval
Learning Outcomes
- Implement hybrid search by combining dense vector retrieval with BM25 sparse retrieval using Reciprocal Rank Fusion
- Apply cross-encoder re-ranking to improve precision on a shortlist of candidates retrieved by a bi-encoder
- Generate multiple query variants and use HyDE (Hypothetical Document Embeddings) to bridge the semantic gap between short queries and long documents
- Build a parent-child retrieval pipeline that indexes fine-grained chunks but returns full parent sections to the LLM
- Measure retrieval quality with precision@k, recall@k, and MRR after adding each technique to the pipeline
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why basic similarity search fails and what we'll fix |
| Step 1–2 | 12 min | Hybrid search: BM25 + dense + RRF fusion |
| Step 3 | 10 min | Cross-encoder re-ranking |
| Step 4 | 8 min | Query expansion and multi-query retrieval |
| Step 5 | 8 min | HyDE — hypothetical document embeddings |
| Step 6–7 | 10 min | Parent-child retrieval and pipeline assembly |
| Wrap-up | 4 min | Measuring quality at each stage, key takeaways |
Before You Begin
Pre-work:
- Complete Lesson 5: Building a Basic RAG Pipeline — this lesson extends that pipeline directly
- Read Lesson 4: Chunking Strategies for context on chunk size trade-offs
- Understand embedding similarity and vector stores from Lesson 2: Embeddings Explained and Lesson 3: Vector Databases
Shopping List:
- Python 3.10+ virtual environment
pip install rank-bm25 sentence-transformers langchain langchain-community chromadb openai- An OpenAI API key (or any LLM API) for HyDE query generation
- The document corpus and Chroma collection from Lesson 5
Dense vector retrieval is excellent at semantic paraphrase — "car" and "automobile" land close together in embedding space. But it struggles with rare tokens, product codes, and exact-match queries where the text must contain a specific string. BM25 is the opposite: it excels at exact matches but knows nothing about meaning.
The vocabulary mismatch problem. A user searches for LDAP CVE-2023-1234. Your documentation uses the exact phrase. A dense retriever may confidently return paragraphs about "directory authentication" that never mention the CVE number. BM25 finds the document immediately.
Hybrid search runs both retrievers in parallel and merges results with Reciprocal Rank Fusion (RRF).
RRF assigns each document a score based on its rank position in each list, not its raw similarity score. This solves score normalization — BM25 and cosine similarity live on incompatible scales, but ranks are universal.
The formula for a single list:
RRF_score(doc) = sum over lists: 1 / (rank_in_list + k)
where k is a constant (typically 60) that dampens the outsized influence of top ranks. Documents ranked first in every list win; documents absent from a list contribute zero.
| Query type | BM25 recall | Dense recall | Hybrid recall |
|---|---|---|---|
| Exact product codes / error numbers | High | Low | High |
| Conceptual / paraphrase queries | Low | High | High |
| Mixed (code + description) | Medium | Medium | Higher than either |
We will use the rank-bm25 library for sparse retrieval and LangChain's EnsembleRetriever to fuse the two ranked lists with RRF.
Install dependencies:
pip install rank-bm25 langchain langchain-community chromadb
Build the hybrid retriever:
import re
from rank_bm25 import BM25Okapi
from langchain.retrievers import BM25Retriever, EnsembleRetriever
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings
# --- 1. Load your existing Chroma vector store from Lesson 5 ---
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma(
persist_directory="./chroma_db",
embedding_function=embeddings,
)
# --- 2. Load the same raw docs (needed for BM25) ---
# Replace with your actual document list from Lesson 5
from langchain_community.document_loaders import DirectoryLoader, TextLoader
loader = DirectoryLoader("./docs", glob="**/*.md", loader_cls=TextLoader)
documents = loader.load()
# --- 3. Build BM25 retriever from the same chunks ---
# BM25Retriever expects LangChain Document objects
bm25_retriever = BM25Retriever.from_documents(documents)
bm25_retriever.k = 10 # fetch top-10 candidates from BM25
# --- 4. Build dense retriever from Chroma ---
dense_retriever = vectorstore.as_retriever(search_kwargs={"k": 10})
# --- 5. Combine with EnsembleRetriever (uses RRF under the hood) ---
# weights=[0.5, 0.5] gives equal weight; tune based on your query mix
hybrid_retriever = EnsembleRetriever(
retrievers=[bm25_retriever, dense_retriever],
weights=[0.5, 0.5],
)
# --- 6. Test it ---
results = hybrid_retriever.invoke("LDAP authentication CVE-2023-1234")
for doc in results[:3]:
print(doc.page_content[:200])
print("---")
weights=[0.5, 0.5]. If your queries are mostly conceptual, shift toward [0.3, 0.7] (more dense). If you have lots of exact-match queries (codes, names, error IDs), shift toward [0.7, 0.3]. Measure with your golden test set from Lesson 7 before locking in a ratio.Measuring improvement. For a quick sanity check before you have a full evaluation harness, run your 10 hardest queries (the ones where your Lesson 5 pipeline returned wrong results) and count how many are now correct. A formal eval setup is covered in Lesson 7: Evaluation & Quality.
After hybrid search, you have a candidate set of perhaps 15–20 documents. Many are plausible but not actually relevant. A cross-encoder re-ranker takes each (query, document) pair, passes both through a single model with full cross-attention, and outputs a calibrated relevance score.
Why cross-encoders are more accurate than bi-encoders. Bi-encoders embed the query and document independently, then measure distance. Cross-encoders see query and document together, so they can attend to exact query terms anywhere in the document. This makes them far more accurate — but too slow to run over millions of documents. The standard pattern is:
query → bi-encoder retrieval (top 50) → cross-encoder rerank (top 10) → LLM
Install sentence-transformers:
pip install sentence-transformers
Re-ranking pipeline:
import torch
from sentence_transformers import CrossEncoder
# cross-encoder/ms-marco-MiniLM-L6-v2: 74.30 NDCG@10, ~1800 docs/sec
# cross-encoder/ms-marco-MiniLM-L12-v2: 74.31 NDCG@10, ~960 docs/sec (marginally better, 2x slower)
# For production latency constraints, L6-v2 is the default choice
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
def rerank(query: str, docs: list, top_k: int = 5) -> list:
"""
Re-score and sort candidate documents by cross-encoder relevance.
Args:
query: The user's original query string.
docs: LangChain Document objects (or any objects with .page_content).
top_k: Number of top documents to return.
Returns:
List of (score, doc) tuples, sorted descending by score.
"""
pairs = [(query, doc.page_content) for doc in docs]
scores = model.predict(pairs) # returns list of floats
ranked = sorted(zip(scores, docs), key=lambda x: x[0], reverse=True)
return ranked[:top_k]
# In your pipeline:
query = "How does the LDAP connector handle certificate validation?"
candidates = hybrid_retriever.invoke(query) # ~20 docs from Step 2
top_docs = rerank(query, candidates, top_k=5)
context = "\n\n".join(doc.page_content for _, doc in top_docs)
MS-MARCO model performance reference:
| Model | NDCG@10 | Throughput |
|---|---|---|
| cross-encoder/ms-marco-TinyBERT-L2-v2 | 69.84 | ~9000 docs/sec |
| cross-encoder/ms-marco-MiniLM-L6-v2 | 74.30 | ~1800 docs/sec |
| cross-encoder/ms-marco-MiniLM-L12-v2 | 74.31 | ~960 docs/sec |
The L6 and L12 models are nearly identical in quality. Choose L6 unless you have a compelling reason to take the 2× latency hit.
Users write short, ambiguous queries. "connection timeout" could mean a database driver, a load balancer, a Kubernetes readiness probe, or a browser idle session. A single query embedding misses relevant documents that use different terminology.
Query expansion generates several alternative phrasings of the query, retrieves candidates for each, deduplicates, and merges before passing to the re-ranker.
LangChain's MultiQueryRetriever implements this pattern: it prompts an LLM to produce N reformulations (default 3), runs each against your retriever, and merges the unique results.
from langchain.retrievers.multi_query import MultiQueryRetriever
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
# Wrap our hybrid_retriever with multi-query expansion
multi_query_retriever = MultiQueryRetriever.from_llm(
retriever=hybrid_retriever,
llm=llm,
)
# Optional: enable logging to inspect generated queries
import logging
logging.getLogger("langchain.retrievers.multi_query").setLevel(logging.INFO)
results = multi_query_retriever.invoke(
"connection timeout handling in the data pipeline"
)
print(f"Retrieved {len(results)} unique documents across all query variants")
When you run this, the logger will show the three auto-generated alternatives, for example:
Generated queries:
- "data pipeline socket timeout configuration"
- "how to handle connection errors in ETL workflows"
- "retry logic for database connection failures"
Each variant retrieves candidates independently, and the union is deduplicated by document ID before passing downstream.
PROMPT = PromptTemplate(input_variables=['question'], template='You are an expert in cloud infrastructure. Generate 3 alternative phrasings of: {question}'). Pass it as prompt=PROMPT to from_llm.HyDE (Hypothetical Document Embeddings) is a query transformation technique introduced in a 2022 paper by Gao et al. The insight: a short question and a long answer live in different regions of embedding space. A question like "What is RAG?" embeds near other questions; the relevant documentation embeds near other documentation. HyDE closes this gap by:
- Prompting an LLM to write a short hypothetical answer to the query (without access to the knowledge base)
- Embedding the hypothetical answer (not the query)
- Searching for documents similar to that hypothetical answer embedding
The hypothetical answer is never shown to the user — it is only used as a search vector. The LLM generates it based on its parametric knowledge, which may be incomplete or wrong, but that does not matter: you need the shape of an answer, not a correct answer.
from openai import OpenAI
from langchain_openai import OpenAIEmbeddings
openai_client = OpenAI()
embeddings_model = OpenAIEmbeddings(model="text-embedding-3-small")
def hyde_retrieve(query: str, vectorstore, top_k: int = 10) -> list:
"""
Retrieve documents using Hypothetical Document Embeddings (HyDE).
1. Generate a hypothetical answer from the LLM.
2. Embed the hypothetical answer.
3. Search the vector store with that embedding.
"""
# Step 1: Generate a plausible hypothetical answer
response = openai_client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "system",
"content": (
"Write a concise, factual paragraph that directly answers "
"the question below. This is a hypothetical passage from "
"technical documentation. Be specific and use domain terminology."
),
},
{"role": "user", "content": query},
],
temperature=0.0,
max_tokens=200,
)
hypothetical_doc = response.choices[0].message.content
# Step 2: Embed the hypothetical document
hypo_embedding = embeddings_model.embed_query(hypothetical_doc)
# Step 3: Search using the hypothetical embedding directly
# Chroma supports passing a pre-computed embedding vector
results = vectorstore._collection.query(
query_embeddings=[hypo_embedding],
n_results=top_k,
)
return results
# Example
query = "How does HNSW indexing affect recall at high load?"
docs = hyde_retrieve(query, vectorstore, top_k=10)
When HyDE helps most. HyDE is particularly effective when:
- Queries are short but documents are long and detailed
- The domain has precise terminology that users rarely use in their queries
- You are doing zero-shot retrieval on a new domain
When HyDE is less useful. If your queries are already long and specific (e.g., users paste error messages), HyDE adds LLM latency for no gain. Use query expansion instead.
HypotheticalDocumentEmbedder component that handles this automatically.Fine-grained chunks (128–256 tokens) are excellent for embedding precision — short passages have focused meaning, and their embeddings are accurate. But returning 128-token snippets to the LLM robs it of context: the answer might require the surrounding section.
Parent-child retrieval stores two levels of chunks:
- Child nodes (small, e.g. 128 tokens) — indexed in the vector store, used for search
- Parent nodes (large, e.g. 512 tokens or full section) — stored in a document store, returned to the LLM
At query time: embed the query, find the most similar child chunks, look up their parent IDs, return the parent chunks to the LLM. The LLM gets richer context; the retrieval stays precise.
LlamaIndex's AutoMergingRetriever implements this with automatic promotion: if enough child nodes from the same parent are retrieved, the system replaces them with the parent.
from llama_index.core import VectorStoreIndex, StorageContext
from llama_index.core.node_parser import HierarchicalNodeParser, get_leaf_nodes
from llama_index.core.storage.docstore import SimpleDocumentStore
from llama_index.core.retrievers import AutoMergingRetriever
from llama_index.core.query_engine import RetrieverQueryEngine
from llama_index.core import SimpleDirectoryReader
# --- 1. Load documents ---
docs = SimpleDirectoryReader("./docs").load_data()
# --- 2. Build hierarchical chunks ---
# Default: chunk_sizes=[2048, 512, 128]
# 2048 = grandparent, 512 = parent, 128 = child (leaf)
node_parser = HierarchicalNodeParser.from_defaults(
chunk_sizes=[512, 128]
)
nodes = node_parser.get_nodes_from_documents(docs)
leaf_nodes = get_leaf_nodes(nodes)
# --- 3. Store ALL nodes in docstore (parent nodes must be accessible) ---
docstore = SimpleDocumentStore()
docstore.add_documents(nodes)
storage_context = StorageContext.from_defaults(docstore=docstore)
# --- 4. Index ONLY leaf nodes in vector store ---
# Only fine-grained chunks are searchable; parents are in docstore only
base_index = VectorStoreIndex(
leaf_nodes,
storage_context=storage_context,
)
base_retriever = base_index.as_retriever(similarity_top_k=12)
# --- 5. Wrap with AutoMergingRetriever ---
# If >= 50% of a parent's children are retrieved, it promotes to the parent
retriever = AutoMergingRetriever(
base_retriever,
storage_context,
verbose=True,
)
# --- 6. Build a query engine ---
query_engine = RetrieverQueryEngine.from_args(retriever)
response = query_engine.query(
"Explain the certificate pinning configuration options"
)
print(response)
[2048, 512, 128] hierarchy works well for long-form technical docs (API references, runbooks). For shorter content like FAQ entries or news articles, try [512, 128]. The rule of thumb: child chunks should be short enough that their embedding is unambiguous; parent chunks should be long enough to fully answer a question.ParentDocumentRetriever that implements the same idea with a flat two-level hierarchy. It is easier to configure if you are already using LangChain's abstractions, but does not support the auto-merging promotion that LlamaIndex provides.Comparing approaches side by side:
| Technique | Retrieves | Returns to LLM | Key benefit |
|---|---|---|---|
| Standard chunking | 512-token chunks | Same chunks | Simplicity |
| Parent-child | 128-token leaf chunks | 512-token parent chunks | Better context for generation |
| Auto-merging | 128-token leaf chunks | Parent if majority matched | Adaptive context window |
Now that you have five techniques individually, the question is: which ones to stack, and in what order? The answer depends on your latency budget and the failure modes you observed in your baseline pipeline.
Recommended order of adoption (by impact-to-effort ratio):
- Hybrid search — always on; adds ~5 ms
- Cross-encoder re-ranking — always on for precision-critical use cases; adds 100–300 ms
- Query expansion — add when queries are ambiguous or short; adds 1 full LLM call
- Parent-child retrieval — add when context truncation is a problem (LLM gives partial answers)
- HyDE — add when there is a large semantic gap between query style and document style
Here is a complete pipeline that stacks hybrid search, query expansion, and cross-encoder re-ranking:
from langchain.retrievers import BM25Retriever, EnsembleRetriever
from langchain.retrievers.multi_query import MultiQueryRetriever
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_community.vectorstores import Chroma
from sentence_transformers import CrossEncoder
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
# ---------------------------------------------------------------
# Component setup (replace with your actual values from Lesson 5)
# ---------------------------------------------------------------
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma(persist_directory="./chroma_db", embedding_function=embeddings)
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
# ---------------------------------------------------------------
# Layer 1: Hybrid retriever (BM25 + dense vector)
# ---------------------------------------------------------------
bm25_retriever = BM25Retriever.from_documents(raw_documents)
bm25_retriever.k = 15
dense_retriever = vectorstore.as_retriever(search_kwargs={"k": 15})
hybrid_retriever = EnsembleRetriever(
retrievers=[bm25_retriever, dense_retriever],
weights=[0.5, 0.5],
)
# ---------------------------------------------------------------
# Layer 2: Query expansion wrapping the hybrid retriever
# ---------------------------------------------------------------
multi_query_retriever = MultiQueryRetriever.from_llm(
retriever=hybrid_retriever,
llm=llm,
)
# ---------------------------------------------------------------
# Layer 3: Cross-encoder re-ranking function
# ---------------------------------------------------------------
def retrieve_and_rerank(query: str, top_k: int = 5) -> list:
candidates = multi_query_retriever.invoke(query)
if not candidates:
return []
pairs = [(query, doc.page_content) for doc in candidates]
scores = reranker.predict(pairs)
ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
return [doc for _, doc in ranked[:top_k]]
# ---------------------------------------------------------------
# Layer 4: Generation with retrieved context
# ---------------------------------------------------------------
prompt = ChatPromptTemplate.from_messages([
(
"system",
"You are a technical assistant. Answer the question using ONLY "
"the provided context. If the context does not contain enough "
"information, say so explicitly.\n\nContext:\n{context}",
),
("human", "{question}"),
])
def format_docs(docs):
return "\n\n---\n\n".join(doc.page_content for doc in docs)
def rag_chain(query: str) -> str:
docs = retrieve_and_rerank(query)
context = format_docs(docs)
messages = prompt.format_messages(context=context, question=query)
return llm.invoke(messages).content
# Usage
answer = rag_chain("How does the LDAP connector handle certificate validation errors?")
print(answer)
Measuring improvement at each step. Before you move to full evaluation (Lesson 7), a quick precision measurement on a small golden set tells you whether each layer is helping:
# Minimal precision@5 measurement
golden_set = [
{
"query": "LDAP certificate validation errors",
"relevant_doc_ids": ["doc_42", "doc_107"],
},
# ... add 10-20 queries
]
def precision_at_k(retrieved_ids, relevant_ids, k=5):
top_k = set(retrieved_ids[:k])
relevant = set(relevant_ids)
return len(top_k & relevant) / k
# Evaluate baseline vs advanced pipeline
for item in golden_set:
baseline_ids = [d.metadata.get("doc_id") for d in dense_retriever.invoke(item["query"])[:5]]
advanced_ids = [d.metadata.get("doc_id") for d in retrieve_and_rerank(item["query"])]
baseline_p = precision_at_k(baseline_ids, item["relevant_doc_ids"])
advanced_p = precision_at_k(advanced_ids, item["relevant_doc_ids"])
print(f"Query: {item['query'][:50]}")
print(f" Baseline P@5: {baseline_p:.2f} | Advanced P@5: {advanced_p:.2f}")
Questions & Answers
ms-marco-MiniLM-L12-v2 to ms-marco-MiniLM-L6-v2 — nearly identical quality, 2× faster. (3) Run the re-ranker on GPU (even a modest one). (4) Cache re-ranking results keyed on (query hash, doc ID set hash) — repeated queries are free. (5) Run the re-ranker asynchronously while streaming the LLM response so the user perceives less latency.ParentDocumentRetriever always returns the parent chunk — it is a flat two-level hierarchy. LlamaIndex's AutoMergingRetriever supports deeper hierarchies (three levels) and has the auto-promotion logic: if a majority of a parent's children are retrieved, it replaces them with the parent. For simple cases, LangChain's version is easier to set up. For complex, long-form documentation where you want adaptive context depth, LlamaIndex's version is more capable. If you covered tool use in the Agentic AI course, note that both integrations work well as tool-enabled retrievers in agent pipelines — see the Agentic AI course for that pattern.rank-bm25 + EnsembleRetriever approach is the right choice when you are using Chroma, pgvector, or any store that does not support hybrid queries natively, or when you want full control over the fusion logic in Python.Key Takeaways
-
Hybrid search is the minimum viable retrieval baseline — combining BM25 sparse retrieval with dense vectors via RRF costs almost nothing in latency and reliably improves recall for exact-match and semantic queries alike.
-
Cross-encoder re-ranking raises precision at the cost of latency — scoring 20 candidates with
cross-encoder/ms-marco-MiniLM-L6-v2takes ~100–300 ms and consistently removes the off-topic documents that dense retrieval includes. -
Query expansion catches what the user did not say — LangChain's
MultiQueryRetrievergenerates alternative phrasings automatically; apply it when users write short, ambiguous queries and you observe low recall on diverse terminology. -
HyDE closes the query-to-document semantic gap — embedding a hypothetical answer rather than the raw query improves retrieval when the linguistic style of queries differs significantly from the style of your documents; factual accuracy of the hypothetical is irrelevant.
-
Parent-child retrieval gives the LLM context without sacrificing index precision — index small leaf chunks (128 tokens) for accurate retrieval, return parent sections (512 tokens) for rich generation context; LlamaIndex's
AutoMergingRetrieverautomates the promotion logic. -
Add techniques one at a time and measure after each — measure precision@k, recall@k, or MRR on a small golden test set before stacking the next technique; this is the only reliable way to isolate what is helping and diagnose regressions.
Next Steps: Lesson 7: Evaluation & Quality