Multi-Modal RAG
Learning Outcomes
- Embed images and text into a shared vector space using CLIP-style models and query across modalities
- Extract tables from PDFs and convert them to a form that a retrieval pipeline can index and search
- Chunk source code using Abstract Syntax Tree boundaries so retrievers return complete, compilable units
- Store multi-modal embeddings in a vector database that supports separate named vector spaces per modality
- Build a query router that inspects incoming questions and dispatches them to the correct modality pipeline
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why text-only RAG breaks on real documents |
| Explain | 7 min | Multi-modal embedding fundamentals — CLIP and the shared latent space |
| Demo | 10 min | Embedding and retrieving images with sentence-transformers and Chroma |
| Demo | 10 min | Table extraction from PDFs with PyMuPDF4LLM and Unstructured |
| Demo | 8 min | Syntax-aware code chunking with ASTChunk |
| Explain | 7 min | Named-vector storage in Qdrant and modality routing |
| Wrap-up | 5 min | Architecture review and key takeaways |
Before You Begin
Pre-work:
- Complete Lesson 5: Building a Basic RAG Pipeline — this lesson extends that end-to-end pipeline
- Complete Lesson 4: Chunking Strategies — especially the section on document-structure-aware chunking
- Skim Lesson 9: Production RAG for context on the operational implications of adding modalities
- If you have taken the Agentic AI course, the concept of tool-use routing is directly analogous to modality routing — see Agentic AI for a refresher
Shopping List:
- Python 3.10 or later
pip install "sentence-transformers[image]" chromadb pymupdf4llm unstructured[pdf] astchunk qdrant-client Pillow- A PDF that mixes text, images, and tables (an annual report or academic paper works well)
- A directory of Python or TypeScript source files for the code-retrieval exercise
Most enterprise documents are not clean prose. A product datasheet embeds specs in a table. A research paper buries key results in a chart. A codebase expresses architecture in function signatures and type annotations. When your RAG pipeline treats the entire document as a flat string of tokens, you lose the structure that carries meaning.
The failures show up predictably:
| Document type | Text-only failure mode |
|---|---|
| PDF with tables | Columns merged or row/column relationships scrambled |
| Slide deck with diagrams | Captions extracted without the image they describe |
| Source code repository | Functions split mid-body, losing the call signature |
| Technical manual with figures | Figure references appear in text but figure content is missing |
Multi-modal RAG solves this by giving each content type its own ingestion path, its own embedding strategy, and — optionally — its own vector index. A single query can then be routed to the right index (or fanned out across all of them) before the results are merged for the LLM.
The three modalities we cover in this lesson:
- Images — embed with a model that shares a vector space with text (CLIP-style), so a text query can retrieve an image directly
- Tables — extract structure-preserving representations (HTML or Markdown), then embed the serialised form and/or attach a verbatim copy for the LLM to reason over
- Code — chunk at AST node boundaries (functions, classes, methods) rather than character counts, and embed with a code-aware model
What CLIP introduced. OpenAI's CLIP (Contrastive Language-Image Pre-Training) was trained on hundreds of millions of image–caption pairs. The training objective is contrastive: push the embedding of an image and its matching caption close together in vector space, and push non-matching pairs apart. The result is a single embedding space where "a golden retriever running on a beach" and a matching photo land near each other — close enough that cosine similarity is a useful similarity signal.
That shared space enables cross-modal retrieval: you embed a text query and find the nearest image vectors. No image-to-text translation step is needed at query time.
Beyond CLIP. More recent models extend this idea with larger backbones and richer training data. The sentence-transformers library (v5.4+) ships with first-class support for multi-modal models including:
| Model | Modalities | Embedding dim |
|---|---|---|
clip-ViT-B-32 (classic) |
text + image | 512 |
BAAI/bge-visualized-base-en-v1.5 |
text + image (+ combined) | 768 |
nomic-ai/nomic-embed-multimodal-3b |
text + image | 768 |
Qwen/Qwen3-VL-Embedding-2B |
text + image + video | 2048 |
For most RAG use cases clip-ViT-B-32 or bge-visualized-base-en-v1.5 give a good balance of speed and quality. The Qwen and Nomic models reach higher retrieval accuracy at the cost of inference time.
Modality gap. An important subtlety: text-to-text cosine similarity scores are typically higher than text-to-image scores even when the content matches, because the two modalities occupy overlapping but distinct regions of the shared space. When building a unified ranked list that mixes text and image results, normalise scores per-modality before merging.
from sentence_transformers import SentenceTransformer
from PIL import Image
# Install with: pip install "sentence-transformers[image]"
model = SentenceTransformer("clip-ViT-B-32")
# Embed a text query
query_emb = model.encode("a bar chart showing quarterly revenue growth")
# Embed an image (accepts file path, PIL.Image, or URL)
img_emb = model.encode(Image.open("charts/q3_revenue.png"))
# Both embeddings live in the same 512-d space — cosine similarity is valid
similarity = model.similarity(query_emb.reshape(1, -1), img_emb.reshape(1, -1))
print(f"similarity: {similarity.item():.4f}")
clip-ViT-B-32 is well-tested and widely supported. For multilingual content or higher retrieval accuracy on technical images, upgrade to nomic-embed-multimodal-3b or BAAI/bge-visualized-m3 (1024-d, multilingual). Benchmark on your actual documents before committing.Chroma ships with an OpenCLIPEmbeddingFunction built in, so you can add images to a collection without manually computing embeddings.
import chromadb
from chromadb.utils.embedding_functions import OpenCLIPEmbeddingFunction
from chromadb.utils.data_loaders import ImageLoader
client = chromadb.PersistentClient(path="./chroma_multimodal")
embedding_fn = OpenCLIPEmbeddingFunction() # uses ViT-B-32 by default
data_loader = ImageLoader() # reads files at query time
collection = client.get_or_create_collection(
name="document_images",
embedding_function=embedding_fn,
data_loader=data_loader,
)
# Index images by URI; Chroma reads the file, embeds it, stores the vector
collection.add(
ids=["img_001", "img_002", "img_003"],
uris=[
"assets/revenue_chart_q3.png",
"assets/architecture_diagram.png",
"assets/product_photo.jpg",
],
metadatas=[
{"source": "annual_report_2025.pdf", "page": 12, "caption": "Q3 revenue breakdown"},
{"source": "design_doc.pdf", "page": 4, "caption": "System architecture"},
{"source": "catalog.pdf", "page": 8, "caption": "Product SKU-447"},
],
)
# Query with plain text — Chroma embeds the query and searches the image vectors
results = collection.query(
query_texts=["quarterly revenue performance"],
n_results=3,
include=["uris", "metadatas", "distances"],
)
for uri, meta, dist in zip(
results["uris"][0],
results["metadatas"][0],
results["distances"][0],
):
print(f"{dist:.4f} {uri} ({meta['caption']})")
Extracting images from PDFs. Before indexing you need the images as files. PyMuPDF handles this efficiently:
import fitz # pip install pymupdf
def extract_images_from_pdf(pdf_path: str, out_dir: str) -> list[dict]:
"""Return a list of dicts with keys: path, page, xref."""
import os
os.makedirs(out_dir, exist_ok=True)
doc = fitz.open(pdf_path)
items = []
for page_num, page in enumerate(doc):
for img in page.get_images(full=True):
xref = img[0]
base = doc.extract_image(xref)
ext = base["ext"] # "png" or "jpeg"
fname = f"page{page_num:03d}_img{xref}.{ext}"
fpath = os.path.join(out_dir, fname)
with open(fpath, "wb") as f:
f.write(base["image"])
items.append({"path": fpath, "page": page_num, "xref": xref})
return items
Tables are the hardest content type for naive text extraction. A PDF renderer draws table cells as positioned text fragments; a plain page.get_text() call concatenates them left-to-right, collapsing the two-dimensional structure into a meaningless string. The fix is layout-aware parsing.
Option A — PyMuPDF4LLM (fast, no OCR dependency)
pymupdf4llm wraps PyMuPDF's native table detection and outputs Markdown with GitHub-compatible table syntax. It interleaves tables with surrounding text in reading order, which preserves the prose context that explains what the table means.
import pymupdf4llm
# Returns a single Markdown string with tables rendered as | col | col | rows
md_text = pymupdf4llm.to_markdown("annual_report_2025.pdf")
# For per-page control:
pages_md = pymupdf4llm.to_markdown(
"annual_report_2025.pdf",
pages=[10, 11, 12], # zero-indexed page numbers
write_images=True, # also extract images to disk
image_path="./assets",
)
The Markdown table can be embedded directly (a table is just a string from the embedder's perspective), but you should also store it verbatim in metadata so the LLM receives the exact cell values, not a vector-averaged approximation.
Option B — Unstructured (richer extraction, OCR fallback)
Unstructured's partition_pdf function uses a computer-vision pipeline to detect table bounding boxes and returns each table as an Element object with text_as_html in its metadata.
from unstructured.partition.pdf import partition_pdf
elements = partition_pdf(
filename="annual_report_2025.pdf",
strategy="hi_res", # triggers layout detection model
infer_table_structure=True,
)
table_elements = [e for e in elements if e.category == "Table"]
for elem in table_elements:
html_table = elem.metadata.text_as_html # full HTML <table> markup
plain_text = str(elem) # flattened text for embedding
source_page = elem.metadata.page_number
# Embed the plain text; store HTML verbatim for the LLM
print(f"Page {source_page}: {plain_text[:120]}...")
What to embed vs. what to store. The embedding model sees the serialised text representation of the table (Markdown or plain text). The LLM context must receive the structured version (HTML or Markdown) so it can read column headers and cell values correctly. Store both:
def table_to_document(plain: str, html: str, meta: dict) -> dict:
return {
"text_for_embedding": plain, # fed to embed()
"content_for_llm": html, # passed to LLM as context
"metadata": meta,
}
strategy='hi_res' runs a YOLOX object detection model on every page. On a CPU this takes several seconds per page. For ingestion pipelines that process thousands of documents, run this step offline in a batch job and cache the extracted elements — do not run it on the hot path of a user query. See Lesson 9: Production RAG for caching patterns.pymupdf4llm when your PDFs are machine-generated (most enterprise reports, papers). Use Unstructured when you are processing scanned documents or PDFs where text is embedded as images — Unstructured falls back to OCR automatically.Code is a modality where arbitrary token-count chunking is especially destructive. A chunk boundary falling inside a function body produces a fragment that is syntactically invalid and semantically orphaned — the retriever may score it highly for a query about that function, but the LLM receives half a function and cannot reason about it correctly.
Abstract Syntax Tree chunking solves this by parsing source code into an AST and splitting at node boundaries: a chunk is always one or more complete top-level definitions (functions, classes, methods). Every chunk is syntactically valid on its own.
ASTChunk wraps tree-sitter and supports Python, Java, C#, and TypeScript out of the box:
from astchunk import ASTChunkBuilder
builder = ASTChunkBuilder(
max_chunk_size=150, # lines per chunk (not characters)
language="python",
metadata_template="default", # prepends file path + class hierarchy
)
with open("src/retrieval/pipeline.py") as f:
source = f.read()
chunks = builder.chunkify(source)
for chunk in chunks:
print("--- content ---")
print(chunk["content"][:200])
print("--- metadata ---")
print(chunk["metadata"])
Each chunk["metadata"] includes the file path and the containing class/function hierarchy (e.g., "src/retrieval/pipeline.py::RAGPipeline::retrieve"), which you should embed alongside the code text so that queries about class methods match the right chunk:
from sentence_transformers import SentenceTransformer
code_model = SentenceTransformer("BAAI/bge-base-en-v1.5") # or a code-specific model
def embed_code_chunk(chunk: dict) -> tuple[list[float], dict]:
# Prepend the metadata path so the embedder sees full context
text = chunk["metadata"] + "\n\n" + chunk["content"]
emb = code_model.encode(text).tolist()
return emb, {
"metadata_path": chunk["metadata"],
"language": "python",
"content": chunk["content"],
}
Choosing a code embedding model. General-purpose models like bge-base-en-v1.5 handle code adequately because code contains meaningful English identifiers and docstrings. For higher precision on code-to-code retrieval (finding similar implementations), consider code-specific models — search HuggingFace for code-embedding or CodeBERT-family models. The trade-off is that a separate model adds deployment complexity. For most enterprise knowledge bases (indexing documentation + code together), a single general model is simpler and nearly as accurate.
| Chunking method | Chunk validity | Retrieves context? | Speed |
|---|---|---|---|
| Fixed character count | Often invalid | Sometimes | Fast |
| Recursive character split | Sometimes invalid | Mostly | Fast |
| AST node boundaries | Always valid | Yes — full definitions | Moderate |
chunk_expansion which prepends the containing class signature to every method chunk — giving the LLM class-level context without duplicating the full class body.When you have three distinct embedding models (text, image, code), you have three vector spaces with different dimensionalities and distance semantics. Qdrant's named vectors feature lets you store all three in a single collection, each in its own indexed space.
from qdrant_client import QdrantClient
from qdrant_client.models import (
VectorParams,
Distance,
PointStruct,
)
client = QdrantClient(path="./qdrant_multimodal")
# Create a collection with three named vector spaces
client.recreate_collection(
collection_name="knowledge_base",
vectors_config={
"text": VectorParams(size=768, distance=Distance.COSINE),
"image": VectorParams(size=512, distance=Distance.COSINE),
"code": VectorParams(size=768, distance=Distance.COSINE),
},
)
Upsert a point that has vectors for only the relevant modalities — Qdrant handles sparse population gracefully:
from uuid import uuid4
# A text chunk (no image or code vector)
client.upsert(
collection_name="knowledge_base",
points=[
PointStruct(
id=str(uuid4()),
vector={"text": text_embedding},
payload={
"modality": "text",
"content": "The gross margin improved by 3.2 percentage points...",
"source": "annual_report_2025.pdf",
"page": 7,
},
)
],
)
# An image chunk (image vector + caption stored in payload)
client.upsert(
collection_name="knowledge_base",
points=[
PointStruct(
id=str(uuid4()),
vector={"image": image_embedding},
payload={
"modality": "image",
"uri": "assets/q3_revenue_chart.png",
"caption": "Q3 revenue breakdown by segment",
"source": "annual_report_2025.pdf",
"page": 12,
},
)
],
)
Query a specific named vector by specifying using:
from qdrant_client.models import SearchRequest
# Search only the image space with a text query embedding
hits = client.search(
collection_name="knowledge_base",
query_vector=("image", query_image_embedding), # (name, vector)
limit=5,
with_payload=True,
)
for hit in hits:
print(f"{hit.score:.4f} {hit.payload['modality']} {hit.payload.get('caption', '')}")
Fan-out search. When you want to search all modalities simultaneously, issue three searches in parallel and merge:
import asyncio
from qdrant_client.async_qdrant_client import AsyncQdrantClient
async def multi_modal_search(
query_text_emb: list[float],
query_image_emb: list[float],
k: int = 5,
) -> list[dict]:
aclient = AsyncQdrantClient(path="./qdrant_multimodal")
text_task = aclient.search("knowledge_base", ("text", query_text_emb), limit=k)
image_task = aclient.search("knowledge_base", ("image", query_image_emb), limit=k)
code_task = aclient.search("knowledge_base", ("code", query_text_emb), limit=k)
text_hits, image_hits, code_hits = await asyncio.gather(
text_task, image_task, code_task
)
# Normalise scores before merging (text and image distances are on different scales)
def normalise(hits):
if not hits:
return []
top = hits[0].score
return [{"score": h.score / top, "payload": h.payload} for h in hits]
merged = normalise(text_hits) + normalise(image_hits) + normalise(code_hits)
merged.sort(key=lambda x: x["score"], reverse=True)
return merged[:k]
Not every query benefits from searching every modality. "Show me the revenue chart from Q3" is clearly an image query. "What does the retrieve() method return?" is clearly a code query. "Summarise the executive summary" is a text query. A modality router classifies the incoming question and dispatches it to one or more vector spaces, reducing latency and noise in the retrieved context.
Simple keyword/heuristic router. Cheap, deterministic, easy to debug:
import re
def classify_query_modality(query: str) -> list[str]:
"""Returns a list of modalities to search: 'text', 'image', 'code'."""
q = query.lower()
image_signals = [
"chart", "graph", "figure", "diagram", "plot", "image",
"photo", "screenshot", "visualis", "show me",
]
code_signals = [
"function", "method", "class", "def ", "import", "returns",
"signature", "implement", "how does.*work", r"\(\)", "code",
]
modalities = []
if any(s in q for s in image_signals):
modalities.append("image")
if any(re.search(s, q) for s in code_signals):
modalities.append("code")
if not modalities or "text" not in modalities:
modalities.insert(0, "text") # always include text as a fallback
return list(dict.fromkeys(modalities)) # deduplicate, preserve order
LLM-based router. More accurate, especially for ambiguous queries. Pass the query to a fast model with a short system prompt:
import json
import anthropic
_client = anthropic.Anthropic()
def llm_route_query(query: str) -> list[str]:
"""Ask the LLM which modalities to search."""
resp = _client.messages.create(
model="claude-haiku-4-5",
max_tokens=64,
system=(
"You route search queries to the correct modality index. "
"Reply with a JSON array containing one or more of: "
"'text', 'image', 'code'. No explanation, just the array."
),
messages=[{"role": "user", "content": query}],
)
raw = resp.content[0].text.strip()
try:
modalities = json.loads(raw)
return [m for m in modalities if m in ("text", "image", "code")]
except json.JSONDecodeError:
return ["text"] # safe fallback
Full pipeline — putting it together:
def multi_modal_rag(query: str, k: int = 4) -> str:
# 1. Route
modalities = classify_query_modality(query)
# 2. Embed query for each relevant modality
text_model = SentenceTransformer("BAAI/bge-base-en-v1.5")
image_model = SentenceTransformer("clip-ViT-B-32")
query_embeddings = {}
if "text" in modalities or "code" in modalities:
query_embeddings["text"] = text_model.encode(query).tolist()
if "image" in modalities:
query_embeddings["image"] = image_model.encode(query).tolist()
# 3. Search named vector spaces
context_items = []
for modality in modalities:
vec_name = "code" if modality == "code" else modality
emb_key = "text" if modality == "code" else modality
hits = client.search(
collection_name="knowledge_base",
query_vector=(vec_name, query_embeddings[emb_key]),
limit=k,
with_payload=True,
)
context_items.extend(hits)
# 4. Deduplicate and sort by score
seen_ids = set()
unique_items = []
for hit in sorted(context_items, key=lambda h: h.score, reverse=True):
if hit.id not in seen_ids:
seen_ids.add(hit.id)
unique_items.append(hit)
# 5. Build LLM prompt
context_parts = []
for item in unique_items[:k]:
payload = item.payload
if payload["modality"] == "image":
context_parts.append(f"[Image: {payload.get('caption', payload['uri'])}]")
elif payload["modality"] == "code":
context_parts.append(f"```python\n{payload['content']}\n```")
else:
context_parts.append(payload["content"])
context = "\n\n---\n\n".join(context_parts)
# Pass context + query to your LLM (not shown — same pattern as Lesson 5)
return context
text as a fallback modality, and set a minimum score threshold so low-confidence results are dropped rather than passed to the LLM. Evaluation with precision@k per modality (covered in Lesson 7) will surface routing failures quickly.Multi-modal RAG introduces evaluation complexity because you now have three retrieval pipelines, each with its own quality metrics.
Per-modality retrieval metrics (see Lesson 7: Evaluation & Quality for the full evaluation framework):
| Modality | Key metric | How to measure |
|---|---|---|
| Text | Recall@5, MRR | Ground-truth QA set with known source documents |
| Image | Recall@5 | Annotated query–image pairs (can bootstrap with LLM-generated captions) |
| Code | Recall@5, Pass@1 | Code search benchmarks, or custom test: given a query, does the retrieved function solve the problem? |
Building an image evaluation set. The absence of labelled query–image pairs is the main evaluation bottleneck. A practical bootstrapping approach:
# Generate synthetic queries from image captions using an LLM
# Then manually verify a random sample (~50 pairs is enough to track trends)
captions = [item.payload["caption"] for item in all_image_items]
def generate_image_queries(caption: str, n: int = 3) -> list[str]:
"""Ask the LLM to write n natural-language queries for this image."""
resp = _client.messages.create(
model="claude-haiku-4-5",
max_tokens=128,
messages=[{
"role": "user",
"content": (
f"Caption: {caption}\n\n"
f"Write {n} short search queries that a user might type "
"to find this image. Return one per line."
),
}],
)
return [line.strip() for line in resp.content[0].text.strip().splitlines() if line.strip()]
Architecture decision points. When designing a multi-modal RAG system, these are the choices that most affect quality and operational cost:
| Decision | Options | When to choose each |
|---|---|---|
| Embedding model | CLIP ViT-B/32 vs. larger VLM | ViT-B/32 for speed; larger model if image queries are core to the product |
| Table representation | Markdown (pymupdf4llm) vs. HTML (Unstructured) | Markdown for text embedding; HTML for LLM context |
| Vector storage | Named vectors in one collection vs. separate collections | One collection for small-to-medium corpora; separate for independent scaling |
| Routing | Heuristic vs. LLM | Heuristic for latency-sensitive; LLM for accuracy-sensitive |
| Image in LLM context | Always include image URIs vs. include only if image score is high | Score threshold avoids irrelevant images inflating token usage |
Questions & Answers
function_definition or equivalent for your target language. The tree-sitter Python bindings (pip install tree-sitter tree-sitter-languages) give you access to all grammars without building from source. Alternatively, use LangChain's Language.detect() splitter which has built-in language detection and per-language split patterns as a fallback — it is not AST-based but respects function-level indentation in most languages.k per modality: searching for 3 results per modality instead of 10 is 3x faster at the Qdrant layer. Fourth, place your Qdrant instance in the same region or on the same host as your API server — network round-trips dominate at the Qdrant layer for small result sets. Finally, consider whether image search is worth it at all for your query distribution — if fewer than 10% of queries benefit from image retrieval, add it as an async enrichment after returning the initial text-only answer.Key Takeaways
- CLIP-style models create a shared vector space — text and image embeddings are directly comparable with cosine similarity, enabling text-to-image retrieval without any translation step. Models like
clip-ViT-B-32andBAAI/bge-visualized-base-en-v1.5give you this out of the box via sentence-transformers. - Tables need structure-preserving extraction — plain text extraction destroys the two-dimensional structure of tables. Use
pymupdf4llmfor machine-generated PDFs (fast, Markdown output) or Unstructured withstrategy='hi_res'for scanned or complex layouts. Always store the structured representation (Markdown or HTML) in metadata separately from the embedding text. - AST-aware chunking makes code retrievable — splitting code at AST node boundaries (functions, classes, methods) produces syntactically valid, semantically complete chunks. ASTChunk wraps tree-sitter and handles Python, Java, C#, and TypeScript; prepend the class hierarchy metadata to each chunk to improve retrieval accuracy.
- Named vector spaces allow one collection per corpus — Qdrant's named-vector feature stores text, image, and code embeddings in separate indexed spaces within a single collection, avoiding the complexity of cross-collection joins while allowing each modality to use its own dimensionality and distance metric.
- Query routing reduces noise — dispatching a query only to the relevant modality indexes improves retrieval precision and reduces latency. Start with a keyword heuristic router; upgrade to an LLM-based router when routing accuracy becomes a measurable bottleneck.
- Evaluate each modality independently — precision@k and recall@k for images and code require purpose-built test sets. Bootstrap image evaluation from LLM-generated caption queries; bootstrap code evaluation from known function-lookup tasks in your codebase. Track metrics per modality to identify which retrieval path degrades first in production.
Next Steps: Back to all RAG & Knowledge lessons