Multi-Modal RAG

50 min advanced Lesson 10

Learning Outcomes

  • Embed images and text into a shared vector space using CLIP-style models and query across modalities
  • Extract tables from PDFs and convert them to a form that a retrieval pipeline can index and search
  • Chunk source code using Abstract Syntax Tree boundaries so retrievers return complete, compilable units
  • Store multi-modal embeddings in a vector database that supports separate named vector spaces per modality
  • Build a query router that inspects incoming questions and dispatches them to the correct modality pipeline

Lesson Plan

Segment Duration Topic
Intro 3 min Why text-only RAG breaks on real documents
Explain 7 min Multi-modal embedding fundamentals — CLIP and the shared latent space
Demo 10 min Embedding and retrieving images with sentence-transformers and Chroma
Demo 10 min Table extraction from PDFs with PyMuPDF4LLM and Unstructured
Demo 8 min Syntax-aware code chunking with ASTChunk
Explain 7 min Named-vector storage in Qdrant and modality routing
Wrap-up 5 min Architecture review and key takeaways

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10 or later
  • pip install "sentence-transformers[image]" chromadb pymupdf4llm unstructured[pdf] astchunk qdrant-client Pillow
  • A PDF that mixes text, images, and tables (an annual report or academic paper works well)
  • A directory of Python or TypeScript source files for the code-retrieval exercise

1 Why Text-Only RAG Fails on Real Documents

Most enterprise documents are not clean prose. A product datasheet embeds specs in a table. A research paper buries key results in a chart. A codebase expresses architecture in function signatures and type annotations. When your RAG pipeline treats the entire document as a flat string of tokens, you lose the structure that carries meaning.

The failures show up predictably:

Document type Text-only failure mode
PDF with tables Columns merged or row/column relationships scrambled
Slide deck with diagrams Captions extracted without the image they describe
Source code repository Functions split mid-body, losing the call signature
Technical manual with figures Figure references appear in text but figure content is missing

Multi-modal RAG solves this by giving each content type its own ingestion path, its own embedding strategy, and — optionally — its own vector index. A single query can then be routed to the right index (or fanned out across all of them) before the results are merged for the LLM.

The three modalities we cover in this lesson:

  1. Images — embed with a model that shares a vector space with text (CLIP-style), so a text query can retrieve an image directly
  2. Tables — extract structure-preserving representations (HTML or Markdown), then embed the serialised form and/or attach a verbatim copy for the LLM to reason over
  3. Code — chunk at AST node boundaries (functions, classes, methods) rather than character counts, and embed with a code-aware model
NOTE
Scope
Audio and video retrieval follow similar patterns but require specialised pre-processing (transcription, frame sampling). They are outside scope for this lesson. The same architectural principles — separate embedding models, named vector spaces, routed queries — apply directly.

2 Multi-Modal Embeddings — The Shared Latent Space

What CLIP introduced. OpenAI's CLIP (Contrastive Language-Image Pre-Training) was trained on hundreds of millions of image–caption pairs. The training objective is contrastive: push the embedding of an image and its matching caption close together in vector space, and push non-matching pairs apart. The result is a single embedding space where "a golden retriever running on a beach" and a matching photo land near each other — close enough that cosine similarity is a useful similarity signal.

That shared space enables cross-modal retrieval: you embed a text query and find the nearest image vectors. No image-to-text translation step is needed at query time.

Beyond CLIP. More recent models extend this idea with larger backbones and richer training data. The sentence-transformers library (v5.4+) ships with first-class support for multi-modal models including:

Model Modalities Embedding dim
clip-ViT-B-32 (classic) text + image 512
BAAI/bge-visualized-base-en-v1.5 text + image (+ combined) 768
nomic-ai/nomic-embed-multimodal-3b text + image 768
Qwen/Qwen3-VL-Embedding-2B text + image + video 2048

For most RAG use cases clip-ViT-B-32 or bge-visualized-base-en-v1.5 give a good balance of speed and quality. The Qwen and Nomic models reach higher retrieval accuracy at the cost of inference time.

Modality gap. An important subtlety: text-to-text cosine similarity scores are typically higher than text-to-image scores even when the content matches, because the two modalities occupy overlapping but distinct regions of the shared space. When building a unified ranked list that mixes text and image results, normalise scores per-modality before merging.

from sentence_transformers import SentenceTransformer
from PIL import Image

# Install with: pip install "sentence-transformers[image]"
model = SentenceTransformer("clip-ViT-B-32")

# Embed a text query
query_emb = model.encode("a bar chart showing quarterly revenue growth")

# Embed an image (accepts file path, PIL.Image, or URL)
img_emb = model.encode(Image.open("charts/q3_revenue.png"))

# Both embeddings live in the same 512-d space — cosine similarity is valid
similarity = model.similarity(query_emb.reshape(1, -1), img_emb.reshape(1, -1))
print(f"similarity: {similarity.item():.4f}")
TIP
Choosing a model
For English-only corpora where speed matters, clip-ViT-B-32 is well-tested and widely supported. For multilingual content or higher retrieval accuracy on technical images, upgrade to nomic-embed-multimodal-3b or BAAI/bge-visualized-m3 (1024-d, multilingual). Benchmark on your actual documents before committing.
WARNING
Image resolution
CLIP-style models were trained on images resized to 224×224 pixels. Very high resolution images do not add retrieval quality — the model can not see fine detail. For charts and tables containing small text, a better approach is OCR or layout-aware extraction (covered in Step 3), not higher-resolution embedding.

3 Image Indexing — Embedding and Storing in Chroma

Chroma ships with an OpenCLIPEmbeddingFunction built in, so you can add images to a collection without manually computing embeddings.

import chromadb
from chromadb.utils.embedding_functions import OpenCLIPEmbeddingFunction
from chromadb.utils.data_loaders import ImageLoader

client = chromadb.PersistentClient(path="./chroma_multimodal")

embedding_fn = OpenCLIPEmbeddingFunction()   # uses ViT-B-32 by default
data_loader  = ImageLoader()                 # reads files at query time

collection = client.get_or_create_collection(
    name="document_images",
    embedding_function=embedding_fn,
    data_loader=data_loader,
)

# Index images by URI; Chroma reads the file, embeds it, stores the vector
collection.add(
    ids=["img_001", "img_002", "img_003"],
    uris=[
        "assets/revenue_chart_q3.png",
        "assets/architecture_diagram.png",
        "assets/product_photo.jpg",
    ],
    metadatas=[
        {"source": "annual_report_2025.pdf", "page": 12, "caption": "Q3 revenue breakdown"},
        {"source": "design_doc.pdf",         "page": 4,  "caption": "System architecture"},
        {"source": "catalog.pdf",            "page": 8,  "caption": "Product SKU-447"},
    ],
)

# Query with plain text — Chroma embeds the query and searches the image vectors
results = collection.query(
    query_texts=["quarterly revenue performance"],
    n_results=3,
    include=["uris", "metadatas", "distances"],
)

for uri, meta, dist in zip(
    results["uris"][0],
    results["metadatas"][0],
    results["distances"][0],
):
    print(f"{dist:.4f}  {uri}  ({meta['caption']})")

Extracting images from PDFs. Before indexing you need the images as files. PyMuPDF handles this efficiently:

import fitz  # pip install pymupdf

def extract_images_from_pdf(pdf_path: str, out_dir: str) -> list[dict]:
    """Return a list of dicts with keys: path, page, xref."""
    import os
    os.makedirs(out_dir, exist_ok=True)
    doc   = fitz.open(pdf_path)
    items = []
    for page_num, page in enumerate(doc):
        for img in page.get_images(full=True):
            xref   = img[0]
            base   = doc.extract_image(xref)
            ext    = base["ext"]            # "png" or "jpeg"
            fname  = f"page{page_num:03d}_img{xref}.{ext}"
            fpath  = os.path.join(out_dir, fname)
            with open(fpath, "wb") as f:
                f.write(base["image"])
            items.append({"path": fpath, "page": page_num, "xref": xref})
    return items
NOTE
What gets retrieved
At query time Chroma returns the URI, metadata, and distance. You then pass the image file to your multi-modal LLM (e.g., Claude claude-sonnet-4-6, GPT-4o) as a base64-encoded image alongside the text context. The LLM synthesises an answer that references both text chunks and image content. This is the same image-in-context pattern used by vision-capable models — retrieval just handles which images to include.

4 Table Extraction — Structure-Preserving PDF Parsing

Tables are the hardest content type for naive text extraction. A PDF renderer draws table cells as positioned text fragments; a plain page.get_text() call concatenates them left-to-right, collapsing the two-dimensional structure into a meaningless string. The fix is layout-aware parsing.

Option A — PyMuPDF4LLM (fast, no OCR dependency)

pymupdf4llm wraps PyMuPDF's native table detection and outputs Markdown with GitHub-compatible table syntax. It interleaves tables with surrounding text in reading order, which preserves the prose context that explains what the table means.

import pymupdf4llm

# Returns a single Markdown string with tables rendered as | col | col | rows
md_text = pymupdf4llm.to_markdown("annual_report_2025.pdf")

# For per-page control:
pages_md = pymupdf4llm.to_markdown(
    "annual_report_2025.pdf",
    pages=[10, 11, 12],       # zero-indexed page numbers
    write_images=True,        # also extract images to disk
    image_path="./assets",
)

The Markdown table can be embedded directly (a table is just a string from the embedder's perspective), but you should also store it verbatim in metadata so the LLM receives the exact cell values, not a vector-averaged approximation.

Option B — Unstructured (richer extraction, OCR fallback)

Unstructured's partition_pdf function uses a computer-vision pipeline to detect table bounding boxes and returns each table as an Element object with text_as_html in its metadata.

from unstructured.partition.pdf import partition_pdf

elements = partition_pdf(
    filename="annual_report_2025.pdf",
    strategy="hi_res",          # triggers layout detection model
    infer_table_structure=True,
)

table_elements = [e for e in elements if e.category == "Table"]

for elem in table_elements:
    html_table = elem.metadata.text_as_html   # full HTML <table> markup
    plain_text = str(elem)                    # flattened text for embedding
    source_page = elem.metadata.page_number

    # Embed the plain text; store HTML verbatim for the LLM
    print(f"Page {source_page}: {plain_text[:120]}...")

What to embed vs. what to store. The embedding model sees the serialised text representation of the table (Markdown or plain text). The LLM context must receive the structured version (HTML or Markdown) so it can read column headers and cell values correctly. Store both:

def table_to_document(plain: str, html: str, meta: dict) -> dict:
    return {
        "text_for_embedding": plain,      # fed to embed()
        "content_for_llm": html,          # passed to LLM as context
        "metadata": meta,
    }
WARNING
Strategy hi_res is slow
Unstructured's strategy='hi_res' runs a YOLOX object detection model on every page. On a CPU this takes several seconds per page. For ingestion pipelines that process thousands of documents, run this step offline in a batch job and cache the extracted elements — do not run it on the hot path of a user query. See Lesson 9: Production RAG for caching patterns.
TIP
When to use which
Use pymupdf4llm when your PDFs are machine-generated (most enterprise reports, papers). Use Unstructured when you are processing scanned documents or PDFs where text is embedded as images — Unstructured falls back to OCR automatically.

5 Syntax-Aware Code Chunking with ASTChunk

Code is a modality where arbitrary token-count chunking is especially destructive. A chunk boundary falling inside a function body produces a fragment that is syntactically invalid and semantically orphaned — the retriever may score it highly for a query about that function, but the LLM receives half a function and cannot reason about it correctly.

Abstract Syntax Tree chunking solves this by parsing source code into an AST and splitting at node boundaries: a chunk is always one or more complete top-level definitions (functions, classes, methods). Every chunk is syntactically valid on its own.

ASTChunk wraps tree-sitter and supports Python, Java, C#, and TypeScript out of the box:

from astchunk import ASTChunkBuilder

builder = ASTChunkBuilder(
    max_chunk_size=150,         # lines per chunk (not characters)
    language="python",
    metadata_template="default",  # prepends file path + class hierarchy
)

with open("src/retrieval/pipeline.py") as f:
    source = f.read()

chunks = builder.chunkify(source)

for chunk in chunks:
    print("--- content ---")
    print(chunk["content"][:200])
    print("--- metadata ---")
    print(chunk["metadata"])

Each chunk["metadata"] includes the file path and the containing class/function hierarchy (e.g., "src/retrieval/pipeline.py::RAGPipeline::retrieve"), which you should embed alongside the code text so that queries about class methods match the right chunk:

from sentence_transformers import SentenceTransformer

code_model = SentenceTransformer("BAAI/bge-base-en-v1.5")  # or a code-specific model

def embed_code_chunk(chunk: dict) -> tuple[list[float], dict]:
    # Prepend the metadata path so the embedder sees full context
    text = chunk["metadata"] + "\n\n" + chunk["content"]
    emb  = code_model.encode(text).tolist()
    return emb, {
        "metadata_path": chunk["metadata"],
        "language": "python",
        "content": chunk["content"],
    }

Choosing a code embedding model. General-purpose models like bge-base-en-v1.5 handle code adequately because code contains meaningful English identifiers and docstrings. For higher precision on code-to-code retrieval (finding similar implementations), consider code-specific models — search HuggingFace for code-embedding or CodeBERT-family models. The trade-off is that a separate model adds deployment complexity. For most enterprise knowledge bases (indexing documentation + code together), a single general model is simpler and nearly as accurate.

Chunking method Chunk validity Retrieves context? Speed
Fixed character count Often invalid Sometimes Fast
Recursive character split Sometimes invalid Mostly Fast
AST node boundaries Always valid Yes — full definitions Moderate
NOTE
Overlap with AST chunking
Classic overlap (duplicate N tokens at chunk boundaries) makes less sense for AST chunking because chunks are already complete units. Instead, ASTChunk supports chunk_expansion which prepends the containing class signature to every method chunk — giving the LLM class-level context without duplicating the full class body.

6 Named-Vector Storage in Qdrant

When you have three distinct embedding models (text, image, code), you have three vector spaces with different dimensionalities and distance semantics. Qdrant's named vectors feature lets you store all three in a single collection, each in its own indexed space.

from qdrant_client import QdrantClient
from qdrant_client.models import (
    VectorParams,
    Distance,
    PointStruct,
)

client = QdrantClient(path="./qdrant_multimodal")

# Create a collection with three named vector spaces
client.recreate_collection(
    collection_name="knowledge_base",
    vectors_config={
        "text": VectorParams(size=768,  distance=Distance.COSINE),
        "image": VectorParams(size=512,  distance=Distance.COSINE),
        "code": VectorParams(size=768,  distance=Distance.COSINE),
    },
)

Upsert a point that has vectors for only the relevant modalities — Qdrant handles sparse population gracefully:

from uuid import uuid4

# A text chunk (no image or code vector)
client.upsert(
    collection_name="knowledge_base",
    points=[
        PointStruct(
            id=str(uuid4()),
            vector={"text": text_embedding},
            payload={
                "modality": "text",
                "content": "The gross margin improved by 3.2 percentage points...",
                "source": "annual_report_2025.pdf",
                "page": 7,
            },
        )
    ],
)

# An image chunk (image vector + caption stored in payload)
client.upsert(
    collection_name="knowledge_base",
    points=[
        PointStruct(
            id=str(uuid4()),
            vector={"image": image_embedding},
            payload={
                "modality": "image",
                "uri": "assets/q3_revenue_chart.png",
                "caption": "Q3 revenue breakdown by segment",
                "source": "annual_report_2025.pdf",
                "page": 12,
            },
        )
    ],
)

Query a specific named vector by specifying using:

from qdrant_client.models import SearchRequest

# Search only the image space with a text query embedding
hits = client.search(
    collection_name="knowledge_base",
    query_vector=("image", query_image_embedding),  # (name, vector)
    limit=5,
    with_payload=True,
)

for hit in hits:
    print(f"{hit.score:.4f}  {hit.payload['modality']}  {hit.payload.get('caption', '')}")

Fan-out search. When you want to search all modalities simultaneously, issue three searches in parallel and merge:

import asyncio
from qdrant_client.async_qdrant_client import AsyncQdrantClient

async def multi_modal_search(
    query_text_emb: list[float],
    query_image_emb: list[float],
    k: int = 5,
) -> list[dict]:
    aclient = AsyncQdrantClient(path="./qdrant_multimodal")

    text_task  = aclient.search("knowledge_base", ("text",  query_text_emb),  limit=k)
    image_task = aclient.search("knowledge_base", ("image", query_image_emb), limit=k)
    code_task  = aclient.search("knowledge_base", ("code",  query_text_emb),  limit=k)

    text_hits, image_hits, code_hits = await asyncio.gather(
        text_task, image_task, code_task
    )

    # Normalise scores before merging (text and image distances are on different scales)
    def normalise(hits):
        if not hits:
            return []
        top = hits[0].score
        return [{"score": h.score / top, "payload": h.payload} for h in hits]

    merged = normalise(text_hits) + normalise(image_hits) + normalise(code_hits)
    merged.sort(key=lambda x: x["score"], reverse=True)
    return merged[:k]
TIP
Named vectors vs. separate collections
Named vectors are preferable when a single logical document can have all three modalities (e.g., a PDF page with text, a table, and a figure). Separate collections make sense when modalities are entirely independent corpora — e.g., a code repository indexed separately from a documentation corpus. Named vectors simplify metadata joins; separate collections simplify scaling each modality independently.

7 Query Routing — Dispatching to the Right Modality

Not every query benefits from searching every modality. "Show me the revenue chart from Q3" is clearly an image query. "What does the retrieve() method return?" is clearly a code query. "Summarise the executive summary" is a text query. A modality router classifies the incoming question and dispatches it to one or more vector spaces, reducing latency and noise in the retrieved context.

Simple keyword/heuristic router. Cheap, deterministic, easy to debug:

import re

def classify_query_modality(query: str) -> list[str]:
    """Returns a list of modalities to search: 'text', 'image', 'code'."""
    q = query.lower()

    image_signals = [
        "chart", "graph", "figure", "diagram", "plot", "image",
        "photo", "screenshot", "visualis", "show me",
    ]
    code_signals = [
        "function", "method", "class", "def ", "import", "returns",
        "signature", "implement", "how does.*work", r"\(\)", "code",
    ]

    modalities = []
    if any(s in q for s in image_signals):
        modalities.append("image")
    if any(re.search(s, q) for s in code_signals):
        modalities.append("code")
    if not modalities or "text" not in modalities:
        modalities.insert(0, "text")   # always include text as a fallback

    return list(dict.fromkeys(modalities))  # deduplicate, preserve order

LLM-based router. More accurate, especially for ambiguous queries. Pass the query to a fast model with a short system prompt:

import json
import anthropic

_client = anthropic.Anthropic()

def llm_route_query(query: str) -> list[str]:
    """Ask the LLM which modalities to search."""
    resp = _client.messages.create(
        model="claude-haiku-4-5",
        max_tokens=64,
        system=(
            "You route search queries to the correct modality index. "
            "Reply with a JSON array containing one or more of: "
            "'text', 'image', 'code'. No explanation, just the array."
        ),
        messages=[{"role": "user", "content": query}],
    )
    raw = resp.content[0].text.strip()
    try:
        modalities = json.loads(raw)
        return [m for m in modalities if m in ("text", "image", "code")]
    except json.JSONDecodeError:
        return ["text"]    # safe fallback

Full pipeline — putting it together:

def multi_modal_rag(query: str, k: int = 4) -> str:
    # 1. Route
    modalities = classify_query_modality(query)

    # 2. Embed query for each relevant modality
    text_model  = SentenceTransformer("BAAI/bge-base-en-v1.5")
    image_model = SentenceTransformer("clip-ViT-B-32")

    query_embeddings = {}
    if "text" in modalities or "code" in modalities:
        query_embeddings["text"] = text_model.encode(query).tolist()
    if "image" in modalities:
        query_embeddings["image"] = image_model.encode(query).tolist()

    # 3. Search named vector spaces
    context_items = []
    for modality in modalities:
        vec_name = "code" if modality == "code" else modality
        emb_key  = "text" if modality == "code" else modality
        hits = client.search(
            collection_name="knowledge_base",
            query_vector=(vec_name, query_embeddings[emb_key]),
            limit=k,
            with_payload=True,
        )
        context_items.extend(hits)

    # 4. Deduplicate and sort by score
    seen_ids = set()
    unique_items = []
    for hit in sorted(context_items, key=lambda h: h.score, reverse=True):
        if hit.id not in seen_ids:
            seen_ids.add(hit.id)
            unique_items.append(hit)

    # 5. Build LLM prompt
    context_parts = []
    for item in unique_items[:k]:
        payload = item.payload
        if payload["modality"] == "image":
            context_parts.append(f"[Image: {payload.get('caption', payload['uri'])}]")
        elif payload["modality"] == "code":
            context_parts.append(f"```python\n{payload['content']}\n```")
        else:
            context_parts.append(payload["content"])

    context = "\n\n---\n\n".join(context_parts)
    # Pass context + query to your LLM (not shown — same pattern as Lesson 5)
    return context
WARNING
Router errors cascade
A misrouted query that only searches the image index for a code question returns zero relevant results — the LLM then hallucinates. Defensive pattern: always include text as a fallback modality, and set a minimum score threshold so low-confidence results are dropped rather than passed to the LLM. Evaluation with precision@k per modality (covered in Lesson 7) will surface routing failures quickly.
TIP
Cost of LLM routing
The LLM-based router adds one fast-model API call per user query. At the scale of most internal knowledge bases (tens to hundreds of queries per minute), this is negligible. If you are running at higher throughput, cache routing decisions by embedding the query and storing the modality label — similar queries will reuse cached routes.

8 Evaluation and Architecture Trade-offs

Multi-modal RAG introduces evaluation complexity because you now have three retrieval pipelines, each with its own quality metrics.

Per-modality retrieval metrics (see Lesson 7: Evaluation & Quality for the full evaluation framework):

Modality Key metric How to measure
Text Recall@5, MRR Ground-truth QA set with known source documents
Image Recall@5 Annotated query–image pairs (can bootstrap with LLM-generated captions)
Code Recall@5, Pass@1 Code search benchmarks, or custom test: given a query, does the retrieved function solve the problem?

Building an image evaluation set. The absence of labelled query–image pairs is the main evaluation bottleneck. A practical bootstrapping approach:

# Generate synthetic queries from image captions using an LLM
# Then manually verify a random sample (~50 pairs is enough to track trends)

captions = [item.payload["caption"] for item in all_image_items]

def generate_image_queries(caption: str, n: int = 3) -> list[str]:
    """Ask the LLM to write n natural-language queries for this image."""
    resp = _client.messages.create(
        model="claude-haiku-4-5",
        max_tokens=128,
        messages=[{
            "role": "user",
            "content": (
                f"Caption: {caption}\n\n"
                f"Write {n} short search queries that a user might type "
                "to find this image. Return one per line."
            ),
        }],
    )
    return [line.strip() for line in resp.content[0].text.strip().splitlines() if line.strip()]

Architecture decision points. When designing a multi-modal RAG system, these are the choices that most affect quality and operational cost:

Decision Options When to choose each
Embedding model CLIP ViT-B/32 vs. larger VLM ViT-B/32 for speed; larger model if image queries are core to the product
Table representation Markdown (pymupdf4llm) vs. HTML (Unstructured) Markdown for text embedding; HTML for LLM context
Vector storage Named vectors in one collection vs. separate collections One collection for small-to-medium corpora; separate for independent scaling
Routing Heuristic vs. LLM Heuristic for latency-sensitive; LLM for accuracy-sensitive
Image in LLM context Always include image URIs vs. include only if image score is high Score threshold avoids irrelevant images inflating token usage
NOTE
Start simple
You do not need all three modalities on day one. The most common path: (1) add table extraction to an existing text RAG pipeline first — tables are the most common failure mode, (2) add code chunking if the corpus includes source files, (3) add image embeddings last, since they require a vision-capable LLM to synthesise answers. Each addition is independent; the routing layer just needs a new branch.

Questions & Answers

Q: My image retrieval recall is low — a text query that clearly describes an image fails to retrieve it. What's wrong?
Three common causes. First, the modality gap: text-to-image similarity is inherently lower than text-to-text, so your score threshold may be too high. Lower it, or normalise scores within each modality before thresholding. Second, the image may be ambiguous — a diagram without a clear visual subject embeds poorly. Fix: store a rich caption or alt-text in the payload and also embed the caption text into the text index. A text query then retrieves the caption chunk, which carries the image URI in its metadata. Third, the image quality or resolution may be causing issues — CLIP was trained on photographic images and struggles with sparse line diagrams. For technical diagrams, embedding the caption + any labels you can OCR off the figure is more reliable than embedding the image pixel values.
Q: Table extraction produces garbled output — columns are merged and row values are out of order. Neither pymupdf4llm nor Unstructured fixes it. What next?
Some PDFs store table data as positioned text with no underlying table structure — the PDF renderer draws the characters, but there is no semantic grouping. For these documents, layout-aware OCR is the only reliable path: pass each page as a rendered image to a vision-capable LLM (GPT-4o, Claude with image input) with a prompt like "Extract the table on this page as Markdown, preserving all column headers and row values exactly." This is expensive per page, so gate it: only invoke the LLM when heuristic extraction confidence is low (e.g., when pymupdf4llm's table bounding box has fewer than 3 columns or 2 rows).
Q: ASTChunk doesn't support the language I need (Rust, Go, Ruby). How do I get syntax-aware chunking?
Tree-sitter itself supports over 100 languages via language-specific grammars. You can write a thin wrapper that (1) loads the right tree-sitter grammar for your language, (2) parses the file, (3) walks the AST to find top-level function and class nodes, and (4) extracts the source slice for each node. The key query is function_definition or equivalent for your target language. The tree-sitter Python bindings (pip install tree-sitter tree-sitter-languages) give you access to all grammars without building from source. Alternatively, use LangChain's Language.detect() splitter which has built-in language detection and per-language split patterns as a fallback — it is not AST-based but respects function-level indentation in most languages.
Q: Latency for multi-modal queries is too high — three vector searches in parallel plus routing still takes over 500ms. How do I bring this down?
Several levers. First, use heuristic routing instead of LLM routing — it is sub-millisecond. Second, cache embeddings: if the same query (or a near-duplicate) has been seen recently, reuse the cached embedding vector. A Redis LMDB or in-process LRU cache keyed by query text handles this. Third, reduce k per modality: searching for 3 results per modality instead of 10 is 3x faster at the Qdrant layer. Fourth, place your Qdrant instance in the same region or on the same host as your API server — network round-trips dominate at the Qdrant layer for small result sets. Finally, consider whether image search is worth it at all for your query distribution — if fewer than 10% of queries benefit from image retrieval, add it as an async enrichment after returning the initial text-only answer.
Q: When I pass retrieved images to the LLM the answers become longer and less focused. Is multi-modal context always worth including?
Not always. Image context is valuable when the answer cannot be fully expressed in text — charts with data points, architectural diagrams, screenshots. It is actively harmful when the image is marginally related: the LLM describes the image rather than answering the question, inflating token usage and reducing faithfulness. Use a score threshold: only include an image in the LLM prompt when its retrieval distance is above a minimum confidence. Start at the 75th percentile of your observed score distribution and adjust based on RAGAS faithfulness scores (covered in Lesson 7). You can also ask the LLM router whether a given query genuinely requires an image as a second-pass filter.

Key Takeaways

  1. CLIP-style models create a shared vector space — text and image embeddings are directly comparable with cosine similarity, enabling text-to-image retrieval without any translation step. Models like clip-ViT-B-32 and BAAI/bge-visualized-base-en-v1.5 give you this out of the box via sentence-transformers.
  2. Tables need structure-preserving extraction — plain text extraction destroys the two-dimensional structure of tables. Use pymupdf4llm for machine-generated PDFs (fast, Markdown output) or Unstructured with strategy='hi_res' for scanned or complex layouts. Always store the structured representation (Markdown or HTML) in metadata separately from the embedding text.
  3. AST-aware chunking makes code retrievable — splitting code at AST node boundaries (functions, classes, methods) produces syntactically valid, semantically complete chunks. ASTChunk wraps tree-sitter and handles Python, Java, C#, and TypeScript; prepend the class hierarchy metadata to each chunk to improve retrieval accuracy.
  4. Named vector spaces allow one collection per corpus — Qdrant's named-vector feature stores text, image, and code embeddings in separate indexed spaces within a single collection, avoiding the complexity of cross-collection joins while allowing each modality to use its own dimensionality and distance metric.
  5. Query routing reduces noise — dispatching a query only to the relevant modality indexes improves retrieval precision and reduces latency. Start with a keyword heuristic router; upgrade to an LLM-based router when routing accuracy becomes a measurable bottleneck.
  6. Evaluate each modality independently — precision@k and recall@k for images and code require purpose-built test sets. Bootstrap image evaluation from LLM-generated caption queries; bootstrap code evaluation from known function-lookup tasks in your codebase. Track metrics per modality to identify which retrieval path degrades first in production.

Next Steps: Back to all RAG & Knowledge lessons