Chunking Strategies

50 min intermediate Lesson 4

Learning Outcomes

  • Explain why chunking decisions affect retrieval quality more than almost any other pipeline choice
  • Implement fixed-size, recursive, semantic, and structure-aware chunking strategies in Python
  • Evaluate the precision and recall trade-offs produced by different chunk sizes and overlap settings
  • Enrich chunks with metadata — source path, section title, parent context — to improve post-retrieval grounding
  • Select the right chunking strategy for a given document type and query workload

Lesson Plan

Segment Duration Topic
Intro 3 min Why chunking is the highest-leverage pipeline decision
Explain 7 min Mental model: retrieval granularity vs. context completeness
Demo 8 min Fixed-size chunking and its failure modes
Demo 10 min Recursive and semantic chunking
Demo 8 min Structure-aware chunking for HTML, Markdown, code
Demo 8 min Metadata enrichment and contextual headers
Lab 4 min Measuring retrieval quality across strategies
Wrap-up 2 min Decision framework and preview of Lesson 5

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with a virtual environment
  • pip install chromadb langchain langchain-community sentence-transformers tiktoken
  • A handful of sample documents — a PDF technical manual, a Markdown README, and an HTML page work well; we provide a sample corpus below
  • Optional: pip install unstructured for advanced structure detection

1 Why Chunking Is the Highest-Leverage Decision

Before writing a single line of chunking code, it is worth understanding the fundamental tension you are navigating.

The retrieval granularity problem:

A RAG system embeds each chunk independently and retrieves the top-k chunks most similar to the query embedding. This means the chunk is the atomic unit of retrieval. If a chunk is too large, one query can only surface it when a majority of its sentences are relevant — meaning many relevant passages are buried inside chunks that score poorly. If a chunk is too small, retrieved text may lack the surrounding context the LLM needs to produce a coherent answer.

This tension does not resolve neatly. It shifts depending on the query pattern:

Query type Prefers Because
Fact lookup ("What is the max retry count?") Small chunks (128–256 tokens) One sentence holds the answer; large chunks dilute the signal
Conceptual ("Explain the authentication flow") Larger chunks (512–1024 tokens) Answer spans multiple paragraphs; small chunks fracture the explanation
Comparative ("How do approaches A and B differ?") Carefully positioned chunks Both halves of the comparison must be retrievable

Most real query workloads are mixed, which is why advanced strategies (parent-child retrieval, covered in Lesson 6) often outperform any single fixed strategy.

Two concrete failure modes to recognise:

DOCUMENT EXCERPT:
  "The retry limit defaults to 3. When this limit is exceeded,
   the system raises a MaxRetriesExceeded exception and logs
   the event to the audit trail."

QUERY: "What exception is raised on too many retries?"

--- FAILURE: chunks split at sentence boundary ---
Chunk A: "The retry limit defaults to 3."
Chunk B: "When this limit is exceeded, the system raises a
          MaxRetriesExceeded exception and logs the event
          to the audit trail."

  Chunk A scores high (contains "retry" and "limit") but
  does NOT contain the exception name.
  Chunk B scores lower (no "retry limit" context) — it may
  not even appear in top-k retrieval.

--- SUCCESS: boundary respects sentence grouping ---
Chunk: "The retry limit defaults to 3. When this limit is
        exceeded, the system raises a MaxRetriesExceeded
        exception and logs the event to the audit trail."
  A single retrieval surfaces both the condition and the answer.
NOTE
First-Principles Intuition
Think of each chunk as a self-contained flash card. The embedding measures how relevant that card is to the question. A card split mid-sentence is incomplete on both sides.
WARNING
The Token Budget Is Real
Your LLM has a context window. If you retrieve 10 chunks of 1 000 tokens each, you consume 10 000 tokens before the question and system prompt. Generous chunk sizes compound quickly — keep top-k × chunk_size well within the model's context limit, typically below half of it.

2 Fixed-Size Chunking — Simple but Brittle

Fixed-size chunking splits text every N tokens or characters, with an optional overlap. It is the simplest strategy and a useful baseline, but it is blind to document structure.

Implementing fixed-size chunking with LangChain:

from langchain.text_splitter import CharacterTextSplitter, TokenTextSplitter
from pathlib import Path

raw_text = Path("docs/manual.md").read_text()

# Character-based: split every 1000 chars, 200-char overlap
char_splitter = CharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200,
    separator="\n",           # prefer splitting on newlines
    length_function=len,
)
char_chunks = char_splitter.split_text(raw_text)

# Token-based: split every 512 tokens (matches most embedding models)
token_splitter = TokenTextSplitter(
    chunk_size=512,
    chunk_overlap=64,
    encoding_name="cl100k_base",   # tiktoken encoding for OpenAI/many models
)
token_chunks = token_splitter.split_text(raw_text)

print(f"Character-based: {len(char_chunks)} chunks")
print(f"Token-based:     {len(token_chunks)} chunks")
print(f"\nFirst token chunk ({len(token_chunks[0].split())} words):")
print(token_chunks[0][:300])

Understanding overlap:

Overlap re-includes the last N tokens of the previous chunk at the start of the next one. This reduces the chance of slicing an answer across a boundary — but it increases storage and embedding costs proportionally.

# Measuring overlap cost
base_chunks = len(token_chunks)
overlap_tokens = 64 * (base_chunks - 1)           # overlap added per boundary
total_tokens_stored = 512 * base_chunks            # approximate
print(f"Overlap adds ~{overlap_tokens} extra tokens ({overlap_tokens / total_tokens_stored:.0%} overhead)")

When fixed-size is appropriate:

Situation Verdict
Plain prose with no headers Acceptable — structure is linear anyway
FAQ documents with short Q&A pairs Poor — splits questions from answers
Legal text with numbered clauses Poor — severs clause relationships
Logs or append-only event streams Good — each event is atomic
Initial prototyping on any corpus Good enough — validates the rest of the pipeline fast
WARNING
Character vs. Token Counting
Most embedding models have token limits, not character limits. A 1 000-character chunk can be anywhere from 200 to 400 tokens depending on language and punctuation. Always use token-based splitting when you need precise control over what fits in an embedding model's context window.

3 Recursive Character Splitting — Respecting Natural Boundaries

LangChain's RecursiveCharacterTextSplitter is the most widely used general-purpose splitter. It attempts to split on a priority-ordered list of separators, falling back to harder splits only when softer ones would exceed the chunk size.

Default separator hierarchy:

["\n\n", "\n", " ", ""]

This means: try to split on paragraph breaks first; if a paragraph is too long, split on line breaks; if a line is too long, split on spaces; as a last resort, split mid-word.

Implementation:

from langchain.text_splitter import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,           # in tokens when using from_tiktoken_encoder
    chunk_overlap=64,
    length_function=len,      # swap for token counter below
)

# Preferred: use tiktoken for accurate token counting
from langchain.text_splitter import RecursiveCharacterTextSplitter
import tiktoken

enc = tiktoken.get_encoding("cl100k_base")

def token_len(text: str) -> int:
    return len(enc.encode(text))

splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,
    chunk_overlap=64,
    length_function=token_len,
    separators=["\n\n", "\n", ". ", " ", ""],
)

chunks = splitter.split_text(raw_text)

Customising separators for your document type:

# Markdown — respect heading hierarchy before paragraphs
md_splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,
    chunk_overlap=64,
    length_function=token_len,
    separators=["## ", "### ", "\n\n", "\n", " "],
)

# Python source code — respect function/class boundaries
code_splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,
    chunk_overlap=32,
    length_function=token_len,
    separators=["\nclass ", "\ndef ", "\n\n", "\n", " "],
)

Comparing fixed vs. recursive on the same document:

fixed = TokenTextSplitter(chunk_size=512, chunk_overlap=64)
recursive = RecursiveCharacterTextSplitter(
    chunk_size=512, chunk_overlap=64, length_function=token_len
)

fixed_chunks = fixed.split_text(raw_text)
recursive_chunks = recursive.split_text(raw_text)

# Measure how often chunks end mid-sentence
def ends_mid_sentence(chunk: str) -> bool:
    return not chunk.rstrip().endswith((".", "!", "?", "```", ":"))

fixed_broken = sum(1 for c in fixed_chunks if ends_mid_sentence(c))
recursive_broken = sum(1 for c in recursive_chunks if ends_mid_sentence(c))

print(f"Fixed — mid-sentence breaks: {fixed_broken}/{len(fixed_chunks)} "
      f"({fixed_broken/len(fixed_chunks):.0%})")
print(f"Recursive — mid-sentence breaks: {recursive_broken}/{len(recursive_chunks)} "
      f"({recursive_broken/len(recursive_chunks):.0%})")

On a typical technical document, recursive splitting reduces mid-sentence breaks by 60–80% versus fixed splitting with the same token budget.

TIP
Start Here for General Text
For most unstructured prose corpora, RecursiveCharacterTextSplitter with 512 tokens and 64-token overlap is a solid default. Switch to a more specialised strategy only after measuring that retrieval quality is actually suffering.

4 Semantic Chunking — Splitting at Topic Boundaries

Recursive splitting respects surface syntax (whitespace, punctuation) but is blind to meaning. Two sentences on completely different topics may end up in the same chunk simply because they appear adjacent. Semantic chunking solves this by measuring the embedding distance between consecutive sentences and inserting a boundary whenever the semantic similarity drops sharply — signalling a topic shift.

The algorithm:

  1. Split the document into individual sentences (using a sentence tokeniser).
  2. Embed each sentence (or a sliding window of sentences for stability).
  3. Compute cosine similarity between each adjacent pair.
  4. Insert a chunk boundary wherever similarity drops below a threshold — a local minimum in the similarity curve.
  5. Group sentences between boundaries into chunks; if a group still exceeds the max token budget, apply a secondary recursive split.

Implementation using LangChain's SemanticChunker:

from langchain_experimental.text_splitter import SemanticChunker
from langchain_community.embeddings import HuggingFaceEmbeddings

embed_model = HuggingFaceEmbeddings(
    model_name="BAAI/bge-small-en-v1.5",
    model_kwargs={"device": "cpu"},
    encode_kwargs={"normalize_embeddings": True},
)

# percentile: insert boundary when similarity is in the lowest N-th percentile
semantic_splitter = SemanticChunker(
    embeddings=embed_model,
    breakpoint_threshold_type="percentile",
    breakpoint_threshold_amount=80,   # 80th percentile drop = boundary
)

semantic_chunks = semantic_splitter.split_text(raw_text)

print(f"Semantic chunks: {len(semantic_chunks)}")
for i, chunk in enumerate(semantic_chunks[:3]):
    toks = token_len(chunk)
    print(f"  Chunk {i+1}: {toks} tokens — {chunk[:80].strip()!r}...")

Threshold strategies:

Strategy Parameter Effect
percentile 70–90 Boundaries where similarity is in the bottom N%; higher = more chunks
standard_deviation 1.0–2.0 Boundary when drop exceeds N std devs below mean
interquartile — Boundary below Q1 minus 1.5×IQR; robust to outliers
gradient — Rate of change rather than absolute level; catches rapid shifts

When semantic chunking wins and when it struggles:

WINS:
  Long mixed-topic documents (wikis, incident reports, research PDFs)
  Documents where paragraph breaks do not align with topic shifts

STRUGGLES:
  Very short documents (not enough signal for similarity curves)
  Highly technical text where adjacent sentences look dissimilar
  even when they're part of the same argument (low baseline similarity)
  Documents with dense cross-references (similarity stays high throughout)
NOTE
Embedding Cost at Indexing Time
Semantic chunking embeds every sentence in the document to find boundaries — typically 5–15x more embedding calls than the final chunk count. For a one-time index build this is fine; for continuous ingestion pipelines, factor the cost into your architecture. Pre-splitting by section heading and then applying semantic chunking within each section reduces this overhead.
WARNING
Same Embeddings, Same Boundaries
Use the same embedding model for semantic chunking and for query-time retrieval. If you chunk with bge-small-en-v1.5 but query with text-embedding-3-large, the chunk boundaries that looked natural in one embedding space may not correspond to retrieval-relevant units in the other.

5 Structure-Aware Chunking for Markdown, HTML, and Code

Many real document corpora are not plain prose — they are Markdown documentation, HTML pages, or source code. These formats encode structure explicitly. Ignoring that structure and falling back to character/token splitting throws away information that makes it trivial to find natural chunk boundaries.

Markdown-aware chunking:

LangChain provides MarkdownHeaderTextSplitter, which splits first at heading boundaries and then applies recursive splitting within each section. Heading levels become chunk metadata automatically.

from langchain.text_splitter import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter

headers_to_split_on = [
    ("#", "h1"),
    ("##", "h2"),
    ("###", "h3"),
]

md_header_splitter = MarkdownHeaderTextSplitter(
    headers_to_split_on=headers_to_split_on,
    strip_headers=False,    # keep header text in the chunk body
)

header_splits = md_header_splitter.split_text(raw_text)

# Each split now carries heading metadata:
# split.metadata == {"h1": "Installation", "h2": "Quick Start", "h3": "macOS"}

# Apply token-level splitting within each section
secondary_splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,
    chunk_overlap=64,
    length_function=token_len,
)

final_chunks = secondary_splitter.split_documents(header_splits)

print(f"Sections after header split: {len(header_splits)}")
print(f"Final chunks after token split: {len(final_chunks)}")
print(f"\nSample metadata: {final_chunks[2].metadata}")

HTML-aware chunking:

from langchain_community.document_transformers import Html2TextTransformer
from langchain_community.document_loaders import AsyncHtmlLoader

# Load HTML and convert to clean Markdown-like text
loader = AsyncHtmlLoader(["https://example.com/docs/api"])
html_docs = loader.load()

html2text = Html2TextTransformer()
docs_transformed = html2text.transform_documents(html_docs)

# Now apply recursive splitting to the cleaned text
chunks = secondary_splitter.split_documents(docs_transformed)

Code-aware chunking with tree-sitter:

For source code, respecting function and class boundaries matters far more than token counts. The unstructured library provides language-aware splitting:

from langchain.text_splitter import Language, RecursiveCharacterTextSplitter

python_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.PYTHON,
    chunk_size=512,
    chunk_overlap=32,
    length_function=token_len,
)

# Language.PYTHON uses these separators in order:
# ["\nclass ", "\ndef ", "\n\tdef ", "\n\n", "\n", " ", ""]

code_text = Path("src/retriever.py").read_text()
code_chunks = python_splitter.split_text(code_text)

# Supported languages: PYTHON, JS, TS, RUST, GO, JAVA, C, CPP, MARKDOWN, HTML, SOL, ...

Comparison of structure-aware vs. naive on a Markdown document:

Metric Recursive (naive) Markdown-header-aware
Chunks crossing section boundaries ~35% ~4%
Average heading context in chunk None Full h1/h2/h3 breadcrumb
Retrieval hit rate on section-specific queries Baseline +18–25% (typical)
TIP
Always Parse Structure First, Split Tokens Second
The two-pass approach — parse document structure into sections, then apply token splitting within sections — gives you the best of both worlds: clean semantic boundaries from the document's own structure, and predictable token budgets from the secondary split. This pattern works for Markdown, HTML, and any format with consistent heading conventions.

6 Metadata Enrichment — Making Chunks Self-Describing

A chunk without metadata is an orphan — the LLM can read it but cannot tell the user where it came from or filter it by provenance. Metadata enrichment attaches contextual information to every chunk at index time, enabling three things: better post-retrieval citations, metadata-filtered similarity search, and hybrid ranking signals.

Core metadata fields to attach to every chunk:

from dataclasses import dataclass, field
from typing import Any

@dataclass
class ChunkMetadata:
    source_path: str          # absolute or relative path to originating file
    source_url: str           # URL if loaded from the web
    doc_title: str            # document-level title
    section_heading: str      # most specific heading above this chunk
    chunk_index: int          # position within the document (0-based)
    total_chunks: int         # total chunks in this document
    char_start: int           # character offset in original document
    char_end: int
    ingested_at: str          # ISO-8601 timestamp
    doc_version: str          # e.g., git SHA or document modified date
    extra: dict[str, Any] = field(default_factory=dict)

Attaching metadata during chunking:

from langchain.schema import Document
from datetime import datetime, timezone
from pathlib import Path

def chunk_document_with_metadata(
    file_path: str,
    splitter,
    doc_title: str = "",
    doc_version: str = "",
) -> list[Document]:
    path = Path(file_path)
    raw_text = path.read_text()

    # Split into LangChain Documents
    raw_chunks = splitter.create_documents([raw_text])

    ingested_at = datetime.now(timezone.utc).isoformat()

    enriched = []
    for i, chunk in enumerate(raw_chunks):
        chunk.metadata.update({
            "source_path": str(path),
            "doc_title": doc_title or path.stem,
            "chunk_index": i,
            "total_chunks": len(raw_chunks),
            "ingested_at": ingested_at,
            "doc_version": doc_version,
        })
        enriched.append(chunk)
    return enriched

Contextual headers — prepending section context to every chunk:

One of the highest-leverage metadata enrichment techniques is prepending a brief context header to the chunk text itself before embedding. This means the embedding captures both the content and its provenance, dramatically improving retrieval for queries that reference section names.

def add_contextual_header(chunk: Document) -> Document:
    """Prepend doc title and section heading to chunk text before embedding."""
    meta = chunk.metadata
    header_parts = []
    if meta.get("doc_title"):
        header_parts.append(f"Document: {meta['doc_title']}")
    if meta.get("h1"):
        header_parts.append(f"Section: {meta['h1']}")
    if meta.get("h2"):
        header_parts.append(f"Subsection: {meta['h2']}")

    if header_parts:
        header = " | ".join(header_parts)
        enriched_text = f"[{header}]\n\n{chunk.page_content}"
        return Document(page_content=enriched_text, metadata=chunk.metadata)
    return chunk

# Apply before embedding
chunks_for_embedding = [add_contextual_header(c) for c in enriched_chunks]

Metadata-filtered retrieval in Chroma:

import chromadb
from chromadb.utils.embedding_functions import SentenceTransformerEmbeddingFunction

client = chromadb.PersistentClient(path="./chroma_db")
embed_fn = SentenceTransformerEmbeddingFunction(model_name="BAAI/bge-small-en-v1.5")

collection = client.get_or_create_collection(
    name="rag_chunks",
    embedding_function=embed_fn,
)

# Upsert chunks with full metadata
collection.upsert(
    ids=[f"chunk_{i}" for i in range(len(chunks_for_embedding))],
    documents=[c.page_content for c in chunks_for_embedding],
    metadatas=[c.metadata for c in chunks_for_embedding],
)

# Retrieval filtered by source document
results = collection.query(
    query_texts=["What is the retry limit?"],
    n_results=5,
    where={"source_path": {"$eq": "docs/manual.md"}},
)
NOTE
Metadata Is Cheap, Omitting It Is Expensive
Storing metadata adds negligible storage. Omitting it means you cannot do source-filtered retrieval, cannot generate citations the user can verify, and cannot diagnose which documents are causing retrieval failures. Always attach at minimum: source_path, doc_title, ingested_at, and chunk_index.
TIP
Contextual Headers and Retrieval Quality
Prepending [Document: X | Section: Y] before the chunk content before embedding is a low-cost technique that consistently improves retrieval on section-specific queries by 10–20 percentage points. The embedding captures both content and structure, so a query about a specific section surfaces chunks from that section even when the content words alone are ambiguous.

7 Measuring Retrieval Quality Across Chunking Strategies

Choosing a chunking strategy without measuring its effect is guesswork. This step walks through building a lightweight evaluation harness that compares strategies on the same corpus and query set.

Core retrieval metrics (defined from first principles):

  • Precision@k — of the top-k chunks retrieved, what fraction are relevant? Measures whether retrieved results are signal or noise.
  • Recall@k — of all relevant chunks in the corpus, what fraction appear in top-k? Measures whether relevant chunks are being found at all.
  • MRR (Mean Reciprocal Rank) — the mean of 1/rank of the first relevant chunk across queries. Measures how quickly a relevant result appears.
from sentence_transformers import SentenceTransformer, util
import numpy as np
from typing import Callable

# A golden eval set: list of (query, list_of_relevant_source_passages)
EVAL_SET = [
    {
        "query": "What is the maximum retry count?",
        "relevant_passages": [
            "The retry limit defaults to 3. When this limit is exceeded",
        ],
    },
    {
        "query": "How does the authentication token expire?",
        "relevant_passages": [
            "Tokens expire after 24 hours. After expiry, the client must",
            "The token expiry policy is configurable via AUTH_TOKEN_TTL",
        ],
    },
    # ... add 20-50 queries for meaningful signal
]

def precision_at_k(retrieved: list[str], relevant: list[str], k: int) -> float:
    top_k = retrieved[:k]
    hits = sum(
        1 for r in top_k
        if any(rel.lower() in r.lower() for rel in relevant)
    )
    return hits / k

def recall_at_k(retrieved: list[str], relevant: list[str], k: int) -> float:
    top_k = retrieved[:k]
    found = sum(
        1 for rel in relevant
        if any(rel.lower() in r.lower() for r in top_k)
    )
    return found / len(relevant) if relevant else 0.0

def mrr(retrieved: list[str], relevant: list[str]) -> float:
    for rank, r in enumerate(retrieved, start=1):
        if any(rel.lower() in r.lower() for rel in relevant):
            return 1.0 / rank
    return 0.0

def evaluate_chunking_strategy(
    strategy_name: str,
    collection,
    eval_set: list[dict],
    k: int = 5,
) -> dict:
    p_scores, r_scores, mrr_scores = [], [], []

    for item in eval_set:
        results = collection.query(
            query_texts=[item["query"]],
            n_results=k,
        )
        retrieved_docs = results["documents"][0]

        p_scores.append(precision_at_k(retrieved_docs, item["relevant_passages"], k))
        r_scores.append(recall_at_k(retrieved_docs, item["relevant_passages"], k))
        mrr_scores.append(mrr(retrieved_docs, item["relevant_passages"]))

    return {
        "strategy": strategy_name,
        "precision@5": np.mean(p_scores),
        "recall@5": np.mean(r_scores),
        "mrr": np.mean(mrr_scores),
    }

Running a head-to-head comparison:

import chromadb
from chromadb.utils.embedding_functions import SentenceTransformerEmbeddingFunction

client = chromadb.EphemeralClient()
embed_fn = SentenceTransformerEmbeddingFunction(model_name="BAAI/bge-small-en-v1.5")

strategies = {
    "fixed_512": TokenTextSplitter(chunk_size=512, chunk_overlap=0),
    "fixed_512_overlap": TokenTextSplitter(chunk_size=512, chunk_overlap=64),
    "recursive_512": RecursiveCharacterTextSplitter(
        chunk_size=512, chunk_overlap=64, length_function=token_len
    ),
    "recursive_256": RecursiveCharacterTextSplitter(
        chunk_size=256, chunk_overlap=32, length_function=token_len
    ),
    "md_header_aware": md_header_splitter,  # from Step 5
}

results_table = []
for name, splitter in strategies.items():
    # Build a fresh collection for each strategy
    coll = client.get_or_create_collection(name=name, embedding_function=embed_fn)

    chunks = splitter.split_text(raw_text)
    coll.upsert(
        ids=[f"{name}_{i}" for i in range(len(chunks))],
        documents=chunks,
    )

    metrics = evaluate_chunking_strategy(name, coll, EVAL_SET, k=5)
    results_table.append(metrics)
    print(f"{name:30s}  P@5={metrics['precision@5']:.3f}  "
          f"R@5={metrics['recall@5']:.3f}  MRR={metrics['mrr']:.3f}")

What typical results show:

The precise numbers depend on your corpus and queries, but patterns that appear consistently:

Observation Explanation
Overlap improves recall, slightly hurts precision Duplicate context from overlap gives another chance to surface the right passage, but also adds near-duplicate noise to top-k
Smaller chunks raise precision on fact-lookup queries One-sentence answers are not diluted by surrounding text
Structure-aware chunking dominates on section-specific queries Heading metadata boosts cosine similarity when query mentions a section name
Semantic chunking has high variance at small corpus sizes Threshold tuning matters; test at least three threshold levels
TIP
Build Your Eval Set Before Choosing a Strategy
Resist the urge to pick a chunking strategy and then design queries that play to its strengths. Build a representative set of 30–50 real queries from your actual use case first, label the relevant passages, then run every strategy through the same harness. The numbers tell you the truth; intuition often doesn't.

8 Choosing a Strategy — Decision Framework

The four strategies covered in this lesson are not mutually exclusive — a production pipeline often applies different strategies to different document types and then stores everything in the same vector collection (keyed by doc_type metadata).

Decision tree:

Is your corpus structured (Markdown, HTML, code with headers)?
├── YES → Use structure-aware chunking (Step 5) as the primary split,
│          then token-level secondary split within each section.
│          Add heading metadata to every chunk.
└── NO → Is it long-form prose with clear topic shifts?
         ├── YES → Try semantic chunking (Step 4).
         │          Tune threshold type and amount with your eval set.
         └── NO → Use RecursiveCharacterTextSplitter (Step 3).
                   Start at 512 tokens / 64 overlap and adjust
                   based on eval metrics.

Are your queries mostly fact-lookup (one-sentence answers)?
→ Skew smaller: 128–256 tokens

Are your queries mostly conceptual / multi-step?
→ Skew larger: 512–1024 tokens, or use parent-child retrieval (Lesson 6)

Does your LLM context window constrain you?
→ top_k × avg_chunk_tokens must fit comfortably; reduce chunk size or top_k

Quick-reference summary table:

Strategy Best for Weakness Typical chunk size
Fixed-size Logs, uniform streams, quick prototyping Blindly splits sentences and paragraphs 256–512 tokens
Recursive character General prose, mixed documents No semantic awareness 256–1024 tokens
Semantic Long mixed-topic documents Embedding cost; threshold tuning needed Variable
Structure-aware Markdown, HTML, source code Requires consistent formatting Variable by section

Putting it all together — a multi-strategy ingestion pipeline:

from pathlib import Path

STRATEGY_BY_EXT = {
    ".md": "markdown_header",
    ".py": "code_python",
    ".html": "html",
    ".txt": "recursive",
    ".pdf": "recursive",   # after PDF text extraction
}

def get_splitter(ext: str):
    if ext == ".md":
        return md_header_splitter        # from Step 5, secondary token split applied after
    elif ext == ".py":
        return RecursiveCharacterTextSplitter.from_language(
            Language.PYTHON, chunk_size=512, chunk_overlap=32, length_function=token_len
        )
    elif ext == ".html":
        return secondary_splitter        # after Html2TextTransformer
    else:
        return RecursiveCharacterTextSplitter(
            chunk_size=512, chunk_overlap=64, length_function=token_len
        )

def ingest_corpus(corpus_dir: str, collection) -> int:
    corpus = Path(corpus_dir)
    total = 0
    for doc_path in corpus.rglob("*"):
        if not doc_path.is_file():
            continue
        ext = doc_path.suffix.lower()
        if ext not in STRATEGY_BY_EXT:
            continue

        splitter = get_splitter(ext)
        chunks = chunk_document_with_metadata(
            str(doc_path),
            splitter,
            doc_version="v1.0",
        )
        chunks = [add_contextual_header(c) for c in chunks]

        collection.upsert(
            ids=[f"{doc_path.stem}_{c.metadata['chunk_index']}" for c in chunks],
            documents=[c.page_content for c in chunks],
            metadatas=[c.metadata for c in chunks],
        )
        total += len(chunks)
        print(f"  {doc_path.name}: {len(chunks)} chunks ({ext})")

    return total

total_chunks = ingest_corpus("./corpus", collection)
print(f"\nTotal chunks indexed: {total_chunks}")
NOTE
Revisit Chunking After Lesson 7
The evaluation harness you build in Lesson 7: Evaluation and Quality will give you the most rigorous signal about whether your chunking strategy is the bottleneck. Many teams discover that retrieval failures they blamed on embeddings or prompts were actually chunking problems. Measure before you optimise.
WARNING
Re-Chunking Means Re-Indexing
Changing your chunk size or strategy after you have already indexed a large corpus means discarding all existing embeddings and re-processing everything. This is expensive in time, API calls, and (if using a managed vector DB) write units. Design your chunking strategy carefully before committing to a large index build — prototype on a 5–10% sample first.

Questions & Answers

Q: I have a 10 000-document corpus. I tried 512-token chunks and retrieval quality is poor. Should I switch to semantic chunking across the whole corpus, or is there something cheaper to try first?
Start with the cheapest interventions first. First, add contextual headers (Step 6) — this often improves precision by 10–20 points at zero extra chunking cost. Second, check whether your corpus has consistent structure (Markdown headings, HTML sections) that you're currently ignoring; switching from recursive to structure-aware splitting is low cost and high reward for structured docs. Semantic chunking is the right next step only if those fail — it requires re-embedding every sentence in the corpus to find topic boundaries, which for 10 000 documents is a significant compute investment. Build a 500-document sample corpus, run your eval harness on it with all strategies, and let the numbers guide you before committing to a full re-index.
Q: My documents include tables. Chunking seems to split rows away from their header row, and the LLM can't interpret the retrieved table fragment. How do I handle this?
Tables are a known hard case for text-based chunking. The most robust approaches are: (1) detect table regions before chunking (the unstructured library can identify table elements in PDFs and HTML), extract each table as a single atomic unit, and never split it — instead serialise the table to Markdown or CSV and store it as one chunk regardless of token count; (2) add a prose summary of each table as a parallel chunk that describes what the table contains — the summary embeds and retrieves well, and you return the full table text alongside the summary. For very large tables that genuinely cannot fit in a single chunk, split column-wise (keep header row in every chunk) rather than row-wise. Multi-modal RAG for structured tables is covered in Lesson 10.
Q: Should I store the contextual header text in the chunk that gets shown to the LLM, or only in the version that gets embedded?
You have two options and both are legitimate. Option 1 — embed with the header, store and return the full text including the header: the LLM sees the header and can use it for citation, but the response may repeat the header verbatim. Option 2 — embed the header-enriched text, but store only the original content without the header in a separate display_text metadata field, and return that to the LLM: cleaner LLM output but requires your retrieval layer to return the right field. Most production systems use Option 1 for simplicity and prompt the LLM to use the header context for citations without repeating it. The important thing is that the header is present when embeddings are computed — it should never be added only at query time.
Q: What is the right overlap amount? I see 10%, 20%, and fixed values like 64 tokens recommended in different places.
There is no universally correct number. Overlap exists solely to reduce the penalty for answers that straddle a chunk boundary. The right amount is empirically the minimum that prevents measurable recall loss at your boundary rate. In practice: 10–15% of chunk size (so 50–75 tokens for a 512-token chunk) is a sensible default. Beyond 20–25%, you are mostly storing duplicate text that raises index size and embedding cost without proportional recall improvement. Measure recall@k on your eval set with overlap = 0, 32, 64, and 128 tokens; you typically see a diminishing-returns curve where the gains from 0→64 are large and 64→128 are small. Use the 64-token point unless your eval shows otherwise.
Q: My documents are versioned — the same article gets updated every month. Do I need to re-chunk the whole corpus on every update, or can I do incremental updates?
You can do incremental updates if you track document identity correctly. Store a doc_id (e.g., a stable slug or UUID for the article, independent of its version) and a doc_version (e.g., the last-modified timestamp or git SHA) in every chunk's metadata. On update: (1) query your vector DB for all chunks where doc_id equals the updated document's ID; (2) delete those chunk IDs; (3) re-chunk and re-embed the new version; (4) upsert the new chunks with the same doc_id and updated doc_version. This is strictly cheaper than a full re-index. The main risk is stale chunks surviving if your delete step fails — production strategies for handling this are covered in detail in Lesson 9: Production RAG.

Key Takeaways

  1. Chunking determines retrieval granularity — a chunk is the atomic unit of retrieval; boundaries that split answers across chunks directly reduce precision and recall, regardless of how good your embedding model is.
  2. Start recursive, evolve toward structure-aware — RecursiveCharacterTextSplitter is the right default for unstructured text; switch to Markdown/HTML/code-aware splitting as soon as your corpus has consistent structure, and apply token splitting as a secondary pass within sections.
  3. Size is a trade-off, not a setting — smaller chunks (128–256 tokens) favour fact-lookup precision; larger chunks (512–1024 tokens) favour contextual and multi-step queries; measure with your actual query distribution before deciding.
  4. Metadata enrichment multiplies chunk value — attaching source path, section heading, document title, and timestamp to every chunk enables citation, filtered retrieval, and incremental updates; contextual headers prepended before embedding improve retrieval accuracy for section-specific queries at near-zero cost.
  5. Measure before you commit — build a 30–50 query golden eval set and compute precision@k, recall@k, and MRR across strategies before indexing a large corpus; changing strategy after a large index build means re-embedding everything from scratch.
  6. Multi-strategy pipelines are normal — Markdown files, Python source, and PDF prose in the same corpus warrant different splitters; route by file extension in your ingestion pipeline and tag every chunk with its doc_type so you can diagnose retrieval failures by document type.

Next Steps: Lesson 5: Building a Basic RAG Pipeline