Chunking Strategies
Learning Outcomes
- Explain why chunking decisions affect retrieval quality more than almost any other pipeline choice
- Implement fixed-size, recursive, semantic, and structure-aware chunking strategies in Python
- Evaluate the precision and recall trade-offs produced by different chunk sizes and overlap settings
- Enrich chunks with metadata — source path, section title, parent context — to improve post-retrieval grounding
- Select the right chunking strategy for a given document type and query workload
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why chunking is the highest-leverage pipeline decision |
| Explain | 7 min | Mental model: retrieval granularity vs. context completeness |
| Demo | 8 min | Fixed-size chunking and its failure modes |
| Demo | 10 min | Recursive and semantic chunking |
| Demo | 8 min | Structure-aware chunking for HTML, Markdown, code |
| Demo | 8 min | Metadata enrichment and contextual headers |
| Lab | 4 min | Measuring retrieval quality across strategies |
| Wrap-up | 2 min | Decision framework and preview of Lesson 5 |
Before You Begin
Pre-work:
- Complete Lesson 2: Embeddings Explained — you need to understand cosine similarity before measuring retrieval quality
- Complete Lesson 3: Vector Databases — we use Chroma to store and query chunks throughout this lesson
- If you took the Agentic AI course, the concept of tool-mediated retrieval covered there connects directly to the retrieval step here
Shopping List:
- Python 3.10+ with a virtual environment
pip install chromadb langchain langchain-community sentence-transformers tiktoken- A handful of sample documents — a PDF technical manual, a Markdown README, and an HTML page work well; we provide a sample corpus below
- Optional:
pip install unstructuredfor advanced structure detection
Before writing a single line of chunking code, it is worth understanding the fundamental tension you are navigating.
The retrieval granularity problem:
A RAG system embeds each chunk independently and retrieves the top-k chunks most similar to the query embedding. This means the chunk is the atomic unit of retrieval. If a chunk is too large, one query can only surface it when a majority of its sentences are relevant — meaning many relevant passages are buried inside chunks that score poorly. If a chunk is too small, retrieved text may lack the surrounding context the LLM needs to produce a coherent answer.
This tension does not resolve neatly. It shifts depending on the query pattern:
| Query type | Prefers | Because |
|---|---|---|
| Fact lookup ("What is the max retry count?") | Small chunks (128–256 tokens) | One sentence holds the answer; large chunks dilute the signal |
| Conceptual ("Explain the authentication flow") | Larger chunks (512–1024 tokens) | Answer spans multiple paragraphs; small chunks fracture the explanation |
| Comparative ("How do approaches A and B differ?") | Carefully positioned chunks | Both halves of the comparison must be retrievable |
Most real query workloads are mixed, which is why advanced strategies (parent-child retrieval, covered in Lesson 6) often outperform any single fixed strategy.
Two concrete failure modes to recognise:
DOCUMENT EXCERPT:
"The retry limit defaults to 3. When this limit is exceeded,
the system raises a MaxRetriesExceeded exception and logs
the event to the audit trail."
QUERY: "What exception is raised on too many retries?"
--- FAILURE: chunks split at sentence boundary ---
Chunk A: "The retry limit defaults to 3."
Chunk B: "When this limit is exceeded, the system raises a
MaxRetriesExceeded exception and logs the event
to the audit trail."
Chunk A scores high (contains "retry" and "limit") but
does NOT contain the exception name.
Chunk B scores lower (no "retry limit" context) — it may
not even appear in top-k retrieval.
--- SUCCESS: boundary respects sentence grouping ---
Chunk: "The retry limit defaults to 3. When this limit is
exceeded, the system raises a MaxRetriesExceeded
exception and logs the event to the audit trail."
A single retrieval surfaces both the condition and the answer.
Fixed-size chunking splits text every N tokens or characters, with an optional overlap. It is the simplest strategy and a useful baseline, but it is blind to document structure.
Implementing fixed-size chunking with LangChain:
from langchain.text_splitter import CharacterTextSplitter, TokenTextSplitter
from pathlib import Path
raw_text = Path("docs/manual.md").read_text()
# Character-based: split every 1000 chars, 200-char overlap
char_splitter = CharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
separator="\n", # prefer splitting on newlines
length_function=len,
)
char_chunks = char_splitter.split_text(raw_text)
# Token-based: split every 512 tokens (matches most embedding models)
token_splitter = TokenTextSplitter(
chunk_size=512,
chunk_overlap=64,
encoding_name="cl100k_base", # tiktoken encoding for OpenAI/many models
)
token_chunks = token_splitter.split_text(raw_text)
print(f"Character-based: {len(char_chunks)} chunks")
print(f"Token-based: {len(token_chunks)} chunks")
print(f"\nFirst token chunk ({len(token_chunks[0].split())} words):")
print(token_chunks[0][:300])
Understanding overlap:
Overlap re-includes the last N tokens of the previous chunk at the start of the next one. This reduces the chance of slicing an answer across a boundary — but it increases storage and embedding costs proportionally.
# Measuring overlap cost
base_chunks = len(token_chunks)
overlap_tokens = 64 * (base_chunks - 1) # overlap added per boundary
total_tokens_stored = 512 * base_chunks # approximate
print(f"Overlap adds ~{overlap_tokens} extra tokens ({overlap_tokens / total_tokens_stored:.0%} overhead)")
When fixed-size is appropriate:
| Situation | Verdict |
|---|---|
| Plain prose with no headers | Acceptable — structure is linear anyway |
| FAQ documents with short Q&A pairs | Poor — splits questions from answers |
| Legal text with numbered clauses | Poor — severs clause relationships |
| Logs or append-only event streams | Good — each event is atomic |
| Initial prototyping on any corpus | Good enough — validates the rest of the pipeline fast |
LangChain's RecursiveCharacterTextSplitter is the most widely used general-purpose splitter. It attempts to split on a priority-ordered list of separators, falling back to harder splits only when softer ones would exceed the chunk size.
Default separator hierarchy:
["\n\n", "\n", " ", ""]
This means: try to split on paragraph breaks first; if a paragraph is too long, split on line breaks; if a line is too long, split on spaces; as a last resort, split mid-word.
Implementation:
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=512, # in tokens when using from_tiktoken_encoder
chunk_overlap=64,
length_function=len, # swap for token counter below
)
# Preferred: use tiktoken for accurate token counting
from langchain.text_splitter import RecursiveCharacterTextSplitter
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
def token_len(text: str) -> int:
return len(enc.encode(text))
splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
length_function=token_len,
separators=["\n\n", "\n", ". ", " ", ""],
)
chunks = splitter.split_text(raw_text)
Customising separators for your document type:
# Markdown — respect heading hierarchy before paragraphs
md_splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
length_function=token_len,
separators=["## ", "### ", "\n\n", "\n", " "],
)
# Python source code — respect function/class boundaries
code_splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=32,
length_function=token_len,
separators=["\nclass ", "\ndef ", "\n\n", "\n", " "],
)
Comparing fixed vs. recursive on the same document:
fixed = TokenTextSplitter(chunk_size=512, chunk_overlap=64)
recursive = RecursiveCharacterTextSplitter(
chunk_size=512, chunk_overlap=64, length_function=token_len
)
fixed_chunks = fixed.split_text(raw_text)
recursive_chunks = recursive.split_text(raw_text)
# Measure how often chunks end mid-sentence
def ends_mid_sentence(chunk: str) -> bool:
return not chunk.rstrip().endswith((".", "!", "?", "```", ":"))
fixed_broken = sum(1 for c in fixed_chunks if ends_mid_sentence(c))
recursive_broken = sum(1 for c in recursive_chunks if ends_mid_sentence(c))
print(f"Fixed — mid-sentence breaks: {fixed_broken}/{len(fixed_chunks)} "
f"({fixed_broken/len(fixed_chunks):.0%})")
print(f"Recursive — mid-sentence breaks: {recursive_broken}/{len(recursive_chunks)} "
f"({recursive_broken/len(recursive_chunks):.0%})")
On a typical technical document, recursive splitting reduces mid-sentence breaks by 60–80% versus fixed splitting with the same token budget.
RecursiveCharacterTextSplitter with 512 tokens and 64-token overlap is a solid default. Switch to a more specialised strategy only after measuring that retrieval quality is actually suffering.Recursive splitting respects surface syntax (whitespace, punctuation) but is blind to meaning. Two sentences on completely different topics may end up in the same chunk simply because they appear adjacent. Semantic chunking solves this by measuring the embedding distance between consecutive sentences and inserting a boundary whenever the semantic similarity drops sharply — signalling a topic shift.
The algorithm:
- Split the document into individual sentences (using a sentence tokeniser).
- Embed each sentence (or a sliding window of sentences for stability).
- Compute cosine similarity between each adjacent pair.
- Insert a chunk boundary wherever similarity drops below a threshold — a local minimum in the similarity curve.
- Group sentences between boundaries into chunks; if a group still exceeds the max token budget, apply a secondary recursive split.
Implementation using LangChain's SemanticChunker:
from langchain_experimental.text_splitter import SemanticChunker
from langchain_community.embeddings import HuggingFaceEmbeddings
embed_model = HuggingFaceEmbeddings(
model_name="BAAI/bge-small-en-v1.5",
model_kwargs={"device": "cpu"},
encode_kwargs={"normalize_embeddings": True},
)
# percentile: insert boundary when similarity is in the lowest N-th percentile
semantic_splitter = SemanticChunker(
embeddings=embed_model,
breakpoint_threshold_type="percentile",
breakpoint_threshold_amount=80, # 80th percentile drop = boundary
)
semantic_chunks = semantic_splitter.split_text(raw_text)
print(f"Semantic chunks: {len(semantic_chunks)}")
for i, chunk in enumerate(semantic_chunks[:3]):
toks = token_len(chunk)
print(f" Chunk {i+1}: {toks} tokens — {chunk[:80].strip()!r}...")
Threshold strategies:
| Strategy | Parameter | Effect |
|---|---|---|
percentile |
70–90 | Boundaries where similarity is in the bottom N%; higher = more chunks |
standard_deviation |
1.0–2.0 | Boundary when drop exceeds N std devs below mean |
interquartile |
— | Boundary below Q1 minus 1.5×IQR; robust to outliers |
gradient |
— | Rate of change rather than absolute level; catches rapid shifts |
When semantic chunking wins and when it struggles:
WINS:
Long mixed-topic documents (wikis, incident reports, research PDFs)
Documents where paragraph breaks do not align with topic shifts
STRUGGLES:
Very short documents (not enough signal for similarity curves)
Highly technical text where adjacent sentences look dissimilar
even when they're part of the same argument (low baseline similarity)
Documents with dense cross-references (similarity stays high throughout)
bge-small-en-v1.5 but query with text-embedding-3-large, the chunk boundaries that looked natural in one embedding space may not correspond to retrieval-relevant units in the other.Many real document corpora are not plain prose — they are Markdown documentation, HTML pages, or source code. These formats encode structure explicitly. Ignoring that structure and falling back to character/token splitting throws away information that makes it trivial to find natural chunk boundaries.
Markdown-aware chunking:
LangChain provides MarkdownHeaderTextSplitter, which splits first at heading boundaries and then applies recursive splitting within each section. Heading levels become chunk metadata automatically.
from langchain.text_splitter import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
headers_to_split_on = [
("#", "h1"),
("##", "h2"),
("###", "h3"),
]
md_header_splitter = MarkdownHeaderTextSplitter(
headers_to_split_on=headers_to_split_on,
strip_headers=False, # keep header text in the chunk body
)
header_splits = md_header_splitter.split_text(raw_text)
# Each split now carries heading metadata:
# split.metadata == {"h1": "Installation", "h2": "Quick Start", "h3": "macOS"}
# Apply token-level splitting within each section
secondary_splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
length_function=token_len,
)
final_chunks = secondary_splitter.split_documents(header_splits)
print(f"Sections after header split: {len(header_splits)}")
print(f"Final chunks after token split: {len(final_chunks)}")
print(f"\nSample metadata: {final_chunks[2].metadata}")
HTML-aware chunking:
from langchain_community.document_transformers import Html2TextTransformer
from langchain_community.document_loaders import AsyncHtmlLoader
# Load HTML and convert to clean Markdown-like text
loader = AsyncHtmlLoader(["https://example.com/docs/api"])
html_docs = loader.load()
html2text = Html2TextTransformer()
docs_transformed = html2text.transform_documents(html_docs)
# Now apply recursive splitting to the cleaned text
chunks = secondary_splitter.split_documents(docs_transformed)
Code-aware chunking with tree-sitter:
For source code, respecting function and class boundaries matters far more than token counts. The unstructured library provides language-aware splitting:
from langchain.text_splitter import Language, RecursiveCharacterTextSplitter
python_splitter = RecursiveCharacterTextSplitter.from_language(
language=Language.PYTHON,
chunk_size=512,
chunk_overlap=32,
length_function=token_len,
)
# Language.PYTHON uses these separators in order:
# ["\nclass ", "\ndef ", "\n\tdef ", "\n\n", "\n", " ", ""]
code_text = Path("src/retriever.py").read_text()
code_chunks = python_splitter.split_text(code_text)
# Supported languages: PYTHON, JS, TS, RUST, GO, JAVA, C, CPP, MARKDOWN, HTML, SOL, ...
Comparison of structure-aware vs. naive on a Markdown document:
| Metric | Recursive (naive) | Markdown-header-aware |
|---|---|---|
| Chunks crossing section boundaries | ~35% | ~4% |
| Average heading context in chunk | None | Full h1/h2/h3 breadcrumb |
| Retrieval hit rate on section-specific queries | Baseline | +18–25% (typical) |
A chunk without metadata is an orphan — the LLM can read it but cannot tell the user where it came from or filter it by provenance. Metadata enrichment attaches contextual information to every chunk at index time, enabling three things: better post-retrieval citations, metadata-filtered similarity search, and hybrid ranking signals.
Core metadata fields to attach to every chunk:
from dataclasses import dataclass, field
from typing import Any
@dataclass
class ChunkMetadata:
source_path: str # absolute or relative path to originating file
source_url: str # URL if loaded from the web
doc_title: str # document-level title
section_heading: str # most specific heading above this chunk
chunk_index: int # position within the document (0-based)
total_chunks: int # total chunks in this document
char_start: int # character offset in original document
char_end: int
ingested_at: str # ISO-8601 timestamp
doc_version: str # e.g., git SHA or document modified date
extra: dict[str, Any] = field(default_factory=dict)
Attaching metadata during chunking:
from langchain.schema import Document
from datetime import datetime, timezone
from pathlib import Path
def chunk_document_with_metadata(
file_path: str,
splitter,
doc_title: str = "",
doc_version: str = "",
) -> list[Document]:
path = Path(file_path)
raw_text = path.read_text()
# Split into LangChain Documents
raw_chunks = splitter.create_documents([raw_text])
ingested_at = datetime.now(timezone.utc).isoformat()
enriched = []
for i, chunk in enumerate(raw_chunks):
chunk.metadata.update({
"source_path": str(path),
"doc_title": doc_title or path.stem,
"chunk_index": i,
"total_chunks": len(raw_chunks),
"ingested_at": ingested_at,
"doc_version": doc_version,
})
enriched.append(chunk)
return enriched
Contextual headers — prepending section context to every chunk:
One of the highest-leverage metadata enrichment techniques is prepending a brief context header to the chunk text itself before embedding. This means the embedding captures both the content and its provenance, dramatically improving retrieval for queries that reference section names.
def add_contextual_header(chunk: Document) -> Document:
"""Prepend doc title and section heading to chunk text before embedding."""
meta = chunk.metadata
header_parts = []
if meta.get("doc_title"):
header_parts.append(f"Document: {meta['doc_title']}")
if meta.get("h1"):
header_parts.append(f"Section: {meta['h1']}")
if meta.get("h2"):
header_parts.append(f"Subsection: {meta['h2']}")
if header_parts:
header = " | ".join(header_parts)
enriched_text = f"[{header}]\n\n{chunk.page_content}"
return Document(page_content=enriched_text, metadata=chunk.metadata)
return chunk
# Apply before embedding
chunks_for_embedding = [add_contextual_header(c) for c in enriched_chunks]
Metadata-filtered retrieval in Chroma:
import chromadb
from chromadb.utils.embedding_functions import SentenceTransformerEmbeddingFunction
client = chromadb.PersistentClient(path="./chroma_db")
embed_fn = SentenceTransformerEmbeddingFunction(model_name="BAAI/bge-small-en-v1.5")
collection = client.get_or_create_collection(
name="rag_chunks",
embedding_function=embed_fn,
)
# Upsert chunks with full metadata
collection.upsert(
ids=[f"chunk_{i}" for i in range(len(chunks_for_embedding))],
documents=[c.page_content for c in chunks_for_embedding],
metadatas=[c.metadata for c in chunks_for_embedding],
)
# Retrieval filtered by source document
results = collection.query(
query_texts=["What is the retry limit?"],
n_results=5,
where={"source_path": {"$eq": "docs/manual.md"}},
)
source_path, doc_title, ingested_at, and chunk_index.[Document: X | Section: Y] before the chunk content before embedding is a low-cost technique that consistently improves retrieval on section-specific queries by 10–20 percentage points. The embedding captures both content and structure, so a query about a specific section surfaces chunks from that section even when the content words alone are ambiguous.Choosing a chunking strategy without measuring its effect is guesswork. This step walks through building a lightweight evaluation harness that compares strategies on the same corpus and query set.
Core retrieval metrics (defined from first principles):
- Precision@k — of the top-k chunks retrieved, what fraction are relevant? Measures whether retrieved results are signal or noise.
- Recall@k — of all relevant chunks in the corpus, what fraction appear in top-k? Measures whether relevant chunks are being found at all.
- MRR (Mean Reciprocal Rank) — the mean of 1/rank of the first relevant chunk across queries. Measures how quickly a relevant result appears.
from sentence_transformers import SentenceTransformer, util
import numpy as np
from typing import Callable
# A golden eval set: list of (query, list_of_relevant_source_passages)
EVAL_SET = [
{
"query": "What is the maximum retry count?",
"relevant_passages": [
"The retry limit defaults to 3. When this limit is exceeded",
],
},
{
"query": "How does the authentication token expire?",
"relevant_passages": [
"Tokens expire after 24 hours. After expiry, the client must",
"The token expiry policy is configurable via AUTH_TOKEN_TTL",
],
},
# ... add 20-50 queries for meaningful signal
]
def precision_at_k(retrieved: list[str], relevant: list[str], k: int) -> float:
top_k = retrieved[:k]
hits = sum(
1 for r in top_k
if any(rel.lower() in r.lower() for rel in relevant)
)
return hits / k
def recall_at_k(retrieved: list[str], relevant: list[str], k: int) -> float:
top_k = retrieved[:k]
found = sum(
1 for rel in relevant
if any(rel.lower() in r.lower() for r in top_k)
)
return found / len(relevant) if relevant else 0.0
def mrr(retrieved: list[str], relevant: list[str]) -> float:
for rank, r in enumerate(retrieved, start=1):
if any(rel.lower() in r.lower() for rel in relevant):
return 1.0 / rank
return 0.0
def evaluate_chunking_strategy(
strategy_name: str,
collection,
eval_set: list[dict],
k: int = 5,
) -> dict:
p_scores, r_scores, mrr_scores = [], [], []
for item in eval_set:
results = collection.query(
query_texts=[item["query"]],
n_results=k,
)
retrieved_docs = results["documents"][0]
p_scores.append(precision_at_k(retrieved_docs, item["relevant_passages"], k))
r_scores.append(recall_at_k(retrieved_docs, item["relevant_passages"], k))
mrr_scores.append(mrr(retrieved_docs, item["relevant_passages"]))
return {
"strategy": strategy_name,
"precision@5": np.mean(p_scores),
"recall@5": np.mean(r_scores),
"mrr": np.mean(mrr_scores),
}
Running a head-to-head comparison:
import chromadb
from chromadb.utils.embedding_functions import SentenceTransformerEmbeddingFunction
client = chromadb.EphemeralClient()
embed_fn = SentenceTransformerEmbeddingFunction(model_name="BAAI/bge-small-en-v1.5")
strategies = {
"fixed_512": TokenTextSplitter(chunk_size=512, chunk_overlap=0),
"fixed_512_overlap": TokenTextSplitter(chunk_size=512, chunk_overlap=64),
"recursive_512": RecursiveCharacterTextSplitter(
chunk_size=512, chunk_overlap=64, length_function=token_len
),
"recursive_256": RecursiveCharacterTextSplitter(
chunk_size=256, chunk_overlap=32, length_function=token_len
),
"md_header_aware": md_header_splitter, # from Step 5
}
results_table = []
for name, splitter in strategies.items():
# Build a fresh collection for each strategy
coll = client.get_or_create_collection(name=name, embedding_function=embed_fn)
chunks = splitter.split_text(raw_text)
coll.upsert(
ids=[f"{name}_{i}" for i in range(len(chunks))],
documents=chunks,
)
metrics = evaluate_chunking_strategy(name, coll, EVAL_SET, k=5)
results_table.append(metrics)
print(f"{name:30s} P@5={metrics['precision@5']:.3f} "
f"R@5={metrics['recall@5']:.3f} MRR={metrics['mrr']:.3f}")
What typical results show:
The precise numbers depend on your corpus and queries, but patterns that appear consistently:
| Observation | Explanation |
|---|---|
| Overlap improves recall, slightly hurts precision | Duplicate context from overlap gives another chance to surface the right passage, but also adds near-duplicate noise to top-k |
| Smaller chunks raise precision on fact-lookup queries | One-sentence answers are not diluted by surrounding text |
| Structure-aware chunking dominates on section-specific queries | Heading metadata boosts cosine similarity when query mentions a section name |
| Semantic chunking has high variance at small corpus sizes | Threshold tuning matters; test at least three threshold levels |
The four strategies covered in this lesson are not mutually exclusive — a production pipeline often applies different strategies to different document types and then stores everything in the same vector collection (keyed by doc_type metadata).
Decision tree:
Is your corpus structured (Markdown, HTML, code with headers)?
├── YES → Use structure-aware chunking (Step 5) as the primary split,
│ then token-level secondary split within each section.
│ Add heading metadata to every chunk.
└── NO → Is it long-form prose with clear topic shifts?
├── YES → Try semantic chunking (Step 4).
│ Tune threshold type and amount with your eval set.
└── NO → Use RecursiveCharacterTextSplitter (Step 3).
Start at 512 tokens / 64 overlap and adjust
based on eval metrics.
Are your queries mostly fact-lookup (one-sentence answers)?
→ Skew smaller: 128–256 tokens
Are your queries mostly conceptual / multi-step?
→ Skew larger: 512–1024 tokens, or use parent-child retrieval (Lesson 6)
Does your LLM context window constrain you?
→ top_k × avg_chunk_tokens must fit comfortably; reduce chunk size or top_k
Quick-reference summary table:
| Strategy | Best for | Weakness | Typical chunk size |
|---|---|---|---|
| Fixed-size | Logs, uniform streams, quick prototyping | Blindly splits sentences and paragraphs | 256–512 tokens |
| Recursive character | General prose, mixed documents | No semantic awareness | 256–1024 tokens |
| Semantic | Long mixed-topic documents | Embedding cost; threshold tuning needed | Variable |
| Structure-aware | Markdown, HTML, source code | Requires consistent formatting | Variable by section |
Putting it all together — a multi-strategy ingestion pipeline:
from pathlib import Path
STRATEGY_BY_EXT = {
".md": "markdown_header",
".py": "code_python",
".html": "html",
".txt": "recursive",
".pdf": "recursive", # after PDF text extraction
}
def get_splitter(ext: str):
if ext == ".md":
return md_header_splitter # from Step 5, secondary token split applied after
elif ext == ".py":
return RecursiveCharacterTextSplitter.from_language(
Language.PYTHON, chunk_size=512, chunk_overlap=32, length_function=token_len
)
elif ext == ".html":
return secondary_splitter # after Html2TextTransformer
else:
return RecursiveCharacterTextSplitter(
chunk_size=512, chunk_overlap=64, length_function=token_len
)
def ingest_corpus(corpus_dir: str, collection) -> int:
corpus = Path(corpus_dir)
total = 0
for doc_path in corpus.rglob("*"):
if not doc_path.is_file():
continue
ext = doc_path.suffix.lower()
if ext not in STRATEGY_BY_EXT:
continue
splitter = get_splitter(ext)
chunks = chunk_document_with_metadata(
str(doc_path),
splitter,
doc_version="v1.0",
)
chunks = [add_contextual_header(c) for c in chunks]
collection.upsert(
ids=[f"{doc_path.stem}_{c.metadata['chunk_index']}" for c in chunks],
documents=[c.page_content for c in chunks],
metadatas=[c.metadata for c in chunks],
)
total += len(chunks)
print(f" {doc_path.name}: {len(chunks)} chunks ({ext})")
return total
total_chunks = ingest_corpus("./corpus", collection)
print(f"\nTotal chunks indexed: {total_chunks}")
Questions & Answers
unstructured library can identify table elements in PDFs and HTML), extract each table as a single atomic unit, and never split it — instead serialise the table to Markdown or CSV and store it as one chunk regardless of token count; (2) add a prose summary of each table as a parallel chunk that describes what the table contains — the summary embeds and retrieves well, and you return the full table text alongside the summary. For very large tables that genuinely cannot fit in a single chunk, split column-wise (keep header row in every chunk) rather than row-wise. Multi-modal RAG for structured tables is covered in Lesson 10.display_text metadata field, and return that to the LLM: cleaner LLM output but requires your retrieval layer to return the right field. Most production systems use Option 1 for simplicity and prompt the LLM to use the header context for citations without repeating it. The important thing is that the header is present when embeddings are computed — it should never be added only at query time.doc_id (e.g., a stable slug or UUID for the article, independent of its version) and a doc_version (e.g., the last-modified timestamp or git SHA) in every chunk's metadata. On update: (1) query your vector DB for all chunks where doc_id equals the updated document's ID; (2) delete those chunk IDs; (3) re-chunk and re-embed the new version; (4) upsert the new chunks with the same doc_id and updated doc_version. This is strictly cheaper than a full re-index. The main risk is stale chunks surviving if your delete step fails — production strategies for handling this are covered in detail in Lesson 9: Production RAG.Key Takeaways
- Chunking determines retrieval granularity — a chunk is the atomic unit of retrieval; boundaries that split answers across chunks directly reduce precision and recall, regardless of how good your embedding model is.
- Start recursive, evolve toward structure-aware —
RecursiveCharacterTextSplitteris the right default for unstructured text; switch to Markdown/HTML/code-aware splitting as soon as your corpus has consistent structure, and apply token splitting as a secondary pass within sections. - Size is a trade-off, not a setting — smaller chunks (128–256 tokens) favour fact-lookup precision; larger chunks (512–1024 tokens) favour contextual and multi-step queries; measure with your actual query distribution before deciding.
- Metadata enrichment multiplies chunk value — attaching source path, section heading, document title, and timestamp to every chunk enables citation, filtered retrieval, and incremental updates; contextual headers prepended before embedding improve retrieval accuracy for section-specific queries at near-zero cost.
- Measure before you commit — build a 30–50 query golden eval set and compute precision@k, recall@k, and MRR across strategies before indexing a large corpus; changing strategy after a large index build means re-embedding everything from scratch.
- Multi-strategy pipelines are normal — Markdown files, Python source, and PDF prose in the same corpus warrant different splitters; route by file extension in your ingestion pipeline and tag every chunk with its
doc_typeso you can diagnose retrieval failures by document type.
Next Steps: Lesson 5: Building a Basic RAG Pipeline