Local RAG — Your Own Knowledge Base
Learning Outcomes
- Explain the RAG pipeline end to end — ingest, chunk, embed, store, retrieve, generate — and why each stage matters
- Ingest mixed documents (PDF, Markdown, code) and chunk them with size and overlap tuned for retrieval
- Generate embeddings entirely locally using an Ollama embedding model
- Stand up a vector store with both ChromaDB (embedded) and Qdrant (Docker) and load your vectors
- Wire retrieval into a local chat model and measure answer quality with and without context
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What RAG is and why local RAG is worth the effort |
| Explain | 7 min | The pipeline: ingest, chunk, embed, store, retrieve, generate |
| Build | 8 min | Ingestion and chunking strategy |
| Build | 8 min | Local embeddings with Ollama |
| Build | 8 min | ChromaDB — the embedded vector store |
| Build | 8 min | Qdrant — the production-grade alternative |
| Build | 9 min | Retrieval + generation with a local model |
| Wrap-up | 4 min | Evaluation, gotchas, and key takeaways |
Before You Begin
Pre-work:
- Have Ollama installed and a chat model pulled — see Lesson 3: Ollama
- Be comfortable choosing a model — see Lesson 4: Model Selection
- Know how to drive models from the shell — see Lesson 5: Running Models from the CLI
- Know whether RAG or fine-tuning fits your problem — see Lesson 7: Fine-Tuning Basics
Shopping List:
- Ollama running locally (
ollama serveactive, default port 11434) - Python 3.10+ and a virtual environment
- Docker (only for the Qdrant step)
- A folder of your own documents to index — PDFs, Markdown notes, or a code repo
- Roughly 1 GB of free disk: ~274 MB for the embedding model, the rest for a chat model and the vector index
RAG (Retrieval-Augmented Generation) fetches the most relevant snippets from your own documents and pastes them into the prompt before the model answers. The model then answers grounded in your data, not just what it memorised in training. Doing it locally matters: the alternative is uploading your contracts, source code, or notes to a hosted embedding API — which defeats the point of local models.
The pipeline has six stages, and every later step maps to one of them:
documents ──► ingest ──► chunk ──► embed ──► store (vector DB)
│
query ──► embed ──► similarity search ────────────┘
│
top-k chunks ──► prompt ──► local LLM ──► answer
A few terms, defined once:
| Term | Plain-English meaning |
|---|---|
| Embedding | A vector of floats capturing the meaning of text, so similar text lands nearby |
| Chunk | A bite-sized slice of a document — what you embed and retrieve |
| Vector store | A database that finds the nearest vectors to a query vector, fast |
| top-k | Chunks retrieved per query (commonly 3–8) |
Create an isolated environment and install the libraries we will use across the lesson:
On macOS and Linux:
python3 -m venv rag-env
source rag-env/bin/activate
pip install ollama chromadb qdrant-client pypdf
On Windows (PowerShell):
python -m venv rag-env
rag-env\Scripts\Activate.ps1
pip install ollama chromadb qdrant-client pypdf
Now pull a dedicated embedding model — separate from your chat model, it only turns text into vectors. We will use nomic-embed-text: small (~274 MB), fast on CPU, 768-dimensional vectors:
ollama pull nomic-embed-text
Confirm Ollama's embeddings endpoint responds. It exposes embeddings at /api/embed on port 11434:
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": "Llamas are members of the camelid family"
}'
You should get back a JSON object containing an embeddings array of 768 floats.
nomic-embed-text emits 768 dimensions, mxbai-embed-large emits 1024. Your vector store collection MUST be created with the exact same size, or inserts will fail. Keep the embedding model and the chat model separate too: the embedder runs on every chunk and query, so favour something small and quick.Retrieval quality is decided here, not at the model. Chunks too big bury the relevant sentence in noise; too small and you lose context. A solid prose default is ~800 characters per chunk with ~100 characters of overlap, so ideas straddling a boundary are not lost.
Create ingest.py to load PDFs, Markdown, and code, then split into overlapping chunks:
import os
from pypdf import PdfReader
def load_file(path):
if path.endswith(".pdf"):
reader = PdfReader(path)
return "\n".join(page.extract_text() or "" for page in reader.pages)
with open(path, encoding="utf-8", errors="ignore") as f:
return f.read()
def chunk_text(text, size=800, overlap=100):
chunks, start = [], 0
while start < len(text):
end = start + size
chunk = text[start:end].strip()
if chunk:
chunks.append(chunk)
start += size - overlap # slide forward, keeping an overlap
return chunks
def load_corpus(root):
docs = []
for dirpath, _, files in os.walk(root):
for name in files:
if name.endswith((".pdf", ".md", ".txt", ".py", ".js")):
path = os.path.join(dirpath, name)
for i, chunk in enumerate(chunk_text(load_file(path))):
docs.append({"id": f"{path}::{i}", "text": chunk,
"source": path})
return docs
if __name__ == "__main__":
corpus = load_corpus("./docs")
print(f"Loaded {len(corpus)} chunks")
Run it against a folder of your own files:
mkdir -p docs # drop your PDFs / markdown / code in here
python ingest.py
Each chunk carries a stable id (path::index) and its source. The id lets you re-ingest a changed file and overwrite exactly its chunks; the source lets you cite where an answer came from.
| Content type | Suggested chunk size | Why |
|---|---|---|
| Prose / docs | 600–1000 chars | Paragraph-ish units retrieve cleanly |
| Code | Split on function/class | Keep logical units whole, not mid-statement |
| Chat logs / FAQs | One Q&A per chunk | Each entry is self-contained |
| Dense tables | One row group per chunk | Avoid splitting a row from its header |
Now turn every chunk into a vector. The ollama Python package calls the same local endpoint you tested with curl — no API key, no network egress.
Add embed.py:
import ollama
EMBED_MODEL = "nomic-embed-text"
def embed_one(text):
resp = ollama.embed(model=EMBED_MODEL, input=text)
return resp["embeddings"][0] # a list of 768 floats
def embed_many(texts, batch=32):
vectors = []
for i in range(0, len(texts), batch):
resp = ollama.embed(model=EMBED_MODEL, input=texts[i:i + batch])
vectors.extend(resp["embeddings"])
return vectors
if __name__ == "__main__":
v = embed_one("Local RAG keeps your data on your own machine.")
print(f"Vector length: {len(v)}") # -> 768
python embed.py
# Vector length: 768
The embed_many helper batches chunks. Embedding is the most repetitive stage — a few thousand chunks is fine on CPU, but batching cuts round-trips and keeps the GPU (if present) busier.
nomic-embed-text and mxbai-embed-large are not comparable, so distances become meaningless. Some models are also asymmetric, expecting different prefixes for documents vs queries (check the model card); the wrong prefix quietly tanks accuracy while everything still appears to run.ChromaDB is the fastest way to get a working store: it runs in your Python process and persists to a local folder via SQLite — no server, no Docker. Perfect for prototypes and single-user tools.
Create store_chroma.py:
import chromadb
from ingest import load_corpus
from embed import embed_many
client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_or_create_collection(
name="knowledge",
metadata={"hnsw:space": "cosine"}, # cosine similarity for text
)
corpus = load_corpus("./docs")
vectors = embed_many([c["text"] for c in corpus])
collection.add(
ids=[c["id"] for c in corpus],
embeddings=vectors,
documents=[c["text"] for c in corpus],
metadatas=[{"source": c["source"]} for c in corpus],
)
print(f"Stored {collection.count()} chunks in ./chroma_db")
python store_chroma.py
Three things to notice:
PersistentClient(path=...)writes to disk and reloads next run — your index survives restarts.- We pass
embeddingsdirectly because we computed them with Ollama. Chroma can embed for you, but supplying your own vectors guarantees the chat side and the index use the same model. hnsw:space: cosinesets the distance metric. HNSW is the approximate-nearest-neighbour index that makes search fast.
Query it to sanity-check:
from embed import embed_one
q = embed_one("How do I keep my data private?")
res = collection.query(query_embeddings=[q], n_results=3)
for doc, meta in zip(res["documents"][0], res["metadatas"][0]):
print(meta["source"], "->", doc[:80])
When you need a shared, networked, high-throughput store, Qdrant is the upgrade. It runs as a server (via Docker) and is what you would put behind the API from Lesson 10: Production Deployment.
Start Qdrant — it exposes REST on 6333 and gRPC on 6334, and the -v mount persists data:
docker run -p 6333:6333 -p 6334:6334 \
-v "$(pwd)/qdrant_storage:/qdrant/storage:z" \
qdrant/qdrant
A dashboard lives at http://localhost:6333/dashboard. Now create store_qdrant.py:
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct
from ingest import load_corpus
from embed import embed_many
client = QdrantClient(url="http://localhost:6333")
client.recreate_collection(
collection_name="knowledge",
vectors_config=VectorParams(size=768, distance=Distance.COSINE),
)
corpus = load_corpus("./docs")
vectors = embed_many([c["text"] for c in corpus])
client.upsert(
collection_name="knowledge",
points=[
PointStruct(id=i, vector=vectors[i],
payload={"text": c["text"], "source": c["source"]})
for i, c in enumerate(corpus)
],
)
print(client.count(collection_name="knowledge"))
python store_qdrant.py
size=768 MUST match nomic-embed-text. Search mirrors Chroma:
from embed import embed_one
hits = client.query_points(
collection_name="knowledge",
query=embed_one("How do I keep my data private?"),
limit=3,
).points
for h in hits:
print(round(h.score, 3), h.payload["source"])
How the two stores compare, so you can pick deliberately:
| Concern | ChromaDB (embedded) | Qdrant (server) |
|---|---|---|
| Setup | pip install, zero infra |
Docker container |
| Process model | In-process library | Networked service |
| Concurrency | Single process | Many clients |
| Best for | Prototypes, desktop apps | Teams, APIs, large corpora |
recreate_collection drops the collection first — convenient while iterating, destructive in production. Switch to create_collection guarded by an existence check once your index is real.Now connect the dots: embed the question, retrieve top-k chunks, stuff them into a prompt, and let a local chat model answer. This ask.py uses the Chroma store and a chat model from earlier lessons (swap in whatever you pulled):
import chromadb
import ollama
from embed import embed_one
CHAT_MODEL = "llama3.1:8b" # any chat model you have pulled
client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_collection("knowledge")
def retrieve(question, k=4):
res = collection.query(query_embeddings=[embed_one(question)],
n_results=k)
return list(zip(res["documents"][0], res["metadatas"][0]))
def answer(question):
chunks = retrieve(question)
context = "\n\n---\n\n".join(
f"[Source: {m['source']}]\n{doc}" for doc, m in chunks)
prompt = (
"Answer the question using ONLY the context below. "
"If the answer is not in the context, say you do not know. "
"Cite the source path for each fact.\n\n"
f"Context:\n{context}\n\nQuestion: {question}"
)
resp = ollama.chat(model=CHAT_MODEL,
messages=[{"role": "user", "content": prompt}])
return resp["message"]["content"]
if __name__ == "__main__":
print(answer("How do I keep my data private with local RAG?"))
python ask.py
The system instruction — answer only from context, admit when you do not know, cite sources — suppresses hallucination. Without it the model fills gaps from training data, exactly what RAG exists to prevent. To feel the difference, ask the same question with no retrieval (just send the bare question): the model gives a generic answer, while the RAG version quotes your documents and names the file.
k=4 and 800-char chunks you spend roughly 800–1200 tokens — fine. Crank k too high and you slow generation and dilute the signal: more context is not better context. For a precision jump, retrieve k=20 then use a small cross-encoder reranker to pick the best 4 — search broad, read narrow.Questions & Answers
path::index), you re-ingest only changed files and upsert their chunks — Chroma and Qdrant both overwrite by id. Track file modification times or content hashes and skip unchanged files. Full re-embeds are only needed when you switch embedding models or chunking strategy.Key Takeaways
- RAG grounds answers in your data — it retrieves relevant chunks and feeds them into the prompt, so the model cites your documents instead of guessing from training data.
- Chunking decides quality — tune size (~800 chars) and overlap (~100 chars) before blaming the embedding model; bad chunks cannot be rescued downstream.
- Keep embedding and chat models separate — a small fast embedder (
nomic-embed-text, 768 dims) runs on every chunk and query; save the big model for the final answer. - Match dimensions everywhere — the embedding model fixes the vector size, and your collection must be created with that exact number or inserts fail.
- ChromaDB for prototypes, Qdrant for scale — embedded Chroma needs zero infra and persists to SQLite; Qdrant runs as a networked Docker service for teams and APIs.
- Everything stays local — Ollama on localhost, an in-process or containerised vector store, and your files on disk — no data ever leaves the machine.
Next Steps: Lesson 9: Performance Tuning