Local RAG — Your Own Knowledge Base

55 min advanced Lesson 8

Learning Outcomes

  • Explain the RAG pipeline end to end — ingest, chunk, embed, store, retrieve, generate — and why each stage matters
  • Ingest mixed documents (PDF, Markdown, code) and chunk them with size and overlap tuned for retrieval
  • Generate embeddings entirely locally using an Ollama embedding model
  • Stand up a vector store with both ChromaDB (embedded) and Qdrant (Docker) and load your vectors
  • Wire retrieval into a local chat model and measure answer quality with and without context

Lesson Plan

Segment Duration Topic
Intro 3 min What RAG is and why local RAG is worth the effort
Explain 7 min The pipeline: ingest, chunk, embed, store, retrieve, generate
Build 8 min Ingestion and chunking strategy
Build 8 min Local embeddings with Ollama
Build 8 min ChromaDB — the embedded vector store
Build 8 min Qdrant — the production-grade alternative
Build 9 min Retrieval + generation with a local model
Wrap-up 4 min Evaluation, gotchas, and key takeaways

Before You Begin

Pre-work:

Shopping List:

  • Ollama running locally (ollama serve active, default port 11434)
  • Python 3.10+ and a virtual environment
  • Docker (only for the Qdrant step)
  • A folder of your own documents to index — PDFs, Markdown notes, or a code repo
  • Roughly 1 GB of free disk: ~274 MB for the embedding model, the rest for a chat model and the vector index

1 Understand the RAG Pipeline

RAG (Retrieval-Augmented Generation) fetches the most relevant snippets from your own documents and pastes them into the prompt before the model answers. The model then answers grounded in your data, not just what it memorised in training. Doing it locally matters: the alternative is uploading your contracts, source code, or notes to a hosted embedding API — which defeats the point of local models.

The pipeline has six stages, and every later step maps to one of them:

documents ──► ingest ──► chunk ──► embed ──► store (vector DB)
                                                  │
query ──► embed ──► similarity search ────────────┘
                              │
                       top-k chunks ──► prompt ──► local LLM ──► answer

A few terms, defined once:

Term Plain-English meaning
Embedding A vector of floats capturing the meaning of text, so similar text lands nearby
Chunk A bite-sized slice of a document — what you embed and retrieve
Vector store A database that finds the nearest vectors to a query vector, fast
top-k Chunks retrieved per query (commonly 3–8)
NOTE
Key Insight
RAG does not change the model's weights — it is prompt engineering at scale, assembling the best context for each question. That makes it cheaper and faster to iterate on than fine-tuning. Use RAG when knowledge changes often or must be cited; fine-tune when you need a new behaviour or style.

2 Set Up the Environment and Pull the Embedding Model

Create an isolated environment and install the libraries we will use across the lesson:

On macOS and Linux:

python3 -m venv rag-env
source rag-env/bin/activate
pip install ollama chromadb qdrant-client pypdf

On Windows (PowerShell):

python -m venv rag-env
rag-env\Scripts\Activate.ps1
pip install ollama chromadb qdrant-client pypdf

Now pull a dedicated embedding model — separate from your chat model, it only turns text into vectors. We will use nomic-embed-text: small (~274 MB), fast on CPU, 768-dimensional vectors:

ollama pull nomic-embed-text

Confirm Ollama's embeddings endpoint responds. It exposes embeddings at /api/embed on port 11434:

curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": "Llamas are members of the camelid family"
}'

You should get back a JSON object containing an embeddings array of 768 floats.

WARNING
Match your dimensions
The vector size is fixed by the embedding model — nomic-embed-text emits 768 dimensions, mxbai-embed-large emits 1024. Your vector store collection MUST be created with the exact same size, or inserts will fail. Keep the embedding model and the chat model separate too: the embedder runs on every chunk and query, so favour something small and quick.

3 Ingest and Chunk Your Documents

Retrieval quality is decided here, not at the model. Chunks too big bury the relevant sentence in noise; too small and you lose context. A solid prose default is ~800 characters per chunk with ~100 characters of overlap, so ideas straddling a boundary are not lost.

Create ingest.py to load PDFs, Markdown, and code, then split into overlapping chunks:

import os
from pypdf import PdfReader

def load_file(path):
    if path.endswith(".pdf"):
        reader = PdfReader(path)
        return "\n".join(page.extract_text() or "" for page in reader.pages)
    with open(path, encoding="utf-8", errors="ignore") as f:
        return f.read()

def chunk_text(text, size=800, overlap=100):
    chunks, start = [], 0
    while start < len(text):
        end = start + size
        chunk = text[start:end].strip()
        if chunk:
            chunks.append(chunk)
        start += size - overlap   # slide forward, keeping an overlap
    return chunks

def load_corpus(root):
    docs = []
    for dirpath, _, files in os.walk(root):
        for name in files:
            if name.endswith((".pdf", ".md", ".txt", ".py", ".js")):
                path = os.path.join(dirpath, name)
                for i, chunk in enumerate(chunk_text(load_file(path))):
                    docs.append({"id": f"{path}::{i}", "text": chunk,
                                 "source": path})
    return docs

if __name__ == "__main__":
    corpus = load_corpus("./docs")
    print(f"Loaded {len(corpus)} chunks")

Run it against a folder of your own files:

mkdir -p docs   # drop your PDFs / markdown / code in here
python ingest.py

Each chunk carries a stable id (path::index) and its source. The id lets you re-ingest a changed file and overwrite exactly its chunks; the source lets you cite where an answer came from.

Content type Suggested chunk size Why
Prose / docs 600–1000 chars Paragraph-ish units retrieve cleanly
Code Split on function/class Keep logical units whole, not mid-statement
Chat logs / FAQs One Q&A per chunk Each entry is self-contained
Dense tables One row group per chunk Avoid splitting a row from its header
WARNING
Garbage in, garbage out
PDF extraction is messy — headers, footers, and column layouts leak in. Print a few chunks before embedding; if they are full of page numbers and broken words, fix extraction first. Once it works, try splitting on Markdown headings or paragraph breaks so each chunk is a coherent thought — that often beats any embedding-model upgrade.

4 Generate Embeddings Locally with Ollama

Now turn every chunk into a vector. The ollama Python package calls the same local endpoint you tested with curl — no API key, no network egress.

Add embed.py:

import ollama

EMBED_MODEL = "nomic-embed-text"

def embed_one(text):
    resp = ollama.embed(model=EMBED_MODEL, input=text)
    return resp["embeddings"][0]   # a list of 768 floats

def embed_many(texts, batch=32):
    vectors = []
    for i in range(0, len(texts), batch):
        resp = ollama.embed(model=EMBED_MODEL, input=texts[i:i + batch])
        vectors.extend(resp["embeddings"])
    return vectors

if __name__ == "__main__":
    v = embed_one("Local RAG keeps your data on your own machine.")
    print(f"Vector length: {len(v)}")   # -> 768
python embed.py
# Vector length: 768

The embed_many helper batches chunks. Embedding is the most repetitive stage — a few thousand chunks is fine on CPU, but batching cuts round-trips and keeps the GPU (if present) busier.

WARNING
One model, one index
Never mix embedding models in a collection — vectors from nomic-embed-text and mxbai-embed-large are not comparable, so distances become meaningless. Some models are also asymmetric, expecting different prefixes for documents vs queries (check the model card); the wrong prefix quietly tanks accuracy while everything still appears to run.

5 Store Vectors in ChromaDB (Embedded)

ChromaDB is the fastest way to get a working store: it runs in your Python process and persists to a local folder via SQLite — no server, no Docker. Perfect for prototypes and single-user tools.

Create store_chroma.py:

import chromadb
from ingest import load_corpus
from embed import embed_many

client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_or_create_collection(
    name="knowledge",
    metadata={"hnsw:space": "cosine"},   # cosine similarity for text
)

corpus = load_corpus("./docs")
vectors = embed_many([c["text"] for c in corpus])

collection.add(
    ids=[c["id"] for c in corpus],
    embeddings=vectors,
    documents=[c["text"] for c in corpus],
    metadatas=[{"source": c["source"]} for c in corpus],
)
print(f"Stored {collection.count()} chunks in ./chroma_db")
python store_chroma.py

Three things to notice:

  • PersistentClient(path=...) writes to disk and reloads next run — your index survives restarts.
  • We pass embeddings directly because we computed them with Ollama. Chroma can embed for you, but supplying your own vectors guarantees the chat side and the index use the same model.
  • hnsw:space: cosine sets the distance metric. HNSW is the approximate-nearest-neighbour index that makes search fast.

Query it to sanity-check:

from embed import embed_one
q = embed_one("How do I keep my data private?")
res = collection.query(query_embeddings=[q], n_results=3)
for doc, meta in zip(res["documents"][0], res["metadatas"][0]):
    print(meta["source"], "->", doc[:80])
TIP
When Chroma is enough
For a personal knowledge base or a desktop tool shipping to one user, embedded ChromaDB is genuinely production-ready. You only outgrow it when multiple processes or machines need to hit the same index concurrently.

6 Store Vectors in Qdrant (Server)

When you need a shared, networked, high-throughput store, Qdrant is the upgrade. It runs as a server (via Docker) and is what you would put behind the API from Lesson 10: Production Deployment.

Start Qdrant — it exposes REST on 6333 and gRPC on 6334, and the -v mount persists data:

docker run -p 6333:6333 -p 6334:6334 \
    -v "$(pwd)/qdrant_storage:/qdrant/storage:z" \
    qdrant/qdrant

A dashboard lives at http://localhost:6333/dashboard. Now create store_qdrant.py:

from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct
from ingest import load_corpus
from embed import embed_many

client = QdrantClient(url="http://localhost:6333")

client.recreate_collection(
    collection_name="knowledge",
    vectors_config=VectorParams(size=768, distance=Distance.COSINE),
)

corpus = load_corpus("./docs")
vectors = embed_many([c["text"] for c in corpus])

client.upsert(
    collection_name="knowledge",
    points=[
        PointStruct(id=i, vector=vectors[i],
                    payload={"text": c["text"], "source": c["source"]})
        for i, c in enumerate(corpus)
    ],
)
print(client.count(collection_name="knowledge"))
python store_qdrant.py

size=768 MUST match nomic-embed-text. Search mirrors Chroma:

from embed import embed_one
hits = client.query_points(
    collection_name="knowledge",
    query=embed_one("How do I keep my data private?"),
    limit=3,
).points
for h in hits:
    print(round(h.score, 3), h.payload["source"])

How the two stores compare, so you can pick deliberately:

Concern ChromaDB (embedded) Qdrant (server)
Setup pip install, zero infra Docker container
Process model In-process library Networked service
Concurrency Single process Many clients
Best for Prototypes, desktop apps Teams, APIs, large corpora
WARNING
recreate_collection wipes data
recreate_collection drops the collection first — convenient while iterating, destructive in production. Switch to create_collection guarded by an existence check once your index is real.

7 Wire Retrieval into a Local LLM

Now connect the dots: embed the question, retrieve top-k chunks, stuff them into a prompt, and let a local chat model answer. This ask.py uses the Chroma store and a chat model from earlier lessons (swap in whatever you pulled):

import chromadb
import ollama
from embed import embed_one

CHAT_MODEL = "llama3.1:8b"   # any chat model you have pulled

client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_collection("knowledge")

def retrieve(question, k=4):
    res = collection.query(query_embeddings=[embed_one(question)],
                           n_results=k)
    return list(zip(res["documents"][0], res["metadatas"][0]))

def answer(question):
    chunks = retrieve(question)
    context = "\n\n---\n\n".join(
        f"[Source: {m['source']}]\n{doc}" for doc, m in chunks)
    prompt = (
        "Answer the question using ONLY the context below. "
        "If the answer is not in the context, say you do not know. "
        "Cite the source path for each fact.\n\n"
        f"Context:\n{context}\n\nQuestion: {question}"
    )
    resp = ollama.chat(model=CHAT_MODEL,
                       messages=[{"role": "user", "content": prompt}])
    return resp["message"]["content"]

if __name__ == "__main__":
    print(answer("How do I keep my data private with local RAG?"))
python ask.py

The system instruction — answer only from context, admit when you do not know, cite sources — suppresses hallucination. Without it the model fills gaps from training data, exactly what RAG exists to prevent. To feel the difference, ask the same question with no retrieval (just send the bare question): the model gives a generic answer, while the RAG version quotes your documents and names the file.

NOTE
Context budget
Retrieved chunks compete for the model's context window. With k=4 and 800-char chunks you spend roughly 800–1200 tokens — fine. Crank k too high and you slow generation and dilute the signal: more context is not better context. For a precision jump, retrieve k=20 then use a small cross-encoder reranker to pick the best 4 — search broad, read narrow.

Questions & Answers

Q: My answers cite the wrong document or miss the obvious one. Is the embedding model bad?
Usually not — it is chunking or k. First print the retrieved chunks for a failing query: if the right text is not even in the top-k, the problem is upstream (chunks too large, bad PDF extraction, or k too low). Only if the correct chunk exists but ranks poorly should you suspect the embedding model. Fix retrieval before touching the model.
Q: My documents change daily. Do I re-embed everything every time?
No. Because each chunk has a stable id (path::index), you re-ingest only changed files and upsert their chunks — Chroma and Qdrant both overwrite by id. Track file modification times or content hashes and skip unchanged files. Full re-embeds are only needed when you switch embedding models or chunking strategy.
Q: How big can my corpus get before this falls over on local hardware?
Embedded ChromaDB comfortably handles tens to low-hundreds of thousands of chunks. Beyond that, or once multiple processes need concurrent access, move to Qdrant, which uses an on-disk HNSW index built for millions of vectors. The bottleneck is rarely storage — 768 floats is ~3 KB per chunk — it is the one-time embedding pass, which is CPU/GPU-bound.
Q: Is anything actually sent off my machine in this setup?
Nothing. Ollama embeddings hit localhost:11434, ChromaDB is an in-process library, and Qdrant runs in a local Docker container. Every vector and document stays on disk under your control. That is the whole reason to build RAG locally instead of using a hosted embedding API.
Q: Should I use cosine or dot-product distance?
For text embeddings, cosine similarity is the safe default — it compares direction (meaning) and ignores vector magnitude. Dot product can be faster and is fine when your embeddings are already normalised to unit length. If unsure, use cosine; the speed difference is negligible at local scale.

Key Takeaways

  1. RAG grounds answers in your data — it retrieves relevant chunks and feeds them into the prompt, so the model cites your documents instead of guessing from training data.
  2. Chunking decides quality — tune size (~800 chars) and overlap (~100 chars) before blaming the embedding model; bad chunks cannot be rescued downstream.
  3. Keep embedding and chat models separate — a small fast embedder (nomic-embed-text, 768 dims) runs on every chunk and query; save the big model for the final answer.
  4. Match dimensions everywhere — the embedding model fixes the vector size, and your collection must be created with that exact number or inserts fail.
  5. ChromaDB for prototypes, Qdrant for scale — embedded Chroma needs zero infra and persists to SQLite; Qdrant runs as a networked Docker service for teams and APIs.
  6. Everything stays local — Ollama on localhost, an in-process or containerised vector store, and your files on disk — no data ever leaves the machine.

Next Steps: Lesson 9: Performance Tuning