What Is RAG?

35 min beginner Lesson 1

Learning Outcomes

  • Explain why LLMs need retrieval augmentation and which three problems it addresses
  • Trace data flow through a RAG pipeline from query to grounded answer
  • Distinguish RAG from fine-tuning and from naive prompt stuffing
  • Identify where RAG fits — and where it does not — in a knowledge architecture
  • Run a minimal Python example that retrieves a document and conditions an LLM on it

Lesson Plan

Segment Duration Topic
Intro 3 min Why LLMs fail at factual questions
Explain 5 min Cutoffs, hallucination, private data
Demo 5 min Pure-LLM vs RAG-augmented answer
Explain 5 min The pipeline — indexing and query phases
Demo 7 min Minimal Python RAG in 25 lines
Explain 7 min RAG vs fine-tuning vs prompt stuffing
Wrap-up 3 min When to use RAG; course preview

Before You Begin

Pre-work:

  • Comfortable with Python and pip
  • At least one successful LLM API call (OpenAI, Anthropic, or any chat-completion endpoint)
  • If you completed the Agentic AI course, RAG is a specific pattern within the tool-use picture you already know

Shopping List:

  • Python 3.10 or later, an OpenAI API key (gpt-4o-mini used in examples)
  • Install once: pip install openai sentence-transformers chromadb

1 The Three Problems RAG Solves

Every LLM has the same three limitations when answering factual questions:

Problem Root Cause Symptom
Knowledge cutoff Facts after training date absent from weights Refuses or confabulates recent information
Hallucination Weak weight representation of specific facts Fluent-sounding answer with wrong specifics
Private data Internal docs never in any training corpus No representation of them at all

RAG addresses all three by the same mechanism: at inference time, relevant source text is retrieved from a corpus you control and placed into the model's context window. The model shifts from "recall a fact from weights" to "read provided text and reason over it" — a task LLMs are far more reliable at.

NOTE
Mental Model
Think of the LLM as a capable analyst and RAG as their briefing document. They read it and reason over it — no memorisation required.
WARNING
What RAG Does Not Fix
RAG does not eliminate hallucination. If retrieved documents are irrelevant, the model can still confabulate. Retrieval quality is the first lever.

2 The RAG Pipeline — Indexing and Query Phases

A RAG system has two phases — indexing (offline, run once then refreshed) and querying (runtime):

INDEXING  Documents -> Chunk -> Embed -> Vector Store
QUERYING  Query -> Embed -> ANN search -> Top-k chunks -> LLM -> Answer

Indexing: ingest documents, split into chunks (L4), encode each chunk as a dense vector (L2), store vectors + metadata in a vector database (L3). Query: embed the question with the same encoder, run approximate-nearest-neighbour (ANN) search, inject the top-k chunks into the prompt, and generate.

NOTE
Why Two Phases?
Embedding is slow and costly. Indexing offline keeps query latency low — typically under 500 ms. You pay the embedding cost once per document, not once per query.
Stage Library / Tool Lesson
Chunk LangChain RecursiveCharacterTextSplitter, LlamaIndex L4
Embed sentence-transformers, OpenAI Embeddings L2
Store + search Chroma, pgvector, Pinecone, Qdrant, Weaviate L3
Augment + generate Any chat-completion client L5
Evaluate RAGAS, custom harness L7

3 Pure LLM vs RAG — Side by Side

The example uses a fictional internal policy — the kind of text absent from any public training corpus.

Source document: "P0 incidents must be acknowledged within 15 minutes. If no acknowledgement in 15 minutes, escalation to VP of Engineering triggers."

Without RAG: The model admits ignorance and returns a generic answer — "generally 15–30 minutes." Useless.

With RAG: The retrieved chunk is injected into the prompt. The LLM responds: "P0 incidents must be acknowledged within 15 minutes. After that window, escalation to the VP of Engineering triggers automatically. (Source: Incident Response Policy v3.2, §4.1)"

Specific, accurate, cited — because the model is reading the document, not guessing.

TIP
The Citation Test
Instruct the model to cite the source chunk. If a claim cannot be traced to a retrieved document, treat it as potentially hallucinated. Faithfulness evaluation is covered in Lesson 7.

4 A Minimal RAG Pipeline in Python

Every essential RAG component in one file. The course replaces each part with production-grade alternatives and measures quality at each step.

import os
from openai import OpenAI
import chromadb
from sentence_transformers import SentenceTransformer

encoder = SentenceTransformer("all-MiniLM-L6-v2")
collection = chromadb.Client().create_collection("demo")

# Index
docs = [
    "P0 incidents must be acknowledged within 15 minutes.",
    "P1 incidents must be acknowledged within 60 minutes.",
]
collection.add(documents=docs,
               embeddings=encoder.encode(docs).tolist(),
               ids=["d0", "d1"])

# Retrieve
query = "What is the SLA for a P0 incident?"
results = collection.query(
    query_embeddings=encoder.encode([query]).tolist(), n_results=1)
chunk = results["documents"][0][0]

# Generate
llm = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
resp = llm.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content":
               f"Context: {chunk}\n\nQuestion: {query}\nAnswer in one sentence."}],
)
print(resp.choices[0].message.content)
export OPENAI_API_KEY="sk-..."
python rag_demo.py
# A P0 incident must be acknowledged within 15 minutes, per the context.
WARNING
Keep Encoders Consistent
The encoder at index time and query time must be identical. Mixing models produces incompatible vector spaces — retrieval quality collapses silently.
NOTE
Why sentence-transformers?
Runs pre-trained encoder models locally — no API key, no cost. Lesson 2 covers when to use local vs hosted APIs.

5 RAG vs Fine-Tuning vs Prompt Stuffing

Three approaches compete when you need an LLM grounded in specific knowledge:

Approach Best For Key Drawback
RAG Large, updated, or private corpora Retrieval can miss; adds latency
Fine-tuning Style, tone, output format Does not memorise facts; expensive; stale immediately
Prompt stuffing Very small corpora under ~50 pages Token cost per query; does not scale

Fine-tuning is the most common misapplication. The model learns language patterns, not lookup tables — fine-tuned models still hallucinate specific policy clauses and version numbers. Fine-tune for style; use RAG for facts.

WARNING
The Fine-Tuning Trap
Test first: does a RAG-augmented general model beat the fine-tuned model on factual accuracy? Almost always yes. Fine-tuning is expensive and stale the moment training ends.

6 Where RAG Fits in a Knowledge Architecture

RAG sits inside a three-layer knowledge architecture:

  • Application layer — chat UI, API endpoint, team-chat bot
  • RAG layer — query, retrieve, augment, generate (where this course lives)
  • Knowledge layer — vector store, metadata DB, source documents

Knowledge layer hygiene matters most. Stale documents and inconsistent ingestion are the leading causes of poor RAG output — no retrieval tuning compensates for a corrupt index.

Skip RAG when: the answer requires live data (use a direct API call), the question needs multi-hop reasoning over structured data (Lesson 8 covers graph + RAG), or the full corpus fits in one context window at low volume (prompt stuffing is simpler).

NOTE
Agentic RAG
If you completed the Agentic AI course, you know how agents call external tools. Agentic RAG extends that: the agent decides whether to retrieve and issues multiple retrieval calls if needed. This course covers mechanics; orchestration is in the Agentic AI course.
TIP
Start With the Query Log
Collect 50 real queries before building. Categorise: document lookup, live data, or SQL. RAG typically fits 60–70% of questions.

Questions & Answers

Q: My LLM has a huge context window. Does that make RAG obsolete?
Not at production scale. Stuffing a large corpus into context has three problems: cost (per-token per query), latency (larger contexts are slower), and the "lost in the middle" effect (models attend less reliably to facts buried mid-context). RAG places only the most relevant 2–5% of your corpus in context per query — typically 10–50x cheaper and faster at scale.
Q: Won't RAG answers drift as the corpus changes?
That is RAG's advantage. Re-index changed chunks and retrieval immediately reflects the new content — no retraining. Fine-tuned models are frozen; every update needs another training run. RAG's staleness is bounded by your re-indexing frequency.
Q: How do I enforce document-level access control?
Store access-control metadata (role, team, classification) alongside each chunk and filter at query time. Most stores (Pinecone, Qdrant, Weaviate, pgvector) support pre-filter. Never rely on the LLM to enforce access control — filter before retrieval, not after generation. Production patterns are in Lesson 9.
Q: How many chunks should I retrieve per query?
Typically 3–10. More chunks improve recall but dilute generation quality and increase token cost. Start at 4–5 and tune on precision@k and recall@k. Re-ranking (scoring more candidates with a cross-encoder before generation) is covered in Lesson 6.
Q: Can RAG work over structured database tables, not just text?
Yes with caveats. For simple lookups serialise rows as text chunks and embed them. For aggregations and joins, text-to-SQL is more reliable: the LLM generates SQL from the question and executes it. Combine both — RAG for free-text, text-to-SQL for structured data. Knowledge graph integration is in Lesson 8.

Key Takeaways

  1. Three problems, one mechanism — RAG solves knowledge cutoffs, hallucination on specifics, and private-data access by retrieving source text at inference time rather than relying on model weights.
  2. Two phases — indexing (embed offline, pay once per document) and querying (retrieve at runtime, pay per query) have separate cost and latency profiles.
  3. The LLM reads, not recalls — reasoning over provided text is substantially more reliable than weight-based recall; that is why grounded answers are far more accurate.
  4. Fine-tuning does not fix factual recall — it teaches language patterns, not lookup tables; use RAG for facts, fine-tuning for style.
  5. Retrieval quality sets the ceiling — no LLM recovers from irrelevant retrieved chunks; every subsequent lesson improves retrieval first, then generation, then operations.
  6. Evaluate from day one — a 20-item golden test set checked on every change catches regressions before they reach users.

Next Steps: Lesson 2: Embeddings Explained