What Is RAG?
Learning Outcomes
- Explain why LLMs need retrieval augmentation and which three problems it addresses
- Trace data flow through a RAG pipeline from query to grounded answer
- Distinguish RAG from fine-tuning and from naive prompt stuffing
- Identify where RAG fits — and where it does not — in a knowledge architecture
- Run a minimal Python example that retrieves a document and conditions an LLM on it
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why LLMs fail at factual questions |
| Explain | 5 min | Cutoffs, hallucination, private data |
| Demo | 5 min | Pure-LLM vs RAG-augmented answer |
| Explain | 5 min | The pipeline — indexing and query phases |
| Demo | 7 min | Minimal Python RAG in 25 lines |
| Explain | 7 min | RAG vs fine-tuning vs prompt stuffing |
| Wrap-up | 3 min | When to use RAG; course preview |
Before You Begin
Pre-work:
- Comfortable with Python and pip
- At least one successful LLM API call (OpenAI, Anthropic, or any chat-completion endpoint)
- If you completed the Agentic AI course, RAG is a specific pattern within the tool-use picture you already know
Shopping List:
- Python 3.10 or later, an OpenAI API key (
gpt-4o-miniused in examples) - Install once:
pip install openai sentence-transformers chromadb
Every LLM has the same three limitations when answering factual questions:
| Problem | Root Cause | Symptom |
|---|---|---|
| Knowledge cutoff | Facts after training date absent from weights | Refuses or confabulates recent information |
| Hallucination | Weak weight representation of specific facts | Fluent-sounding answer with wrong specifics |
| Private data | Internal docs never in any training corpus | No representation of them at all |
RAG addresses all three by the same mechanism: at inference time, relevant source text is retrieved from a corpus you control and placed into the model's context window. The model shifts from "recall a fact from weights" to "read provided text and reason over it" — a task LLMs are far more reliable at.
A RAG system has two phases — indexing (offline, run once then refreshed) and querying (runtime):
INDEXING Documents -> Chunk -> Embed -> Vector Store
QUERYING Query -> Embed -> ANN search -> Top-k chunks -> LLM -> Answer
Indexing: ingest documents, split into chunks (L4), encode each chunk as a dense vector (L2), store vectors + metadata in a vector database (L3). Query: embed the question with the same encoder, run approximate-nearest-neighbour (ANN) search, inject the top-k chunks into the prompt, and generate.
| Stage | Library / Tool | Lesson |
|---|---|---|
| Chunk | LangChain RecursiveCharacterTextSplitter, LlamaIndex |
L4 |
| Embed | sentence-transformers, OpenAI Embeddings |
L2 |
| Store + search | Chroma, pgvector, Pinecone, Qdrant, Weaviate | L3 |
| Augment + generate | Any chat-completion client | L5 |
| Evaluate | RAGAS, custom harness | L7 |
The example uses a fictional internal policy — the kind of text absent from any public training corpus.
Source document: "P0 incidents must be acknowledged within 15 minutes. If no acknowledgement in 15 minutes, escalation to VP of Engineering triggers."
Without RAG: The model admits ignorance and returns a generic answer — "generally 15–30 minutes." Useless.
With RAG: The retrieved chunk is injected into the prompt. The LLM responds: "P0 incidents must be acknowledged within 15 minutes. After that window, escalation to the VP of Engineering triggers automatically. (Source: Incident Response Policy v3.2, §4.1)"
Specific, accurate, cited — because the model is reading the document, not guessing.
Every essential RAG component in one file. The course replaces each part with production-grade alternatives and measures quality at each step.
import os
from openai import OpenAI
import chromadb
from sentence_transformers import SentenceTransformer
encoder = SentenceTransformer("all-MiniLM-L6-v2")
collection = chromadb.Client().create_collection("demo")
# Index
docs = [
"P0 incidents must be acknowledged within 15 minutes.",
"P1 incidents must be acknowledged within 60 minutes.",
]
collection.add(documents=docs,
embeddings=encoder.encode(docs).tolist(),
ids=["d0", "d1"])
# Retrieve
query = "What is the SLA for a P0 incident?"
results = collection.query(
query_embeddings=encoder.encode([query]).tolist(), n_results=1)
chunk = results["documents"][0][0]
# Generate
llm = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
resp = llm.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content":
f"Context: {chunk}\n\nQuestion: {query}\nAnswer in one sentence."}],
)
print(resp.choices[0].message.content)
export OPENAI_API_KEY="sk-..."
python rag_demo.py
# A P0 incident must be acknowledged within 15 minutes, per the context.
Three approaches compete when you need an LLM grounded in specific knowledge:
| Approach | Best For | Key Drawback |
|---|---|---|
| RAG | Large, updated, or private corpora | Retrieval can miss; adds latency |
| Fine-tuning | Style, tone, output format | Does not memorise facts; expensive; stale immediately |
| Prompt stuffing | Very small corpora under ~50 pages | Token cost per query; does not scale |
Fine-tuning is the most common misapplication. The model learns language patterns, not lookup tables — fine-tuned models still hallucinate specific policy clauses and version numbers. Fine-tune for style; use RAG for facts.
RAG sits inside a three-layer knowledge architecture:
- Application layer — chat UI, API endpoint, team-chat bot
- RAG layer — query, retrieve, augment, generate (where this course lives)
- Knowledge layer — vector store, metadata DB, source documents
Knowledge layer hygiene matters most. Stale documents and inconsistent ingestion are the leading causes of poor RAG output — no retrieval tuning compensates for a corrupt index.
Skip RAG when: the answer requires live data (use a direct API call), the question needs multi-hop reasoning over structured data (Lesson 8 covers graph + RAG), or the full corpus fits in one context window at low volume (prompt stuffing is simpler).
Questions & Answers
Key Takeaways
- Three problems, one mechanism — RAG solves knowledge cutoffs, hallucination on specifics, and private-data access by retrieving source text at inference time rather than relying on model weights.
- Two phases — indexing (embed offline, pay once per document) and querying (retrieve at runtime, pay per query) have separate cost and latency profiles.
- The LLM reads, not recalls — reasoning over provided text is substantially more reliable than weight-based recall; that is why grounded answers are far more accurate.
- Fine-tuning does not fix factual recall — it teaches language patterns, not lookup tables; use RAG for facts, fine-tuning for style.
- Retrieval quality sets the ceiling — no LLM recovers from irrelevant retrieved chunks; every subsequent lesson improves retrieval first, then generation, then operations.
- Evaluate from day one — a 20-item golden test set checked on every change catches regressions before they reach users.
Next Steps: Lesson 2: Embeddings Explained