Knowledge Graphs + RAG
Learning Outcomes
- Explain why vector search alone fails on multi-hop and relationship-dependent questions
- Extract entities and relationships from documents to construct a knowledge graph
- Query a knowledge graph programmatically to traverse typed relationships
- Combine graph traversal with vector retrieval in a hybrid retrieval architecture
- Evaluate when a knowledge graph layer adds enough value to justify the engineering cost
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why vector search has a structural blind spot |
| Explain | 7 min | Knowledge graphs — entities, edges, and properties |
| Demo | 8 min | Extracting a graph from raw documents with an LLM |
| Demo | 9 min | Storing and querying the graph with NetworkX and Neo4j |
| Demo | 10 min | Graph-based retrieval for multi-hop questions |
| Demo | 8 min | Hybrid graph + vector pipeline end-to-end |
| Explain | 2 min | When to reach for a graph vs. when not to |
| Wrap-up | 3 min | Key takeaways and what comes next |
Before You Begin
Pre-work:
- Complete Lesson 5: Building a Basic RAG Pipeline — you need a working RAG pipeline to extend
- Complete Lesson 6: Advanced Retrieval — we build on hybrid search and re-ranking concepts
- Complete Lesson 7: Evaluation & Quality — we use RAGAS-style metrics to validate the graph layer
- Familiarity with LLM-as-tool patterns is helpful; if you took the Agentic AI course, the entity-extraction step will look familiar — it is the same function-calling pattern covered there
Shopping List:
- Python 3.10 or later
pip install networkx spacy sentence-transformers openai chromadb(or your preferred vector store from Lesson 3)python -m spacy download en_core_web_sm(lightweight NER model)- Optional, for production graphs: a running Neo4j instance (Community Edition is free) and
pip install neo4j - An OpenAI-compatible API key for LLM-assisted entity extraction in Step 3
Vector similarity finds chunks whose surface meaning is close to the query. That works well for direct lookup questions: "What does the retry_policy configuration field do?" But it breaks down on a class of questions that require following a chain of relationships rather than recognising textual proximity.
Consider a corpus of engineering documents. The following questions all require traversal, not similarity:
| Question type | Example | Why vector search struggles |
|---|---|---|
| Multi-hop | "Who owns the service that exposes the payments API?" | "Owner" and "payments API" may never appear in the same chunk |
| Aggregation over relationships | "Which teams depend on the auth service?" | Requires collecting every document that mentions a dependency |
| Hierarchical | "What sub-components belong to the checkout domain?" | Hierarchy is implicit in headings and prose, not encoded in any single chunk |
| Contradiction detection | "Does any document say feature X is deprecated?" | Contradiction only visible after collecting all statements about X |
A knowledge graph (KG) makes these relationships explicit. In a KG:
- Entities are named things: services, teams, people, components, features
- Edges are typed, directed relationships:
OWNS,DEPENDS_ON,PART_OF,DEPRECATED_BY - Properties are scalar attributes on entities or edges: version numbers, dates, status flags
Once the graph exists, multi-hop questions become graph traversal problems. "Who owns the service that exposes the payments API?" becomes: find the node labelled payments_api, follow the EXPOSED_BY edge to a Service node, then follow its OWNED_BY edge to a Team node.
The architecture you will build by the end of this lesson looks like this:
User query
│
├─── Entity detector ────► Graph traversal ──► Graph context
│ │
└─── Query embedder ────► Vector search ──► Chunk context
│
Context merger
│
LLM generation
│
Grounded answer
The entity detector decides whether the query names known graph entities. If it does, graph traversal runs alongside — or instead of — vector search. The results are merged before generation.
Before writing code, it helps to be precise about the data model. Most knowledge graphs follow one of two storage paradigms: property graphs (nodes and edges can carry arbitrary key-value properties) and RDF triples (subject-predicate-object). For RAG-adjacent work, property graphs are nearly always the right choice — they map naturally to JSON and are supported by Neo4j, NetworkX, Amazon Neptune, and Memgraph.
A property graph schema for a software documentation corpus might look like this:
# Node labels and their core properties
(:Service) {name, status, version, description}
(:Team) {name, chat_channel, oncall_rotation}
(:API) {name, endpoint, protocol, slo_p99_ms}
(:Document) {url, title, ingested_at, chunk_count}
# Relationship types
(:Team)-[:OWNS]->(:Service)
(:Service)-[:EXPOSES]->(:API)
(:Service)-[:DEPENDS_ON]->(:Service)
(:API)-[:DOCUMENTED_IN]->(:Document)
(:Document)-[:MENTIONS_ENTITY]->(:Service | :Team | :API)
The MENTIONS_ENTITY edges link graph nodes back to the document chunks that describe them — this is the bridge that lets you retrieve prose context for any entity you find through graph traversal.
Choosing relationship types is the most consequential design decision. A useful heuristic: model only relationships you expect to query directionally. "Team A owns Service B" is worth modelling. "Document A mentions the word 'latency'" is not — vector search handles that better.
The table below maps common document corpora to the relationship types worth extracting:
| Document corpus | High-value relationship types |
|---|---|
| Engineering runbooks | OWNS, DEPENDS_ON, ALERTS_ON, ESCALATES_TO |
| API documentation | EXPOSES, CONSUMES, RETURNS, AUTHENTICATED_BY |
| HR / org charts | REPORTS_TO, MEMBER_OF, RESPONSIBLE_FOR |
| Research papers | CITES, EXTENDS, CONTRADICTS, EVALUATES |
| Product specs | PART_OF, REQUIRES, SUPERSEDES, TESTED_BY |
Graph construction has two stages: entity recognition (which named things appear in this text?) and relationship extraction (how do those entities relate to each other?).
Stage 1 — Named entity recognition with spaCy
For well-defined entity types such as proper nouns, product names, and person names, a lightweight spaCy NER model is fast and accurate enough:
import spacy
nlp = spacy.load("en_core_web_sm")
def extract_entities_spacy(text: str) -> list[dict]:
doc = nlp(text)
entities = []
for ent in doc.ents:
if ent.label_ in {"ORG", "PRODUCT", "PERSON", "GPE", "WORK_OF_ART"}:
entities.append({
"text": ent.text,
"label": ent.label_,
"start": ent.start_char,
"end": ent.end_char,
})
return entities
spaCy is fast (thousands of documents per second on a CPU) but only recognises entity types it was trained on. For domain-specific entities — service names, internal project codenames, custom configuration keys — you need the next stage.
Stage 2 — LLM-assisted relationship extraction
Use an LLM to extract typed (subject, predicate, object) triples from each chunk. Structure the output so it is directly loadable into your graph store:
import json
from openai import OpenAI
client = OpenAI()
EXTRACTION_SYSTEM = """You are a knowledge-graph extraction engine.
Given a passage of technical documentation, extract all relationships
as JSON triples. Use only these relationship types:
OWNS, DEPENDS_ON, EXPOSES, PART_OF, SUPERSEDES, DOCUMENTED_IN.
Output a JSON array of objects with keys:
subject (entity name, normalised to snake_case)
predicate (one of the allowed types)
object (entity name, normalised to snake_case)
subject_type (Service | Team | API | Feature | Person | Other)
object_type (Service | Team | API | Feature | Person | Other)
Return ONLY valid JSON. No prose, no markdown fences."""
def extract_triples(chunk_text: str) -> list[dict]:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": EXTRACTION_SYSTEM},
{"role": "user", "content": chunk_text},
],
temperature=0,
response_format={"type": "json_object"},
)
raw = response.choices[0].message.content
data = json.loads(raw)
# The model returns {"triples": [...]} or directly [...]
return data.get("triples", data) if isinstance(data, dict) else data
Stage 3 — Building the in-memory graph
Load extracted triples into NetworkX for development and experimentation:
import networkx as nx
def build_graph(all_triples: list[dict]) -> nx.DiGraph:
G = nx.DiGraph()
for triple in all_triples:
subj = triple["subject"]
obj = triple["object"]
pred = triple["predicate"]
# Add nodes with type metadata
G.add_node(subj, entity_type=triple.get("subject_type", "Other"))
G.add_node(obj, entity_type=triple.get("object_type", "Other"))
# Add edge; allow multiple edges with different predicates
G.add_edge(subj, obj, relation=pred)
return G
Inspect what you have before moving to retrieval:
G = build_graph(triples)
print(f"Nodes: {G.number_of_nodes()}, Edges: {G.number_of_edges()}")
# Sample: what does the payments_service depend on?
deps = [
(u, v, d["relation"])
for u, v, d in G.out_edges("payments_service", data=True)
if d["relation"] == "DEPENDS_ON"
]
print(deps)
# [('payments_service', 'auth_service', 'DEPENDS_ON'),
# ('payments_service', 'ledger_service', 'DEPENDS_ON')]
NetworkX is ideal for experimentation, but it lives entirely in memory and cannot be queried by multiple processes. For anything beyond a prototype, persist the graph.
Option A — NetworkX + GEXF file (prototype, single process)
import networkx as nx
# Save
nx.write_gexf(G, "knowledge_graph.gexf")
# Load
G = nx.read_gexf("knowledge_graph.gexf")
Option B — Neo4j (recommended for production)
Neo4j is a property graph database with a declarative query language called Cypher. The Community Edition runs locally with no licence fee.
from neo4j import GraphDatabase
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
def upsert_triple(tx, subj, subj_type, pred, obj, obj_type):
# MERGE prevents duplicate nodes/edges on re-ingestion
query = """
MERGE (s {name: $subj})
SET s.entity_type = $subj_type
MERGE (o {name: $obj})
SET o.entity_type = $obj_type
MERGE (s)-[r:RELATIONSHIP {type: $pred}]->(o)
RETURN s, r, o
"""
tx.run(query, subj=subj, subj_type=subj_type,
pred=pred, obj=obj, obj_type=obj_type)
def load_triples_to_neo4j(triples: list[dict]):
with driver.session() as session:
for triple in triples:
session.execute_write(
upsert_triple,
triple["subject"], triple.get("subject_type", "Other"),
triple["predicate"],
triple["object"], triple.get("object_type", "Other"),
)
Cypher queries read naturally once you see the arrow-based syntax:
-- Which teams own services that depend on auth_service?
MATCH (t:Team)-[:OWNS]->(s:Service)-[:DEPENDS_ON]->(dep {name: "auth_service"})
RETURN t.name AS team, s.name AS service
-- All services exposed by any team in the payments domain
MATCH (t:Team)-[:OWNS]->(s:Service)-[:EXPOSES]->(api:API)
WHERE t.name CONTAINS "payments"
RETURN t.name, s.name, api.name
The equivalent NetworkX traversal for the first query:
def teams_depending_on(G: nx.DiGraph, service_name: str) -> list[str]:
teams = []
# Find all services that depend on service_name
dependents = [
u for u, v, d in G.in_edges(service_name, data=True)
if d["relation"] == "DEPENDS_ON"
]
for svc in dependents:
# Find teams that own those services
owners = [
u for u, v, d in G.in_edges(svc, data=True)
if d["relation"] == "OWNS"
]
teams.extend(owners)
return list(set(teams))
Once the graph is queryable, you need a retrieval function that converts a natural-language question into a graph traversal, then returns the relevant context as text that can be passed to the LLM.
The pattern has three sub-steps: entity detection, traversal, and context materialisation.
Sub-step 1 — Detect graph entities in the query
def detect_entities_in_query(query: str, G: nx.DiGraph) -> list[str]:
"""Return graph node names mentioned in the query (substring match + normalise)."""
query_slug = query.lower().replace(" ", "_").replace("-", "_")
matched = []
for node in G.nodes():
# Check if the node name appears (as slug) in the query
if node in query_slug or node.replace("_", " ") in query.lower():
matched.append(node)
return matched
For production, replace the substring match with an embedding-based lookup against node names — this handles paraphrasing and abbreviations. A simple version:
from sentence_transformers import SentenceTransformer, util
import torch
model = SentenceTransformer("all-MiniLM-L6-v2")
# Pre-compute embeddings for all node names once at startup
node_names = list(G.nodes())
node_embeds = model.encode(node_names, convert_to_tensor=True)
def detect_entities_semantic(query: str, top_k: int = 3) -> list[str]:
q_embed = model.encode(query, convert_to_tensor=True)
scores = util.cos_sim(q_embed, node_embeds)[0]
top_idx = torch.topk(scores, k=top_k).indices.tolist()
# Only return matches above a relevance threshold
return [node_names[i] for i in top_idx if scores[i] > 0.6]
Sub-step 2 — Traverse the graph
def graph_context_for_entities(
G: nx.DiGraph,
seed_entities: list[str],
max_hops: int = 2,
max_edges: int = 30,
) -> list[dict]:
"""
BFS from seed entities up to max_hops.
Returns a list of (subject, predicate, object) dicts.
"""
visited_edges = []
frontier = set(seed_entities)
visited_nodes = set(seed_entities)
for _ in range(max_hops):
next_frontier = set()
for node in frontier:
for u, v, data in G.out_edges(node, data=True):
if len(visited_edges) >= max_edges:
break
visited_edges.append({
"subject": u,
"predicate": data["relation"],
"object": v,
})
if v not in visited_nodes:
next_frontier.add(v)
visited_nodes.add(v)
for u, v, data in G.in_edges(node, data=True):
if len(visited_edges) >= max_edges:
break
visited_edges.append({
"subject": u,
"predicate": data["relation"],
"object": v,
})
if u not in visited_nodes:
next_frontier.add(u)
visited_nodes.add(u)
frontier = next_frontier
if not frontier:
break
return visited_edges
Sub-step 3 — Materialise as text context
def triples_to_text(triples: list[dict]) -> str:
lines = ["Knowledge graph facts:"]
for t in triples:
lines.append(f" - {t['subject']} {t['predicate']} {t['object']}")
return "\n".join(lines)
A representative output for a query about "payments_service":
Knowledge graph facts:
- payments_service DEPENDS_ON auth_service
- payments_service DEPENDS_ON ledger_service
- payments_service EXPOSES payments_api
- platform_team OWNS payments_service
- payments_api DOCUMENTED_IN runbook_payments_v3
This structured context goes into the LLM prompt alongside any retrieved vector chunks, and the model can answer relational questions accurately even if no single document chunk contained all the facts.
max_hops=2 and only increase it if you have specific query patterns that require deeper chains. Measure retrieval quality (recall@k from Lesson 7) before and after each increase.The complete hybrid pipeline merges graph context and vector chunk context before passing either to the LLM.
from chromadb import Client as ChromaClient
from openai import OpenAI
openai_client = OpenAI()
chroma_client = ChromaClient()
collection = chroma_client.get_collection("docs")
def embed_query(query: str) -> list[float]:
resp = openai_client.embeddings.create(
model="text-embedding-3-small",
input=query,
)
return resp.data[0].embedding
def vector_retrieve(query: str, k: int = 5) -> list[str]:
embedding = embed_query(query)
results = collection.query(
query_embeddings=[embedding],
n_results=k,
include=["documents"],
)
return results["documents"][0] # list of chunk strings
def hybrid_rag_answer(
query: str,
G: nx.DiGraph,
use_graph: bool = True,
vector_k: int = 5,
) -> str:
# --- Vector retrieval ---
chunks = vector_retrieve(query, k=vector_k)
vector_context = "\n\n---\n\n".join(chunks)
# --- Graph retrieval (conditional) ---
graph_context = ""
if use_graph:
seed_entities = detect_entities_semantic(query)
if seed_entities:
triples = graph_context_for_entities(G, seed_entities, max_hops=2)
graph_context = triples_to_text(triples)
# --- Assemble prompt ---
system_prompt = (
"You are a technical assistant. Answer the user's question using ONLY "
"the provided context. If the answer requires following a relationship "
"chain, use the Knowledge Graph Facts section. Cite sources where possible."
)
context_block = ""
if graph_context:
context_block += f"{graph_context}\n\n"
context_block += f"Retrieved document chunks:\n{vector_context}"
response = openai_client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": f"Context:\n{context_block}\n\nQuestion: {query}"},
],
temperature=0.1,
)
return response.choices[0].message.content
Routing logic — when to skip graph retrieval
Not every query benefits from graph traversal. Calling the entity detector on every query adds latency for no gain on purely semantic questions. A lightweight router:
RELATIONAL_SIGNALS = [
"who owns", "who manages", "which team", "depends on",
"part of", "related to", "connected to", "upstream",
"downstream", "responsible for", "reports to",
]
def query_needs_graph(query: str) -> bool:
q_lower = query.lower()
return any(signal in q_lower for signal in RELATIONAL_SIGNALS)
| Query | Router decision | Reason |
|---|---|---|
| "What does the retry_policy field do?" | Vector only | Semantic lookup |
| "Which team owns the service that exposes the billing API?" | Graph + Vector | Explicit ownership chain |
| "How do I configure TLS for the proxy?" | Vector only | Procedural lookup |
| "What services depend on the identity service?" | Graph + Vector | Dependency traversal |
| "What is the SLO for the payments API?" | Graph + Vector | Property via entity |
Before committing to the engineering overhead of a knowledge graph, measure whether it actually helps. Use the evaluation framework from Lesson 7 with a graph-specific test set.
Build a multi-hop golden set
Create 15–30 questions that require at least two relationship hops to answer correctly, alongside ground-truth answers and the entities involved. Example rows:
| Question | Ground-truth answer | Seed entities | Hops required |
|---|---|---|---|
| "Which team is responsible for the service that the checkout API calls for fraud checks?" | fraud_team | checkout_api | 2 |
| "What documentation covers the database that stores user sessions?" | session_db_runbook | user_sessions_db | 1 |
| "List all services that a downstream failure in the ledger service could affect." | [payments_service, billing_service] | ledger_service | 2 |
Measure recall and answer correctness
def evaluate_multi_hop(
test_cases: list[dict],
G: nx.DiGraph,
use_graph: bool,
) -> dict:
correct = 0
total = len(test_cases)
for case in test_cases:
answer = hybrid_rag_answer(
case["question"], G, use_graph=use_graph
)
# Simple string-match correctness — replace with LLM-judge for nuance
expected = case["ground_truth"].lower()
if expected in answer.lower():
correct += 1
return {
"accuracy": correct / total,
"correct": correct,
"total": total,
"mode": "graph+vector" if use_graph else "vector_only",
}
# Compare modes
results_graph = evaluate_multi_hop(test_cases, G, use_graph=True)
results_vector = evaluate_multi_hop(test_cases, G, use_graph=False)
print(results_graph) # e.g. {'accuracy': 0.78, 'mode': 'graph+vector'}
print(results_vector) # e.g. {'accuracy': 0.31, 'mode': 'vector_only'}
What good looks like
On a well-constructed test set of genuinely multi-hop questions, vector-only RAG typically scores below 40% accuracy. Adding the graph layer should push this above 70%. If your improvement is modest, the most common causes are:
- Entity normalisation failures (duplicate nodes for the same entity)
- Missing relationship types (the relevant predicate was never extracted)
- Graph construction coverage gaps (the source documents for key facts were not processed)
- Max hops too low (the answer requires three hops but you stop at two)
A knowledge graph is a meaningful engineering investment. Before building one, apply this decision framework:
Strong signals that a graph is worth it
- Your corpus has a clear, stable entity taxonomy (services, teams, APIs, products)
- You receive frequent questions that require following ownership or dependency chains
- Consistency matters: the same entity must always retrieve the same set of facts, not whichever chunk happened to score highest in ANN search
- You need to answer aggregation questions ("how many services depend on X?", "which teams are affected by Y outage?")
Signals that vector search is sufficient
- Questions are predominantly "find me prose about topic X"
- Your corpus has no strong relational structure (e.g., a collection of independent blog posts)
- Entity taxonomy is unclear or changes frequently
- Team capacity does not allow maintaining extraction pipelines as documents evolve
The maintenance cost is real
Every time your documents change, the graph needs re-extraction. You need:
- A re-ingestion pipeline that detects changed documents and updates affected triples
- Conflict resolution when two documents disagree on a relationship
- Version tracking (covered in Lesson 9: Production RAG)
For most teams, a pragmatic middle path is to maintain a manually curated "core graph" (key ownership and dependency relationships edited by humans) alongside LLM-extracted auxiliary facts from documents. The core graph is small (hundreds of edges), authoritative, and cheap to keep current. The auxiliary layer provides depth but its quality is measured and tolerated.
| Approach | Accuracy | Maintenance | Best for |
|---|---|---|---|
| Vector only | Moderate on relational Qs | Low | General-purpose Q&A |
| Fully automated KG | High ceiling, variable floor | High | Large, stable, well-structured corpora |
| Manual core + LLM auxiliary | High and reliable | Medium | Most production teams |
Questions & Answers
AUTH_SVC=auth_service, AUTHENTICATION_MICROSERVICE=auth_service; (3) for long-tail variants, use embedding similarity over node names to detect candidates with cosine similarity above 0.92, then manually approve merges. The synonym dictionary is the leverage point — once you have it, re-processing the whole graph is a batch job that takes minutes.Key Takeaways
-
Vector search has a structural blind spot — it cannot follow chains of relationships across documents. Questions involving ownership, dependencies, hierarchies, or multi-hop chains require a complementary retrieval mechanism.
-
Knowledge graphs make relationships explicit — entities become nodes, relationships become typed directed edges, and multi-hop questions become graph traversal queries that return precise, repeatable answers.
-
LLM-assisted extraction is the practical path — use spaCy for known entity types and a small LLM for typed relationship extraction, with strict entity normalisation applied before loading into the graph.
-
The hybrid architecture routes queries intelligently — a simple keyword or embedding-based router decides whether a query needs graph traversal, vector search, or both, keeping latency low for purely semantic queries.
-
Measure before committing — build a multi-hop golden test set and verify that adding the graph layer improves multi-hop accuracy by at least 20 percentage points over vector-only before investing in a production graph pipeline.
-
Maintenance is the real cost — graph edges must be kept current as documents change. A manually curated core graph covering your highest-value entities, supplemented by automated extraction, is more reliable and cheaper to operate than a fully automated graph at scale.
Next Steps: Lesson 9: Production RAG