Knowledge Graphs + RAG

50 min advanced Lesson 8

Learning Outcomes

  • Explain why vector search alone fails on multi-hop and relationship-dependent questions
  • Extract entities and relationships from documents to construct a knowledge graph
  • Query a knowledge graph programmatically to traverse typed relationships
  • Combine graph traversal with vector retrieval in a hybrid retrieval architecture
  • Evaluate when a knowledge graph layer adds enough value to justify the engineering cost

Lesson Plan

Segment Duration Topic
Intro 3 min Why vector search has a structural blind spot
Explain 7 min Knowledge graphs — entities, edges, and properties
Demo 8 min Extracting a graph from raw documents with an LLM
Demo 9 min Storing and querying the graph with NetworkX and Neo4j
Demo 10 min Graph-based retrieval for multi-hop questions
Demo 8 min Hybrid graph + vector pipeline end-to-end
Explain 2 min When to reach for a graph vs. when not to
Wrap-up 3 min Key takeaways and what comes next

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10 or later
  • pip install networkx spacy sentence-transformers openai chromadb (or your preferred vector store from Lesson 3)
  • python -m spacy download en_core_web_sm (lightweight NER model)
  • Optional, for production graphs: a running Neo4j instance (Community Edition is free) and pip install neo4j
  • An OpenAI-compatible API key for LLM-assisted entity extraction in Step 3

1 The Structural Blind Spot of Vector Search

Vector similarity finds chunks whose surface meaning is close to the query. That works well for direct lookup questions: "What does the retry_policy configuration field do?" But it breaks down on a class of questions that require following a chain of relationships rather than recognising textual proximity.

Consider a corpus of engineering documents. The following questions all require traversal, not similarity:

Question type Example Why vector search struggles
Multi-hop "Who owns the service that exposes the payments API?" "Owner" and "payments API" may never appear in the same chunk
Aggregation over relationships "Which teams depend on the auth service?" Requires collecting every document that mentions a dependency
Hierarchical "What sub-components belong to the checkout domain?" Hierarchy is implicit in headings and prose, not encoded in any single chunk
Contradiction detection "Does any document say feature X is deprecated?" Contradiction only visible after collecting all statements about X

A knowledge graph (KG) makes these relationships explicit. In a KG:

  • Entities are named things: services, teams, people, components, features
  • Edges are typed, directed relationships: OWNS, DEPENDS_ON, PART_OF, DEPRECATED_BY
  • Properties are scalar attributes on entities or edges: version numbers, dates, status flags

Once the graph exists, multi-hop questions become graph traversal problems. "Who owns the service that exposes the payments API?" becomes: find the node labelled payments_api, follow the EXPOSED_BY edge to a Service node, then follow its OWNED_BY edge to a Team node.

NOTE
Key Insight
A knowledge graph does not replace your vector store — it complements it. Vector search handles the 'find content about topic X' part. The graph handles the 'follow the chain from X to Y to Z' part. Most real queries need both.

The architecture you will build by the end of this lesson looks like this:

User query
    │
    ├─── Entity detector ────► Graph traversal ──► Graph context
    │                               │
    └─── Query embedder ────► Vector search  ──► Chunk context
                                    │
                              Context merger
                                    │
                              LLM generation
                                    │
                              Grounded answer

The entity detector decides whether the query names known graph entities. If it does, graph traversal runs alongside — or instead of — vector search. The results are merged before generation.

TIP
Start Small
You do not need to extract every possible relationship from day one. Start with the two or three relationship types that your known failure-mode queries require, then expand incrementally.

2 Knowledge Graph Fundamentals — Entities, Edges, and Properties

Before writing code, it helps to be precise about the data model. Most knowledge graphs follow one of two storage paradigms: property graphs (nodes and edges can carry arbitrary key-value properties) and RDF triples (subject-predicate-object). For RAG-adjacent work, property graphs are nearly always the right choice — they map naturally to JSON and are supported by Neo4j, NetworkX, Amazon Neptune, and Memgraph.

A property graph schema for a software documentation corpus might look like this:

# Node labels and their core properties
(:Service)      {name, status, version, description}
(:Team)         {name, chat_channel, oncall_rotation}
(:API)          {name, endpoint, protocol, slo_p99_ms}
(:Document)     {url, title, ingested_at, chunk_count}

# Relationship types
(:Team)-[:OWNS]->(:Service)
(:Service)-[:EXPOSES]->(:API)
(:Service)-[:DEPENDS_ON]->(:Service)
(:API)-[:DOCUMENTED_IN]->(:Document)
(:Document)-[:MENTIONS_ENTITY]->(:Service | :Team | :API)

The MENTIONS_ENTITY edges link graph nodes back to the document chunks that describe them — this is the bridge that lets you retrieve prose context for any entity you find through graph traversal.

Choosing relationship types is the most consequential design decision. A useful heuristic: model only relationships you expect to query directionally. "Team A owns Service B" is worth modelling. "Document A mentions the word 'latency'" is not — vector search handles that better.

WARNING
Schema Sprawl
It is tempting to extract every relationship an LLM can name. Resist this. Graphs with dozens of poorly-defined relationship types are expensive to maintain and hard to query reliably. Five well-defined types beat twenty vague ones.

The table below maps common document corpora to the relationship types worth extracting:

Document corpus High-value relationship types
Engineering runbooks OWNS, DEPENDS_ON, ALERTS_ON, ESCALATES_TO
API documentation EXPOSES, CONSUMES, RETURNS, AUTHENTICATED_BY
HR / org charts REPORTS_TO, MEMBER_OF, RESPONSIBLE_FOR
Research papers CITES, EXTENDS, CONTRADICTS, EVALUATES
Product specs PART_OF, REQUIRES, SUPERSEDES, TESTED_BY

3 Extracting a Graph from Documents

Graph construction has two stages: entity recognition (which named things appear in this text?) and relationship extraction (how do those entities relate to each other?).

Stage 1 — Named entity recognition with spaCy

For well-defined entity types such as proper nouns, product names, and person names, a lightweight spaCy NER model is fast and accurate enough:

import spacy

nlp = spacy.load("en_core_web_sm")

def extract_entities_spacy(text: str) -> list[dict]:
    doc = nlp(text)
    entities = []
    for ent in doc.ents:
        if ent.label_ in {"ORG", "PRODUCT", "PERSON", "GPE", "WORK_OF_ART"}:
            entities.append({
                "text": ent.text,
                "label": ent.label_,
                "start": ent.start_char,
                "end": ent.end_char,
            })
    return entities

spaCy is fast (thousands of documents per second on a CPU) but only recognises entity types it was trained on. For domain-specific entities — service names, internal project codenames, custom configuration keys — you need the next stage.

Stage 2 — LLM-assisted relationship extraction

Use an LLM to extract typed (subject, predicate, object) triples from each chunk. Structure the output so it is directly loadable into your graph store:

import json
from openai import OpenAI

client = OpenAI()

EXTRACTION_SYSTEM = """You are a knowledge-graph extraction engine.
Given a passage of technical documentation, extract all relationships
as JSON triples. Use only these relationship types:
OWNS, DEPENDS_ON, EXPOSES, PART_OF, SUPERSEDES, DOCUMENTED_IN.

Output a JSON array of objects with keys:
  subject  (entity name, normalised to snake_case)
  predicate (one of the allowed types)
  object    (entity name, normalised to snake_case)
  subject_type  (Service | Team | API | Feature | Person | Other)
  object_type   (Service | Team | API | Feature | Person | Other)

Return ONLY valid JSON. No prose, no markdown fences."""

def extract_triples(chunk_text: str) -> list[dict]:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": EXTRACTION_SYSTEM},
            {"role": "user", "content": chunk_text},
        ],
        temperature=0,
        response_format={"type": "json_object"},
    )
    raw = response.choices[0].message.content
    data = json.loads(raw)
    # The model returns {"triples": [...]} or directly [...]
    return data.get("triples", data) if isinstance(data, dict) else data
TIP
Batch for Cost
LLM extraction is the expensive step. Process chunks in batches of 5-10 using the Batch API (or your provider's equivalent) to cut costs by roughly 50% with no quality penalty.

Stage 3 — Building the in-memory graph

Load extracted triples into NetworkX for development and experimentation:

import networkx as nx

def build_graph(all_triples: list[dict]) -> nx.DiGraph:
    G = nx.DiGraph()
    for triple in all_triples:
        subj = triple["subject"]
        obj  = triple["object"]
        pred = triple["predicate"]
        # Add nodes with type metadata
        G.add_node(subj, entity_type=triple.get("subject_type", "Other"))
        G.add_node(obj,  entity_type=triple.get("object_type", "Other"))
        # Add edge; allow multiple edges with different predicates
        G.add_edge(subj, obj, relation=pred)
    return G

Inspect what you have before moving to retrieval:

G = build_graph(triples)
print(f"Nodes: {G.number_of_nodes()}, Edges: {G.number_of_edges()}")

# Sample: what does the payments_service depend on?
deps = [
    (u, v, d["relation"])
    for u, v, d in G.out_edges("payments_service", data=True)
    if d["relation"] == "DEPENDS_ON"
]
print(deps)
# [('payments_service', 'auth_service', 'DEPENDS_ON'),
#  ('payments_service', 'ledger_service', 'DEPENDS_ON')]
WARNING
Entity Normalisation
The LLM will name the same entity differently across chunks: 'Auth Service', 'auth-service', 'the authentication service'. Apply lowercase + slug normalisation before adding to the graph, or you will end up with dozens of disconnected duplicate nodes.

4 Persisting and Querying the Graph

NetworkX is ideal for experimentation, but it lives entirely in memory and cannot be queried by multiple processes. For anything beyond a prototype, persist the graph.

Option A — NetworkX + GEXF file (prototype, single process)

import networkx as nx

# Save
nx.write_gexf(G, "knowledge_graph.gexf")

# Load
G = nx.read_gexf("knowledge_graph.gexf")

Option B — Neo4j (recommended for production)

Neo4j is a property graph database with a declarative query language called Cypher. The Community Edition runs locally with no licence fee.

from neo4j import GraphDatabase

driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

def upsert_triple(tx, subj, subj_type, pred, obj, obj_type):
    # MERGE prevents duplicate nodes/edges on re-ingestion
    query = """
    MERGE (s {name: $subj})
    SET s.entity_type = $subj_type
    MERGE (o {name: $obj})
    SET o.entity_type = $obj_type
    MERGE (s)-[r:RELATIONSHIP {type: $pred}]->(o)
    RETURN s, r, o
    """
    tx.run(query, subj=subj, subj_type=subj_type,
           pred=pred, obj=obj, obj_type=obj_type)

def load_triples_to_neo4j(triples: list[dict]):
    with driver.session() as session:
        for triple in triples:
            session.execute_write(
                upsert_triple,
                triple["subject"], triple.get("subject_type", "Other"),
                triple["predicate"],
                triple["object"],  triple.get("object_type", "Other"),
            )

Cypher queries read naturally once you see the arrow-based syntax:

-- Which teams own services that depend on auth_service?
MATCH (t:Team)-[:OWNS]->(s:Service)-[:DEPENDS_ON]->(dep {name: "auth_service"})
RETURN t.name AS team, s.name AS service

-- All services exposed by any team in the payments domain
MATCH (t:Team)-[:OWNS]->(s:Service)-[:EXPOSES]->(api:API)
WHERE t.name CONTAINS "payments"
RETURN t.name, s.name, api.name

The equivalent NetworkX traversal for the first query:

def teams_depending_on(G: nx.DiGraph, service_name: str) -> list[str]:
    teams = []
    # Find all services that depend on service_name
    dependents = [
        u for u, v, d in G.in_edges(service_name, data=True)
        if d["relation"] == "DEPENDS_ON"
    ]
    for svc in dependents:
        # Find teams that own those services
        owners = [
            u for u, v, d in G.in_edges(svc, data=True)
            if d["relation"] == "OWNS"
        ]
        teams.extend(owners)
    return list(set(teams))
NOTE
When to Use Neo4j
Use Neo4j when: you have more than ~50k edges, you need concurrent read/write access, or you need complex multi-hop queries that would require nested loops in NetworkX. For corpora under a few hundred documents, NetworkX is perfectly adequate.

5 Graph-Based Retrieval for Multi-Hop Questions

Once the graph is queryable, you need a retrieval function that converts a natural-language question into a graph traversal, then returns the relevant context as text that can be passed to the LLM.

The pattern has three sub-steps: entity detection, traversal, and context materialisation.

Sub-step 1 — Detect graph entities in the query

def detect_entities_in_query(query: str, G: nx.DiGraph) -> list[str]:
    """Return graph node names mentioned in the query (substring match + normalise)."""
    query_slug = query.lower().replace(" ", "_").replace("-", "_")
    matched = []
    for node in G.nodes():
        # Check if the node name appears (as slug) in the query
        if node in query_slug or node.replace("_", " ") in query.lower():
            matched.append(node)
    return matched

For production, replace the substring match with an embedding-based lookup against node names — this handles paraphrasing and abbreviations. A simple version:

from sentence_transformers import SentenceTransformer, util
import torch

model = SentenceTransformer("all-MiniLM-L6-v2")

# Pre-compute embeddings for all node names once at startup
node_names  = list(G.nodes())
node_embeds = model.encode(node_names, convert_to_tensor=True)

def detect_entities_semantic(query: str, top_k: int = 3) -> list[str]:
    q_embed = model.encode(query, convert_to_tensor=True)
    scores  = util.cos_sim(q_embed, node_embeds)[0]
    top_idx = torch.topk(scores, k=top_k).indices.tolist()
    # Only return matches above a relevance threshold
    return [node_names[i] for i in top_idx if scores[i] > 0.6]

Sub-step 2 — Traverse the graph

def graph_context_for_entities(
    G: nx.DiGraph,
    seed_entities: list[str],
    max_hops: int = 2,
    max_edges: int = 30,
) -> list[dict]:
    """
    BFS from seed entities up to max_hops.
    Returns a list of (subject, predicate, object) dicts.
    """
    visited_edges = []
    frontier = set(seed_entities)
    visited_nodes = set(seed_entities)

    for _ in range(max_hops):
        next_frontier = set()
        for node in frontier:
            for u, v, data in G.out_edges(node, data=True):
                if len(visited_edges) >= max_edges:
                    break
                visited_edges.append({
                    "subject":   u,
                    "predicate": data["relation"],
                    "object":    v,
                })
                if v not in visited_nodes:
                    next_frontier.add(v)
                    visited_nodes.add(v)
            for u, v, data in G.in_edges(node, data=True):
                if len(visited_edges) >= max_edges:
                    break
                visited_edges.append({
                    "subject":   u,
                    "predicate": data["relation"],
                    "object":    v,
                })
                if u not in visited_nodes:
                    next_frontier.add(u)
                    visited_nodes.add(u)
        frontier = next_frontier
        if not frontier:
            break

    return visited_edges

Sub-step 3 — Materialise as text context

def triples_to_text(triples: list[dict]) -> str:
    lines = ["Knowledge graph facts:"]
    for t in triples:
        lines.append(f"  - {t['subject']} {t['predicate']} {t['object']}")
    return "\n".join(lines)

A representative output for a query about "payments_service":

Knowledge graph facts:
  - payments_service DEPENDS_ON auth_service
  - payments_service DEPENDS_ON ledger_service
  - payments_service EXPOSES payments_api
  - platform_team OWNS payments_service
  - payments_api DOCUMENTED_IN runbook_payments_v3

This structured context goes into the LLM prompt alongside any retrieved vector chunks, and the model can answer relational questions accurately even if no single document chunk contained all the facts.

TIP
Cap Traversal Depth
Beyond two hops, traversal often pulls in loosely-related noise. Start with max_hops=2 and only increase it if you have specific query patterns that require deeper chains. Measure retrieval quality (recall@k from Lesson 7) before and after each increase.

6 Hybrid Graph + Vector Retrieval Pipeline

The complete hybrid pipeline merges graph context and vector chunk context before passing either to the LLM.

from chromadb import Client as ChromaClient
from openai import OpenAI

openai_client = OpenAI()
chroma_client = ChromaClient()
collection = chroma_client.get_collection("docs")

def embed_query(query: str) -> list[float]:
    resp = openai_client.embeddings.create(
        model="text-embedding-3-small",
        input=query,
    )
    return resp.data[0].embedding

def vector_retrieve(query: str, k: int = 5) -> list[str]:
    embedding = embed_query(query)
    results = collection.query(
        query_embeddings=[embedding],
        n_results=k,
        include=["documents"],
    )
    return results["documents"][0]  # list of chunk strings

def hybrid_rag_answer(
    query: str,
    G: nx.DiGraph,
    use_graph: bool = True,
    vector_k: int = 5,
) -> str:
    # --- Vector retrieval ---
    chunks = vector_retrieve(query, k=vector_k)
    vector_context = "\n\n---\n\n".join(chunks)

    # --- Graph retrieval (conditional) ---
    graph_context = ""
    if use_graph:
        seed_entities = detect_entities_semantic(query)
        if seed_entities:
            triples = graph_context_for_entities(G, seed_entities, max_hops=2)
            graph_context = triples_to_text(triples)

    # --- Assemble prompt ---
    system_prompt = (
        "You are a technical assistant. Answer the user's question using ONLY "
        "the provided context. If the answer requires following a relationship "
        "chain, use the Knowledge Graph Facts section. Cite sources where possible."
    )
    context_block = ""
    if graph_context:
        context_block += f"{graph_context}\n\n"
    context_block += f"Retrieved document chunks:\n{vector_context}"

    response = openai_client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": system_prompt},
            {"role": "user",   "content": f"Context:\n{context_block}\n\nQuestion: {query}"},
        ],
        temperature=0.1,
    )
    return response.choices[0].message.content

Routing logic — when to skip graph retrieval

Not every query benefits from graph traversal. Calling the entity detector on every query adds latency for no gain on purely semantic questions. A lightweight router:

RELATIONAL_SIGNALS = [
    "who owns", "who manages", "which team", "depends on",
    "part of", "related to", "connected to", "upstream",
    "downstream", "responsible for", "reports to",
]

def query_needs_graph(query: str) -> bool:
    q_lower = query.lower()
    return any(signal in q_lower for signal in RELATIONAL_SIGNALS)
Query Router decision Reason
"What does the retry_policy field do?" Vector only Semantic lookup
"Which team owns the service that exposes the billing API?" Graph + Vector Explicit ownership chain
"How do I configure TLS for the proxy?" Vector only Procedural lookup
"What services depend on the identity service?" Graph + Vector Dependency traversal
"What is the SLO for the payments API?" Graph + Vector Property via entity
NOTE
Latency Budget
Graph traversal on an in-memory NetworkX graph with hundreds of nodes takes under 5 ms. Neo4j adds a network round-trip (typically 10–30 ms). Entity embedding lookup against a few thousand node names adds another 20–50 ms. The combined graph path rarely exceeds 100 ms — well within a typical 2-second RAG latency budget.

7 Evaluating the Graph Layer

Before committing to the engineering overhead of a knowledge graph, measure whether it actually helps. Use the evaluation framework from Lesson 7 with a graph-specific test set.

Build a multi-hop golden set

Create 15–30 questions that require at least two relationship hops to answer correctly, alongside ground-truth answers and the entities involved. Example rows:

Question Ground-truth answer Seed entities Hops required
"Which team is responsible for the service that the checkout API calls for fraud checks?" fraud_team checkout_api 2
"What documentation covers the database that stores user sessions?" session_db_runbook user_sessions_db 1
"List all services that a downstream failure in the ledger service could affect." [payments_service, billing_service] ledger_service 2

Measure recall and answer correctness

def evaluate_multi_hop(
    test_cases: list[dict],
    G: nx.DiGraph,
    use_graph: bool,
) -> dict:
    correct = 0
    total = len(test_cases)

    for case in test_cases:
        answer = hybrid_rag_answer(
            case["question"], G, use_graph=use_graph
        )
        # Simple string-match correctness — replace with LLM-judge for nuance
        expected = case["ground_truth"].lower()
        if expected in answer.lower():
            correct += 1

    return {
        "accuracy": correct / total,
        "correct": correct,
        "total": total,
        "mode": "graph+vector" if use_graph else "vector_only",
    }

# Compare modes
results_graph  = evaluate_multi_hop(test_cases, G, use_graph=True)
results_vector = evaluate_multi_hop(test_cases, G, use_graph=False)

print(results_graph)   # e.g. {'accuracy': 0.78, 'mode': 'graph+vector'}
print(results_vector)  # e.g. {'accuracy': 0.31, 'mode': 'vector_only'}

What good looks like

On a well-constructed test set of genuinely multi-hop questions, vector-only RAG typically scores below 40% accuracy. Adding the graph layer should push this above 70%. If your improvement is modest, the most common causes are:

  1. Entity normalisation failures (duplicate nodes for the same entity)
  2. Missing relationship types (the relevant predicate was never extracted)
  3. Graph construction coverage gaps (the source documents for key facts were not processed)
  4. Max hops too low (the answer requires three hops but you stop at two)
TIP
LLM-as-Judge for Graph Eval
String matching is brittle for multi-entity answers. Use an LLM judge to compare the generated answer against ground truth: prompt it with 'Does answer A correctly identify the same entities and relationships as the ground truth B? Reply YES or NO with a brief reason.' This scales better than manual review.
WARNING
Graph Quality Gates
Before enabling the graph path in production, verify: (1) entity normalisation produces fewer than 5% duplicate nodes on a random sample, (2) relationship extraction precision is above 80% (manually label 50 triples), and (3) multi-hop accuracy on your golden set improves by at least 20 percentage points over vector-only. If any gate fails, fix extraction first.

8 When to Use a Knowledge Graph — and When Not To

A knowledge graph is a meaningful engineering investment. Before building one, apply this decision framework:

Strong signals that a graph is worth it

  • Your corpus has a clear, stable entity taxonomy (services, teams, APIs, products)
  • You receive frequent questions that require following ownership or dependency chains
  • Consistency matters: the same entity must always retrieve the same set of facts, not whichever chunk happened to score highest in ANN search
  • You need to answer aggregation questions ("how many services depend on X?", "which teams are affected by Y outage?")

Signals that vector search is sufficient

  • Questions are predominantly "find me prose about topic X"
  • Your corpus has no strong relational structure (e.g., a collection of independent blog posts)
  • Entity taxonomy is unclear or changes frequently
  • Team capacity does not allow maintaining extraction pipelines as documents evolve

The maintenance cost is real

Every time your documents change, the graph needs re-extraction. You need:

  • A re-ingestion pipeline that detects changed documents and updates affected triples
  • Conflict resolution when two documents disagree on a relationship
  • Version tracking (covered in Lesson 9: Production RAG)

For most teams, a pragmatic middle path is to maintain a manually curated "core graph" (key ownership and dependency relationships edited by humans) alongside LLM-extracted auxiliary facts from documents. The core graph is small (hundreds of edges), authoritative, and cheap to keep current. The auxiliary layer provides depth but its quality is measured and tolerated.

Approach Accuracy Maintenance Best for
Vector only Moderate on relational Qs Low General-purpose Q&A
Fully automated KG High ceiling, variable floor High Large, stable, well-structured corpora
Manual core + LLM auxiliary High and reliable Medium Most production teams
NOTE
GraphRAG
Microsoft's open-source GraphRAG framework automates much of the extraction and community-detection pipeline described in this lesson. It is worth evaluating for corpora above a few thousand documents — it handles entity normalisation and hierarchical summarisation in a way that would take weeks to build from scratch.

Questions & Answers

Q: LLM entity extraction sounds expensive at scale. How do I make it practical for a corpus of 50,000 documents?
Three strategies compound well. First, run extraction only on chunks that a lightweight NER model (spaCy or a fine-tuned BERT NER) flags as entity-dense — this filters out pure prose from perhaps 60% of chunks. Second, use a small, cheap model (GPT-4o-mini, Haiku, or a locally-hosted 7B instruction model) rather than a frontier model — relationship extraction is a structured output task that smaller models handle reliably. Third, use the provider's batch API to cut per-token costs in half. Combined, extraction across 50k documents typically costs tens of dollars rather than hundreds. Cache all extracted triples so you only pay for re-extraction on document changes.
Q: My entity normalisation is messy — the same service appears as "AuthService", "auth-service", "auth_svc", and "the authentication microservice". How do I merge these?
This is the single hardest problem in knowledge graph construction. A layered approach works best: (1) apply a deterministic slug function first (lowercase, strip punctuation, replace spaces and hyphens with underscores); (2) build a synonym dictionary manually for your known high-value entities — AUTH_SVC=auth_service, AUTHENTICATION_MICROSERVICE=auth_service; (3) for long-tail variants, use embedding similarity over node names to detect candidates with cosine similarity above 0.92, then manually approve merges. The synonym dictionary is the leverage point — once you have it, re-processing the whole graph is a batch job that takes minutes.
Q: How do I keep the graph consistent when source documents are updated or deleted?
Track a mapping from each document (by URL or content hash) to the set of triples it produced. When a document changes, delete all triples that originated solely from that document, then re-extract from the new version. Use MERGE semantics in Neo4j (or equivalent upsert logic) so triples that are also supported by other documents survive the deletion. For deletions, cascade: removing a document removes its triples, which may disconnect nodes — detect and prune zero-degree nodes unless they are also referenced by other documents. This is the same problem as index maintenance in production RAG, which Lesson 9 covers in full.
Q: Should I store the graph in the same database as my vectors, or keep them separate?
Separate systems are usually the right call, for a simple reason: purpose-built vector databases (Chroma, Qdrant, Pinecone, pgvector) are optimised for ANN search but have no concept of graph traversal, while graph databases (Neo4j, Memgraph) have no native ANN index. Weaviate is a notable exception — it natively combines property graph semantics with vector search and is worth evaluating if you want a single system. Otherwise, operate them as independent services and merge results in application code, as shown in Step 6. The coordination logic is a few dozen lines and pays for itself in operational simplicity.
Q: I tried graph extraction on my corpus and the precision of extracted relationships is only around 60%. Is that good enough?
Sixty percent precision means four in ten extracted triples are incorrect. That is too low to trust graph-based answers, because even a single wrong relationship can send traversal down a completely wrong path. The floor for production use is around 80% precision. The fastest path to improvement: tighten the extraction prompt by providing five or six few-shot examples from your specific domain, remove ambiguous relationship types that the model consistently confuses (e.g., DEPENDS_ON vs USES vs CALLS are often conflated — merge them into a single type if your queries do not distinguish them), and add a post-extraction validation pass that checks basic sanity (entity types on both ends of each relationship type match the schema).

Key Takeaways

  1. Vector search has a structural blind spot — it cannot follow chains of relationships across documents. Questions involving ownership, dependencies, hierarchies, or multi-hop chains require a complementary retrieval mechanism.

  2. Knowledge graphs make relationships explicit — entities become nodes, relationships become typed directed edges, and multi-hop questions become graph traversal queries that return precise, repeatable answers.

  3. LLM-assisted extraction is the practical path — use spaCy for known entity types and a small LLM for typed relationship extraction, with strict entity normalisation applied before loading into the graph.

  4. The hybrid architecture routes queries intelligently — a simple keyword or embedding-based router decides whether a query needs graph traversal, vector search, or both, keeping latency low for purely semantic queries.

  5. Measure before committing — build a multi-hop golden test set and verify that adding the graph layer improves multi-hop accuracy by at least 20 percentage points over vector-only before investing in a production graph pipeline.

  6. Maintenance is the real cost — graph edges must be kept current as documents change. A manually curated core graph covering your highest-value entities, supplemented by automated extraction, is more reliable and cheaper to operate than a fully automated graph at scale.

Next Steps: Lesson 9: Production RAG