Skip to content

MODULE 36 — GraphRAG & Knowledge Graph Architecture

36.1 Why Flat Vector Search Fails for Enterprise Knowledge

Vector search is outstanding at one thing: given a query, retrieve the chunks of text most semantically similar to it. For a large fraction of enterprise questions, that is enough. "What is our vacation policy?" retrieves the HR policy chunk, the answer is in that chunk, done.

The failure emerges when the answer to a question is not in any single chunk but exists only between chunks — specifically, in the relationships connecting entities across them.

The Specific Failure Pattern

FLAT RAG FAILURE: THE CONNECTION GAP

Query: "Which contracts expose us to force majeure in APAC?"

Vector search retrieves:
  Chunk A (similarity 0.87): Contract #4421 — APAC Master Agreement
    "...the parties agree to standard commercial terms..."
  Chunk B (similarity 0.85): Force Majeure Policy Document
    "...force majeure events include natural disasters, war,
     government actions, and pandemic conditions..."
  Chunk C (similarity 0.82): Contract #4421 — Clause 14.3
    "...clause 14.3 (b) incorporates Exhibit C by reference..."
  Chunk D (similarity 0.79): Exhibit C — Governing Terms
    "...Section 8 of Exhibit C defines force majeure conditions
     applicable to this agreement..."

THE GAP: The answer requires knowing:
  Contract #4421 CONTAINS Clause 14.3
  Clause 14.3 REFERENCES Exhibit C
  Exhibit C CONTAINS Section 8
  Section 8 DEFINES force_majeure
  Contract #4421 APPLIES_TO Region APAC

None of the four retrieved chunks contains this connection chain.
Each chunk is individually relevant. The answer lives in the traversal.

Result: LLM synthesizes a hallucinated answer or says "I cannot determine this."
         The correct answer was in the corpus. Vector search could not find it.

This is not a prompt engineering problem. It is a structural mismatch between the retrieval mechanism and the information architecture of the question.

Three Query Types That Defeat Vector RAG

Multi-hop relationship queries require traversing a chain of references before the answer emerges. "Which customers have contracts with our subsidiary in Singapore?" requires: Customer → Contract → ContractingEntity → Subsidiary → Jurisdiction. Vector similarity finds contracts; it cannot traverse entity chains.

Entity-network queries ask about the combined state of multiple entities that share a relationship. "Which customers have contracts expiring in 90 days AND open high-severity tickets?" Vector search may retrieve customer records and ticket records independently, but cannot enforce the conjunction across two separate entity types — it has no concept of entity identity that spans chunks.

Corpus-wide thematic queries ask about patterns across an entire corpus that no single document captures. "What are the common themes across all our incident reports from Q4?" This cannot be answered by retrieving the top-k most similar chunks — it requires a global view of the corpus, not local chunk similarity. Vector search returns the most similar documents; it cannot aggregate thematic structure across thousands of them.

The Fundamental Mismatch

Vector similarity measures: how similar is this text to the query?

Enterprise knowledge questions often ask: what is the relationship between these entities across this document set?

These are categorically different operations. Cosine similarity in an embedding space does not model directed relationships, entity identity across documents, multi-hop traversal, or corpus-wide structural patterns. It models local semantic proximity. That is a powerful primitive — but it is not a graph.

The architectural implication: enterprises with entity-rich corpora (legal, financial, biomedical, supply chain, compliance) will eventually encounter the class of queries that flat vector RAG structurally cannot answer. The question is whether the architect builds for this before or after users notice.


36.2 What GraphRAG Actually Is

The Microsoft GraphRAG Paper (2024)

In April 2024, Microsoft Research published "From Local to Global: A Graph RAG Approach to Query-Focused Summarization" (Edge et al., 2024). The accompanying open-source library (graphrag on PyPI, github.com/microsoft/graphrag) followed shortly after. This was not a minor paper — it defined a new architectural pattern that has become the reference implementation for knowledge-graph-augmented retrieval.

The paper's core insight: community detection applied to an entity-relationship graph extracted from a corpus can generate hierarchical summaries that answer corpus-wide questions no single document contains. This solved the third failure category from Section 36.1 (corpus-wide thematic queries) with a tractable approach.

Two Retrieval Modes: Local and Global

Local search is entity-anchored retrieval. Given a query mentioning specific entities, local search: 1. Identifies the entities referenced in the query 2. Retrieves those entity reports (structured summaries generated during indexing) 3. Expands to 1-2 hop graph neighbors (related entities, their relationships) 4. Retrieves text chunks associated with those entities 5. Assembles the combined context and generates an answer

Local search excels at: "Tell me about Acme Corp's contracts," "What incidents involved the payment service?", "Who are the key decision-makers at Customer X?"

Global search is community-based retrieval. Given a query asking about themes or patterns across the entire corpus, global search: 1. Traverses the community hierarchy (communities generated by Leiden algorithm during indexing) 2. For each community, retrieves the pre-generated community summary 3. Scores each community summary for relevance to the query 4. Aggregates relevant summaries into a final answer

Global search excels at: "What are the dominant themes in our incident corpus?", "What risks appear most frequently across our vendor contracts?", "What topics have we written about most in technical documentation?"

GRAPHRAG LOCAL vs. GLOBAL SEARCH

LOCAL SEARCH (entity-anchored):
  Query ──► Entity Extraction ──► Entity Reports
                                        │
                               Graph Neighbor Expansion
                                        │
                               Associated Text Chunks
                                        │
                               Context Assembly ──► LLM ──► Answer

  Depth: 1-3 hops, anchored to specific entities
  Best for: "What do we know about X?" / "How are X and Y related?"

GLOBAL SEARCH (community-based):
  Query ──► Community Hierarchy Traversal
                    │
          Community Summaries (pre-built at index time)
                    │
          Relevance Scoring per Community
                    │
          Relevant Summary Aggregation ──► LLM ──► Answer

  Depth: Corpus-wide, no entity anchor required
  Best for: "What themes exist across the corpus?"

What GraphRAG Adds Over Vector RAG

  • Entity reports: structured summaries of each extracted entity, built at index time, including all mentions and relationships across the corpus
  • Relationship context: every retrieved entity carries its edge context — what it connects to and how
  • Community summaries: hierarchical abstractions over clusters of related entities, enabling corpus-wide reasoning
  • Temporal and provenance links: extracted relationships carry source chunk references, enabling answer grounding

What GraphRAG Is NOT

GraphRAG is not a replacement for vector RAG. It is an augmentation. For semantic similarity queries ("find documents about X"), vector search remains faster and cheaper. GraphRAG's value is additive: it handles the query types that vector RAG cannot, and for entity-anchored queries, enriches the context that vector RAG would have retrieved anyway.

Deploying GraphRAG does not mean removing your vector store. It means adding a graph layer and a router that directs queries to the right retrieval path.

The GraphRAG Pipeline

GRAPHRAG INDEXING PIPELINE

Document Corpus
    │
    ▼
[1] Chunking (fixed or semantic)
    │  Produces: text chunks with source references
    ▼
[2] Entity & Relationship Extraction (LLM call per chunk)
    │  Prompt: "Extract entities and relationships from this text.
    │           Entity types: Person, Organization, Contract, ...
    │           Relationship types: WORKS_FOR, CONTAINS, REFERENCES..."
    │  Produces: (entity, type, properties) tuples
    │            (entity_a, relationship, entity_b, properties) tuples
    ▼
[3] Entity Resolution
    │  "Apple" + "Apple Inc." + "AAPL" → same node
    │  Produces: deduplicated entity graph
    ▼
[4] Graph Construction
    │  Nodes: resolved entities with properties
    │  Edges: typed, directed relationships with confidence scores
    │  Produces: property graph in Neo4j / FalkorDB / NetworkX
    ▼
[5] Community Detection (Leiden algorithm)
    │  Groups densely connected entities into communities
    │  Builds hierarchy: Level 0 (finest) → Level N (coarsest)
    │  Produces: community assignments per entity
    ▼
[6] Community Summary Generation (LLM call per community)
    │  Summarizes what each community represents
    │  Produces: community reports (structured natural language)
    ▼
[7] Index
    │  Graph DB: entity nodes + relationship edges
    │  Vector store: chunk embeddings + entity report embeddings
    │  Community store: community summaries by level
    └──► Ready for local and global search

The Cost Reality

GraphRAG indexing is expensive because it requires LLM calls at steps 2 and 6.

For a 10,000-document corpus, assuming average 5 chunks per document (50,000 chunks), with an efficient-tier model at $0.10/1M input tokens and 800 tokens per extraction prompt (input tokens only; output tokens add more, see 36.7 for a full estimate):

  • Step 2 (extraction): 50,000 chunks × 800 tokens = 40M tokens × $0.10/1M → ~$4
  • Step 6 (community summaries): depends on community count, typically 500-2,000 communities × 1,500 tokens = 0.75-3M tokens × $0.10/1M → ~$0.08-0.30

For a 100,000-document corpus, these numbers scale linearly: ~$40 in LLM input-token cost for the extraction pass alone (400M tokens × $0.10/1M), before output tokens, before community summaries, before re-indexing when new documents arrive. With a mid-tier model ($2.00/1M input tokens), multiply by 20x: ~$800 for 100K documents (400M tokens × $2.00/1M).

Illustrative, October 2026 prices (efficient tier such as GPT-6 Luna; mid-tier such as Claude Sonnet 5.5 or GPT-6 Sol). Prices change quickly; see Appendix G, G.3 for current figures.

This cost is not a reason to avoid GraphRAG. It is a reason to evaluate whether your query patterns justify it before committing, and to design incremental indexing so you pay it once per document, not per rebuild.


36.3 Knowledge Graph Fundamentals for AI Architects

The Three Primitives

A property graph is built from three primitives. Architects who have not worked with graph databases need these internalized before building anything.

Nodes (entities) are typed objects in your domain. Every entity has a label (its type) and a set of properties (its attributes).

Node examples in an enterprise legal corpus:

  (:Person {name: "Sarah Chen", title: "VP Procurement", email: "..."})
  (:Organization {name: "Acme Corp", type: "customer", industry: "manufacturing"})
  (:Contract {id: "CON-4421", value: 2400000, expiry: "2025-03-31", status: "active"})
  (:Clause {id: "CL-4421-14", type: "force_majeure", text: "..."})
  (:Incident {id: "INC-8821", severity: "P1", status: "open", created: "2024-11-01"})
  (:Region {name: "APAC", countries: ["SG", "JP", "AU", "IN"]})

Edges (relationships) are directed, typed connections between nodes. An edge always has a direction, a type, and optionally properties.

Relationship examples:

  (Sarah Chen)-[:SIGNS]->(CON-4421)
  (Acme Corp)-[:PARTY_TO]->(CON-4421)
  (CON-4421)-[:CONTAINS]->(CL-4421-14)
  (CL-4421-14)-[:TYPE]->(force_majeure_clause_type)
  (CON-4421)-[:APPLIES_TO]->(APAC)
  (Acme Corp)-[:HAS_INCIDENT]->(INC-8821)

Relationship with properties:
  (Acme Corp)-[:PARTY_TO {role: "customer", signed_date: "2023-01-15"}]->(CON-4421)

Properties are key-value attributes on either nodes or edges. Properties on relationships are underused and valuable — they carry confidence scores, timestamps, source references, and metadata about the relationship itself.

Property Graph vs. RDF

Two data models compete in the knowledge graph space:

Property graphs (Neo4j, FalkorDB, Amazon Neptune property graph, Memgraph) store entities as nodes with arbitrary key-value properties and relationships as first-class objects with their own properties. Query language: Cypher (Neo4j) or Gremlin (TinkerPop). This is the right model for application development and AI retrieval — it is programmer-friendly, schema-flexible, and performant for traversal queries.

RDF (Resource Description Framework) represents everything as triples: (subject, predicate, object). No native property bags — properties are modeled as additional triples. Query language: SPARQL. This is the right model for data interchange, semantic web standards, and interoperability between organizations. It is less ergonomic for developers and slower for deep traversals. Use RDF when interoperability is the primary requirement; use property graphs for everything else.

When You Need a Graph Instead of Tables

A relational database is the right tool when data relationships are fixed, known, and two or three levels deep. It is the wrong tool when:

  • Relationship depth is variable and unbounded. Finding all upstream dependencies of a component requires a recursive CTE in SQL; it requires a single Cypher query in a graph. At 5+ hops, the SQL performance degrades non-linearly; the graph traversal does not.
  • Relationship structure itself is the data. In an org chart, a knowledge network, or a supply chain, the shape of the graph is what you are analyzing — not just the attribute values of individual records.
  • Schema evolves rapidly. Adding a new relationship type between existing entity types in a property graph requires no migration. The same change in a relational schema requires ALTER TABLE, index rebuilds, and application code changes.
  • Queries traverse many entity types. Joining seven tables in SQL for a single query is a warning sign. A graph traversal across seven entity types is ergonomic and readable Cypher.

The architectural test: if the answer to your most important business questions requires joining more than four tables, or if the join depth changes based on the data, your domain may be a graph problem.

Sample Enterprise Knowledge Graph Fragment

ENTERPRISE KNOWLEDGE GRAPH: LEGAL & CUSTOMER CORPUS

  [Acme Corp]──────────────PARTY_TO──────────────[Contract #4421]
      │                                                  │
      │                                           CONTAINS│
    HAS_INCIDENT                                         │
      │                                          [Clause 14.3]
      ▼                                                  │
  [Incident INC-8821]                            REFERENCES│
   severity: P1                                           │
   status: open                                   [Exhibit C]
   created: 2024-11-01                                    │
                                                 CONTAINS │
  [Sarah Chen]──────────SIGNS──────────────────►[Contract #4421]
   title: VP Procurement                                  │
   email: s.chen@acme.com                        APPLIES_TO│
                                                          ▼
  [Widget Corp]──────PARTY_TO──────────────────────────[APAC]
   (different contract,                          [Contract #7823]
    also APAC)                                         │
                                                CONTAINS│
                                             [Clause 8.1]
                                              type: force_majeure

This fragment makes the "Which APAC contracts have force majeure clauses?" query trivially resolvable by graph traversal. The same query is structurally unsolvable by flat vector search against the text of these documents.


36.4 Building a Knowledge Graph from Documents

The Extraction Pipeline

DOCUMENT-TO-GRAPH EXTRACTION PIPELINE

  Raw Documents (PDF, DOCX, HTML, text)
       │
       ▼ [1] Pre-processing & chunking
  Text Chunks (512-1024 tokens, semantic boundaries preferred)
       │
       ▼ [2] Named Entity Recognition (NER)
  Entity candidates: (text_span, entity_type, chunk_ref)
  e.g., ("Acme Corp", Organization, chunk_047)
       │
       ▼ [3] Relationship extraction
  Relationship triples: (entity_a, relationship_type, entity_b, confidence)
  e.g., ("Acme Corp", PARTY_TO, "Contract #4421", 0.94)
       │
       ▼ [4] Entity resolution ◄─── HARDEST STEP
  Canonical entities: ("Apple", "Apple Inc.", "AAPL") → node:apple_inc
       │
       ▼ [5] Graph merge (upsert)
  New entities → CREATE nodes
  Existing entities → MERGE and update properties
  New relationships → CREATE edges
  Confidence below threshold → flag for review
       │
       ▼ [6] Quality checkpoint
  Entity coverage check: known entities in source docs found?
  Relationship precision sample: 5% manual review
  Resolution audit: spot-check for entity merge errors
       │
       ▼
  Production Knowledge Graph

NER Approaches: Pick Based on Domain Complexity

Rule-based NER (spaCy, regex patterns, dictionaries): Fast, cheap, deterministic. Appropriate for well-defined entity types in structured text — product codes, contract IDs, dates, known organization names. Cannot extract novel entities or complex relationships. Use for high-volume preprocessing of common entity types.

Fine-tuned NER (GLiNER, BERT-based models): A single model fine-tuned to extract multiple entity types without per-type models. GLiNER in particular handles arbitrary entity types with a single model via a "zero-shot" approach — you define the types at inference time, not training time. Cost-effective for 10-50 entity types; requires labeled training data for the specific domain.

LLM-prompted extraction: Give an LLM a structured extraction prompt per chunk. Highest accuracy for complex relationships, nuanced entity types, and domains where entities are described implicitly rather than named explicitly. Significantly more expensive than the above. The right choice for high-value complex domains (legal, biomedical, M&A) where extraction quality directly affects answer quality.

Production pattern: rule-based for common entities (fast, cheap) + LLM for complex relationships (targeted, higher cost). The LLM pass runs only on chunks that pass an initial relevance filter, not on all chunks.

Relationship Extraction via LLM

A production extraction prompt is structured, typed, and returns JSON:

SYSTEM: You are an expert at extracting structured relationships from legal documents.
        Extract all relationships between entities. Be precise about relationship types.

ENTITY_TYPES: [Person, Organization, Contract, Clause, Exhibit, Region, Product]
RELATIONSHIP_TYPES: [PARTY_TO, CONTAINS, REFERENCES, APPLIES_TO, SIGNS, GOVERNED_BY,
                     SUPERSEDES, AMENDS, EXPIRES_ON, DEFINED_IN]

USER: Extract entities and relationships from the following text.
      Return JSON only.
      [CHUNK TEXT]

RESPONSE FORMAT:
{
  "entities": [
    {"id": "e1", "text": "Acme Corp", "type": "Organization", "confidence": 0.97},
    {"id": "e2", "text": "Contract #4421", "type": "Contract", "confidence": 0.99}
  ],
  "relationships": [
    {"from": "e1", "type": "PARTY_TO", "to": "e2", "confidence": 0.94,
     "evidence": "Acme Corp is identified as the Customer in this agreement"}
  ]
}

Every relationship carries a confidence score and an evidence span. Low-confidence relationships (below 0.70) are stored but excluded from default traversal queries; they are available for review.

Entity Resolution: The Hard Part

Entity resolution is the process of determining that multiple textual mentions refer to the same real-world entity. This is the step that most teams underestimate and most graph implementations get wrong.

"Apple," "Apple Inc.," "AAPL," "the Cupertino company," "Tim Cook's company," and "Apple Computer" may all refer to the same node in your graph. Whether they do depends on context.

Resolution strategy hierarchy:

  1. Exact match on canonical identifier — contract ID, employee ID, product SKU. No ambiguity. Apply first, always.
  2. Fuzzy string match (Levenshtein distance, token sort ratio) — "Acme Corporation" vs. "Acme Corp" vs. "ACME CORP." Threshold around 0.85 similarity, with manual review queue for 0.70-0.85.
  3. Embedding similarity — embed entity mention + context window, cosine similarity against existing entity embeddings. Catches semantic equivalents that string match misses ("the Cupertino company" → Apple Inc.).
  4. LLM-based semantic resolution — for ambiguous cases where context is required. Given two entity mentions with their surrounding text, ask the LLM whether they refer to the same entity. Expensive; reserve for high-confidence-required decisions.

Resolution errors are the primary source of graph quality failures. A false positive merge (two distinct entities treated as one) is worse than a false negative split (one entity treated as two) because it introduces incorrect relationships. When in doubt, split and review rather than merge.

Incremental Graph Building

The initial graph build is a one-time batch operation. Every document that enters the corpus after that must be processed incrementally. New documents must be resolved against the existing graph, not trigger a full rebuild.

Incremental pipeline: 1. New document arrives → extract entities and relationships (same pipeline as batch) 2. For each extracted entity: attempt resolution against existing graph nodes 3. If resolved → MERGE properties (new mentions, updated attributes), ADD new relationships 4. If not resolved → CREATE new node 5. Run community re-detection only on the affected subgraph neighborhood, not the full corpus 6. Regenerate community summaries only for communities that changed

Full re-indexing should be a scheduled operation (weekly or monthly), not triggered by every new document.

Quality Metrics

Track three numbers in production: - Entity coverage: of the known entities in your domain (from a reference list), what percentage appear correctly in the graph? Below 80% means the extraction is missing entities. - Relationship precision: of a random sample of extracted relationships, what percentage are correct? Below 85% means the LLM extraction needs prompt tuning or confidence threshold adjustment. - Resolution accuracy: of a random sample of entity merge decisions, what percentage are correct? Below 90% is a critical problem — every wrong merge corrupts the graph.


36.5 Graph-Enhanced Retrieval Patterns

Three patterns cover the production space. Choose based on query structure, not on what sounds most sophisticated.

Pattern 1 — Graph-Augmented RAG

GRAPH-AUGMENTED RAG

Query
  │
  ▼
Vector Search ──► Top-K chunks (standard dense retrieval)
  │
  ▼
Entity Extraction from retrieved chunks
  │
  ▼
Graph Neighbor Expansion (1-2 hops from extracted entities)
  │  Cypher: MATCH (e)-[r*1..2]-(neighbor)
  │          WHERE e.name IN $extracted_entities
  │          RETURN neighbor, r, type(r)
  ▼
Entity reports + neighbor context
  │
  ▼
Reranking: score combined set (chunks + graph context) for query relevance
  │
  ▼
LLM generation with enriched context
  │
  ▼
Answer

Decision criteria — use this pattern when: - Your existing vector RAG works for most queries but has gaps on entity-rich questions - You want to add graph enrichment without a full GraphRAG deployment - Entity context demonstrably improves answers in your domain - Query volume is high — this pattern adds moderate overhead per query

Typical latency overhead: +50-150ms for the graph expansion step vs. pure vector RAG

Pattern 2 — Graph-First Retrieval

GRAPH-FIRST RETRIEVAL

Query (natural language)
  │
  ▼
Text2Cypher (LLM converts NL query to graph query)
  │  "Which APAC contracts have force majeure clauses?"
  │  →
  │  MATCH (c:Contract)-[:CONTAINS]->(cl:Clause {type:'force_majeure'})
  │        (c)-[:APPLIES_TO]->(r:Region {name:'APAC'})
  │  RETURN c.id, c.value, c.expiry, cl.text
  ▼
Cypher execution against graph DB
  │
  ▼
Subgraph result (nodes + edges + properties)
  │
  ▼
Subgraph-to-text summarization (LLM converts subgraph to answer)
  │
  ▼
Answer with source references (contract IDs, clause references)

Decision criteria — use this pattern when: - The question is structurally a relationship query (who-to-what, what-contains-what, find-all-X-related-to-Y) - Answers require precision — partial matches are worse than no answer - Domain has well-defined, stable entity and relationship types - You can tolerate ~60-75% first-pass Text2Cypher accuracy with a validation fallback

The Text2Cypher accuracy problem: current LLMs generate correct Cypher for simple 1-2 hop queries with high reliability (~85-90%). Complex multi-hop queries with filters and aggregations drop to 60-70%. Production deployments must include a Cypher validation layer that catches syntax errors, checks referenced label and property names against the schema, and falls back to vector retrieval on failure.

Pattern 3 — Hybrid Router

HYBRID ROUTER

Query
  │
  ▼
Query Classifier
  │  Input: query text
  │  Output: {type: "structural"|"semantic"|"global", confidence: 0.87}
  │  Implementation: fine-tuned classifier OR LLM with few-shot examples
  │
  ├──[structural]──► Pattern 2 (Graph-First)
  │                  "Which contracts...", "Who are the parties to...",
  │                  "What incidents involve...", "Find all X related to Y"
  │
  ├──[semantic]────► Pattern 1 (Vector search, optional graph augment)
  │                  "Explain the meaning of...", "What does policy X say about...",
  │                  "Summarize the approach to..."
  │
  └──[global]──────► GraphRAG global search (community-based)
                     "What are the themes...", "What risks appear across...",
                     "How has X changed over time across the corpus..."

Results merged when classifier confidence < 0.70 (run both paths)
  │
  ▼
Unified answer with retrieval path attribution

Decision criteria — use this pattern when: - Corpus has both entity-rich structured content and semantic narrative content - Query mix is genuinely heterogeneous — users ask both types - Query volume is high enough that routing overhead (classifier call) is justified - Team can maintain both retrieval backends

Routing overhead: 50-100ms for the classification step. At low volume, simpler to run both paths and merge. At high volume, routing saves ~60% of graph traversal cost on semantic queries.


36.6 Graph Database Selection and Infrastructure

Candidate Platforms

Neo4j is the dominant enterprise graph database. Native graph storage engine, Cypher query language, mature ecosystem, GraphRAG integration via LangChain and LlamaIndex connectors. Available as AuraDB (managed, $0.08/GB/hour at time of writing) or self-hosted. Strong choice when team expertise matters — Cypher is widely known, community resources are abundant. The licensing model changed in 2024 (BSL for enterprise features); verify which features require a commercial license.

Amazon Neptune is AWS-managed, supporting both property graph (Gremlin query language) and RDF (SPARQL). No operational overhead for teams already on AWS. Gremlin is less ergonomic than Cypher; the RDF/SPARQL path is viable for standards-compliance use cases. Scale to billions of edges without manual sharding. Choose Neptune when AWS-native deployment, managed operations, and scale are the primary requirements.

FalkorDB is a Redis-compatible, in-memory graph database using sparse matrix graph representation. Significantly lower operational overhead than Neo4j — it runs as a Redis module, which most teams already operate. Excellent developer experience, Cypher-compatible query language. The right choice for teams that need a graph database without a dedicated graph database operator. Performance degrades on datasets that don't fit in memory; plan capacity accordingly.

Azure Cosmos DB for Gremlin is Azure-managed with Gremlin API. Natural choice for Azure-native organizations. Globally distributed with multi-region writes. Less mature GraphRAG integration than Neo4j. Choose when Azure commitment and global distribution outweigh query language preference.

Memgraph is an in-memory graph database with strong real-time streaming capabilities. Kafka integration is native. Appropriate when the graph needs continuous updates from a streaming event pipeline — IoT, financial transactions, operational monitoring. Less mature for large-scale batch indexing. Choose Memgraph when graph freshness (sub-second update latency) is a requirement.

Selection Matrix

GRAPH DATABASE SELECTION MATRIX

                  Neo4j   Neptune  FalkorDB  CosmosDB  Memgraph
                          (AWS)    (Redis)   (Azure)
─────────────────────────────────────────────────────────────────
Scale (>10B edges)  ●●●     ●●●      ●●       ●●●       ●●
Managed ops         ●●      ●●●      ●●●      ●●●       ●●
Dev experience      ●●●     ●●       ●●●      ●●        ●●●
GraphRAG support    ●●●     ●●       ●●●      ●●        ●●
Real-time updates   ●●      ●●       ●●       ●●        ●●●
AWS affinity        ●●      ●●●      ●●       ●         ●●
Azure affinity      ●●      ●        ●●       ●●●       ●●
Cost (relative)     $$$     $$       $        $$        $$
Schema flexibility  ●●●     ●●       ●●●      ●●        ●●●

●●● = excellent  ●● = good  ● = limited

Schema Design: Ontology vs. Flexible Graph

The degree of schema enforcement is a design choice with significant downstream consequences.

Strict ontology (defined node labels, allowed relationship types, required properties): Consistent graph that supports reliable querying, enforces data quality, enables type-checked Cypher queries. High upfront investment in schema design. Appropriate for stable domains (legal contracts, financial instruments) where the entity model is well understood.

Flexible property graph (any label, any relationship type, any properties): Fast to build, easy to extend, accommodates unexpected relationship types discovered during extraction. Results in inconsistent graphs where the same concept appears under different labels ("Vendor" vs. "Supplier" vs. "ThirdPartyProvider"). Appropriate for exploratory phases or rapidly evolving domains.

Production recommendation: start with a flexible graph during the first 8 weeks (discovery phase); define a canonical ontology by week 12 and migrate the graph to it. The flexibility lets you learn what entity types the corpus actually contains before committing to a schema.


36.7 Entity Extraction at Scale: Production Considerations

The Cost Problem

At scale, LLM-based entity extraction is a significant budget line item. A 100,000-document corpus processed with a mid-tier model (illustrative, October 2026 prices; see Appendix G, G.3):

EXTRACTION COST ESTIMATE: 100K DOCUMENTS

Assumptions:
  Average document length: 3,000 tokens
  Chunks per document: 6 (at 512 tokens/chunk)
  Total chunks: 600,000
  Extraction prompt overhead: 400 tokens (system + instructions)
  Per-chunk LLM input: 400 + 512 = 912 tokens
  Per-chunk LLM output: ~300 tokens (entities + relationships JSON)
  Mid-tier model: $2.00/1M input, $10.00/1M output
                  (e.g., Claude Sonnet 5.5, GPT-6 Sol)
  Efficient-tier model: $0.10/1M input, $0.50/1M output
                  (e.g., GPT-6 Luna)

Mid-tier extraction:
  Input cost:  600,000 × 912 tokens = 547M tokens × $2.00/1M → $1,094
  Output cost: 600,000 × 300 tokens = 180M tokens × $10.00/1M → $1,800
  Total extraction cost: ~$2,894

Community summary generation (est. 5,000 communities × 2,000 tokens):
  Input: 10M tokens × $2.00/1M → $20
  Output: 10M tokens × $10.00/1M → $100 (summary generation)
  Total community cost: ~$120

TOTAL INDEXING COST (mid-tier): ~$3,000 for 100,000 documents

Efficient-tier extraction + summaries:
  Extraction: 547M × $0.10/1M + 180M × $0.50/1M = $55 + $90 → ~$145
  Community:  10M × $0.10/1M + 10M × $0.50/1M = $1 + $5 → ~$6
TOTAL INDEXING COST (efficient tier): ~$150 for 100,000 documents

Decision point: use the efficient tier for bulk extraction + the mid-tier for
               ambiguous/high-value documents (hybrid cost optimization)

Hybrid Extraction Architecture

HYBRID EXTRACTION: COST-OPTIMIZED PIPELINE

All Chunks
    │
    ▼
[Rule-based NER pass] — fast, $0, handles:
    ├── Contract IDs (regex: CON-\d{4,})
    ├── Dates (spaCy date NER)
    ├── Known organization names (lookup against entity dictionary)
    ├── Person names (spaCy person NER)
    └── Monetary values (regex + spaCy)
    │
    ▼
[Relevance classifier] — identifies chunks needing LLM extraction
    │  "Does this chunk contain complex relationships or
    │   entity types not covered by rules?"
    │  Implementation: lightweight BERT classifier (not LLM)
    │
    ├──[low complexity]──► Skip LLM, use rule-based output only
    │
    └──[high complexity]──► LLM extraction (efficient-tier model default,
                             mid-tier model for flagged high-value documents)

This hybrid approach reduces LLM extraction calls to 30-40% of chunks, cutting costs by 60-70% while preserving quality on complex content.

Batch vs. Streaming

Initial indexing: batch. Process all existing documents in parallel workers. Use a job queue (Celery, AWS Batch, Kubernetes Jobs) with checkpointing — if a batch fails at chunk 45,000 of 600,000, resume from 45,000, not from the beginning.

Ongoing ingestion: streaming. New documents arrive continuously and must update the graph without triggering a full rebuild.

STREAMING GRAPH UPDATE PIPELINE

New Document
    │
    ▼
[Event trigger] (document uploaded, API webhook, email ingestion)
    │
    ▼
Kafka topic: document.ingested
    │
    ▼
[Extraction service] (horizontally scaled, stateless)
  Consumes document → runs extraction pipeline → emits graph updates
    │
    ▼
Kafka topic: graph.updates
    │
    ▼
[Graph writer service]
  Consumes updates → resolves against existing graph → upserts to graph DB
    │
    ▼
[Community re-detection trigger]
  If entity count changed by >N or new high-degree node added →
  trigger partial community re-detection on affected neighborhood

Graph Freshness and Staleness

A stale graph produces wrong answers as confidently as a fresh one. The graph does not signal its own staleness to the LLM.

Define a graph freshness SLA for your use case: - Legal/contracts: graph must reflect documents within 24 hours of ingestion - Operational/incident: graph must reflect incidents within 1 hour - Knowledge base/documentation: weekly rebuild acceptable

Monitor: track the lag between document ingestion timestamp and graph node creation timestamp. Alert when lag exceeds SLA. A graph where last night's incidents are not yet in the graph will answer incident-related questions incorrectly without any visible signal of the error.


36.8 Multi-Hop Reasoning: Architecture and Limits

What Multi-Hop Reasoning Requires

Multi-hop reasoning requires three components working together:

  1. A connected graph — entities and their relationships correctly captured
  2. A query language that can traverse it — Cypher's variable-length path patterns
  3. A reasoning step that interprets the subgraph — an LLM that can read a subgraph and synthesize an answer from it

All three must be in place. A perfect graph with no traversal capability is useless. A perfect traversal returning a giant subgraph with no synthesis step produces raw graph data, not answers.

Cypher Multi-Hop Example

The canonical enterprise query from Section 36.1, expressed as Cypher:

-- "Which contracts expose us to force majeure in APAC?"

MATCH (org:Organization)-[:PARTY_TO]->(c:Contract)
      -[:CONTAINS]->(cl:Clause)
      -[:APPLIES_TO]->(r:Region {name: 'APAC'})
WHERE cl.type = 'force_majeure'
  AND c.status = 'active'
RETURN org.name         AS customer,
       c.id             AS contract_id,
       c.value          AS contract_value,
       c.expiry         AS expiry_date,
       cl.id            AS clause_id,
       cl.text          AS clause_text
ORDER BY c.expiry ASC

-- "Which customers have active contracts expiring in 90 days AND open P1 incidents?"

MATCH (org:Organization)-[:PARTY_TO]->(c:Contract),
      (org)-[:HAS_INCIDENT]->(inc:Incident)
WHERE c.status = 'active'
  AND c.expiry <= date() + duration({days: 90})
  AND inc.severity = 'P1'
  AND inc.status = 'open'
RETURN org.name, c.id, c.expiry, count(inc) AS open_p1_count
ORDER BY c.expiry ASC

These queries return precise, grounded answers — specific contract IDs, customer names, expiry dates. No hallucination is possible because the answer is the graph traversal result. The LLM's role in this pattern is converting the query to Cypher and summarizing the result, not synthesizing information from semantic similarity.

The Text2Cypher Problem

Natural language to Cypher generation is the primary reliability challenge in graph-first retrieval. Current performance benchmarks:

  • Simple 1-hop queries ("Find all contracts for Acme Corp"): ~90% accuracy
  • 2-hop with single filter ("Find APAC contracts expiring this year"): ~80% accuracy
  • 3-hop with multiple filters and aggregation: ~60-70% accuracy
  • Complex analytical queries with subqueries: ~45-60% accuracy

Required validation layer for production:

TEXT2CYPHER VALIDATION PIPELINE

NL Query ──► LLM (Text2Cypher) ──► Generated Cypher
                                         │
                                 [Syntax validation]
                                  EXPLAIN {query}
                                   catches syntax errors
                                         │
                                 [Schema validation]
                                  All labels in MATCH clauses
                                  exist in graph schema?
                                  All property names valid?
                                         │
                                 [Safety validation]
                                  Query is read-only? (no MERGE/CREATE/DELETE)
                                  No unbounded variable-length paths? ([:*])
                                  Estimated result set size reasonable?
                                         │
                                         ├──[FAIL]──► Fallback to vector search
                                         │            log for Cypher quality review
                                         │
                                         └──[PASS]──► Execute against graph DB

Never execute LLM-generated Cypher without the validation layer. A poorly generated query with MATCH (n)-[*]-(m) (unbounded traversal) can lock a graph database.

Depth Limits and Path Explosion

Variable-length path queries in Cypher (-[*1..N]-) are exponential in the worst case. Practical limits:

  • 1-2 hops: Fast, predictable. Always safe.
  • 3 hops: Generally fine for well-indexed graphs with specific starting nodes.
  • 4-5 hops: Risky without explicit cardinality limits. Add LIMIT clauses.
  • 5+ hops with unspecified start node: Dangerous. Can saturate the graph database.

Production rule: constrain all variable-length paths to a maximum of 3 hops in auto-generated Cypher. Hand-crafted analytical queries may go deeper with explicit planning.

Answer Assembly: Subgraph to Natural Language

A graph query returns structured data — nodes, edges, properties. This is not an answer. An LLM synthesizes the structured result into a natural language response.

MULTI-HOP: END-TO-END FLOW

User query: "Which customers are at risk from force majeure exposure in APAC?"
      │
      ▼
Text2Cypher ──► [Cypher query] ──► Graph execution
                                        │
                            Structured result:
                            [
                              {customer:"Acme Corp", contract:"CON-4421",
                               expiry:"2025-03-31", clause:"Exhibit C §8"},
                              {customer:"Widget Corp", contract:"CON-7823",
                               expiry:"2024-12-15", clause:"Clause 14.7"}
                            ]
                                        │
                                        ▼
                          LLM synthesis prompt:
                          "The following contracts were found to have
                           force majeure exposure in APAC. Summarize
                           the risk for each and flag the most urgent
                           by expiry date."
                                        │
                                        ▼
                          Natural language answer with grounded citations:
                          "Two active contracts carry APAC force majeure
                           exposure. Widget Corp (CON-7823) is most urgent,
                           expiring December 15 2024, with clause 14.7
                           governing force majeure. Acme Corp (CON-4421)
                           expires March 31 2025 via Exhibit C Section 8..."

The LLM is not retrieving — it is narrating. This separation of retrieval (graph) and synthesis (LLM) is the architectural principle that gives graph-first retrieval its accuracy advantage.


36.9 GraphRAG vs. Traditional RAG: The Decision Framework

The decision to invest in GraphRAG is an architectural commitment, not a configuration change. It requires graph database infrastructure, extraction pipelines, entity resolution logic, and ongoing graph maintenance. Make the decision deliberately.

Decision Tree

GRAPHRAG vs. VECTOR RAG DECISION FRAMEWORK

Does your corpus contain entities with meaningful
relationships across documents?
         │
    NO ──┘──► Vector RAG. You're done here.
         │
        YES
         │
         ▼
Are your most important user queries about
relationships between specific entities?
(vs. conceptual/semantic questions)
         │
    NO ──┘──► Vector RAG for semantic queries.
              GraphRAG adds cost without benefit.
         │
        YES
         │
         ▼
Does your corpus have >1,000 documents with
sufficient density to build a meaningful graph?
         │
    NO ──┘──► Vector RAG. Graph too sparse to be useful.
              Revisit when corpus grows.
         │
        YES
         │
         ▼
Do multi-hop relationship queries currently
fail or produce hallucinated answers?
         │
    NO ──┘──► Vector RAG is working. Monitor query
              failure rate. Add GraphRAG when failures
              exceed 10% of queries by volume.
         │
        YES
         │
         ▼
Does your team have capacity to own a graph
database and extraction pipeline in production?
(or budget for a managed GraphRAG service)
         │
    NO ──┘──► Address capacity gap first.
              GraphRAG without operational ownership
              degrades into a stale graph that
              produces confident wrong answers.
         │
        YES
         ▼
        INVEST IN GRAPHRAG
        Start with Graph-Augmented RAG (Pattern 1, §36.5)
        Add Graph-First (Pattern 2) once extraction quality validated
        Add Hybrid Router (Pattern 3) once query volume justifies it

Criteria Reference

Use GraphRAG when: - Questions require connecting entities across multiple documents (the defining use case) - "What is the relationship between X and Y?" questions are frequent - Domain has well-defined entity types: legal, financial, biomedical, supply chain, compliance - Complete, grounded answers matter more than approximate answers — legal, regulatory, clinical - Corpus is large enough to build a meaningful graph (>1,000 documents with entity density) - Team has or can acquire graph database expertise — at minimum one engineer who owns it

Stick with vector RAG when: - Questions are conceptual or semantic, not entity-specific - Single-document answers are usually sufficient - Corpus is small, unstructured, or entity-sparse (creative content, general Q&A, support documentation) - Query patterns are unpredictable — the vector store handles novel queries more gracefully than a schema-bound graph - Operational complexity is a binding constraint

Use both (hybrid) when: - Corpus has both entity-rich structured content and narrative semantic content - Query mix is heterogeneous — structural and semantic queries both appear in production - You can sustain two indexes operationally - Query volume is high enough to justify the routing infrastructure

The Cost Justification

GraphRAG indexing costs 5-20x more than vector RAG indexing. For a 100,000-document corpus: - Vector RAG indexing (embeddings only): ~$15-30 - GraphRAG indexing (extraction + embeddings + summaries): ~$150-3,000

The justification is not cost reduction. It is query failure rate reduction on the specific class of queries that vector RAG structurally cannot answer.

If your users ask entity-relationship queries and currently get hallucinated or incomplete answers 20-30% of the time, and GraphRAG reduces that to 3-5%, the business value of accurate answers in a legal, financial, or compliance context typically far exceeds the indexing cost difference.

The organizational trigger that makes GraphRAG worth the investment: a documented, recurring, business-consequential failure of vector RAG on entity-relationship queries. Not a theoretical possibility — an actual pattern of wrong answers on real queries. When that pattern exists and is costing decisions, time, or compliance risk, GraphRAG's ROI case is straightforward.


36.10 GraphRAG Production Checklist and Governance

The Graph as a Production Data Asset

The knowledge graph is not infrastructure. It is a data asset. Treat it accordingly.

A graph without an owner degrades. Entity resolution decisions change, new relationship types emerge, old entities become stale, community summaries reflect a corpus state from six months ago. A graph without version history cannot be audited. A graph without a freshness SLA silently serves wrong answers.

Every production knowledge graph needs: - A named owner (team or individual) responsible for quality and freshness - Version history — graph schema changes tracked like database schema migrations - A freshness SLA with monitoring and alerting (see §36.7) - A quality review cadence — quarterly sample of extracted relationships for precision

Entity Ontology Governance

The graph schema defines what entity types and relationship types exist. This is a governance decision, not a technical one — it determines what questions the graph can and cannot answer.

Governance questions that must be answered before production: - Who can add new entity types to the ontology? - What is the review process when an entity type needs to be split or merged? - How are deprecated relationship types handled (migrated, flagged, deleted)? - Who owns the canonical list of entities in the domain (the master data)?

Without governance, the ontology fragments: "Vendor," "Supplier," "ThirdParty," and "ExternalProvider" become four separate node types for the same concept because four teams added them without coordination.

Graph Accuracy Monitoring

Extraction quality degrades as the corpus evolves. The extraction prompt that worked well on 2023 contracts may perform poorly on a new document format added in 2025. The entity dictionary may miss new organizations.

Quarterly review protocol: 1. Sample 200 randomly selected extracted relationships 2. Have a domain expert review each: correct, incorrect, or ambiguous 3. Track precision by relationship type — some types degrade faster than others 4. If precision on any type drops below 80%, retune the extraction prompt for that type 5. Re-extract affected documents (not the full corpus)

PII in Knowledge Graphs

Person entities in a knowledge graph may constitute personal data under GDPR and similar regulations. A node (:Person {name: "Sarah Chen", email: "s.chen@acme.com", title: "VP Procurement"}) is personal data. Every edge connecting that node to contracts, incidents, and communications is personal data in context.

Requirements for compliance:

Right to erasure ("right to be forgotten") requires the ability to delete a person node while preserving non-PII content. The challenge: if (:Person {name:"Sarah Chen"})-[:SIGNS]->(Contract), deleting Sarah's node removes the signing relationship. The contract still exists but its signatory is gone from the graph.

Implementation approaches: 1. Pseudonymization: Replace the person node with an anonymized reference (person_id: "P4421") — preserves graph structure, removes PII 2. Soft delete with tombstone: Mark the node {deleted: true, deletion_date: "..."} — prevents retrieval without destroying relationships; requires all queries to filter deleted nodes 3. Attribute erasure: Delete PII properties (name, email, title) from the node while preserving the node's relationship edges — graph structure intact, PII removed

Choose approach 1 (pseudonymization) for most cases — it preserves graph utility while meeting erasure requirements.

Access Control on Graph Queries

Graph traversal can surface sensitive relationships that flat vector search would not. A user with access to customer records and contract records separately might have no access to the combination — but a graph query naturally traverses the join. The graph does not know about your document-level permissions; it knows about graph structure.

Graph access control requirements: - Node-level access control: certain nodes (e.g., M&A target entities, personnel under HR investigation) must be invisible to certain query principals - Edge-level access control: certain relationships (e.g., compensation data edges, health condition edges) require explicit authorization - Query-level audit logging: every graph query must be logged with the principal, the query text, and the nodes/edges accessed

Neo4j Enterprise supports property-level access control via role-based subgraph access. For FalkorDB and Memgraph, access control must be implemented in the application layer (query rewriting to inject WHERE NOT node.restricted = true clauses based on the requesting principal's permissions).

Query Audit Logging

Graph queries can reveal organizational intelligence that raw document access cannot. Knowing that a principal repeatedly queries for the relationships between an acquisition target and a specific executive is itself sensitive information — it reveals intent and knowledge.

Audit log schema for graph queries:

{
  "timestamp": "2024-11-15T14:23:07Z",
  "principal": "user@company.com",
  "agent_id": "contract-analysis-agent-v2",
  "query_type": "cypher",
  "query_hash": "sha256:abc123...",
  "query_text": "MATCH (o:Organization {name:$name})-[:PARTY_TO]->(c:Contract)...",
  "parameters": {"name": "Acme Corp"},
  "nodes_accessed": ["Organization:Acme_Corp", "Contract:CON-4421", "Contract:CON-7823"],
  "relationships_traversed": ["PARTY_TO", "CONTAINS"],
  "result_count": 12,
  "latency_ms": 43
}

Review this log as you would review database query logs. Anomalous access patterns (unusual principals querying sensitive entity clusters, high-frequency traversal of restricted relationship types) are security signals.

Production Readiness Checklist

Extraction Pipeline - [ ] Entity types defined in canonical ontology (not discovered ad hoc)? - [ ] Extraction pipeline produces confidence scores on all relationships? - [ ] Entity resolution strategy documented with thresholds? - [ ] Incremental ingestion pipeline in place (not full rebuild per document)? - [ ] Extraction quality metrics tracked: coverage, precision, resolution accuracy?

Graph Infrastructure - [ ] Graph database selection documented with justification? - [ ] Read replica strategy for query load (graph writes are serial; reads can fan out)? - [ ] Backup and point-in-time recovery configured? - [ ] Graph schema version-controlled (schema migrations tracked)? - [ ] Cypher query validation layer in place (syntax + schema + safety checks)?

Governance - [ ] Graph owner named with explicit accountability? - [ ] Freshness SLA defined and monitoring in place? - [ ] Entity ontology governed: change process and owner defined? - [ ] Quarterly precision review scheduled? - [ ] PII node deletion capability implemented and tested?

Security - [ ] Node-level access control implemented (application layer or DB native)? - [ ] Graph query audit logging enabled? - [ ] Text2Cypher generated queries validated before execution? - [ ] Graph queries are read-only in production (no CREATE/MERGE/DELETE via retrieval path)? - [ ] Sensitive entity clusters identified and access-controlled?


EXERCISE — When Would GraphRAG Help?: Think of three questions that users actually ask about your organization's knowledge that vector RAG consistently fails on. For each: why does vector search fail (which failure mode from §36.1 applies)? What entities and relationship types would a knowledge graph need to answer it? Write the Cypher query that would answer it correctly. Then ask: does your organization currently have all the source data to populate that graph, or are the key relationships implied in documents but never explicitly recorded?

PONDER — The Indexing Cost Question: GraphRAG indexing costs 5-20x more than vector RAG indexing. For a 100,000-document corpus, that might be $150-3,000 in LLM API calls just to build the graph — and you pay it again whenever you need to re-extract with an improved pipeline. What query quality improvement justifies that? How would you measure whether the graph is providing better answers than vector RAG on the same queries — what is your evaluation set, your quality metric, and your threshold for declaring success? What is the organizational trigger in your context — the specific, recurring, business-consequential failure — that would make this investment clearly justified?

WORKSHOP — Knowledge Graph Design: Choose a domain with rich entity relationships (vendor contracts + compliance requirements + audit findings; patient records + treatments + outcomes + guidelines; software services + dependencies + incidents + owners). Design the following six components: (1) the entity ontology — entity types, relationship types, key properties on each, noting which properties are PII; (2) the extraction pipeline — NER approach for each entity type (rule-based vs. GLiNER vs. LLM-prompted, with justification), relationship extraction prompt structure, entity resolution strategy with thresholds; (3) the retrieval pattern — local search, global search, or hybrid router — with a written justification specific to your query mix; (4) the graph database selection with written justification against the selection matrix in §36.6; (5) the freshness and maintenance plan — SLA, monitoring approach, incremental vs. batch strategy, re-extraction triggers; (6) a GDPR and privacy compliance assessment — which nodes constitute personal data, your erasure implementation approach, and your access control design for sensitive entity clusters.


Next: Module 37 — Model Portability: Engineering the Model as a Replaceable Part. (Agentic memory — short-term, long-term, episodic — is covered under Module 21's persistent memory patterns rather than as a standalone module.)