Skip to content

MODULE 4 — RAG Architecture: From Naive to Production-Grade

4.1 Why RAG Exists and What Problem It Actually Solves

LLMs are trained on static snapshots of the world. A model does not know what happened after its training cutoff. It does not know your organization's internal policies. No frontier model, however capable, knows the specific contents of your proprietary documentation, your product catalog, your regulatory filings, or your customer contracts.

RAG — Retrieval-Augmented Generation — is the architectural pattern that bridges the gap between a model's static training knowledge and the dynamic, proprietary information your AI system needs to operate on. At its core it is simple: before asking the model to answer a question, retrieve relevant information from your knowledge stores and include it in the prompt context.

That simplicity is deceptive. Every word in "retrieve relevant information from your knowledge stores" is an architectural decision with significant implications for quality, cost, latency, and reliability. The retrieval must actually find relevant information (not just similar-looking information). The relevance scoring must be calibrated to your use case (not just to general semantic similarity). The knowledge store must contain current, accurate, well-structured information (not just whatever was dumped in at launch). The retrieved information must be presented to the model in a way that produces grounded, accurate answers (not just whatever the model generates when given chunks as context).

Production RAG failure is almost never about the model. It is almost always about one of these retrieval, knowledge management, or context assembly decisions made carelessly early in development.


4.2 The Seven Failure Modes of Naive RAG

Understanding these failure modes is the prerequisite for designing a production-grade system. Each one has a specific architectural cause and a specific architectural fix.


Failure Mode 1: The Chunk Boundary Problem

What happens: A document is split into fixed-size chunks. A sentence or concept that spans a chunk boundary is split across two chunks. When a query retrieves one chunk, the answer is incomplete or misleading because the critical qualifying information is in the adjacent chunk.

Concrete example:

Document: "Early withdrawal penalties are waived provided the
  account holder has maintained the account for a minimum period
  of 24 consecutive months and the withdrawal does not exceed
  30% of the total balance."

Fixed 512-token chunking splits this at "waived provided the":

Chunk 1: "...Early withdrawal penalties are waived provided the"
Chunk 2: "account holder has maintained the account for a minimum
          period of 24 consecutive months..."

Query: "Can I withdraw without penalty?"
Retrieved: Chunk 1 (high similarity to "withdraw without penalty")
LLM reads: "penalties are waived" → tells the user yes.

The condition — 24 months, 30% limit — is in chunk 2, not retrieved.
The answer is wrong. The consequence is real.

Architectural fix: Semantic chunking (split at sentence and paragraph boundaries, not arbitrary token counts), hierarchical chunking (small chunks for retrieval, larger parent chunks for context), and overlap windows (chunks share N tokens with adjacent chunks so boundary content appears in both). No single fix is complete — the best production systems combine semantic boundaries with overlap.


Failure Mode 2: The Stale Knowledge Problem

What happens: A document is updated. The new version is added to the vector store. The old version is not removed. Both versions coexist. The retrieval system returns a mix of current and outdated content. The model synthesizes an answer from both, producing a response that is partially correct and partially outdated — with no indication of which parts are which.

Concrete example:

January 2024: Fee schedule uploaded. Annual fee: $150.
March 2024:   Fee reduced to $95. New version uploaded.
              Old version not removed.

Vector store now contains both versions:
  Chunk A (Jan): "annual maintenance fee is $150"
  Chunk B (Mar): "annual maintenance fee is $95"

Query: "What is my annual fee?"
Both chunks have high similarity. Model receives both.
Model synthesizes: "The annual fee is $95, though some accounts
  may have a $150 fee." — wrong, confusing, and potentially
  actionable by a customer.

Architectural fix: Document lifecycle management with mandatory versioning. When a document is updated, the prior version's chunks are marked status: archived and excluded from retrieval. The vector store is not a simple append store — it is a versioned document management system with lifecycle operations. This is a data engineering problem, not a prompt engineering problem.


Failure Mode 3: The Confident Hallucination on Low Retrieval

What happens: A query does not match any document in the knowledge base well. The retrieval returns low-relevance chunks. There is no confidence threshold. The model receives these irrelevant chunks as "context" and generates a confident answer anyway — either by making up information or by misapplying the low-relevance context.

Why this happens: LLMs are trained to be helpful. They do not spontaneously say "I don't have enough information." When given context (even poor context), they will use it. The model generates plausible-sounding content grounded in irrelevant retrieved material, producing answers that are coherent but wrong.

Architectural fix: A confidence gate. If the maximum relevance score across all retrieved chunks falls below a defined threshold, the system does not attempt to answer. It responds with: "I cannot find information about that in our knowledge base. Let me connect you with someone who can help." The threshold value is a business decision that requires explicit sign-off — not a technical default.


Failure Mode 4: The Embedding Model Version Drift

What happens: At launch, all documents are embedded with Model A. Six months later, an engineer upgrades the embedding model to Model B for improved quality on new documents. Existing vectors remain in Model A's embedding space. Queries are now executed in Model B's embedding space. The query vector and the document vectors are no longer in the same mathematical space. Cosine similarity between a Model B query and a Model A document vector is meaningless. Retrieval quality silently collapses.

Why it's hard to detect: The system continues to return results. They just aren't the right results anymore. There is no error. There is no alert. Users notice responses are worse. Engineers investigate the model, the prompt, the data — not the embedding version mismatch.

Architectural fix: Store the embedding model name and version with every vector as metadata. The retrieval pipeline must use the same model version as the stored vectors. Model upgrades are treated as data migration events: re-embed all existing vectors with the new model before switching queries to use the new model. Never mix embedding model versions in the same retrieval operation.


Failure Mode 5: The Semantic Similarity ≠ Relevance Problem

What happens: Vector similarity search returns chunks that are semantically similar to the query but not actually relevant to answering it. This is a fundamental limitation of bi-encoder retrieval (the type of search used in all standard vector stores).

Concrete example:

Query: "What is the penalty for early account closure?"

Bi-encoder retrieval returns (by cosine similarity):
  1. "Account closure procedures and timelines" (high similarity)
  2. "Penalty structure for missed payments" (medium-high similarity)
  3. "Early withdrawal fees for investment accounts" (medium similarity)

None of these are the actual policy for early account CLOSURE penalties.
The word overlap produces semantic similarity.
But the model receives these as the relevant context.

Why this happens: Bi-encoder models (which produce the vectors in your vector store) embed query and document separately. They capture topical similarity well. They do not directly model the relevance of a document to a specific query the way a human would judge it.

Architectural fix: Cross-encoder re-ranking. A cross-encoder model reads the query AND the candidate chunk together, producing a direct relevance score. It is computationally expensive (cannot be pre-computed — must run at query time) but dramatically more accurate. The production pattern: bi-encoder retrieval for speed (get Top-50 candidates), cross-encoder re-ranking for precision (rerank to Top-5), pass Top-5 to the LLM. The latency cost (50–150ms additional) is justified for any use case where wrong answers have consequences.


Failure Mode 6: The Long Document Context Degradation

What happens: Multiple relevant chunks are retrieved and passed to the LLM. But the most relevant chunk is in the middle of the context window. The model performs significantly worse on information in the middle of a long context than at the beginning or end — a well-documented phenomenon known as "lost in the middle."

Concrete example:

Retrieved chunks passed to LLM (in retrieval rank order):
  [Chunk 1: somewhat relevant]
  [Chunk 2: most relevant - contains the answer]
  [Chunk 3: somewhat relevant]
  [Chunk 4: tangentially relevant]
  [Chunk 5: tangentially relevant]

LLM reads all 5 chunks. Chunk 2 (the answer) is in the middle.
Research shows models systematically underweight middle-context information.
The model answers based on Chunks 1, 3, 4, 5 — not Chunk 2.

Architectural fix: Place the highest-relevance chunks at the beginning and end of the context, not in the middle. After re-ranking, order chunks so the highest-scoring chunks are first and last, with lower-scoring chunks filling the middle. Additionally, limit the number of chunks passed to the LLM — 3–5 high-quality chunks produce better answers than 10 mixed-quality chunks.


Failure Mode 7: No Audit Trail

What happens: The system produces an incorrect answer. Support is escalated. Compliance requests an audit: which document was retrieved, which version, what was the query, what reasoning did the model use? The system has no answer. Each interaction is ephemeral — query in, response out, nothing logged.

Why this matters beyond compliance: Without an audit trail, you cannot debug failures systematically. You cannot build regression tests based on production failures. You cannot detect when retrieval quality is degrading. You are flying blind.

Architectural fix: Log every interaction with: query hash, retrieved chunk IDs and relevance scores, document version metadata, model used, response hash, timestamp, user identifier (pseudonymized). This is not optional for any regulated use case. It is good practice for any production system.


4.3 Chunking Strategy: The Decision That Shapes Everything Downstream

Chunking is the process of splitting source documents into segments that can be individually embedded and retrieved. The chunking strategy affects retrieval precision, context quality, and the chunk boundary problem directly. It is one of the most impactful and least discussed architectural decisions in RAG.

Fixed-Token Chunking

Split every document into chunks of exactly N tokens with M tokens of overlap.

Example: chunk_size=512, overlap=50

Document: [--------512 tokens---------][overlap 50][--------512 tokens---------]
                                                    ↑ overlap region appears
                                                      in both chunks

When appropriate: Homogeneous document types with consistent density of information (short FAQ entries, uniform product descriptions). Fast to implement. Predictable chunk sizes for cost estimation.

Problems: Splits mid-sentence and mid-concept. Unaware of document structure (headers, sections, paragraphs). The overlap mitigates but does not solve the boundary problem.


Semantic / Paragraph-Aware Chunking

Split at natural semantic boundaries: sentence endings, paragraph breaks, section headers. Chunks vary in size but respect the document's own structure.

Document:
  [Section: Account Features]         → Chunk 1
  [Paragraph: Balance requirements]   → Chunk 2
  [Paragraph: Fee schedule]           → Chunk 3
  [Section: Early Closure Policy]     → Chunk 4
    [Subsection: Standard penalty]    → Chunk 5
    [Subsection: Waiver conditions]   → Chunk 6

When appropriate: Narrative documents, policy documents, technical documentation where section structure reflects semantic organization. Significantly reduces the boundary problem. Produces more coherent, self-contained chunks.

Tools: Unstructured.io (document-type-aware parsing), LangChain's RecursiveCharacterTextSplitter with sentence-aware splitting, LlamaIndex's semantic splitter.


Hierarchical / Parent-Child Chunking

Split documents into two levels: small chunks for precise retrieval, larger parent chunks for context quality.

Document
  └── Parent Chunk (2,000 tokens) — the context provided to the LLM
        ├── Child Chunk 1 (256 tokens) — used for retrieval matching
        ├── Child Chunk 2 (256 tokens) — used for retrieval matching
        ├── Child Chunk 3 (256 tokens) — used for retrieval matching
        └── Child Chunk 4 (256 tokens) — used for retrieval matching

How it works: Embed and index the small child chunks. When a child chunk matches a query, retrieve the parent chunk for the LLM context. The small chunks provide precise matching. The parent chunk provides full surrounding context.

Why this helps: A 256-token child chunk has a tighter semantic focus, producing more precise retrieval. But a 256-token context passed to the LLM may lack surrounding information needed for a complete answer. The parent chunk provides that context without sacrificing retrieval precision.

When appropriate: Documents with natural hierarchical structure (sections containing paragraphs, articles with subsections, manuals with chapters and steps). More complex to implement and maintain but meaningfully better retrieval quality for structured documents.


Document-Type-Aware Chunking

Different document types require different chunking strategies. A uniform strategy applied to all documents degrades quality for documents whose structure the strategy does not match.

CHUNKING STRATEGY BY DOCUMENT TYPE

Policy documents (narrative, hierarchical):
  → Semantic chunking at section/paragraph boundaries
  → Hierarchical with section-level parents

FAQ documents (Q&A pairs):
  → Chunk per Q&A pair, never split a question from its answer
  → Each chunk: {question: "...", answer: "..."}

Tables and structured data:
  → Never split mid-row
  → Include column headers in every chunk derived from the table
  → Convert to text representation that preserves relationships

Legal contracts:
  → Clause-level chunking (each clause is a chunk)
  → Include clause numbering as metadata
  → Definitions section linked to all chunks using defined terms

Code documentation:
  → Function/class level chunking
  → Include module context in every chunk
  → API signatures never split from their descriptions

4.4 Embedding Models: Selection, Pinning, and Migration

The embedding model transforms text into vectors. Every vector in your store is a product of a specific embedding model. The retrieval quality ceiling is set by the embedding model — a poor embedding model produces vectors that are bad at capturing semantic similarity for your content domain, and no amount of retrieval engineering compensates for bad embeddings.

Selection Criteria for Enterprise RAG

Domain fit. General-purpose embedding models (for example, OpenAI text-embedding-3-large or Cohere's Embed family; illustrative, see Appendix G for current options) are trained on broad web text. They perform well on general English prose. For specialized domains — medical literature, legal contracts, financial filings, code — domain-specific or domain-fine-tuned embedding models may significantly outperform general models. The way to know: run a retrieval quality evaluation (precision and recall) on 50–100 representative queries from your actual data before committing to an embedding model.

Dimension count. Higher-dimension embeddings (3072 vs 768) capture more nuance but cost more in storage and retrieval latency. The quality improvement from higher dimensions is task-dependent. Run the evaluation before paying the storage and latency premium.

Multilingual requirement. If your document corpus includes multiple languages, use a multilingual embedding model (multilingual-e5-large, Cohere multilingual). A monolingual model produces meaningless vectors for content in other languages.

Vendor lock-in. Using OpenAI's embedding models ties your vector store to OpenAI. The vectors cannot be reused if you change providers — you must re-embed. Open-source embedding models (bge-m3, nomic-embed-text, e5-mistral-7b) can be self-hosted. The vectors are yours regardless of what happens to any vendor relationship.

Embedding Model Pinning (Non-Negotiable)

Every vector in the store must have the embedding model name and version stored as metadata:

Vector metadata (required):
{
  "doc_id": "policy-001",
  "doc_version": "2024-03-01",
  "chunk_id": "policy-001-chunk-07",
  "embedding_model": "text-embedding-3-large",
  "embedding_model_version": "002",
  "embedded_at": "2024-03-15T09:23:00Z",
  "status": "active"
}

Why this is non-negotiable: When you upgrade the embedding model, you must re-embed all existing vectors before using the new model for queries. If you don't know which model version created each vector, you cannot safely perform a partial migration. You also cannot detect when mixed model versions in the store are degrading retrieval quality.

Embedding Model Migration Process

MIGRATION PROCEDURE: Embedding Model Upgrade

Phase 1: Parallel preparation
  - New embedding model selected and evaluated against your data
  - New model begins embedding new documents as they are ingested
  - Existing documents flagged as "pending re-embed"
  - All new document vectors tagged with new model version

Phase 2: Background re-embedding
  - Batch re-embedding pipeline runs for existing documents
  - Creates NEW vectors for existing documents with new model
  - OLD vectors remain active during this phase
  - Progress tracked: X of Y documents re-embedded

Phase 3: Cutover
  - Validate re-embedding is complete (100% or defined threshold)
  - Run retrieval quality evaluation: before vs. after comparison
  - If quality improves or is neutral: switch query pipeline to new model
  - Query pipeline NOW uses new model vectors only

Phase 4: Cleanup
  - Old vectors marked for archival (keep for rollback period: 30 days)
  - After rollback period: old vectors removed
  - Documentation updated: current embedding model version recorded

Never do: Use the new embedding model for queries while old-model vectors exist in the store. Even a single day of mixed-model retrieval produces misleading results that are difficult to debug.


4.5 Hybrid Search: The Production Standard

Purely vector-based retrieval is not sufficient for production RAG. It excels at semantic similarity but underperforms on exact keyword matching, specific names, identifiers, product codes, and domain-specific terminology.

The case for hybrid search:

Query: "What is the penalty for FDIC-insured accounts under Section 4B?"

Vector search behavior:
  Finds documents about "penalties" and "accounts" and "FDIC"
  May miss the exact section 4B reference if the training data
  doesn't encode legal section numbering as semantically meaningful

BM25 keyword search behavior:
  Finds documents containing the exact string "Section 4B"
  Even if the surrounding language is different

Hybrid search:
  Combines both signals
  Returns documents that are both semantically relevant AND
  contain the specific section reference

Reciprocal Rank Fusion (RRF): The standard method for combining dense and sparse retrieval scores. Both retrievers return ranked lists. RRF combines the rank positions (not the raw scores, which are on different scales) using a formula that rewards documents that rank highly in both lists. Simple, effective, and model-agnostic.

Metadata filtering as the third search dimension:

Beyond dense and sparse signals, metadata filters constrain the retrieval space before search begins. This is not just a quality improvement — it is a correctness requirement for multi-domain or multi-tenant systems.

Multi-domain filter example:
  User in the "Retail Banking" division queries the system.
  Without metadata filter: may retrieve "Corporate Banking" policy
    documents that don't apply to retail customers.
  With filter: {division: "retail", status: "active", jurisdiction: "US"}
    Retrieval is constrained to relevant documents before similarity is computed.

Multi-tenant filter example:
  Customer A's queries must never return content from Customer B's
    document space. Tenant ID filter enforces this at the retrieval
    layer, not the prompt layer.

Relying on the LLM to "figure out" which retrieved documents are applicable is wrong. It places a correctness requirement in a probabilistic system. Metadata filtering is deterministic. Use it.


4.6 Re-ranking: From Candidates to Quality

Re-ranking is the quality layer between retrieval and generation. After hybrid search produces 20–50 candidate chunks, a re-ranker scores each chunk specifically for its relevance to the query. The top 3–5 highest-scoring chunks are passed to the LLM.

Why Re-ranking Matters

Bi-encoders (the models that produce vectors) embed query and document independently. The similarity score is an approximation of relevance. Cross-encoders (re-rankers) encode query and document together, allowing the model to directly reason about how well this specific document answers this specific query.

The quality difference is significant. In enterprise RAG evaluations, adding a cross-encoder re-ranker typically improves context precision (the fraction of retrieved chunks that are actually relevant) by 20–40 percentage points. At the cost of 50–150ms additional latency per query.

Re-ranking Tools

Cohere Rerank API: Hosted, no infrastructure, simple API (Rerank 4 Fast and Rerank 4 Pro as of October 2026). Billed per search, where one search is one query with up to 100 documents, and documents longer than 500 tokens are split into chunks that each count toward that limit. Reported list pricing is about $2.00–$2.50 per 1,000 searches as of 2026 (reported; illustrative, verify current pricing in Appendix G or on the vendor page). At 1 million queries a month that is roughly $2,000–$2,500 a month, which is material for high-volume systems.

BGE Reranker (BAAI): Open-source, self-hostable, strong performance across languages. The bge-reranker-v2-m3 variant handles multilingual content. Requires GPU for production throughput.

Jina Reranker: Open-source option with commercial API. Strong at code and technical content.

When to skip re-ranking: Very simple document corpora (under 10,000 highly uniform documents), use cases where the 50ms latency is not acceptable, high-volume classification tasks where retrieval precision is less critical. The default should be to include re-ranking and explicitly justify its removal.


4.7 Advanced RAG Patterns

Beyond the standard retrieve-then-read pipeline, several advanced patterns address specific failure modes or quality requirements.

HyDE (Hypothetical Document Embeddings)

Problem being solved: For complex queries, the query text and the relevant document text may use very different language. A user asking "why did my payment fail?" may be best answered by a document titled "Transaction Processing Error Codes and Resolutions" — the semantic similarity between the query and document title is low even though the document is highly relevant.

How it works: Instead of embedding the query directly, ask the LLM to generate a hypothetical answer to the query. Embed the hypothetical answer. Use that embedding to retrieve documents. The hypothetical answer uses the same vocabulary and structure as a good answer, so its embedding is closer to relevant documents than the raw query embedding.

Query: "Why did my payment fail?"

Without HyDE:
  Embed query → search for similar → may miss technical docs

With HyDE:
  LLM generates: "Payment failures typically occur due to
    insufficient funds (error code NSF), expired card information
    (error code CARD_EXP), or bank authorization holds. Check your
    transaction history for the specific error code..."
  Embed this hypothetical answer → search → retrieves the technical
    error code documentation with high precision

Trade-off: Adds one LLM call per query (cost, latency). Worth it for complex queries with vocabulary mismatch. Not worth it for simple keyword-style queries.


Multi-Query Retrieval (RAG Fusion)

Problem being solved: A single query may be interpretable in multiple ways. Retrieving for only one interpretation misses relevant documents that would be found by other phrasings.

How it works: Generate N variants of the original query (via LLM). Run retrieval for each variant. Combine results with RRF. Pass the combined top-K to the LLM.

Query: "What happens when I close my account?"

LLM generates 3 query variants:
  1. "Account closure process and timeline"
  2. "Fees and penalties for closing an account"
  3. "Impact of account closure on linked services"

Retrieve for all 3, combine with RRF.
The combined result covers: procedure, costs, AND implications.
A single-query retrieval would likely cover only one of the three.

Trade-off: N LLM calls (cheap, small model works for query generation) + N retrieval operations (index operations, not expensive). Material latency addition at large N. N=3 is typically sufficient.


Agentic Retrieval

Problem being solved: Some questions require multiple retrieval steps, where the answer to one retrieval informs the next query.

How it works: An agent orchestrates retrieval iteratively. It retrieves, reads, determines what additional information is needed, retrieves again, and continues until it has sufficient context to answer.

Query: "Compare the fee structures for our basic and premium 
        accounts, and tell me if there's a promotion currently 
        running that would make premium a better deal."

Step 1: Retrieve basic account fee documentation
Step 2: Retrieve premium account fee documentation
Step 3: LLM compares fees, identifies gap in knowledge about promotions
Step 4: Retrieve current promotions
Step 5: LLM synthesizes complete answer

When agentic retrieval is warranted: Multi-part questions, questions that require comparison across documents, questions where the answer to one part determines what to look for next. Not for simple single-document lookups.

Architectural constraint: Agentic retrieval must have a maximum iteration limit. An agent that retrieves indefinitely in search of a perfect answer is a cost and reliability risk. Set a hard maximum of 3–5 retrieval steps per query.

Section 4.10 covers the full family of agentic retrieval patterns, including deep research, along with their cost profile, controls, and evals.


GraphRAG (Knowledge Graph + Vector Retrieval)

Problem being solved: Standard vector search finds similar documents but does not capture relationships between entities. Questions about relationships — "who reports to whom," "which products are affected by this regulation," "what are all the dependencies of this service" — require relationship traversal that vector similarity cannot provide.

How it works: Build a knowledge graph from the document corpus (entity extraction + relationship extraction). At query time, combine vector search (for semantic similarity) with graph traversal (for relationship queries). The LLM receives both the similar documents and the graph neighborhood around relevant entities.

Document corpus: Organizational policies, reporting structures

Standard vector search for "who is responsible for data governance?":
  Finds documents mentioning "data governance" and "responsibility"
  Returns a list of policy statements

GraphRAG for the same query:
  Entity: "Data Governance"
  Graph traversal: Data Governance → responsible_party → [CISO, CDO]
                   CDO → reports_to → CTO
                   CDO → owns → [Data Classification Policy, 
                                  Data Retention Policy]
  Returns: the specific relationships AND the policy documents

The graph traversal answers "who" precisely.
The vector search provides the policy context.

When GraphRAG is warranted: Organizational hierarchies, regulatory dependency mapping, product relationship queries, knowledge bases where entity relationships are as important as document content. Significantly more complex to build and maintain than standard RAG. Not a default — a deliberate choice for specific query types that require relationship traversal.

Tools: Microsoft GraphRAG (open source), Neo4j with vector search, Amazon Neptune Analytics (combined graph + vector).


Contextual Retrieval

Problem being solved: Individual chunks, when extracted from their surrounding document context, lose the contextual information that makes them interpretable. A chunk containing "The maximum limit is $10,000" is ambiguous without knowing it refers to daily transfer limits for basic accounts — but that context is in a different part of the document.

How it works: Before embedding each chunk, prepend a brief AI-generated context summary that situates the chunk within the broader document. The summary (50–100 tokens) is prepended to the chunk text before embedding. This means the vector for each chunk carries contextual information even when retrieved in isolation.

Original chunk (without context):
  "The maximum limit is $10,000 and applies to all standard accounts.
   Exceptions may be granted with branch manager approval."

With contextual retrieval:
  "This chunk is from the Daily Transaction Limits Policy document,
   Section 3: Wire Transfer Limits for Retail Banking Accounts.

   The maximum limit is $10,000 and applies to all standard accounts.
   Exceptions may be granted with branch manager approval."

The contextualized version retrieves correctly even for queries like "what is the wire transfer limit?" where the original chunk would score lower due to missing the term "wire transfer."

Implementation: Generate the context prefix using a fast, cheap LLM at ingestion time (illustrative small/efficient tiers as of October 2026: Claude Haiku 4.5, GPT-6 Luna, Gemini Flash-Lite; see Appendix G). The added cost is a one-time ingestion cost, not a per-query cost. Anthropic's published research (September 2024) found that contextual embeddings alone reduced the top-20 retrieval failure rate by 35%, adding contextual BM25 raised the reduction to 49%, and adding re-ranking on top raised it to 67%.

When to use: Any RAG system where chunks are derived from longer, structured documents where section context matters. Especially valuable for policy documents, technical manuals, legal texts, and any corpus where the meaning of a passage depends on where it appears in the document.


CRAG (Corrective RAG)

Problem being solved: Standard RAG retrieves and uses whatever is returned, regardless of quality. If the retrieval step returns irrelevant or low-quality chunks, the model generates a poor answer from poor context with no mechanism to detect or correct this.

How it works: After retrieval, a lightweight evaluator assesses the quality of each retrieved chunk. Based on the quality assessment: - High quality: Use the retrieved chunks directly - Low quality: Either discard and reformulate the query, or supplement with a web search / broader retrieval - Ambiguous: Strip the irrelevant portions and augment with additional retrieval

Query → Retrieve candidates → Quality evaluator assesses each chunk
  ├── Chunk quality: HIGH → Use chunk, proceed to generation
  ├── Chunk quality: LOW  → Discard chunk
  │                        Reformulate query
  │                        Re-retrieve with new query
  └── Chunk quality: MED  → Use relevant portions of chunk
                            Supplement with additional retrieval
                            Combine for generation

When to add CRAG: Systems with diverse, inconsistent document quality where standard retrieval frequently returns irrelevant results. Systems where users ask open-ended questions that may be out of domain for the knowledge base. The quality evaluator can be a lightweight classifier model (not a full frontier model) — fast and cheap.

Trade-off: Adds latency for the evaluation step and potential re-retrieval. In systems with consistently high retrieval quality (well-curated corpus, tight domain), CRAG adds complexity without significant benefit. Profile your retrieval quality with RAGAS first — if context precision is already above 0.75, CRAG is unlikely to justify its complexity.


4.8 RAG Evaluation: The RAGAS Framework

Most teams validate RAG systems by running some test queries and checking if the answers "look right." This is not evaluation. It is intuition. It does not scale, it does not catch subtle regressions, and it is not reproducible.

RAGAS (Retrieval-Augmented Generation Assessment) is the standard framework for systematic RAG evaluation. It defines four metrics that together characterize the system's quality.

The Four Core Metrics

Context Precision Question: Of the chunks retrieved, what fraction were actually relevant to answering the query?

A system that retrieves 10 chunks, 3 of which are relevant, has context precision of 0.30. Low context precision means the LLM is receiving a lot of irrelevant information — which increases cost, increases latency, and increases the risk that the LLM anchors on irrelevant content.

High precision (good): Retrieved 5 chunks, 4 relevant → precision = 0.80
Low precision (bad):   Retrieved 5 chunks, 1 relevant → precision = 0.20

Context Recall Question: Of the relevant information that exists in the knowledge base, what fraction was actually retrieved?

A system with high precision but low recall is finding relevant chunks when it retrieves but missing many relevant chunks entirely. The LLM answer is based on an incomplete picture.

High recall (good): 8 relevant chunks exist, 7 retrieved → recall = 0.875
Low recall (bad):   8 relevant chunks exist, 2 retrieved → recall = 0.25

Faithfulness Question: Are all the claims in the generated answer actually supported by the retrieved context?

Faithfulness specifically measures hallucination: is the model saying things that are not in the chunks it was given? A score of 1.0 means every factual claim in the answer is grounded in the retrieved context. Scores below 0.80 indicate significant hallucination risk.

Retrieved context: "The annual fee is $95 for basic accounts."

Faithful answer: "Your annual fee is $95."
Unfaithful answer: "Your annual fee is $95, and premium accounts
  are $145." — the premium account figure was not in the context.
  Faithfulness score < 1.0.

Answer Relevance Question: Does the generated answer actually address what was asked?

A perfectly faithful answer that does not address the question is still a bad answer. Answer relevance measures whether the response is on-topic and responsive to the query.

Running RAGAS in Practice

RAGAS EVALUATION SETUP

1. Build your evaluation dataset:
   - 50–200 query/ground_truth_answer pairs
   - Ground truth answers written by domain experts from your docs
   - Cover diverse query types: simple factual, multi-hop,
     edge cases, out-of-scope

2. Run the RAG pipeline on each query:
   - Capture: query, retrieved_contexts[], generated_answer

3. Compute metrics:
   - Context Precision: LLM-as-judge (is each context chunk relevant?)
   - Context Recall: Compare retrieved content against ground truth
   - Faithfulness: LLM-as-judge (is each claim in answer in context?)
   - Answer Relevance: Embed generated answer and query, compute similarity

4. Establish baseline:
   - Record metrics at launch: P=0.72, R=0.65, F=0.88, AR=0.91
   - This is your quality baseline

5. Run on schedule:
   - After every document corpus update
   - After every embedding model change
   - After every prompt change
   - Weekly in production (on sample of real queries)

6. Alert thresholds:
   - Context Precision drops > 15% from baseline → investigate
   - Faithfulness drops below 0.80 → high hallucination risk alert
   - Any metric drops > 20% from baseline → production incident

RAGAS tooling: The ragas Python library automates metric computation. It requires an LLM for the judge-based metrics (faithfulness, context precision) — typically a capable but cost-efficient model (illustrative as of October 2026: Claude Haiku 4.5, GPT-6 Luna, or a Gemini Flash-Lite model; see Appendix G). Validate the judge model against human labels before trusting its scores. The evaluation cost is a small fraction of production inference cost but yields the quality visibility that production-grade systems require.


4.9 RAG in Regulated Environments: Additional Requirements

Standard RAG architecture needs augmentation for regulated industries. The following requirements apply beyond normal production considerations.

Citation with source traceability. Every AI-generated answer that draws on retrieved content must cite the source: document name, version, section or page. In regulated environments (financial services, healthcare, legal), this is not a feature — it is a compliance requirement. The architecture must store the source metadata with every retrieved chunk and the response pipeline must include citation construction.

Answer scope authorization. Not every document in the knowledge base is authorized to be used for every type of query or every user tier. A customer service representative should not receive an answer that draws on internal risk assessment documents. The retrieval layer must enforce document-level access control, not just tenant-level. This requires user/role metadata to be part of the retrieval filter.

Confidence disclosure. When the confidence gate triggers or when RAGAS faithfulness scores indicate uncertain grounding, the user must be informed. "Based on available information" is insufficient. The system should clearly state when an answer is provided with lower confidence and recommend human verification.

Retention of the full interaction for audit. In regulated industries, the full interaction — query, retrieved chunks with version metadata, generated answer, user identifier — must be retained for the compliance retention period (often 7 years in financial services). This is an architectural requirement that affects the audit logging infrastructure, not just the RAG pipeline.


4.10 Agentic RAG and Deep Research

Everything up to this point treats retrieval as one step in a pipeline: one query goes in, the top-K chunks come out, and the model generates once. Agentic RAG turns retrieval into a loop that the model controls. The model decides whether to retrieve, what to search for, which retriever to use, whether the evidence is good enough, and when to stop. The "Agentic Retrieval" pattern in §4.7 is the simplest version. This section covers the full range of patterns, what they cost, and the controls that keep them bounded.

What Changes When the Model Controls Retrieval

In single-shot RAG, the architect makes every retrieval decision at design time. In agentic RAG, four of those decisions move into the model's runtime behavior:

SINGLE-SHOT RAG (the pipeline controls retrieval)
  Query → Retrieve top-K → Re-rank → Generate → Answer
  One retrieval. One generation. Behavior fixed at design time.

AGENTIC RAG (the model controls retrieval)
  Query
    │
    ▼
  1. PLAN: split the question into sub-questions
    │
    ▼
  ┌─► 2. CHOOSE TOOL: vector │ keyword │ SQL │ graph │ web
  │     │
  │     ▼
  │   RETRIEVE → READ → EXTRACT EVIDENCE (with source IDs)
  │     │
  │     ▼
  │   3. EVALUATE: is the evidence sufficient?
  │     ├── Gap found, budget left → write the next query ──┐
  │     │                                                    │
  └─────┴────────────────────────────────────────────────────┘
        ├── Budget exhausted → answer with stated gaps, or escalate
        └── Sufficient → 4. SYNTHESIZE with citations
                              → VERIFY citations → Answer

Query planning and decomposition. The model rewrites a compound question into sub-questions that can each be answered from a single retrieval. "Is our premium account a better deal than basic, given current promotions?" becomes three lookups: basic fees, premium fees, active promotions. Bad decomposition is the most common root cause of agentic RAG failure. If the plan misses a sub-question, no later step recovers it.

Iterative retrieve → read → decide. Each retrieval result shapes the next query. The model reads what it found, notices a gap ("the promotion references a 'qualifying deposit' that is not defined"), and searches for that specific gap.

Self-evaluation of sufficiency. The model judges whether it has enough evidence to answer. This is the hardest decision to get right. Models tend either to stop at the first plausible hit or to keep searching for confirmation they do not need. The stopping criterion must be explicit (see Controls below).

Tool choice between retrievers. Agentic systems often expose several retrievers as separate tools. The model's routing decision is driven almost entirely by the tool names and descriptions, so treat tool descriptions as prompts: version them and include routing cases in your eval set.

Retriever Best for Example query Failure if misrouted
Vector (dense) Conceptual or paraphrased questions "How do we handle fee complaints?" Misses exact identifiers and codes
Keyword (BM25) IDs, product codes, names, section numbers "Penalty under Section 4B" Misses paraphrased content
SQL / structured Counts, aggregates, current values "How many accounts opened in Q3?" Model estimates numbers from prose
Graph Relationships and multi-hop entity traversal "Which products does regulation X affect?" Returns similar text instead of the actual relationships
Web search Public, fresh, external information "What did the regulator publish this week?" Untrusted content enters the loop (injection risk)

Inside the vector and keyword tools, hybrid search (§4.5), metadata filtering, and re-ranking (§4.6) still apply. Agentic control sits on top of a production retrieval stack. It does not replace one. For relationship-heavy queries, Module 36 covers when a graph retriever is worth its build and maintenance cost.

The Four Patterns

Pattern 1: Single-shot RAG. The pipeline described in §4.3–§4.6. One retrieval, one generation. It remains the right default for most enterprise question-answering traffic.

Pattern 2: Corrective and self-reflective RAG. One generation pass with a quality check on retrieval and, in some variants, on the answer itself. Two research patterns define this category:

  • Self-RAG (Asai et al., 2023; ICLR 2024) trains a single model to emit special "reflection tokens." These tokens decide whether to retrieve at all, judge whether each retrieved passage is relevant, check whether each generated claim is supported by a passage, and rate the overall usefulness of the response. The model can retrieve several times or skip retrieval entirely.
  • Corrective RAG / CRAG (Yan et al., 2024) adds a lightweight retrieval evaluator (a fine-tuned T5 model in the paper) that scores retrieved documents for a query. Based on its confidence, CRAG triggers one of three actions: use the documents, discard them and fall back to web search, or combine both. A "decompose-then-recompose" step strips irrelevant content from documents before generation. §4.7 describes the production form of this pattern.

In practice, most teams implement both ideas as prompted LLM-judge steps or small classifiers rather than the fine-tuned models from the papers. The architectural idea is what transfers: check retrieval quality before trusting it.

Pattern 3: Multi-hop agentic retrieval. A single agent runs a ReAct-style loop (Module 6, §6.2), where the answer to one retrieval determines the next query. Example: "Which of our vendors handling EU customer data have contracts that expire before the new DPA terms take effect?" Hop 1 finds vendors that handle EU data. Hop 2 retrieves each vendor's contract end date. Hop 3 retrieves the effective date of the new terms. The final step compares them. No single query could retrieve this answer.

Pattern 4: Deep-research agents. An orchestrator plans the research, starts several sub-agents in parallel, and synthesizes their findings into a cited report. Each sub-agent has its own context window and its own retrieval loop, and returns only condensed findings to the orchestrator.

DEEP RESEARCH ARCHITECTURE (orchestrator–worker)

  User question
       │
       ▼
  ORCHESTRATOR (strong model): plan, assign, track coverage
       │
       ├──► Sub-agent A: sub-question 1 ── own retrieval loop ──┐
       ├──► Sub-agent B: sub-question 2 ── own retrieval loop ──┤ parallel
       └──► Sub-agent C: sub-question 3 ── own retrieval loop ──┘
                                                                │
       ◄──── condensed findings + source IDs (not raw chunks) ──┘
       │
       ▼
  SYNTHESIS: draft the report from the findings
       │
       ▼
  CITATION PASS: attach every claim to a specific source passage;
                 drop or flag claims that have no supporting source

Anthropic's published account of its multi-agent research system (June 2025) gives the clearest public numbers on this pattern. A lead agent running Claude Opus 4 with Claude Sonnet 4 sub-agents outperformed a single-agent Claude Opus 4 by 90.2% on Anthropic's internal research eval. The same system used about 15× the tokens of a chat interaction, compared with about 4× for a single agent. On the BrowseComp benchmark, token usage alone explained 80% of the variance in performance. Anthropic also reported that the pattern works best for breadth-first questions that split into independent strands. It fits poorly when all agents need shared context or when the subtasks depend heavily on each other. The lesson for architects is that deep research mostly buys quality by spending tokens, so it must be budgeted the same way any other spend is.

Decision Table: Which Pattern, When

Pattern LLM calls per question Typical latency Cost vs. single-shot Quality gain Justified when
Single-shot RAG 1–2 1–3 s 1× Baseline Single-document lookups, FAQs, high-volume support
Corrective / Self-RAG 2–4 2–6 s ~1.5–3× Fewer confident wrong answers when retrieval is weak Uneven corpus quality; frequent out-of-domain queries
Multi-hop agentic 3–10 5–30 s ~3–10× Answers comparison and chained questions that single-shot cannot Multi-document comparison; answers that depend on earlier answers
Deep research Dozens to hundreds of tool calls across sub-agents Minutes Often 10×+ (about 15× a chat turn, per Anthropic) Breadth across many sources; a structured, cited report High-value research; asynchronous delivery; the user will wait

Latency, call counts, and multipliers other than the Anthropic figures are planning estimates, not benchmarks. Measure your own system.

Why agentic RAG costs 3–10×, not just "a few more calls." Each iteration is a full LLM call, and the context grows on every iteration because earlier evidence is carried forward. Token cost therefore grows faster than the call count. A worked estimate:

Single-shot:   1 call × 6K input tokens                      =  6K tokens

Multi-hop (4 retrieval rounds + synthesis, evidence accumulates):
  Round 1: 3K   Round 2: 7K   Round 3: 11K   Round 4: 15K
  Synthesis: 15K
  Total input                                               ≈ 51K tokens
                                                            ≈ 8.5× single-shot

Compressing evidence between rounds (keep extracted facts and source IDs, drop raw chunks) is the main lever for keeping the multiplier near 3× rather than 10×. This is the same context-management discipline covered in Module 6 (Failure Mode 2: context window overflow, and Failure Mode 6: cost explosion).

The routing rule. Do not make every query agentic. Put a classifier or router in front: simple lookups go to single-shot RAG, compound questions go to multi-hop, and explicit research requests go to deep research. If 80% of traffic is simple, routing keeps the blended cost close to the single-shot cost.

Controls: Keeping the Loop Bounded

AGENTIC RAG CONTROL PLANE

  ITERATION CAP
    Single agent: hard max of 3–5 retrieval rounds per question
    Deep research: max N sub-agents, max M tool calls per sub-agent
  RETRIEVAL BUDGET
    Token budget and $ cap per question; wall-clock timeout
    On exhaustion: answer with the gaps stated, or escalate.
    Never fail silently.
  STOPPING CRITERION ("enough evidence")
    Stop when every sub-question in the plan has at least one
    supporting passage from a Tier 1 or Tier 2 source,
    OR when two consecutive rounds add no new evidence.
  REDUNDANT-QUERY GUARD
    Reject a new query that is a near-duplicate of an earlier query
    in the same trajectory, or that returns the same chunk set
  CITATION ENFORCEMENT
    Every factual claim maps to a source ID and passage.
    A verifier pass removes or flags unsupported claims.
  SOURCE-TRUST TIERS
    Every retrieved passage carries a trust tier into the context.

Source-trust tiers. Agentic RAG mixes sources of very different reliability in one context. Make the difference explicit.

Tier Sources Treatment
Tier 1 Curated, versioned, access-controlled internal content Can be cited as authoritative
Tier 2 Internal but unreviewed content (wikis, tickets, email) Cite with a caveat; prefer corroboration from Tier 1
Tier 3 Public web pages, third-party and user-uploaded documents Untrusted data. Never followed as instructions. Consequential claims need corroboration.

Indirect prompt injection is a bigger risk in a loop. In single-shot RAG, a malicious instruction hidden in a retrieved document can corrupt one answer. In agentic RAG, retrieved content also shapes the agent's next actions: which queries it writes and which tools it calls. A web page can instruct the agent to search for internal data and then encode it in an outbound search query or URL, which turns a bad answer into data exfiltration. Architectural mitigations:

  • Label Tier 3 content as data in the context, and never let it change the plan or the tool permissions.
  • Separate the agent that reads the open web from the agent that holds sensitive internal data and write tools. Pass only extracted, validated findings between them.
  • Filter outbound queries and URLs for internal identifiers and sensitive data before they leave the perimeter.
  • Give research agents read-only tools. A research loop has no reason to send email or modify records.

Module 9 covers prompt injection defenses in depth.

Evals Specific to Agentic RAG

The RAGAS metrics in §4.8 still evaluate the final answer. Agentic RAG also needs trajectory-level evals, because two systems can produce the same answer at a 5× cost difference, or reach the right answer through an unsafe path. Log the full trajectory (plan, every query, tool choice, retrieved chunk IDs, stop reason) using the audit pattern from Failure Mode 7 in §4.2.

Metric Definition How to measure Example target
Per-step retrieval recall Fraction of planned sub-questions where the step retrieved at least one gold-relevant passage Gold passages per sub-question in the eval set ≥ 0.80
Redundant-query rate Fraction of queries in a trajectory that duplicate an earlier query or return the same chunk set Embedding similarity between queries, plus overlap of chunk IDs < 10%
Citation faithfulness Fraction of cited claims where the cited passage actually supports the claim LLM judge plus a human-reviewed sample ≥ 0.95 in regulated use
Cost per answered question Total spend ÷ questions answered (excluding refusals and escalations) Gateway cost logs joined to outcomes Set against the business value of an answer
Budget-exhaustion rate Fraction of questions that hit the iteration or token cap Stop-reason field in the trajectory log Track the trend; spikes signal corpus gaps or poor plans
Premature-stop rate Fraction of answers missing a sub-question that the evidence could have answered Compare the answer against gold sub-question coverage < 5%

Targets are illustrative starting points. Set your own baselines.

Cost per answered question is the metric that decides whether agentic RAG survives a budget review. A system that costs $0.40 per answer but resolves questions that previously took an analyst 30 minutes is cheap. A system that costs $0.40 per answer to replace a $0.04 single-shot lookup with no measurable quality gain is not.


4.11 The RAG Architecture Decision Checklist

Before any RAG system goes to production:

Knowledge Architecture - [ ] Chunking strategy defined and matched to document type? - [ ] Semantic or hierarchical chunking used (not just fixed-token)? - [ ] Embedding model selected and documented (name + version)? - [ ] Embedding model version stored as metadata with every vector? - [ ] Hybrid search (dense + sparse) implemented? - [ ] Metadata filtering defined for scope/tenant/access control?

Quality Architecture - [ ] Cross-encoder re-ranking implemented? - [ ] Confidence threshold defined and signed off by business owner? - [ ] Fallback behavior when confidence gate triggers defined? - [ ] RAGAS baseline evaluation completed? - [ ] Eval suite in place for regression testing?

Lifecycle Architecture - [ ] Document version management: old chunks archived on update? - [ ] Re-ingestion pipeline exists for document updates? - [ ] Embedding model migration process documented? - [ ] Document freshness monitoring in place?

Audit Architecture - [ ] Full interaction logging: query, chunks, versions, response? - [ ] Citations included in responses? - [ ] Retention period defined and infrastructure in place?

Agentic Retrieval Architecture (if agentic RAG or deep research is used) - [ ] Router in place so only compound or research queries take the agentic path? - [ ] Hard iteration cap, token budget, $ cap, and timeout set per question? - [ ] Explicit stopping criterion ("enough evidence") defined and tested? - [ ] Retriever tool descriptions versioned and routing cases in the eval set? - [ ] Source-trust tier attached to every retrieved passage? - [ ] Web and third-party content treated as untrusted data; research agents limited to read-only tools? - [ ] Citation verification pass removes or flags unsupported claims? - [ ] Full trajectory logged (plan, queries, tool choices, chunk IDs, stop reason)? - [ ] Per-step recall, redundant-query rate, citation faithfulness, and cost per answered question baselined?


EXERCISE — Failure Mode Diagnosis: You have a customer-facing RAG system. Users report that it occasionally gives incorrect information about account fees. Map the 7 failure modes against this symptom. Which failure modes could produce this symptom? For each plausible failure mode: what evidence would confirm it, and what is the architectural fix?

PONDER — The Confidence Threshold Decision: At what retrieval confidence score should your RAG system refuse to answer and escalate to a human? Write down a number. Now: who in your organization validated that number? What was the reasoning? What is the consequence if it is set too high (too many escalations) vs. too low (too many wrong answers)? Is that trade-off documented?

WORKSHOP — Chunking Strategy Design: Take a real document type from your organization (policy document, contract, FAQ, technical manual). Design the chunking strategy: boundary rules, chunk size target, overlap, metadata fields per chunk, parent-child structure if applicable. Run the strategy against 10 representative documents. Identify 3 places where the strategy produces bad chunks. Refine and document the final strategy.

WORKSHOP — RAGAS Baseline: Build a 30-case evaluation dataset for a RAG system: 10 simple factual queries, 10 multi-document queries, 5 edge cases, 5 out-of-scope queries. For each: write the ground truth answer. Run the current RAG system. Compute context precision and faithfulness manually for 10 cases. What do the numbers tell you about the system's current quality? What architectural change would most improve the lowest-scoring metric?


Next: Module 5 — AI Data Architecture