Skip to content

ONE-PAGER SAMPLES

(Format demonstration — Modules 4 & 6)


MODULE 4 — RAG Architecture

One-Page Essence

THE ESSENCE RAG quality is determined by retrieval, not by the model. If the right content isn't retrieved, no model can save the answer.

THE CORE INSIGHT Naive RAG (embed → vector search → stuff into prompt) fails in production in seven predictable ways. The difference between a demo and a production RAG system is not the LLM — it is the retrieval pipeline: how you chunk, whether you use hybrid search, whether you re-rank, and whether you gate on confidence. Most "the AI gave a wrong answer" problems are actually "the right content was never retrieved" problems.

THE KEY FRAMEWORK — The Production RAG Pipeline

Query → Query processing → Hybrid search (dense + BM25) →
Score fusion → Cross-encoder re-rank → Confidence gate →
  PASS: generate with citations
  FAIL: escalate / "I don't know"
Each stage is independently tunable. The confidence gate is the safety mechanism most teams skip.

THE DECISION RULE - Chunk size: start at 512 tokens with 10-15% overlap; tune by measuring, not guessing - Always use hybrid search (dense + keyword) in production — pure vector search misses exact-match terms - Add re-ranking when precision matters — it's the highest-ROI quality upgrade - Set a confidence threshold via calibration, and route low-confidence queries to humans

RED FLAGS (you've got it wrong if…) - You're tuning the prompt to fix what is actually a retrieval problem - Pure vector search, no keyword search, in production - No confidence gate — the system answers every query no matter how weak the retrieval - The knowledge base has no version management; stale docs silently degrade quality - You can't measure faithfulness, so "quality" is a vibe, not a number

THE ONE QUESTION "When the system gives a wrong answer — do you check whether the right content was retrieved before you touch the prompt?" If retrieval isn't the first place you look, you'll waste weeks tuning generation to compensate for a retrieval gap.

---

MODULE 6 — Agentic Systems

One-Page Essence

THE ESSENCE An agent is a loop with tools and a stopping condition. Most agent failures are missing stopping conditions and missing guardrails — architecture, not intelligence.

THE CORE INSIGHT An agent that can take actions is an agent that can take wrong actions. The architectural work in agentic systems is almost entirely about constraint: what tools the agent can use, when, with what limits, and what it cannot do without a human. The intelligence comes from the model; the reliability comes from the constraints you put around it. A capable model with no guardrails is a liability, not an asset.

THE KEY FRAMEWORK — The Agent Loop + Controls

GOAL → [reason → select tool → act → observe] → repeat → STOP
        ▲                                          │
        └──── controls wrap every iteration ───────┘

Required controls (all code-enforced, not prompt-enforced):
  • Hard iteration limit       • Per-task cost ceiling
  • Phase-scoped tool manifest • Human approval gate for irreversible acts
  • Full execution logging     • Idempotency on side-effecting tools

THE DECISION RULE - Ask "should this even be an agent?" first — if a deterministic workflow works, use that - Give the agent the minimum tools needed for the current phase, not all tools always - Reversible actions can be autonomous; irreversible actions require human approval — in code - Set the iteration limit and cost ceiling before the first run, enforce in code

RED FLAGS (you've got it wrong if…) - "Be careful" lives in the prompt instead of constraints living in the code - The agent has access to irreversible actions (delete, pay, send) without an approval gate - No iteration limit — the agent can loop until the budget is exhausted - 10+ tools available in a single phase — tool selection accuracy collapses - You can't replay or trace why the agent made a given decision

THE ONE QUESTION "What is the worst thing this agent can do before a human sees it — and is that limit enforced in code or just requested in the prompt?" If the answer is "requested in the prompt," it's not a limit. The model can reason around any instruction; it cannot reason around code.


Format notes: each one-pager is ~1 page (400-500 words), self-contained, shareable as a standalone artifact. The six-section structure is fixed across all modules so they form a consistent set.