MODULE 6 — Agentic Systems Architecture¶
Currency note: Product names and features are accurate as of October 2026. Volatile details (model versions, pricing, platform feature status) live in Appendix G.
6.1 What an Agent Actually Is (Architecturally)¶
The word "agent" is overloaded to the point of meaninglessness in 2025–2026. Marketing uses it to describe any feature with AI. Developers use it to mean a script that calls an LLM. Architects need a precise definition.
An AI agent is a system that: 1. Receives a goal, not just a query — it is given an outcome to achieve, not a single question to answer 2. Plans the steps to achieve that goal (either reactively or upfront) 3. Takes actions via tool calls that have effects in the real world 4. Observes the results of those actions 5. Adapts based on what it observes — replanning when steps fail or produce unexpected results 6. Terminates when the goal is achieved or when it cannot proceed
The architectural distinction that matters: a chatbot answers a question. An agent changes state in the world. The consequences of agent actions are real — emails sent, records created, APIs called, transactions initiated. This is why agent architecture requires a fundamentally different safety model than RAG or prompt engineering.
A Gartner August 2025 forecast projects that up to 40% of enterprise applications will feature task-specific AI agents by end of 2026, up from less than 5% in 2025. But a March 2026 enterprise architecture study (AaiNova, reported) found that while 79% of organizations report some AI agent adoption, only 11% are in production and just 2% have deployed at full scale. The architecture is where most deployments fail — not the AI capability.
Gartner (June 2025) also projects that over 40% of agentic AI projects will be cancelled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. The architect's job is to ensure the systems you design are in the 60% that survive.
6.2 The Agent Loop Patterns¶
Every agent is built around a loop: perceive → reason → act → observe → repeat. The design of this loop — how it reasons, when it plans, how it recovers — determines the system's reliability, cost, and safety profile.
Pattern 1: ReAct (Reason + Act)¶
The foundational pattern. At each step, the agent reasons about what to do next, executes one action, observes the result, and reasons again.
REACT LOOP
[Goal received]
│
▼
[THINK] — What is the current state? What should I do next?
│
▼
[ACT] — Call one tool with specific parameters
│
▼
[OBSERVE] — What did the tool return? Did it succeed?
│
├── Goal achieved? → Return result
├── Need more steps? → Back to THINK
└── Error encountered? → Reason about error, retry or adapt
When ReAct is appropriate: - Tasks with unpredictable intermediate states where the next step depends on what was discovered - Debugging workflows, exploratory data analysis, customer support with variable paths - Tasks where the tool call results are needed to determine subsequent actions
ReAct failure mode — the reasoning loop: The agent gets stuck in a loop. Step 1 produces output A. Step 2 depends on A. Step 2 fails. The agent reasons that it should retry step 1 with a different approach. Step 1 produces output B. Step 2 still fails. The agent retries step 1 again. This continues until the context window fills or the cost budget is exhausted. Without a hard iteration limit, ReAct loops are a cost and reliability risk.
Architectural constraint: Every ReAct agent must have a hard maximum iteration count enforced in code, not in the prompt. "Stop after 10 steps" in the system prompt can be reasoned around by a sufficiently confused LLM. if iteration_count > MAX_ITERATIONS: raise AgentIterationLimitExceeded() cannot.
Pattern 2: Plan-and-Execute¶
Separates the planning phase from the execution phase. A planner model generates the complete plan upfront. An executor model (often cheaper and faster) executes each step.
PLAN-AND-EXECUTE
[Goal received]
│
▼
[PLANNER] — Frontier reasoning model at high effort (e.g., GPT-6 class,
Claude Opus 5.5 with adaptive thinking; see Appendix G)
Generates: Directed Acyclic Graph of subtasks
{
step_1: {task: "retrieve customer profile", tool: "get_customer", dependencies: []},
step_2: {task: "check account eligibility", tool: "check_eligibility", dependencies: ["step_1"]},
step_3: {task: "calculate offer", tool: "calculate_offer", dependencies: ["step_1", "step_2"]},
step_4: {task: "generate recommendation", tool: "llm_generate", dependencies: ["step_3"]}
}
│
▼
[EXECUTOR] — Cheaper model or deterministic code
Executes each step in dependency order
Handles tool calls, error management, retries
│
▼
[RE-PLANNER] (triggered on failure only)
If step fails: pass failure context back to planner
Planner generates revised plan from current state
Executor continues with revised plan
Why Plan-and-Execute beats ReAct for many enterprise tasks:
Cost efficiency. The planner makes one expensive LLM call (the reasoning model). Execution uses a cheap model or deterministic code. Benchmarks show Plan-and-Execute achieves ~92% task completion with 3.6x speedup vs pure ReAct for structured multi-step workflows.
Predictability. The plan is visible before execution begins. It can be validated, logged, and presented to a human for approval before any irreversible action is taken. With ReAct, you don't know what the agent will do until it does it.
Parallelism. Steps with no dependencies can execute in parallel. ReAct is inherently sequential. A Plan-and-Execute workflow with 8 steps, 5 of which are independent, can execute in 3 sequential rounds instead of 8.
When Plan-and-Execute is appropriate: - Well-defined workflows with known tool sets and predictable step sequences - Tasks where human review of the plan before execution is a requirement - High-volume automation where the same workflow pattern repeats
When Plan-and-Execute is not appropriate: - Highly dynamic tasks where the correct next step is genuinely unknowable until the prior step completes - Exploratory tasks where the problem definition itself may change during execution
Pattern 3: Hierarchical / State Machine (Enterprise Grade)¶
The production pattern for regulated, high-stakes, or high-complexity agent workflows. The agent does not operate as a free-form loop — it moves through defined states with explicit transitions, and the AI is autonomous within a state but cannot transition without meeting defined criteria.
STATE MACHINE AGENT ARCHITECTURE
States: [DATA_COLLECTION] → [ANALYSIS] → [HUMAN_REVIEW] → [ACTION] → [COMPLETE]
TRANSITIONS:
DATA_COLLECTION → ANALYSIS:
Condition: all required data sources returned valid results
On failure: route to ERROR_STATE with missing data report
ANALYSIS → HUMAN_REVIEW:
Condition: always (human review is mandatory for this workflow)
Human sees: analysis results, confidence scores, recommended action
Human can: approve, modify, reject
HUMAN_REVIEW → ACTION:
Condition: human explicitly approved
Timeout: if no response in 4 hours, escalate to supervisor
HUMAN_REVIEW → DATA_COLLECTION:
Condition: human requests additional data
ACTION → COMPLETE:
Condition: action executed successfully, confirmation received
Any state → ERROR_STATE:
Condition: unrecoverable error, iteration limit reached, cost limit reached
Why this is the architecture for regulated environments:
The state machine makes the agent's behavior auditable by design. Every state transition is logged with the condition that triggered it. Every human interaction is recorded with the approver's identity and the decision made. The agent cannot skip states — the transition conditions are enforced by the orchestration framework, not by the LLM's judgment.
The LLM's judgment is used within states (how to analyze the data, how to interpret results) but not for state transitions (whether to take action, whether human review is needed). State transitions are deterministic code. This is the fundamental principle that makes agentic systems governable.
Tools: LangGraph is the primary framework for state machine agent architectures. Its graph-based model maps directly to this pattern — nodes are states, edges are transitions with conditions. Temporal (workflow engine) provides even stronger guarantees for long-running workflows with persistence, retry, and audit.
Pattern 4: Reflection and Self-Critique¶
An agent that evaluates its own outputs and iterates to improve quality before returning results.
REFLECTION PATTERN
[Initial output generated]
│
▼
[CRITIC LLM] — Evaluates output against criteria:
- Is the answer complete?
- Are all claims supported by evidence?
- Are there logical inconsistencies?
- Does it meet the quality bar?
│
├── Passes criteria → Return output
└── Fails criteria → [REVISE] → [CRITIC] again (max 3 iterations)
When reflection adds value: Complex generation tasks (contract drafting, compliance analysis, code review) where quality is more important than latency and a second-pass evaluation catches genuine errors. When reflection wastes cost: Simple generation tasks, time-sensitive responses, tasks where the critic LLM has the same blind spots as the generator LLM (they may agree on wrong answers).
Pattern 5: Long-Running Agents and Durable Execution¶
Standard agent architectures assume a task completes in seconds to minutes. Enterprise workflows often span hours, days, or weeks — a contract review that waits for legal sign-off, a compliance check that waits for a third-party verification, an onboarding workflow that spans multiple business days.
Long-running agents require a fundamentally different execution model:
DURABLE EXECUTION PATTERN
Problem: A standard agent loop running for 3 days will:
- Lose context on server restart
- Fail on memory pressure
- Accumulate enormous context costs
- Have no recovery path on infrastructure failure
Solution: Durable execution with persistent state
[Task starts] → [Execute Step 1] → [Checkpoint state to durable store]
│
[Server restart / pause] │ State persisted
│
[Task resumes] → [Load state from store] → [Execute Step 2] → [Checkpoint]
│
[Human approval required] → [Suspend task] ─────────────────────────┘
[State: WAITING_FOR_APPROVAL]
[Resume when approval received]
What durable execution requires: - Task state is serializable — every piece of state can be written to and read from storage - Each step is idempotent — resuming from a checkpoint cannot re-execute a completed step - The workflow engine handles scheduling, retries, and timeouts — not the agent's LLM reasoning - Human wait states are first-class — the agent can suspend indefinitely waiting for human input
Tools: Temporal (strongest durable execution guarantees, production-proven), LangGraph with Redis/Postgres persistence (lighter weight for simpler workflows), AWS Step Functions (managed, event-driven long-running workflows).
6.3 Agent Memory Architecture¶
An agent without memory is stateless — it starts from zero on every invocation. An agent with memory can build context over time, learn from prior interactions, and make decisions informed by history. The memory architecture determines what the agent knows, how it knows it, and what persists when the session ends.
The Four Memory Tiers¶
AGENT MEMORY ARCHITECTURE
TIER 1: IN-CONTEXT MEMORY (working memory)
────────────────────────────────────────────────────────────
What: Everything in the current context window
Scope: Current session, current task
Persistence: None — lost when the context ends
Cost: Every token in context costs money at inference time
Capacity: Model-dependent (32K–2M tokens depending on model)
Use for: Current task state, recent tool results, conversation
history for the current session
ARCHITECTURAL CONCERN:
In-context memory grows with every turn and tool call.
At 10 turns × 500 tokens/turn = 5,000 tokens of history.
At 50 turns: 25,000 tokens — meaningful cost overhead.
Manage context size explicitly (summarization, selective inclusion).
Section 6.4 covers the techniques and a budget worksheet.
TIER 2: EXTERNAL SHORT-TERM MEMORY (session memory)
────────────────────────────────────────────────────────────
What: Summaries, key facts, task state for current session
Scope: Current session or task (hours to days)
Persistence: Session duration — stored in Redis, key-value store
Cost: Storage cost (negligible), retrieval latency (fast)
Capacity: Unlimited in principle, but use judiciously
Use for: Multi-turn conversation summaries, task checkpoints,
intermediate results for long-running workflows
Pattern: After every N turns, summarize the conversation into
structured key facts. Store summary in session store.
On next turn: inject summary (200 tokens) instead of full
history (5,000 tokens). Same information, 96% cost reduction.
TIER 3: EXTERNAL LONG-TERM MEMORY (persistent memory)
────────────────────────────────────────────────────────────
What: User preferences, learned facts, historical outcomes
Scope: Persistent across sessions (weeks to indefinitely)
Persistence: Database, vector store (for semantic retrieval)
Cost: Storage + retrieval query cost
Capacity: Effectively unlimited
Use for: User preferences ("always send reports as CSV"),
historical context ("last quarter's review found X"),
personalization ("this user prefers concise answers")
ARCHITECTURAL CONCERN:
Long-term memory raises significant privacy questions.
What is stored? For how long? Who can access it? Can users
delete it? In regulated environments, long-term memory about
users may be subject to data retention and deletion requirements.
Design the privacy model before implementing long-term memory.
TIER 4: EPISODIC MEMORY (experience memory)
────────────────────────────────────────────────────────────
What: Records of past task executions and their outcomes
Scope: Persistent — the agent's track record
Persistence: Append-only log with outcome metadata
Use for: "The last 3 times I tried approach X for this
type of task, it failed. Try approach Y instead."
Few-shot examples from real past successes.
Quality improvement over time.
Most production systems do not implement episodic memory yet.
It requires careful design to prevent the agent from "learning"
incorrect patterns from failed or atypical past executions.
Memory Architecture Decision Matrix¶
QUESTION: Does this information need to persist?
NO → In-context (Tier 1) only
QUESTION: Does it persist across turns in the same session?
YES, session-scoped → External short-term (Tier 2)
QUESTION: Does it persist across sessions for the same user/task?
YES → External long-term (Tier 3)
Is privacy analysis done? If no: do not implement until done.
QUESTION: Do past execution outcomes improve future performance?
YES, and quality evaluation is in place → Episodic (Tier 4)
Quality evaluation NOT in place: do not implement Tier 4.
Bad episodic memories make agents worse, not better.
6.4 Context Engineering: Managing the Agent's Working Memory¶
Section 6.3 treats the context window as Tier 1 memory. This section covers how to manage it. Module 32 §32.5 explains why context engineering now matters more than prompt engineering. This section gives the mechanics: the budget, the techniques, and the signals that tell you which technique you need.
Context Is a Budget, Not a Container¶
A large context window tells you how much the model will accept. It does not tell you how much the model will use well. Two findings matter here:
- Context rot. Chroma's July 2025 technical report Context Rot (Hong, Troynikov, Huber) tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3. Performance became less reliable as input length grew, even on deliberately simple tasks. The decline was not uniform. It was steeper when the question and the relevant passage had low semantic similarity, and it got worse when related but irrelevant "distractor" passages were present. On a conversational memory benchmark (LongMemEval), models did significantly better with a focused input than with a full-length input padded with irrelevant history.
- The attention budget. Anthropic's September 2025 engineering guide Effective Context Engineering for AI Agents describes the same effect: as the number of tokens grows, the model's ability to recall information from the context goes down. Every token you add draws on a finite attention budget.
The architectural consequence: set a design ceiling well below the advertised window. As of October 2026, frontier models offer windows of 1M tokens (see Appendix G). That does not mean an agent should run at 900K. Pick a design ceiling, such as 200K, and treat everything above it as an incident, not a feature. Then decide what fills that budget on purpose.
ANATOMY OF AN AGENT CONTEXT (order matters, see "cache-aware layout" below)
STABLE PREFIX (changes rarely; cached) VOLATILE TAIL (changes every turn)
┌───────────┬────────────┬─────────────┐ ┌─────────────┬─────────────┬─────────────┐
│ Tool defs │ System │ Skills index│ │ Task state /│ Conversation│ Latest tool │
│ (or tool │ prompt + │ + long-term │ │ notes (re- │ + tool-call │ results + │
│ search) │ invariant │ memory │ │ injected) │ history │ user turn │
│ │ constraints│ summary │ │ │ (compacted) │ │
└───────────┴────────────┴─────────────┘ └─────────────┴─────────────┴─────────────┘
▲ + room for
cache breakpoint output
The Techniques¶
| Technique | What it does | Use when | Main trade-off |
|---|---|---|---|
| Compaction | Summarizes older turns into a shorter block | Long sessions that must continue past the ceiling | Lossy; adds a summarization call; rewrites the cache |
| Context editing (clearing) | Deletes stale tool results or old reasoning blocks | Tool-heavy loops where old results are no longer needed | Lossless only if the data can be fetched again |
| Just-in-time retrieval | Keeps references in context and fetches content through tools | Large corpora, codebases, document sets | More tool calls and latency per step |
| Tool search / deferred loading | Loads tool definitions only when needed | More than about 10 tools, or tool definitions over about 10K tokens | Search can miss the right tool |
| Sub-agent isolation | Sends reading-heavy subtasks to sub-agents that return summaries | Broad research or exploration subtasks | Multiplies token cost; lossy handoff |
| Structured note-taking | Agent writes progress to external files and reloads them | Multi-hour tasks, tasks that span context resets | Notes go stale or get polluted |
| Cache-aware layout | Orders content stable-first so the prefix is reused | Always | Constrains how you edit history |
1. Compaction (summarize older turns)
Compaction replaces the older part of the conversation with a summary and continues from there. Failure Mode 2 (§6.7) recommends it at 70% utilization. You can write your own summarizer, or use a provider's server-side version. As of October 2026, Anthropic's API offers server-side compaction in beta in two forms:
- Compaction at a token threshold (beta header
compact-2026-01-12): when input reaches a trigger you set (default 150,000 input tokens, minimum 50,000), the API writes a summary and returns it as acompactionblock. Optionally, it can pause after compaction so you can re-insert recent turns word for word. - Compaction on demand (beta header
compact-2026-09-04): your code asks for the summary when it chooses. This can run in the background and can keep recent turns verbatim.
The integration detail that matters: pass the compaction block back, not the text inside it. With threshold compaction you append the full response content, and the API drops everything before the block on the next request. If you copy only the summary text into a new user message, the API no longer recognizes it as a compaction point. On models that preserve and sign reasoning blocks, the signed block is also what keeps later reasoning valid.
Cost note: compaction is an extra model call. On Anthropic's API, the top-level input_tokens and output_tokens exclude the compaction pass. You must sum the usage.iterations array to get the billed total. If your cost meter reads only the top-level fields, compaction cost is invisible.
Failure modes: the summary drops a constraint (the "definition of non-standard clause" in Failure Mode 2). Repeated compaction turns into a summary of a summary, and detail decays each time. Mitigation: never let invariants live only in the compactable region. Keep the goal, scope, and prohibited actions in the system prompt or in a re-injected task-state block. Supply your own summarization instructions that list what must survive. Add compaction cases to the agent's regression suite.
2. Context editing (clear stale content)
Clearing differs from compaction. Compaction rewrites history into a summary. Clearing removes specific items, usually old tool results, and leaves a placeholder. It is the lightest option. Anthropic's guide calls tool-result clearing one of the safest forms of compaction. As of October 2026, Anthropic's context editing beta (context-management-2025-06-27) offers:
clear_tool_uses_20250919: clears old tool results once input passes a trigger (default 100,000 tokens) and keeps the most recent N tool uses (default 3). You can exclude specific tools and set a minimum amount to clear per pass (clear_at_least).clear_thinking_20251015: clears reasoning blocks from older turns and keeps the last N turns.
Use when: the result can be fetched again (a file read, a search, an API GET). Do not use for results that cannot be reproduced, such as a third-party verification response or a one-time token. Write those to notes first.
Failure mode, the re-fetch thrash: the agent no longer sees a result, so it calls the same tool again. That result gets cleared, and the cycle repeats. In your traces this looks exactly like the infinite loop of Failure Mode 1. Mitigation: have the agent record the conclusion from each large result in its notes (technique 6) before the result is cleared. Count repeated identical tool calls per task.
3. Just-in-time retrieval instead of pre-loading
Pre-loading puts all possibly relevant material into context before the agent starts. Just-in-time (JIT) retrieval keeps lightweight references in context, such as file paths, record IDs, stored query names, and URLs. The agent fetches content through tools only when a step needs it. This is how coding agents work with repositories that are far larger than any window. They list, search, and read ranges of files instead of loading the repository.
| Pre-load | Just-in-time | |
|---|---|---|
| Best for | Small, stable, always-needed material (a policy, a schema) | Large or changing corpora; material only some steps need |
| Latency | Lowest per step | One or more extra tool calls per step |
| Context cost | High and constant | Proportional to what is actually used |
| Risk | Context rot and distractors | Agent does not know something exists, so it never fetches it |
Design rules: give the agent a map (an index, a manifest, a directory listing) so it knows what exists. Cap tool-result size and paginate, for example 8K tokens per result with a "next page" cursor, so one fetch cannot flood the context. Use hybrids: pre-load the 2K-token summary and fetch the details just in time.
4. Tool search and deferred tool loading
Tool definitions are context too. Anthropic's documentation (as of October 2026) gives the numbers. A typical setup with five MCP servers (GitHub, Slack, Sentry, Grafana, Splunk) uses about 55K tokens of tool definitions before the agent does any work. Tool-selection accuracy degrades once more than 30–50 tools are exposed. That is the same accuracy collapse that Failure Mode 4 (tool authorization creep) produces.
The fix is to load definitions on demand. In Anthropic's API, you mark tools defer_loading: true and include a tool-search tool (regex or BM25 variants). The model sees only the search tool and your non-deferred tools. It searches the catalog when it needs a capability, and the API expands the matches, 5 by default, into full definitions. Anthropic reports that this typically cuts tool-definition context by more than 85%. Documented limits: up to 10,000 deferred tools per request. The guidance is to keep the 3–5 most-used tools non-deferred. Deferred tools are kept out of the cached prefix, so discovering a tool does not break the prompt cache. MCP toolsets can be deferred per server. Several MCP clients and agent frameworks implement the same on-demand pattern on the client side.
Two cautions. First, deferral saves context, not request size: you still send every definition with each request. Second, deferral is not least privilege. A deferred tool is still callable once the model finds it. Phase-based tool access (Failure Mode 4) still decides which tools are in the manifest at all. Tool search only decides which ones the model reads.
Failure mode: the search misses the right tool and the agent improvises with the wrong one. Mitigation: use consistent name prefixes (crm_, billing_), put keywords in descriptions that match how users phrase tasks, and log which tools are discovered and which are used.
5. Sub-agent context isolation
Give a reading-heavy subtask, such as "survey these 40 documents" or "find every caller of this function," to a sub-agent with a fresh context. It may read 100K tokens, but it returns only a condensed result. Anthropic's guide reports that these results are often 1,000–2,000 tokens. The orchestrator's context stays clean. Module 32 cites the AWS DevOps Agent as an example of this design.
Trade-offs: total tokens go up, not down. Each sub-agent re-reads its own material and needs its own system prompt and tools. A plan with 5 sub-agents that each read 80K tokens consumes 400K input tokens to save the orchestrator about 70K. Budget for it using the cost model in Module 7 §7.8. Failure modes: the handoff summary omits the detail the orchestrator needed. Sub-agents duplicate each other's work. Sub-agents ignore constraints they were never given. Mitigation: give each sub-agent an explicit return schema (findings, evidence references, open questions) and pass the invariant constraint block to every sub-agent. Multi-agent coordination is covered in Module 7.
6. Structured note-taking (external scratchpads and memory files)
The agent writes its progress to a store outside the context, such as a progress.md, a todo list, a JSON task-state record, or a memory-tool directory. It re-reads that store after a compaction, a context reset, or a resume. This is how §6.3's tiers work in practice. Notes for the current task are Tier 2 session memory. Notes promoted across tasks become Tier 3 or Tier 4, and the privacy and quality gates in §6.3 apply. Anthropic's memory tool combines with context editing: the model is warned as it approaches the clearing threshold so it can save what matters first.
Failure modes: notes go stale (they claim a step is done after it was rolled back). Notes become an injection channel: text copied from a malicious document into notes is re-read later as if the agent had written it. Mitigation: use a schema for task state (step IDs, status, evidence pointers), not free-form prose. Make the orchestrator, not the model, the authority on step completion. Treat note content that came from tools as untrusted input (Module 9).
7. Prompt-cache-aware context layout
Prompt caches are exact prefix matches. Anthropic builds the prefix in a fixed order: tools, then system, then messages. A change at any level invalidates that level and everything after it. Changing a tool definition invalidates the whole cache. On Anthropic's API (as of October 2026), standard cache reads cost 0.1x the base input price (some newer models price reads lower). Five-minute cache writes cost 1.25x and one-hour writes cost 2x. The minimum cacheable prefix is 512 to 4,096 tokens depending on the model. OpenAI caches prompts of 1,024 tokens or more automatically, on the same prefix principle. A cache holds one model's internal state, so a fallback to a different model starts with a cold cache. Include that in fallback cost estimates (Module 37).
CACHE-AWARE LAYOUT RULES
1. Stable first: tools → system prompt → skills index → long-term memory
2. Volatile last: timestamps, user turn, latest tool results
3. Never put a timestamp or request ID in the system prompt
4. Treat messages as append-only; do not rewrite earlier turns
5. Edit history only through primitives designed for it
(server-side compaction, server-side clearing) and do it in
large batches (e.g., clear_at_least) so each cache rewrite pays off
Rule 4 is stronger on newer models. As of October 2026, Anthropic models that preserve and sign reasoning blocks across turns (Claude Opus 5.5, Sonnet 5.5, Fable 5.1) check that the prefix before each block is unchanged. Editing an earlier message, rebuilding the system prompt, changing the tool list, or trimming old tool results on the client invalidates every later reasoning block. On accounts created from August 31, 2026, the default response is an HTTP 400 error. The supported paths are append-only history, mid-conversation system messages for new instructions, and the server-side compaction and clearing features above.
The Context Budget Worksheet¶
Fill this in for every agent before you build it. Example: a document-review agent with a 200K design ceiling on a model with a 1M window.
CONTEXT BUDGET — contract-review agent Design ceiling: 200,000 tokens
Slot Budget % Technique / control
────────────────────────────────────────────────────────────────────────
System prompt + invariants 6,000 3% Fixed; cached; never compacted
Tool definitions 8,000 4% Tool search; 4 hot tools loaded
Skills index (names + descr.) 1,000 0.5% Progressive disclosure (§6.5)
Long-term memory / user profile 3,000 1.5% Retrieved once per session
Task state + notes 6,000 3% Re-injected after compaction
Retrieved documents (JIT) 40,000 20% 8K cap per fetch, paginated
Conversation + tool history 96,000 48% Clear tool results >100K total;
compact at 150K total
Output + reasoning working space 24,000 12% max_tokens + reasoning effort
Headroom (never planned) 16,000 8% Alert if consumed
────────────────────────────────────────────────────────────────────────
Total 200,000 100%
Triggers (total input tokens): 100K clear stale tool results
150K compact (summary instructions pinned)
180K checkpoint, hand off to fresh context (§6.2 P5)
Cost check: 40 calls × ~110K avg input = 4.4M input tokens/task.
With ~85% served from cache at 0.1x: ~1.0M full-price-equivalent
tokens, plus cache-write premiums and any compaction calls.
The percentages are not universal. A coding agent spends more on history, and a RAG-heavy analyst spends more on retrieval. What is universal is that every slot has an owner, a cap, and a technique.
Signals: Which Technique Do You Need?¶
| Signal in traces or evals | Likely cause | Technique |
|---|---|---|
| Quality drops late in long tasks; early constraints ignored | Context rot; invariants buried | Pin invariants; compaction; lower ceiling |
| Same tool called repeatedly with identical arguments | Re-fetch thrash after clearing | Note-taking before clearing; fewer, larger clears |
| Over 20% of input tokens are tool definitions | Large catalog loaded every call | Tool search / deferred loading |
| Wrong tool chosen; plausible but irrelevant tool calls | Too many tools visible | Tool search; phase-based manifests |
| Single tool results over 20K tokens | Unbounded fetches | Result caps, pagination, JIT ranged reads |
| Orchestrator context grows with each research subtask | Raw findings returned to parent | Sub-agents with return schema |
| Agent repeats completed work after a resume | State only lived in context | Structured notes; orchestrator-owned task state |
| Cache hit rate below 50% on a multi-turn agent | Volatile content in prefix; history edits | Cache-aware layout; append-only messages |
| Cost per task jumps after compaction ships | Compaction calls not metered | Sum all usage iterations; budget compaction |
6.5 Agent Harnesses: Build the Loop, Use an SDK, or Use a Managed Runtime¶
Everything in §6.2–§6.4 has to live somewhere. That somewhere is the harness: the code that runs the loop, manages the context, executes tools, holds state, and enforces guardrails. Do not confuse the harness with the deployment, which is where the harness runs (your container, a serverless function, a provider's sandbox). The two decisions are related but separate. You can run a hand-written loop on AgentCore, or run a provider's harness on your own infrastructure.
WHAT A HARNESS CONTAINS
┌──────────────────────────────────────────────────────────────────┐
│ LOOP call model → parse tool calls → execute → repeat │
│ CONTEXT MGR budget, compaction, clearing, caching (§6.4) │
│ TOOL EXECUTOR tool registry, MCP clients, sandbox, timeouts │
│ STATE checkpoints, sessions, resume, idempotency keys │
│ GUARDRAILS iteration/cost/time limits, permissions, │
│ approval gates (§6.6), output validation │
│ TELEMETRY traces, token and cost metering (§6.8) │
└──────────────────────────────────────────────────────────────────┘
DEPLOYMENT = where this runs: your service | serverless | managed runtime
The Four Options¶
Option 1: Hand-written loop on a model API. You call the model API and write every component in the box above. A basic loop is short; production-grade context management, guardrails, sandboxing, and resume logic are not. You get full control and full visibility. You also own every bug, and you have to rebuild compaction, caching discipline, and sandboxing yourself.
Option 2: SDK tool runner or orchestration framework. A library runs the loop, and you supply tools and policy. Examples include provider SDK tool runners (Anthropic's client SDKs ship a beta tool runner), LangGraph (graph and state-machine orchestration with checkpointing, the natural fit for §6.2 Pattern 3), the OpenAI Agents SDK, Strands Agents (AWS, open source), and Microsoft Agent Framework (open source; unifies Semantic Kernel and AutoGen; 1.0 released April 2026). Most are multi-provider. You still host them and still own context strategy.
Option 3: Packaged agent harness as a library. This is a complete harness with built-in tools that you embed in your own process. The leading example is the Claude Agent SDK: Claude Code packaged as a Python and TypeScript library. It includes the same agent loop, context management, and built-in tools (file read, write, and edit; shell; search; web), plus hooks, sub-agents, permissions, sessions, MCP, and skills. You get a production-tested harness without giving up your deployment. In exchange, you accept its loop design and its model family.
Option 4: Managed agent runtime. The provider runs the infrastructure. Two variants differ in architecturally important ways:
- Provider runs the loop: Claude Managed Agents (beta as of October 2026, header
managed-agents-2026-04-01). You define a versioned agent (model, system prompt, tools, MCP servers, skills) and an environment (an Anthropic-managed cloud sandbox or a self-hosted sandbox). Then you start sessions. Each session gets its own isolated container. Events stream over SSE. Event history, sandbox state, and outputs are persisted server-side, and prompt caching and compaction are built in. Sessions can be pinned to an agent version. Compliance note: because it stores state server-side, Managed Agents is not eligible for Zero Data Retention or HIPAA BAA coverage as of October 2026. - Provider hosts your loop: you bring a harness from Option 1, 2, or 3, and the platform supplies compute, identity, memory, and governance.
- AWS Bedrock AgentCore (GA October 2025): Runtime, Memory, Identity, Gateway, Observability, Code Interpreter, and Browser. The default microVM runtime supports sessions of up to 8 hours. Runtime instances (EC2-backed, in your account; GA August 2026) support sessions of up to 14 days.
- Microsoft Foundry Agent Service (GA March 2026): Hosted Agents, a managed runtime for your agent code, reached GA in July 2026.
- Google Gemini Enterprise Agent Platform (replaced Vertex AI in April 2026): Agent Runtime (formerly Agent Engine, reported) runs agents continuously for up to 7 days, with Memory Bank, Agent Identity, and Agent Gateway.
Platform specifics and comparisons are in Module 31.
The Decision Table¶
| Criterion | 1. Hand-written loop | 2. SDK / framework | 3. Packaged harness library | 4a. Managed: provider loop | 4b. Managed: hosts your loop |
|---|---|---|---|---|---|
| Control over loop and context | Total | High | Medium (hooks, config) | Low (config only) | Same as the harness you bring |
| Model portability | Highest (your adapter) | High (multi-provider) | Low (one model family) | Lowest | Same as the harness you bring |
| Time to first production agent | Slowest | Medium | Fast | Fastest | Medium |
| Ops burden | Highest (you run everything) | High | Medium (you host it) | Lowest | Low to medium |
| Compliance and data residency | Your controls | Your controls | Your controls | Provider's terms (check ZDR/BAA/region); self-hosted sandbox option | Cloud region and account controls |
| Lock-in surface | Model API only | Framework + model | Harness + model family | Loop, state format, sessions, config | Platform services (memory, identity, gateway) |
| Debugging visibility | Complete | High (framework tracing) | High (hooks, transcripts) | Event stream only; internals opaque | Platform observability + your traces |
| Cost model | Tokens + your compute | Tokens + your compute | Tokens + your compute | Tokens + provider runtime charges (check current pricing) | Tokens + platform runtime charges (consumption-based; see Module 31) |
Heuristics:
- Regulated workflow with a state machine (Pattern 3): use Option 2 (LangGraph or Microsoft Agent Framework), possibly hosted on Option 4b. Deterministic transitions belong in code you own.
- Coding, file, and operations agents that need a sandbox fast: use Option 3 or 4a. Rebuilding a mature file and shell harness is rarely a good use of time.
- Multi-day asynchronous work (Pattern 5): use Option 4a or 4b for execution windows, but keep durable task state in your own store. A 14-day session is not a substitute for checkpoints.
- Strict portability or multi-model routing: use Option 1 or 2 behind the adapter from Module 37 §37.3.
- Pilot that will be judged on time to value: Option 3 or 4a is fine. Record the lock-in deliberately (see below) so the pilot does not become the production architecture by default.
Skills: Packaging Capability Without Paying for It in Context¶
Agent Skills apply progressive disclosure to the harness. A skill is a folder with a SKILL.md file (a name, a description, and instructions), plus optional scripts, reference documents, and templates. At startup the agent loads only each skill's name and description. When a task matches, it reads the full SKILL.md. It loads bundled files or runs bundled scripts only if the instructions call for them. Anthropic's documentation puts the always-loaded metadata at about 100 tokens per skill and the full instructions at under 5K tokens. Fifty installed skills therefore cost about 5K tokens of index, not 250K tokens of instructions.
Anthropic originated the format and released it as an open standard (agentskills.io). As of October 2026, the standard's client showcase lists support in Claude and Claude Code, OpenAI Codex, Gemini CLI, GitHub Copilot, VS Code, Cursor, and many others. Skills attach directly to Claude Managed Agents agent definitions and load in the Claude Agent SDK. That makes skills one of the more portable harness artifacts.
Architect's rules for skills: a skill with scripts is executable code, so review, version, and sign it like code (supply-chain risk, Module 9). Write skill descriptions for retrieval, because they are what the agent matches against, just like tool descriptions. Include skills in the agent version (§6.10: model + prompt + tool manifest + skill set). Check data-retention terms: as of October 2026, Anthropic's Agent Skills are not covered by Zero Data Retention arrangements.
Portability: Own Your State, Even on Someone Else's Runtime¶
Managed runtimes make it easy to let the provider own your agent's memory. Do not do that by accident.
PORTABILITY RULES FOR ANY HARNESS
1. Canonical task state lives in YOUR store (Postgres, DynamoDB, etc.),
in YOUR schema. The runtime's session is a cache of it, not the source.
2. Mirror conversation/event history in a provider-neutral form
(role, content, tool call, tool result, timestamps). Keep provider-specific
blocks (compaction, signed reasoning) as opaque attachments.
3. Tool definitions live in your registry (MCP servers you control),
not only in a runtime's console.
4. Agent config (prompt, tools, skills, limits) is in version control
and deployed to the runtime, never authored only in the runtime.
5. Your eval suite runs outside the runtime, against any harness.
6. Exit test: could you replay a task from your store on a different
harness within one sprint? If not, record it as lock-in.
Each managed-runtime dependency (server-side session history, provider memory services, runtime-specific identity) should be an entry in the lock-in ledger in Module 37 §37.10. The ledger records why you accepted the dependency, what leaving would cost, and when you will review it. Module 37 also covers the adapter and behavioral-contract mechanics that make a harness or model swap testable. For which cloud runtime fits which organization, see Module 31.
6.6 Human-in-the-Loop Design¶
The hardest architectural decision in agentic systems is not where to put humans — it is what criteria trigger human involvement and what information humans need to make a good decision quickly.
The Four Human Touchpoint Types¶
Type 1: Approval Gate (before irreversible action)
Trigger: Agent is about to take an irreversible action
(send email, execute transaction, delete record, publish content)
Human sees: Proposed action + full context + agent reasoning
Human options: Approve / Modify / Reject / Request more information
Timeout behavior: If no response in T hours, escalate to supervisor
Do NOT default-approve on timeout — this is a risk decision
Architecture: The agent generates an "action proposal" object,
not a request to take the action. The action cannot
be executed until the proposal is approved. Approval
is recorded with approver identity and timestamp.
Type 2: Quality Review (before customer-facing output)
Trigger: Agent-generated content will reach an external party
(customer email, report, public content, regulatory filing)
Human sees: Drafted content + sources + confidence scores
Human options: Approve as-is / Edit and approve / Reject and regenerate
Architecture: Content staging area. Agent deposits to staging.
Human reviews and promotes to delivery queue.
No direct path from agent to customer.
Type 3: Confidence-Triggered Escalation
Trigger: Agent confidence score falls below threshold
(threshold defined per workflow, not globally)
Human sees: Query + what the agent knows + what it's unsure about
Human options: Provide missing information / Take over the task /
Tell agent to proceed with best effort (explicit acceptance)
Architecture: Confidence evaluation at each planning step.
Threshold breach routes to human queue automatically.
Agent suspends, preserving full task state, until human responds.
Type 4: Exception Handling
Trigger: Agent encounters an error it cannot recover from
(tool failure, unexpected data, constraint violation)
Human sees: Error context + what was attempted + what failed
Human options: Resolve the underlying issue / Provide alternative approach /
Abandon the task
Architecture: Structured error objects (not free-form LLM error messages).
Task state preserved for resumption after human resolution.
Incident logged for pattern analysis.
The Human-in-the-Loop Quality Problem¶
The most common failure mode in human-in-the-loop design is not the absence of human review points — it is designing review points that give humans insufficient context to make good decisions quickly, causing them to either rubber-stamp agent actions (defeating the purpose) or spend excessive time reviewing (creating a bottleneck that eliminates the efficiency benefit).
What humans need for effective review:
EFFECTIVE APPROVAL REQUEST STRUCTURE
{
"summary": "Agent proposes to send a rebalancing recommendation
email to 847 customers based on Q2 market analysis.",
"what_changes": "847 customers will receive an email recommending
portfolio adjustments. Average recommended change:
8% equity reduction.",
"why_agent_recommends_this": "Q2 returns underperformed benchmark
by 2.3%. Customer risk profiles
indicate current allocation is
outside tolerance bands for 847 accounts.",
"confidence": "High (0.91) — all 847 accounts have current risk
profiles. Market data is from this morning.",
"what_could_go_wrong": "If market conditions change materially
between now and send time, recommendations
may be suboptimal. Data cutoff: 9:00 AM.",
"irreversibility": "HIGH — emails cannot be unsent. Customers
may act on recommendations immediately.",
"data_sources_used": ["portfolio_service v2.1", "market_data 2024-03-15 09:00",
"risk_profile_service v3.4"],
"approve_by": "2024-03-15T14:00:00Z (5 hours from now)",
"if_not_approved_by_deadline": "Task will be cancelled and flagged
for manual follow-up. No action taken."
}
This structure gives the human reviewer exactly what they need: what will change, why, how confident the agent is, what the risk is, and what happens if they don't act. A reviewer can make a quality decision in 60 seconds with this information.
Never: Present the human with a summary-only approval request for an irreversible action. "Agent recommends sending 847 emails. Approve?" — this is not a reviewable decision. It is a rubber stamp.
6.7 Agent Failure Modes: The Complete Taxonomy¶
Understanding how agents fail is more valuable than understanding how they succeed. These failure modes recur across agent implementations regardless of the underlying model or framework.
Failure Mode 1: The Infinite Loop¶
What happens: The agent enters a cycle where it repeatedly attempts the same failing action or oscillates between two states.
Triggers: Tool call fails but error handling tells the agent to retry. The agent reasons it should try a slightly different approach. The slightly different approach also fails. The agent reasons to try again. The pattern repeats.
Detection: Iteration count exceeds expected range. Tool call log shows the same tool called repeatedly with minor variations.
Prevention: Hard iteration limit in code (not prompt). Per-tool call rate limiting — if the same tool is called more than 3 times in one task, trigger human escalation. Exponential backoff on retries.
Failure Mode 2: Context Window Overflow and Silent Truncation¶
What happens: The agent's context window fills. The model's behavior degrades as it loses earlier context. The agent makes decisions based on incomplete information without knowing it has incomplete information.
Why it's insidious: Most model APIs reject a request that exceeds the window. But many agent frameworks and chat harnesses avoid that error by silently dropping the oldest messages to make the request fit. The agent continues operating, unaware that the context of its goal, its prior actions, or the results of earlier steps has been lost. Even with no truncation at all, quality degrades well before the window is full (context rot, §6.4).
Concrete example:
Task: "Review all 50 customer contracts and flag any non-standard clauses."
After processing 30 contracts:
Context window: 60% full
After processing 40 contracts:
Context window: 85% full
The original task goal (earlier in context) begins to be truncated
After processing 45 contracts:
Context window: 95% full
The agent has lost the definition of "non-standard clause"
It begins making inconsistent flagging decisions
After processing 47 contracts:
The original tool call results (contracts 1-20) are truncated
The agent cannot cross-reference earlier contracts for consistency
Final output: Inconsistent and incomplete review.
Nobody knows because the agent returned a complete-looking result.
Prevention: Monitor context utilization throughout the agent loop. When context utilization exceeds 70%, trigger a summarization step — compress prior results into a structured summary and drop the raw content. For long-running tasks, design episodic checkpointing: complete partial work, store results externally, start a fresh context for the next batch with a structured summary of prior work. Never let a framework truncate silently; make overflow an explicit, logged event. §6.4 covers compaction, clearing, and the context budget in detail.
Failure Mode 3: The Irreversible Action Without Confidence¶
What happens: The agent takes an irreversible action (sends email, executes transaction, deletes records) at a moment when its confidence is actually lower than the system's threshold — because confidence was not checked before the action, or because the confidence check was in the prompt rather than enforced in code.
Concrete scenario:
Agent workflow: Generate personalized financial recommendations → Send to customers
Confidence check in prompt: "Only proceed to sending if you are highly confident."
What actually happens:
The agent generates recommendations. The recommendations look good.
The agent is 0.73 confident (threshold was 0.85).
But the prompt says "highly confident" and the agent interprets
its 0.73 as "pretty confident" in natural language.
The agent sends the emails.
247 customers receive recommendations that should have triggered
human review.
Prevention: Confidence evaluation is code, not prompt. The action execution path requires a structured confidence score (a number) that must exceed a hardcoded threshold before the action tool is made available to the agent. The agent cannot call send_email unless confidence_score >= 0.85 has been evaluated and recorded. The evaluation happens in the orchestration layer, not in the LLM's self-assessment.
Failure Mode 4: Tool Authorization Creep¶
What happens: Over time, tools are added to the agent's manifest "temporarily" for specific tasks and never removed. The agent accumulates an ever-growing set of available tools. The principle of least privilege erodes completely.
Why it happens: A developer needs the agent to do something new. They add the necessary tool to the global tool manifest. It works. They move on. Nobody removes it. Six months later, the agent has access to 40 tools with no review of which it actually needs.
Consequence: An agent with 40 tools has a much larger attack surface for both injection-driven misuse and accidental misuse (the LLM reasoning its way to calling a tool it shouldn't in an unexpected context).
Prevention: Tool manifests are versioned and reviewed quarterly. Each tool in the manifest requires a documented justification: "This tool is included because [specific task requires it]." Tools with no documented justification are removed. Phase-based tool loading (from the integration patterns reference) is the technical enforcement — the agent only receives the tools appropriate for its current phase. Tool search with deferred loading (§6.4) fixes the context cost of a large catalog, but not the privilege problem: a deferred tool is still callable.
Failure Mode 5: Non-Idempotent Retry¶
What happens: A step in the agent workflow fails partway through. The orchestration layer retries the entire workflow from the beginning. The step that succeeded in the first run now executes again, creating a duplicate action.
Concrete example:
Workflow: [Create order] → [Send confirmation email] → [Update inventory]
Run 1: Create order (success) → Send email (success) → Update inventory (FAIL)
Retry: Create order (DUPLICATE ORDER CREATED) → ...
The customer now has two orders. The inventory may go negative.
Prevention: Every step that has side effects must be idempotent. Idempotency is achieved through idempotency keys — a unique identifier for each workflow execution that downstream systems use to detect and reject duplicate calls.
Idempotency key pattern:
workflow_id = hash(goal_input + session_id + timestamp_rounded_to_minute)
Every tool call passes the workflow_id as an idempotency key.
The create_order tool:
If an order with this workflow_id exists: return the existing order
If not: create new order with this workflow_id
On retry, create_order detects the existing order and returns it.
The retry continues from the failed step (inventory update),
not from the beginning.
Failure Mode 6: Cost Explosion¶
What happens: An agent task that was designed for a simple use case encounters a complex input. The agent spawns more iterations, makes more tool calls, and accumulates more context than expected. The cost per task spikes by 10x–100x. At scale, this produces a bill shock at the end of the month.
Concrete example:
Normal task: Process a 5-page contract. Expected: 8 LLM calls, $0.04.
Pathological task: Process a 150-page contract with 300 clauses.
The agent makes 156 LLM calls, $1.87 per task.
If 1,000 contracts/day, 1% are long: $18.70/day unexpected cost.
Scale to 10,000 contracts/day with 5% long: $935/day in overages.
Prevention: Per-task cost budget enforced in code. When the running cost of a task exceeds the budget ceiling (MAX_COST_USD = 0.50), the task is suspended and routed to a human for decision: continue with cost override, break the task into smaller pieces, or handle manually. The cost ceiling is configuration, reviewed and adjusted based on actual task distribution.
Failure Mode 7: Instruction Drift Across Long Sessions¶
What happens: Over a long multi-turn agent session, the cumulative context dilutes the original system prompt instructions. The agent's behavior gradually drifts away from its defined role and constraints.
What causes it: The original system prompt is at the beginning of the context window. As the session grows, the ratio of system prompt to total context shrinks. The model's behavior is influenced by the full context, not just the system prompt. Instruction content that was prominent at turn 1 has less relative influence at turn 50.
Prevention: Periodic instruction reinforcement — at every N turns, re-inject the core constraint summary into the context. Not the full system prompt (cost), but the key constraints: scope boundaries, prohibited actions, escalation rules. Tools like LangGraph support this through checkpointing that resets context while preserving task state. Better still, keep invariants in a fixed, cached system prompt or task-state block that compaction never touches (§6.4). On models that sign reasoning blocks, add reinforcement as a new message, not by rewriting the system prompt.
6.8 Agent Observability: What You Must See¶
Standard application monitoring (latency, error rate, throughput) is necessary but radically insufficient for agentic systems. You need visibility into the agent's reasoning and behavior, not just its API calls.
The Agent Observability Stack¶
WHAT TO INSTRUMENT IN EVERY AGENT TASK
Task-level metrics:
├── task_id, workflow_id, task_type
├── start_time, end_time, total_duration_ms
├── total_llm_calls (planned vs. actual — divergence = sign of problems)
├── total_tool_calls (by tool name)
├── total_tokens (input + output + reasoning)
├── total_cost_usd
├── final_state: SUCCESS | HUMAN_ESCALATED | ERROR | TIMEOUT | COST_EXCEEDED
└── human_intervention_count
Per-LLM-call metrics:
├── call_id, task_id, call_sequence_number
├── model, model_version
├── tokens_input, tokens_output, tokens_reasoning
├── tokens_cache_read, tokens_cache_write (cache hit rate)
├── context_utilization (input tokens / design ceiling)
├── context_edits (compaction or clearing applied, tokens removed)
├── latency_ms
├── finish_reason: stop | max_tokens | tool_call | error
└── tool_called (if applicable)
Per-tool-call metrics:
├── tool_name, tool_version
├── call_parameters (sanitized — no PII)
├── success: true | false
├── latency_ms
├── error_type (if failed)
└── retry_count
Reasoning trace:
├── Agent's stated reasoning at each step (the "think" in ReAct)
├── Stored as structured text associated with each LLM call
├── Critical for debugging — shows WHY the agent made each decision
└── Stored separately from audit log (may contain sensitive reasoning)
The Alerts That Matter¶
AGENT HEALTH ALERTS
Cost:
WARN: Any single task exceeds 2x expected cost
ALERT: Any single task exceeds 5x expected cost
ALERT: Daily cost for agent workflow exceeds weekly average by >30%
Behavior:
ALERT: Any task exceeds MAX_ITERATIONS threshold
WARN: P95 iteration count for a workflow type increases >20% WoW
ALERT: Human escalation rate exceeds 15% for any workflow type
(indicates AI quality problem, not human capacity problem)
WARN: Human escalation rate drops to near zero
(may indicate escalation conditions are too lenient)
Reliability:
ALERT: Task failure rate (non-human-escalation) exceeds 5%
ALERT: Tool call failure rate for any tool exceeds 10%
ALERT: Context window utilization P90 exceeds 80% for any workflow
(measure against the design ceiling, not the model's window)
WARN: Prompt cache hit rate drops below 50% on a multi-turn workflow
WARN: Same tool called with identical arguments 3+ times in a task
(re-fetch thrash after context clearing)
Quality:
WARN: Reflection/self-critique agent: revision rate exceeds 40%
(indicates the initial generation quality is poor)
ALERT: LLM-as-judge quality score drops below threshold
(requires automated eval pipeline — Module 7)
6.9 The "Should This Be an Agent?" Test¶
The most important architectural decision is often whether to use an agent at all. Agentic architectures add complexity, cost, latency, and risk. They are justified when the benefits outweigh these costs. They are not justified by default.
DECISION FRAMEWORK: AGENT VS. SIMPLER ARCHITECTURE
Question 1: Is the task structure known upfront?
YES, fully predictable steps → Use a deterministic workflow (not an agent)
NO, steps depend on intermediate results → Agent may be appropriate
Question 2: Is there genuine branching logic driven by content?
YES, the content of tool results determines what to do next → Agent
NO, it's if/else logic a developer could write → Deterministic workflow
Question 3: Does the task require interpretation of ambiguous inputs?
YES → LLM-based agent
NO, inputs are structured → Rule engine or traditional code
Question 4: Can the task be broken into a fixed sequence of LLM calls?
YES → Pipeline of LLM calls (not an agent loop)
NO → Agent
Question 5: What is the consequence of agent error?
LOW (reversible, low-stakes) → Agent acceptable with basic controls
HIGH (irreversible, financial, customer-facing) → Require state machine
pattern with human approval gates, not ReAct
HEURISTIC: If a sufficiently senior developer could write the logic as
code (even complex code), the answer is probably not "agent."
Agents are for tasks where the logic cannot be predetermined
because it depends on content interpretation at each step.
The premature agent problem. Teams reach for agents because they are interesting, not because the task requires them. A workflow that sends a daily report, pulls data from three APIs, and formats it as a PDF does not need an agent. A workflow that reads unstructured customer feedback, determines which feedback themes require action, identifies the appropriate team owner for each theme, drafts a summary, and routes it appropriately — that is a legitimate agent use case.
6.10 Production Readiness Checklist for Agents¶
Before any agent system goes to production:
Safety - [ ] Hard iteration limit enforced in code? - [ ] Hard cost budget per task enforced in code? - [ ] All irreversible actions require human approval (code-enforced, not prompt-enforced)? - [ ] Phase-based tool access — agent cannot call irreversible tools in read-only phases? - [ ] Confidence threshold evaluation in code before any consequential action?
Reliability - [ ] All stateful operations are idempotent with idempotency keys? - [ ] Retry logic does not re-execute completed steps? - [ ] Context window utilization monitored and managed? - [ ] Task state persisted at checkpoints (so partial failure can resume, not restart)? - [ ] Graceful degradation defined: what happens when agent cannot complete?
Context Engineering - [ ] Context budget worksheet completed: design ceiling set below the model's window, every slot capped? - [ ] Invariants (goal, scope, prohibited actions) stored outside the compactable region? - [ ] Compaction and clearing triggers defined, tested in the regression suite, and their cost metered (all usage iterations)? - [ ] Irreproducible tool results written to notes or task state before they can be cleared? - [ ] Tool results capped and paginated; large corpora fetched just in time, not pre-loaded? - [ ] Tool search / deferred loading used if the catalog exceeds ~10 tools or ~10K tokens of definitions? - [ ] Cache-aware layout: stable content first, no timestamps in the prefix, messages append-only? - [ ] Sub-agent cost multiplication estimated, with a return schema for each sub-agent?
Harness and Runtime - [ ] Harness option chosen and recorded (hand-written, SDK/framework, packaged harness, managed runtime) with rationale? - [ ] Managed runtime terms checked: data retention (ZDR/BAA eligibility), region, session limits, pricing? - [ ] Canonical task state and provider-neutral history kept in your own store? - [ ] Agent config, tool registry, and skills in version control, deployed to the runtime (not authored only in a console)? - [ ] Skills reviewed and versioned like code? - [ ] Each runtime dependency recorded in the lock-in ledger (Module 37 §37.10)?
Observability - [ ] Task-level metrics instrumented (cost, duration, iteration count)? - [ ] Reasoning trace captured per LLM call? - [ ] Tool call audit log: input parameters, output, latency, success? - [ ] Human intervention events logged with approver and decision? - [ ] Alert thresholds defined for cost, iteration, escalation, failure rate?
Governance - [ ] Agent identity defined (separate from user identity)? - [ ] Agent authorization scope: what can it do, what requires human approval? - [ ] Tool manifest versioned and reviewed? - [ ] Full audit trail: goal → plan → actions → outcomes — for compliance? - [ ] Rollback procedure for partially-executed tasks defined? - [ ] Agent version defined and pinned? An agent version = model version + prompt version + tool manifest version + skill set version. Any change to any component creates a new agent version. Behavioral changes can be traced to specific version transitions. - [ ] Regression test suite defined for this agent? At least 10 representative tasks with expected outcome characteristics, run on every agent version change before deployment.
EXERCISE — Loop Termination Design: Take a ReAct agent you are designing or have access to. Write the exact code (not prompt instructions) that enforces: maximum 10 iterations, maximum $0.50 cost, maximum 60-second wall-clock time. For each limit: what is the handler when the limit is breached? What state is preserved? What is presented to the human for resolution?
PONDER — The Irreversibility Audit: List every action in your current or planned agent's tool manifest. Categorize each as: fully reversible, partially reversible (with effort), irreversible. For each irreversible action: what is the human approval gate? Is it enforced in code or in the prompt? If it is enforced in the prompt — rewrite it as a code-enforced gate.
WORKSHOP — Agent vs. Workflow Decision: Take five automation use cases from your organization. For each: apply the "Should This Be an Agent?" test from Section 6.9. Determine: deterministic workflow, pipeline of LLM calls, ReAct agent, Plan-and-Execute agent, or state machine agent. Justify each decision. For those that should be agents: identify the single most likely failure mode from Section 6.7 and design the specific mitigation.
WORKSHOP — Human-in-the-Loop Design: Design the complete human review interface for an agentic financial report generation system. The agent pulls data from 5 sources, analyzes it, and proposes a quarterly risk summary to be sent to the board. Use the approval request structure from Section 6.6. What does the reviewer see? What can they do? What is the timeout behavior? What is the audit record?
Next: Module 7 — Multi-Agent Systems