THE AI ARCHITECT'S FIELD CARDS¶
One-Page Essence for All 37 Modules¶
What this is: The distilled cream of each module — the part that survives in memory after the course. Each card is self-contained and follows the same six-section structure: Essence → Core Insight → Key Framework → Decision Rule → Red Flags → The One Question.
How to use: Read the full module to learn it. Keep the field card to remember it. The "One Question" on each card is the diagnostic that tells you whether you've actually internalized the module or just read it.
Currency: Cards contain principles, not prices or model names. Where a card mentions a dated fact (a regulation, a platform name, a market statistic), it was accurate as of October 2026. Current figures live in Appendix G — Current Landscape.
The set: Part 1 (Cards 1-11) · Part 2 (Cards 12-22) · Part 3 (Cards 23-33) · Part 4 (Cards 34-37)
---¶
ACT I — FOUNDATIONS¶
CARD 1 — The AI Architect's Operating Model¶
THE ESSENCE Your job is not to know the most about AI. It is to make sound decisions about AI under uncertainty, and to be the person in the room who asks what happens when it's wrong.
THE CORE INSIGHT AI architecture is different from software architecture because the system is probabilistic, the failure modes are content-level (not crashes), and the technology changes every six weeks. The architect operates at multiple altitudes — strategy, system, component — and the skill is knowing which altitude a given decision belongs to and staying there.
THE KEY FRAMEWORK — The Altitude Map
STRATEGY (why / whether) → business value, build-vs-buy, risk appetite
SYSTEM (what / how) → architecture, data flows, integration, governance
COMPONENT (which / detail) → model, chunking, prompt — delegate-able
The trap: drifting down into component detail when the decision needed strategy.
THE DECISION RULE - Own the decisions that are expensive to reverse; delegate the rest - Match your altitude to the decision — don't solve a strategy problem with a tool choice - Build credibility through judgment shown in hard moments, not through knowing the most
RED FLAGS - You're debating model choice before defining the business problem - You approve everything — nobody trusts the architect who never says "not yet" - You drift into implementation detail and lose the strategic thread
THE ONE QUESTION "Is this a decision I need to own, or one I need to enable someone else to make well?"
---¶
CARD 2 — The Model Ecosystem¶
THE ESSENCE The model is no longer the differentiator. It's a commodity input. Design for substitutability, not for a model.
THE CORE INSIGHT The landscape has fractured into five categories (frontier, open-weight, reasoning, small, and the newer non-generative typed-decision models) with no single winner. Prices drop, new models ship every few weeks, and the right model for a task changes. Reasoning is now mostly a mode of flagship models (an effort or thinking level) rather than a separate family. The architectural move is to abstract the model behind a gateway and route by task, so switching models is configuration, not surgery. Making that true takes more than a gateway (see Card 37).
THE KEY FRAMEWORK — Five Categories + Routing
FRONTIER (closed, best quality) → hardest tasks, non-sensitive data
OPEN-WEIGHT (self-host) → data residency, high volume, control
REASONING (effort/thinking mode) → complex multi-step planning — tune depth per task
SMALL (SLM) → edge, high-volume classification, cost tier
TYPED-DECISION (non-generative) → bounded classify/route/score/gate decisions
(Jev, Laya, Clef — see Module 16 §16.9)
Router sends each task to the cheapest model that meets its quality bar.
THE DECISION RULE - Route by task complexity. Don't use a frontier model for classification - Pin model versions in production; never use floating aliases - Set reasoning depth explicitly per task class, because defaults change between versions - For sensitive data + high volume, evaluate self-hosting open-weight - Abstract every model call through a gateway, and treat today's leaderboard as perishable
RED FLAGS - "Which model" is your central debate (wrong altitude: it's commoditizing) - One model hardcoded across the whole system - Using a Chinese-origin model API for data with sovereignty constraints - Model selection based on a snapshot more than a quarter old
THE ONE QUESTION "If my primary model doubled in price or got deprecated next month, how much would it cost me to switch?"
---¶
CARD 3 — Prompt Engineering as a Discipline¶
THE ESSENCE Prompts are production code. Version them, test them, review them, and govern them — or they become ungovernable liabilities.
THE CORE INSIGHT The hidden risk in most organizations is that prompts powering production systems live in nobody-knows-where, were never reviewed, and change without testing. Treating prompts as software — version control, regression tests, an owner, a change process — is what separates a maintainable AI system from one nobody dares to touch.
THE KEY FRAMEWORK — Prompts as Code
Prompt in version control → eval suite gates the change →
staging validation → production promotion (with approval) → rollback ready
Layered structure: system role | constraints | context | task | output format
THE DECISION RULE - Every production prompt has a version, an owner, and a test suite - Separate user content from instructions structurally (XML delimiters / typed fields) - A prompt change is a code change — it goes through the same rigor
RED FLAGS - Nobody can tell you which prompt version is in production - Prompts contain contradictory or dead instructions accumulated over time - A prompt change ships without running an eval suite
THE ONE QUESTION "Can I roll back a prompt change in under five minutes, and do I have a test that would catch it if it broke something?"
---¶
ACT II — KNOWLEDGE ARCHITECTURE¶
CARD 4 — RAG Architecture¶
THE ESSENCE RAG quality is determined by retrieval, not by the model. If the right content isn't retrieved, no model can save the answer.
THE CORE INSIGHT Naive RAG fails in production in seven predictable ways. The difference between a demo and a production system is the retrieval pipeline — chunking, hybrid search, re-ranking, confidence gating. Most "the AI gave a wrong answer" problems are actually "the right content was never retrieved" problems. When one retrieval isn't enough, retrieval becomes a loop the model controls (agentic RAG, deep research). That buys quality on hard questions at roughly 3–10× the calls, so route only the questions that need it, and cap the loop.
THE KEY FRAMEWORK — The Production RAG Pipeline
Query → hybrid search (dense + BM25) → score fusion → cross-encoder rerank →
confidence gate → PASS: generate with citations / FAIL: escalate
Each stage independently tunable. The confidence gate is the skipped safety net.
Escalation ladder: single-shot → corrective → multi-hop → deep research
(each step: more calls, more latency — route by question difficulty)
THE DECISION RULE - Always hybrid search in production (pure vector misses exact matches) - Add re-ranking when precision matters — highest-ROI quality upgrade - Set the confidence threshold by calibration; route low-confidence to humans - Version-manage the knowledge base; archive superseded docs - Agentic retrieval gets an iteration cap, a budget, a stopping rule, and enforced citations; web content is untrusted input
RED FLAGS - Tuning the prompt to fix what is actually a retrieval problem - No confidence gate — every query gets an answer no matter how weak the match - Quality is a vibe, not a faithfulness number - Every query sent through a deep-research loop "for quality" (cost and latency explode)
THE ONE QUESTION "When the system is wrong — do I check whether the right content was retrieved before I touch the prompt?"
---¶
CARD 5 — AI Data Architecture¶
THE ESSENCE AI runs on unstructured data your organization has never had to make machine-readable before. The data work is 80% of the project and always underestimated.
THE CORE INSIGHT The traditional data stack was built for structured, normalized data. AI needs documents parsed, chunked, embedded, versioned, lineage-tracked, and PII-handled. The organizations that succeed built the data foundation before the AI initiative — not alongside it.
THE KEY FRAMEWORK — The AI Data Pipeline
Source docs → parse (Unstructured.io) → quality check → chunk →
embed → vector store + document registry (version, owner, lineage, PII flags)
Governance wraps every stage; PII handled before content reaches any model.
THE DECISION RULE - Assess data readiness before committing to AI timelines - Every document has a version, an owner, a review date, and a lineage record - Handle PII before content enters the AI pipeline, not after - Never let user-submitted content flow unfiltered into the vector store
RED FLAGS - "We'll build the data pipeline as part of the AI project" (it'll be the blocker) - No document lifecycle — stale docs silently degrade quality - PII handling is an afterthought
THE ONE QUESTION "Is the data this AI needs actually ready — profiled, governed, and integrated — or am I assuming it is?"
---¶
ACT III — AGENTIC SYSTEMS¶
CARD 6 — Agentic Systems¶
THE ESSENCE An agent is a loop with tools and a stopping condition. Most agent failures are missing stopping conditions and missing guardrails — architecture, not intelligence.
THE CORE INSIGHT An agent that can take actions can take wrong actions. The architectural work is almost entirely constraint: which tools, when, with what limits, and what requires a human. Intelligence comes from the model; reliability comes from the constraints around it. The second lever is context: quality drops as context fills ("context rot"), long before the window is full. Treat the context as a budget managed with compaction, clearing stale tool results, just-in-time retrieval, tool search, and sub-agent isolation. Then choose the harness deliberately: your own loop, an SDK, or a managed runtime.
THE KEY FRAMEWORK — Agent Loop + Controls
GOAL → [reason → select tool → act → observe] → repeat → STOP
Controls (all code-enforced): iteration limit, cost ceiling, phase-scoped tools,
human approval for irreversible acts, full logging, idempotency on side effects
Context: stable prefix first, volatile last · design ceiling < window ·
compact / clear / retrieve-on-demand / tool search / sub-agents
Harness: own loop → SDK → managed runtime (more control ◄──► less ops)
THE DECISION RULE - Ask "should this even be an agent?" before building one - Minimum tools per phase, not all tools always - Reversible acts can be autonomous; irreversible acts need human approval — in code - Iteration limit and cost ceiling set before first run - Budget the context window like a resource; set a design ceiling below the advertised window - Keep task state and history in your own canonical form, even on a managed runtime
RED FLAGS - "Be careful" in the prompt instead of constraints in the code - Irreversible actions (delete/pay/send) with no approval gate - No iteration limit; 10+ tools in one phase - Full history and every tool schema sent on every call ("it fits in 1M tokens") - A managed runtime adopted with no record of what leaving it would cost
THE ONE QUESTION "What is the worst thing this agent can do before a human sees it — and is that limit in code or just in the prompt?"
---¶
CARD 7 — Multi-Agent, MCP & A2A¶
THE ESSENCE Most "multi-agent" problems are better solved by one well-designed agent. When you do need many, the hard part is communication, state, and emergent failure — not the agents.
THE CORE INSIGHT Multi-agent systems multiply capability and complexity together. MCP standardizes agent-to-tool connection; A2A standardizes agent-to-agent delegation. But more agents means emergent behavior, trust boundaries, and cascading failures. Production multi-agent systems are small (3-7 agents), tightly scoped, and human-supervised.
THE KEY FRAMEWORK — The Protocol Stack
MCP: agent ↔ tools/data (the "USB-C of AI tools")
A2A: agent ↔ agent (delegation across boundaries)
Orchestrator = deterministic CODE, not an LLM, for routing and sequencing
THE DECISION RULE - Default to single-agent; justify every additional agent - Orchestration logic is code, not model judgment - Scope each agent narrowly with its own tool manifest - Govern MCP servers like any other production dependency
RED FLAGS - An LLM is making your routing/orchestration decisions - "Deploy 100 autonomous agents" — emergent failure you can't debug - No registry or governance for the MCP servers agents can reach
THE ONE QUESTION "Could one well-designed agent do this — and if not, what specifically requires the second one?"
---¶
CARD 8 — Copilot Ecosystem & Skills¶
THE ESSENCE Microsoft 365 Copilot's biggest risk isn't the AI — it's that it surfaces every document the user already had over-broad access to. Oversharing is the threat.
THE CORE INSIGHT Copilot inherits the user's permissions. In organizations where access control drifted for years, Copilot suddenly makes that latent over-access actionable — an employee can now ask for and instantly receive sensitive content they technically could always reach but never would have found. The pre-deployment work is permissions remediation, not AI configuration.
THE KEY FRAMEWORK — The Oversharing Problem
User's effective access (often far broader than intended)
↓ Copilot makes all of it instantly searchable & synthesizable
Pre-deployment: audit & remediate access BEFORE enabling Copilot broadly
THE DECISION RULE - Run an access/oversharing audit before broad Copilot rollout - Govern Copilot Studio agents like production applications - Treat connectors and plugins as new attack surface requiring review
RED FLAGS - Rolling out Copilot before remediating SharePoint/file permissions - Citizen-built Copilot Studio agents with no governance - No inventory of what data Copilot can actually reach
THE ONE QUESTION "Before we turn this on for everyone — do we actually know what each user's Copilot can reach, and is that what we intend?"
---¶
ACT IV — SECURITY & GOVERNANCE¶
CARD 9 — AI Security¶
THE ESSENCE Traditional security assumes the attacker is outside and the input is data. With LLMs, the input is instructions and the attack is in the text. Defense is architectural, not a filter.
THE CORE INSIGHT Prompt injection has no complete solution. In one 2025 study, roleplay-based attacks succeeded about 90% of the time. You cannot patch your way to safety; you design defense-in-depth: structural input/instruction separation, least-privilege tool access, output monitoring, and the assumption that any single layer will be bypassed. Agents widen the surface: tool descriptions, MCP metadata, retrieved pages, and agent memory are all places to hide instructions. The OWASP 2026 list moved Excessive Agency to #3. Use the "lethal trifecta" as a review test: private data + untrusted content + a way to send data out, all in one agent, is an exfiltration path.
THE KEY FRAMEWORK — Defense in Depth
Input filtering → structural instruction/content separation →
least-privilege tool scope → output validation → monitoring
No single layer is trusted. Assume injection will succeed somewhere.
Agent layer: pinned + hashed MCP tool manifests · per-agent identity with
scoped, time-bound delegated tokens · memory writes with provenance ·
sandboxed execution with egress allow-lists · spend mandates for payments
THE DECISION RULE - Treat all retrieved/user content as untrusted instructions - Limit blast radius with phase-scoped tools (an injected agent can do less) - Red team with Garak/PyRIT before production; add every finding to the eval suite - Never put secrets in system prompts — assume they leak - Break the lethal trifecta: no single agent gets private data, untrusted input, and an outbound channel without a human gate - Agents act under their own scoped identity, never a human's full session
RED FLAGS - Relying on a single input filter to stop injection - System prompt contains secrets or internal URLs - No red teaming before a customer-facing launch - MCP servers whose tool definitions can change after approval (rug pull) - Agent memory that untrusted content can write to
THE ONE QUESTION "If a user's input were treated by the model as an instruction, what's the worst it could trigger — and what limits that in code?"
---¶
CARD 10 — Shadow AI & Governance¶
THE ESSENCE ~45% of employees use AI regularly and ~70% of it is outside IT oversight. You don't have a choice about whether AI is in your org — only about whether you govern it.
THE CORE INSIGHT Blocking AI drives it underground; employees paste sensitive data into consumer tools you can't see. The winning posture is enable-and-govern: provide sanctioned tools good enough that nobody needs shadow ones, with policy-as-code guardrails rather than blanket bans. The 2026 version of the problem is agents: citizen-built agents multiply, owners leave, and orphaned agents keep their credentials. Govern the fleet with a registry (owner, tools, data, identity, autonomy level, budget, last review) and a lifecycle from proposal to retirement.
THE KEY FRAMEWORK — The Governance Spectrum
BLOCK ──────────────────────────────────────────── ENABLE
(drives shadow AI) (sanctioned tools + policy-as-code guardrails)
Goal: make the governed path easier than the ungoverned one.
Agent fleet: REGISTRY → autonomy level (L0–L3) → recertify on a cadence →
orphan check when owners leave → retire + revoke credentials
THE DECISION RULE - Detect current shadow AI before writing policy (know the real usage) - Provide a sanctioned tool catalog with a fast review process - Enforce policy as code (OPA), not as PDF nobody reads - Make the compliant path the path of least resistance - Every agent has a named owner, a registry entry, and an expiry or recertification date
RED FLAGS - Your AI policy is "don't use AI" (guarantees shadow usage) - No sanctioned alternative to the consumer tools people want - You don't actually know what AI tools are in use - You can't say how many agents are running, who owns them, or what they can reach
THE ONE QUESTION "Is the sanctioned, governed way to use AI easier than the shadow way — or am I pushing people to work around me?"
---¶
CARD 11 — Compliance & Model Risk¶
THE ESSENCE In regulated industries, an AI system that can't be inventoried, validated, explained, and monitored isn't production-ready — no matter how good its outputs.
THE CORE INSIGHT Model risk management didn't disappear when SR 11-7 was superseded (April 2026). It was replaced by principles-based interagency guidance (Federal Reserve SR 26-2, OCC Bulletin 2026-13, FDIC FIL-15-2026). The new guidance is aimed at banks over $30B in assets and is non-enforceable. It explicitly leaves generative and agentic AI out of scope as "novel and rapidly evolving," pending a planned request for information. For the EU, the AI Act's high-risk deadline for Annex III systems moved to December 2, 2027, but transparency duties still start August 2, 2026. ISO/IEC 42001 gives a certifiable AI management system, which helps but does not by itself prove AI Act conformity. The principles endure: inventory every model, validate proportional to materiality, explain consequential decisions, monitor continuously, and assign a named risk owner.
THE KEY FRAMEWORK — The MRM Pillars
INVENTORY → VALIDATION (proportional to materiality) →
EXPLAINABILITY (for adverse decisions) → ONGOING MONITORING → NAMED OWNER
Tier by materiality: high-stakes models get the most rigor.
Management system: ISO/IEC 42001 (certifiable) + NIST AI RMF; map to EU AI Act dates
THE DECISION RULE - Apply enduring MRM principles to AI even where genAI guidance is still forming. Out of scope does not mean unexamined - Tier models by materiality; match validation rigor to the tier - Adverse decisions need a plain-language explanation capability - Every regulated model has a named risk owner - Use ISO/IEC 42001 as the management-system backbone, but don't claim it equals AI Act compliance
RED FLAGS - Governance framework still built only around superseded SR 11-7 specifics - "GenAI is out of scope of SR 26-2, so it needs no model risk management" - No model inventory — you can't list your AI systems and their owners - A consequential AI decision with no explanation capability
THE ONE QUESTION "Could I show an examiner the inventory, the validation, the monitoring, and the owner for this model today?"
Continued in Part 2 (Cards 12-22), Part 3 (Cards 23-33), and Part 4 (Cards 34-37).