THE AI ARCHITECT CERTIFICATION — COURSE HANDBOOK¶
Part 2 of 3 — Acts IV-VI (Modules 9-21)¶
Continues the learning journey. Same loop: read the card, read the module, answer the checks, do the work, revisit the One Question.
---¶
ACT IV — SECURITY & GOVERNANCE¶
Making AI systems safe, compliant, and accountable.
MODULE 9 — AI Security¶
Essence: With LLMs, the input is instructions and the attack is in the text. Defense is architectural, not a filter.
Learning objectives — you can: 1. Map the AI attack surface and OWASP LLM Top 10 to architectural responses 2. Design defense-in-depth that assumes any single layer is bypassed 3. Run and interpret a red-team exercise with Garak 4. Convert findings into regression eval cases
Knowledge checks:
- Prompt injection is best understood as:
- A) A bug that can be patched
- B) A structural property of LLMs with no complete solution — managed via defense-in-depth
- C) Rare and low-risk
-
D) Solved by content filters
Answer: B. Why: roleplay attacks succeed ~90% of the time against frontier models; you design layers and limit blast radius, you don't "fix" it.
-
The best way to limit the damage of a successful injection on an agent is:
- A) A longer system prompt
- B) Least-privilege, phase-scoped tool access so an injected agent can do little
- C) A disclaimer
-
D) A faster model
Answer: B. Why: minimizing available tools shrinks the blast radius regardless of whether injection succeeds.
-
After a red-team finds a working jailbreak, the most durable response is:
- A) Patch the prompt and move on
- B) Add the attack as a regression eval case so it's tested forever
- C) Ignore it if rare
- D) Disable the feature
Answer: B. Why: every security finding should become a permanent test; otherwise the fix silently regresses.
Apply: LAB 04 (red-team an assistant) + Module 9 WORKSHOP. One Question: If a user's input were treated as an instruction, what's the worst it could trigger, and what limits that in code?
MODULE 10 — Shadow AI & Governance¶
Essence: ~45% of employees use AI regularly and ~70% of it is outside IT oversight. You don't choose whether AI is in your org — only whether you govern it.
Learning objectives — you can: 1. Detect shadow AI usage 2. Design an enable-and-govern posture, not a block posture 3. Implement policy-as-code guardrails 4. Assess governance maturity
Knowledge checks:
- A blanket "no AI tools" policy most likely results in:
- A) No AI usage
- B) Shadow AI — employees using consumer tools you can't see, often with sensitive data
- C) Higher productivity
-
D) Better security
Answer: B. Why: blocking drives usage underground; the data exposure gets worse, not better.
-
The goal of an enable-and-govern posture is to:
- A) Make AI hard to use
- B) Make the sanctioned, governed path easier than the shadow path
- C) Approve every tool
-
D) Eliminate all risk
Answer: B. Why: people take the path of least resistance; make the compliant path that path.
-
Policy-as-code (e.g., OPA) is preferable to a written policy because:
- A) It's cheaper
- B) It's enforced automatically rather than relying on people reading a PDF
- C) It's shorter
- D) It's required by law
Answer: B. Why: enforcement-in-code is consistent; a document nobody reads isn't a control.
Apply: Module 10 — design a tool catalog + review process. One Question: Is the sanctioned way to use AI easier than the shadow way, or am I pushing people to work around me?
MODULE 11 — Compliance & Model Risk¶
Essence: In regulated industries, an AI system that can't be inventoried, validated, explained, and monitored isn't production-ready — no matter how good its outputs.
Learning objectives — you can: 1. Apply enduring model-risk principles even where genAI guidance is still forming 2. Build a model inventory and tier by materiality 3. Design explainability for adverse decisions 4. Establish ongoing monitoring with a named owner
Knowledge checks:
- After SR 11-7 was rescinded (April 2026), model risk management:
- A) No longer applies
- B) Was replaced by a principles-based interagency framework; the core principles endure and genAI is explicitly carved out as still-evolving
- C) Only applies to banks
-
D) Became optional
Answer: B. Why: the specific guidance changed but the discipline (inventory, validation, monitoring, ownership) endures; genAI guidance is still being developed.
-
Validation rigor for a model should be:
- A) The same for all models
- B) Proportional to the model's materiality (tier)
- C) Maximal always
-
D) Decided by the vendor
Answer: B. Why: tiering by materiality focuses rigor where the stakes are highest.
-
For an adverse automated decision about a person, you must be able to:
- A) Explain the model's neural weights
- B) Provide a plain-language explanation of the factors that drove the decision
- C) Nothing
- D) Refer them to the vendor
Answer: B. Why: explainability for adverse decisions is a recurring legal/ethical requirement; it's about understandable factors, not model internals.
Apply: Module 11 — build a model inventory entry + tiering. One Question: Could I show an examiner the inventory, validation, monitoring, and owner for this model today?
---¶
ACT V — PRODUCTION & PLATFORMS¶
Running AI reliably and economically at scale.
MODULE 12 — Observability & Evals¶
Essence: If you can't measure AI quality with a number, you can't improve it, catch regressions, or tell a CFO it works.
Learning objectives — you can: 1. Instrument the three observability signals via OpenTelemetry 2. Build an offline eval suite that gates deployment 3. Design and calibrate an LLM-as-judge 4. Detect drift and attribute cost in production
Knowledge checks:
- Traditional APM is insufficient for AI because:
- A) It's too expensive
- B) An AI system can return HTTP 200 while giving wrong answers — failure is content-level, not system-level
- C) It doesn't support Python
-
D) It's too slow
Answer: B. Why: AI failures are silent at the infrastructure layer; you need quality as a first-class signal.
-
Before trusting an LLM-as-judge, you must:
- A) Use the largest model
- B) Calibrate it against human judgment and confirm acceptable agreement
- C) Run it twice
-
D) Lower its temperature
Answer: B. Why: an uncalibrated judge produces unreliable scores; calibration against humans is required.
-
The relationship between production and evals should be:
- A) One-directional (evals then deploy)
- B) A loop — production failures become new eval cases
- C) Unrelated
- D) Evals only run once
Answer: B. Why: every incident should feed the eval suite so it can't recur undetected.
Apply: LAB 01 (build the eval suite) — you've done this one. One Question: What's my faithfulness number, and would my eval suite catch it if this change made the system worse?
MODULE 13 — Cost Engineering¶
Essence: AI cost is an architecture problem, not a procurement problem. The savings come from routing, caching, and context discipline.
Learning objectives — you can: 1. Build a full cost model across its eight components 2. Apply routing, caching, context management 3. Calculate build-vs-host break-even 4. Enforce cost ceilings in code
Knowledge checks:
- Output tokens vs. input tokens cost roughly:
-
A) The same B) Output costs several times more (about 5-6x on closed frontier models) C) Input costs more D) Both are free
Answer: B. Why: the output/input asymmetry drives architecture toward input-heavy, output-light designs. The exact ratio moves with each price change (Appendix G §G.3); the asymmetry has held.
-
The most reliable way to prevent an agent from running up a huge bill is:
-
A) Hope it stops B) A code-enforced per-task cost ceiling C) A note in the prompt D) Checking the invoice monthly
Answer: B. Why: only a hard, code-enforced ceiling bounds runaway loops.
-
Semantic caching reduces cost by:
- A) Compressing prompts
- B) Serving cached responses for semantically similar (not just identical) queries
- C) Using a smaller model
- D) Removing context
Answer: B. Why: it reuses answers across similar queries, cutting redundant inference.
Apply: LAB 05 (the cost model calculator). One Question: What's the most this system could cost if every safeguard failed, and is that bounded in code?
MODULE 14 — AI Infrastructure¶
Essence: Build-vs-host is driven by data sensitivity and volume, not preference. Above a threshold with sensitive data, self-hosting wins on both cost and compliance.
Learning objectives — you can: 1. Choose API / cloud-VPC / on-prem by sensitivity and volume 2. Configure vLLM for production serving 3. Select a quantization approach 4. Design autoscaling for variable load
Knowledge checks:
- The deployment decision is driven primarily by:
-
A) What's trendy B) Data sensitivity and volume C) The team's favorite cloud D) Model popularity
Answer: B. Why: sensitivity sets the residency requirement; volume sets the break-even — together they decide the tier.
-
The production standard for self-hosted open-weight serving is:
-
A) Ollama B) vLLM C) running the model in a notebook D) a single-threaded server
Answer: B. Why: vLLM (PagedAttention, continuous batching) is built for concurrent production serving; Ollama is for local/dev.
-
Quantization (AWQ/GPTQ/FP8) is used to:
- A) Improve accuracy
- B) Reduce memory and increase throughput at some quality cost
- C) Encrypt the model
- D) Speed up training
Answer: B. Why: it trades a little quality for big memory/throughput gains, fitting models to available hardware.
Apply: Module 14 — model a build-vs-host break-even for a real workload. One Question: At my actual volume and data sensitivity, does the math favor API or self-hosting — or am I deciding by habit?
MODULE 15 — AI Coding Assistants¶
Essence: Real 30-55% productivity gains and a ~29% vulnerability rate in generated code are both true. Governance keeps the gain without the breach.
Learning objectives — you can: 1. Quantify the gain and the security risk 2. Design governance with mandatory security review of AI code 3. Match tool configuration to IP/data requirements 4. Assess the agentic coding tools' specific risks
Knowledge checks:
- AI-generated code should be:
- A) Trusted more because the AI is consistent
- B) Reviewed at least as carefully as human code, because of its vulnerability rate
- C) Merged automatically to save time
-
D) Exempt from security scans
Answer: B. Why: ~29% vuln rate means AI code needs equal-or-stricter review, not less.
-
The biggest governance concern with consumer-grade coding assistants is:
- A) Speed B) IP/data exposure with no enterprise agreement C) Cost D) UI quality
Answer: B. Why: without an enterprise data/IP agreement, proprietary code and secrets can leak.
Apply: Module 15 — draft an AI coding governance policy. One Question: Does AI-generated code get reviewed at least as carefully as human code, or are we trusting it more because it looks confident?
MODULE 16 — Integration Patterns & AI-Native Design¶
Essence: AI-augmented adds AI to a workflow; AI-native rebuilds the workflow around AI. The architectures differ fundamentally.
Learning objectives — you can: 1. Distinguish augmented from native and choose correctly 2. Apply strangler fig and sidecar patterns 3. Design confidence/provenance into responses 4. Identify the integration anti-patterns
Knowledge checks:
- An AI-augmented system differs from an AI-native one mainly in that augmented systems:
- A) Are always better
- B) Have a working non-AI fallback; native systems are designed assuming AI with no "old way"
- C) Don't use AI
-
D) Are cheaper
Answer: B. Why: the presence of a non-AI fallback vs. designing around AI from scratch is the defining difference.
-
The strangler fig pattern is used to:
- A) Replace AI with humans
- B) Incrementally add AI to a legacy system with both running in parallel during transition
- C) Delete legacy code at once
- D) Avoid AI entirely
Answer: B. Why: it's the safe incremental-replacement pattern for adding AI to existing systems.
Apply: Module 16 + the Integration Patterns reference (Artifact 2). One Question: Am I adding AI to this workflow or rebuilding it around AI — and does my architecture match that answer?
MODULE 17 — Platform Engineering for AI¶
Essence: Every team solving gateway/prompt-management/evals from scratch is waste. A platform turns shared infrastructure into a paved road.
Learning objectives — you can: 1. Define the AI platform team's scope 2. Design the gateway as the enforcement point for cost, PII, audit 3. Build a golden path 4. Measure the platform as a product
Knowledge checks:
- The gateway should be built first because it is:
-
A) The cheapest B) The single enforcement point for cost limits, PII scanning, and audit C) The easiest D) Required by vendors
Answer: B. Why: centralizing calls through the gateway is what makes cost/PII/audit enforceable consistently.
-
A platform team is succeeding when:
- A) It ships the most features
- B) Product teams adopt it because it's easier than DIY
- C) Usage is mandated
- D) It has the biggest budget
Answer: B. Why: platform-as-product is measured by voluntary adoption and developer experience.
Apply: Module 17 — assess your org against the platform maturity model. One Question: Do teams use the platform because it's easier than building their own, or only because they're told to?
---¶
ACT VI — LANDSCAPE, VALUE & DIRECTION¶
Seeing the market, proving value, and reading where it's going.
MODULE 18 — AI Startup Landscape¶
Essence: The model isn't the moat. Winning AI companies are defended by proprietary data, workflow embedding, distribution, or switching costs — your build-vs-buy signals.
Learning objectives — you can: 1. Identify the four moat types and apply them to build-vs-buy 2. Map a capability to its category and vendors 3. Run a competitive-intelligence assessment 4. Recognize where the SaaSpocalypse shifts the calculus
Knowledge checks:
- For an AI capability you're considering building, the key build-vs-buy question is:
- A) Can we build it?
- B) What's the moat, and could a funded startup out-build us and sell it to everyone?
- C) Is it fun to build?
-
D) Do we have engineers free?
Answer: B. Why: "can we" is almost always yes; the real question is defensibility and whether a vendor's moat makes buying smarter.
-
The four moats in AI applications are proprietary data, workflow embedding, distribution/brand, and:
- A) Model size B) Switching costs C) Funding D) Team size
Answer: B. Why: accumulated user state creates switching costs — the fourth durable moat.
Apply: Module 18 WORKSHOP — competitive scan of your industry. One Question: What's the moat, and could a funded startup out-build me and sell it to everyone?
MODULE 19 — Business Value¶
Essence: 95% of AI pilots deliver no measurable P&L impact — usually because no metric was defined, data wasn't ready, or no one owned the outcome. Not because the AI failed.
Learning objectives — you can: 1. Construct the five-element business case 2. Identify which pilot-failure pattern threatens an initiative 3. Design a pilot that proves economic feasibility 4. Translate AI metrics into CFO/COO/board language
Knowledge checks:
- A pilot judged successful because "users found it helpful" most likely failed to:
- A) Use enough compute
- B) Define and measure a business outcome metric
- C) Use the right model
-
D) Run long enough
Answer: B. Why: "helpful" is a vanity metric; without a business number the pilot can't justify scaling.
-
The highest AI ROI is most often found in:
- A) Customer-facing flagship features
- B) Back-office automation with clear cost baselines
- C) Sales and marketing
-
D) Executive dashboards
Answer: B. Why: MIT's finding — the boring, measurable back-office work has the highest ROI, the opposite of where budgets go.
-
A complete AI business case must include all of these EXCEPT:
- A) Cost baseline B) A named business owner C) A guarantee of zero errors D) Total implementation cost
Answer: C. Why: no AI system guarantees zero errors; the other three are required elements.
Apply: Module 19 WORKSHOP — build a full business case + the CFO conversation. One Question: What's the cost baseline, the target metric, and the name of the person accountable for moving it?
MODULE 20 — Hype vs. Reality¶
Essence: You're the only person in the room whose job is to be right, not interesting. Separate capability (can do once) from reliability (does consistently in production).
Learning objectives — you can: 1. Distinguish capability from reliability 2. Place a capability in reliable / overhyped / research 3. Apply the vendor claim filter 4. Apply the earned-autonomy model
Knowledge checks:
- A vendor demos an agent completing a complex task flawlessly. Before relying on it you most want to know:
- A) The demo's resolution
- B) The production override/escalation rate on real-world cases
- C) The vendor's funding
-
D) The model it uses
Answer: B. Why: demos show capability; the override rate reveals production reliability — the thing that matters.
-
"Capability" vs. "reliability" means:
- A) The same thing
- B) Can-do-once-in-conditions vs. does-consistently-in-production
- C) Speed vs. accuracy
- D) Cost vs. quality
Answer: B. Why: most AI disappointment comes from mistaking a capable demo for reliable production behavior.
Apply: Module 20 WORKSHOP — run the vendor claim filter on a real claim. One Question: Is this reliable in production conditions like mine, or have I only seen it be capable in a demo?
MODULE 21 — Emerging Patterns (18-Month Horizon)¶
Essence: The model is commoditizing; competition moves up to orchestration, knowledge, and evaluation. Build model-agnostic infrastructure now.
Learning objectives — you can: 1. Explain the model-commodity shift and its implications 2. Decide long-context vs. RAG vs. hybrid 3. Sort 18-month patterns into plan / prototype / monitor 4. Design model-agnostic infrastructure
Knowledge checks:
- The claim "long context will kill RAG" is:
- A) True
- B) Misleading — they coevolve; hybrid (full-load small/stable + retrieve large/dynamic) outperforms either alone
- C) Obviously correct
-
D) Irrelevant
Answer: B. Why: long context and RAG are complementary; the production pattern is hybrid, not replacement.
-
As the model commoditizes, an organization's durable AI advantage shifts to:
- A) Having the newest model
- B) Proprietary data, deep workflow integration, and evaluation infrastructure
- C) The biggest GPU cluster
- D) The largest prompt library
Answer: B. Why: when everyone has good models, the moat is the orchestration/knowledge/eval layer above the model.
Apply: Module 21 WORKSHOP — build an 18-month architecture roadmap. One Question: When the model stops being a differentiator in 18 months, what is my organization's actual AI advantage?
Continued in Part 3 — Practical Layer, Cloud, Responsible AI, Advanced Architecture (Modules 22-37) + Assessment + Capstone.