Skip to content

THE AI ARCHITECT CERTIFICATION — COURSE HANDBOOK

Part 2 of 3 — Acts IV-VI (Modules 9-21)

Continues the learning journey. Same loop: read the card, read the module, answer the checks, do the work, revisit the One Question.

---

ACT IV — SECURITY & GOVERNANCE

Making AI systems safe, compliant, and accountable.


MODULE 9 — AI Security

Essence: With LLMs, the input is instructions and the attack is in the text. Defense is architectural, not a filter.

Learning objectives — you can: 1. Map the AI attack surface and OWASP LLM Top 10 to architectural responses 2. Design defense-in-depth that assumes any single layer is bypassed 3. Run and interpret a red-team exercise with Garak 4. Convert findings into regression eval cases

Knowledge checks:

  1. Prompt injection is best understood as:
  2. A) A bug that can be patched
  3. B) A structural property of LLMs with no complete solution — managed via defense-in-depth
  4. C) Rare and low-risk
  5. D) Solved by content filters

    Answer: B. Why: roleplay attacks succeed ~90% of the time against frontier models; you design layers and limit blast radius, you don't "fix" it.

  6. The best way to limit the damage of a successful injection on an agent is:

  7. A) A longer system prompt
  8. B) Least-privilege, phase-scoped tool access so an injected agent can do little
  9. C) A disclaimer
  10. D) A faster model

    Answer: B. Why: minimizing available tools shrinks the blast radius regardless of whether injection succeeds.

  11. After a red-team finds a working jailbreak, the most durable response is:

  12. A) Patch the prompt and move on
  13. B) Add the attack as a regression eval case so it's tested forever
  14. C) Ignore it if rare
  15. D) Disable the feature

    Answer: B. Why: every security finding should become a permanent test; otherwise the fix silently regresses.

Apply: LAB 04 (red-team an assistant) + Module 9 WORKSHOP. One Question: If a user's input were treated as an instruction, what's the worst it could trigger, and what limits that in code?


MODULE 10 — Shadow AI & Governance

Essence: ~45% of employees use AI regularly and ~70% of it is outside IT oversight. You don't choose whether AI is in your org — only whether you govern it.

Learning objectives — you can: 1. Detect shadow AI usage 2. Design an enable-and-govern posture, not a block posture 3. Implement policy-as-code guardrails 4. Assess governance maturity

Knowledge checks:

  1. A blanket "no AI tools" policy most likely results in:
  2. A) No AI usage
  3. B) Shadow AI — employees using consumer tools you can't see, often with sensitive data
  4. C) Higher productivity
  5. D) Better security

    Answer: B. Why: blocking drives usage underground; the data exposure gets worse, not better.

  6. The goal of an enable-and-govern posture is to:

  7. A) Make AI hard to use
  8. B) Make the sanctioned, governed path easier than the shadow path
  9. C) Approve every tool
  10. D) Eliminate all risk

    Answer: B. Why: people take the path of least resistance; make the compliant path that path.

  11. Policy-as-code (e.g., OPA) is preferable to a written policy because:

  12. A) It's cheaper
  13. B) It's enforced automatically rather than relying on people reading a PDF
  14. C) It's shorter
  15. D) It's required by law

    Answer: B. Why: enforcement-in-code is consistent; a document nobody reads isn't a control.

Apply: Module 10 — design a tool catalog + review process. One Question: Is the sanctioned way to use AI easier than the shadow way, or am I pushing people to work around me?


MODULE 11 — Compliance & Model Risk

Essence: In regulated industries, an AI system that can't be inventoried, validated, explained, and monitored isn't production-ready — no matter how good its outputs.

Learning objectives — you can: 1. Apply enduring model-risk principles even where genAI guidance is still forming 2. Build a model inventory and tier by materiality 3. Design explainability for adverse decisions 4. Establish ongoing monitoring with a named owner

Knowledge checks:

  1. After SR 11-7 was rescinded (April 2026), model risk management:
  2. A) No longer applies
  3. B) Was replaced by a principles-based interagency framework; the core principles endure and genAI is explicitly carved out as still-evolving
  4. C) Only applies to banks
  5. D) Became optional

    Answer: B. Why: the specific guidance changed but the discipline (inventory, validation, monitoring, ownership) endures; genAI guidance is still being developed.

  6. Validation rigor for a model should be:

  7. A) The same for all models
  8. B) Proportional to the model's materiality (tier)
  9. C) Maximal always
  10. D) Decided by the vendor

    Answer: B. Why: tiering by materiality focuses rigor where the stakes are highest.

  11. For an adverse automated decision about a person, you must be able to:

  12. A) Explain the model's neural weights
  13. B) Provide a plain-language explanation of the factors that drove the decision
  14. C) Nothing
  15. D) Refer them to the vendor

    Answer: B. Why: explainability for adverse decisions is a recurring legal/ethical requirement; it's about understandable factors, not model internals.

Apply: Module 11 — build a model inventory entry + tiering. One Question: Could I show an examiner the inventory, validation, monitoring, and owner for this model today?

---

ACT V — PRODUCTION & PLATFORMS

Running AI reliably and economically at scale.


MODULE 12 — Observability & Evals

Essence: If you can't measure AI quality with a number, you can't improve it, catch regressions, or tell a CFO it works.

Learning objectives — you can: 1. Instrument the three observability signals via OpenTelemetry 2. Build an offline eval suite that gates deployment 3. Design and calibrate an LLM-as-judge 4. Detect drift and attribute cost in production

Knowledge checks:

  1. Traditional APM is insufficient for AI because:
  2. A) It's too expensive
  3. B) An AI system can return HTTP 200 while giving wrong answers — failure is content-level, not system-level
  4. C) It doesn't support Python
  5. D) It's too slow

    Answer: B. Why: AI failures are silent at the infrastructure layer; you need quality as a first-class signal.

  6. Before trusting an LLM-as-judge, you must:

  7. A) Use the largest model
  8. B) Calibrate it against human judgment and confirm acceptable agreement
  9. C) Run it twice
  10. D) Lower its temperature

    Answer: B. Why: an uncalibrated judge produces unreliable scores; calibration against humans is required.

  11. The relationship between production and evals should be:

  12. A) One-directional (evals then deploy)
  13. B) A loop — production failures become new eval cases
  14. C) Unrelated
  15. D) Evals only run once

    Answer: B. Why: every incident should feed the eval suite so it can't recur undetected.

Apply: LAB 01 (build the eval suite) — you've done this one. One Question: What's my faithfulness number, and would my eval suite catch it if this change made the system worse?


MODULE 13 — Cost Engineering

Essence: AI cost is an architecture problem, not a procurement problem. The savings come from routing, caching, and context discipline.

Learning objectives — you can: 1. Build a full cost model across its eight components 2. Apply routing, caching, context management 3. Calculate build-vs-host break-even 4. Enforce cost ceilings in code

Knowledge checks:

  1. Output tokens vs. input tokens cost roughly:
  2. A) The same B) Output costs several times more (about 5-6x on closed frontier models) C) Input costs more D) Both are free

    Answer: B. Why: the output/input asymmetry drives architecture toward input-heavy, output-light designs. The exact ratio moves with each price change (Appendix G §G.3); the asymmetry has held.

  3. The most reliable way to prevent an agent from running up a huge bill is:

  4. A) Hope it stops B) A code-enforced per-task cost ceiling C) A note in the prompt D) Checking the invoice monthly

    Answer: B. Why: only a hard, code-enforced ceiling bounds runaway loops.

  5. Semantic caching reduces cost by:

  6. A) Compressing prompts
  7. B) Serving cached responses for semantically similar (not just identical) queries
  8. C) Using a smaller model
  9. D) Removing context

    Answer: B. Why: it reuses answers across similar queries, cutting redundant inference.

Apply: LAB 05 (the cost model calculator). One Question: What's the most this system could cost if every safeguard failed, and is that bounded in code?


MODULE 14 — AI Infrastructure

Essence: Build-vs-host is driven by data sensitivity and volume, not preference. Above a threshold with sensitive data, self-hosting wins on both cost and compliance.

Learning objectives — you can: 1. Choose API / cloud-VPC / on-prem by sensitivity and volume 2. Configure vLLM for production serving 3. Select a quantization approach 4. Design autoscaling for variable load

Knowledge checks:

  1. The deployment decision is driven primarily by:
  2. A) What's trendy B) Data sensitivity and volume C) The team's favorite cloud D) Model popularity

    Answer: B. Why: sensitivity sets the residency requirement; volume sets the break-even — together they decide the tier.

  3. The production standard for self-hosted open-weight serving is:

  4. A) Ollama B) vLLM C) running the model in a notebook D) a single-threaded server

    Answer: B. Why: vLLM (PagedAttention, continuous batching) is built for concurrent production serving; Ollama is for local/dev.

  5. Quantization (AWQ/GPTQ/FP8) is used to:

  6. A) Improve accuracy
  7. B) Reduce memory and increase throughput at some quality cost
  8. C) Encrypt the model
  9. D) Speed up training

    Answer: B. Why: it trades a little quality for big memory/throughput gains, fitting models to available hardware.

Apply: Module 14 — model a build-vs-host break-even for a real workload. One Question: At my actual volume and data sensitivity, does the math favor API or self-hosting — or am I deciding by habit?


MODULE 15 — AI Coding Assistants

Essence: Real 30-55% productivity gains and a ~29% vulnerability rate in generated code are both true. Governance keeps the gain without the breach.

Learning objectives — you can: 1. Quantify the gain and the security risk 2. Design governance with mandatory security review of AI code 3. Match tool configuration to IP/data requirements 4. Assess the agentic coding tools' specific risks

Knowledge checks:

  1. AI-generated code should be:
  2. A) Trusted more because the AI is consistent
  3. B) Reviewed at least as carefully as human code, because of its vulnerability rate
  4. C) Merged automatically to save time
  5. D) Exempt from security scans

    Answer: B. Why: ~29% vuln rate means AI code needs equal-or-stricter review, not less.

  6. The biggest governance concern with consumer-grade coding assistants is:

  7. A) Speed B) IP/data exposure with no enterprise agreement C) Cost D) UI quality

    Answer: B. Why: without an enterprise data/IP agreement, proprietary code and secrets can leak.

Apply: Module 15 — draft an AI coding governance policy. One Question: Does AI-generated code get reviewed at least as carefully as human code, or are we trusting it more because it looks confident?


MODULE 16 — Integration Patterns & AI-Native Design

Essence: AI-augmented adds AI to a workflow; AI-native rebuilds the workflow around AI. The architectures differ fundamentally.

Learning objectives — you can: 1. Distinguish augmented from native and choose correctly 2. Apply strangler fig and sidecar patterns 3. Design confidence/provenance into responses 4. Identify the integration anti-patterns

Knowledge checks:

  1. An AI-augmented system differs from an AI-native one mainly in that augmented systems:
  2. A) Are always better
  3. B) Have a working non-AI fallback; native systems are designed assuming AI with no "old way"
  4. C) Don't use AI
  5. D) Are cheaper

    Answer: B. Why: the presence of a non-AI fallback vs. designing around AI from scratch is the defining difference.

  6. The strangler fig pattern is used to:

  7. A) Replace AI with humans
  8. B) Incrementally add AI to a legacy system with both running in parallel during transition
  9. C) Delete legacy code at once
  10. D) Avoid AI entirely

    Answer: B. Why: it's the safe incremental-replacement pattern for adding AI to existing systems.

Apply: Module 16 + the Integration Patterns reference (Artifact 2). One Question: Am I adding AI to this workflow or rebuilding it around AI — and does my architecture match that answer?


MODULE 17 — Platform Engineering for AI

Essence: Every team solving gateway/prompt-management/evals from scratch is waste. A platform turns shared infrastructure into a paved road.

Learning objectives — you can: 1. Define the AI platform team's scope 2. Design the gateway as the enforcement point for cost, PII, audit 3. Build a golden path 4. Measure the platform as a product

Knowledge checks:

  1. The gateway should be built first because it is:
  2. A) The cheapest B) The single enforcement point for cost limits, PII scanning, and audit C) The easiest D) Required by vendors

    Answer: B. Why: centralizing calls through the gateway is what makes cost/PII/audit enforceable consistently.

  3. A platform team is succeeding when:

  4. A) It ships the most features
  5. B) Product teams adopt it because it's easier than DIY
  6. C) Usage is mandated
  7. D) It has the biggest budget

    Answer: B. Why: platform-as-product is measured by voluntary adoption and developer experience.

Apply: Module 17 — assess your org against the platform maturity model. One Question: Do teams use the platform because it's easier than building their own, or only because they're told to?

---

ACT VI — LANDSCAPE, VALUE & DIRECTION

Seeing the market, proving value, and reading where it's going.


MODULE 18 — AI Startup Landscape

Essence: The model isn't the moat. Winning AI companies are defended by proprietary data, workflow embedding, distribution, or switching costs — your build-vs-buy signals.

Learning objectives — you can: 1. Identify the four moat types and apply them to build-vs-buy 2. Map a capability to its category and vendors 3. Run a competitive-intelligence assessment 4. Recognize where the SaaSpocalypse shifts the calculus

Knowledge checks:

  1. For an AI capability you're considering building, the key build-vs-buy question is:
  2. A) Can we build it?
  3. B) What's the moat, and could a funded startup out-build us and sell it to everyone?
  4. C) Is it fun to build?
  5. D) Do we have engineers free?

    Answer: B. Why: "can we" is almost always yes; the real question is defensibility and whether a vendor's moat makes buying smarter.

  6. The four moats in AI applications are proprietary data, workflow embedding, distribution/brand, and:

  7. A) Model size B) Switching costs C) Funding D) Team size

    Answer: B. Why: accumulated user state creates switching costs — the fourth durable moat.

Apply: Module 18 WORKSHOP — competitive scan of your industry. One Question: What's the moat, and could a funded startup out-build me and sell it to everyone?


MODULE 19 — Business Value

Essence: 95% of AI pilots deliver no measurable P&L impact — usually because no metric was defined, data wasn't ready, or no one owned the outcome. Not because the AI failed.

Learning objectives — you can: 1. Construct the five-element business case 2. Identify which pilot-failure pattern threatens an initiative 3. Design a pilot that proves economic feasibility 4. Translate AI metrics into CFO/COO/board language

Knowledge checks:

  1. A pilot judged successful because "users found it helpful" most likely failed to:
  2. A) Use enough compute
  3. B) Define and measure a business outcome metric
  4. C) Use the right model
  5. D) Run long enough

    Answer: B. Why: "helpful" is a vanity metric; without a business number the pilot can't justify scaling.

  6. The highest AI ROI is most often found in:

  7. A) Customer-facing flagship features
  8. B) Back-office automation with clear cost baselines
  9. C) Sales and marketing
  10. D) Executive dashboards

    Answer: B. Why: MIT's finding — the boring, measurable back-office work has the highest ROI, the opposite of where budgets go.

  11. A complete AI business case must include all of these EXCEPT:

  12. A) Cost baseline B) A named business owner C) A guarantee of zero errors D) Total implementation cost

    Answer: C. Why: no AI system guarantees zero errors; the other three are required elements.

Apply: Module 19 WORKSHOP — build a full business case + the CFO conversation. One Question: What's the cost baseline, the target metric, and the name of the person accountable for moving it?


MODULE 20 — Hype vs. Reality

Essence: You're the only person in the room whose job is to be right, not interesting. Separate capability (can do once) from reliability (does consistently in production).

Learning objectives — you can: 1. Distinguish capability from reliability 2. Place a capability in reliable / overhyped / research 3. Apply the vendor claim filter 4. Apply the earned-autonomy model

Knowledge checks:

  1. A vendor demos an agent completing a complex task flawlessly. Before relying on it you most want to know:
  2. A) The demo's resolution
  3. B) The production override/escalation rate on real-world cases
  4. C) The vendor's funding
  5. D) The model it uses

    Answer: B. Why: demos show capability; the override rate reveals production reliability — the thing that matters.

  6. "Capability" vs. "reliability" means:

  7. A) The same thing
  8. B) Can-do-once-in-conditions vs. does-consistently-in-production
  9. C) Speed vs. accuracy
  10. D) Cost vs. quality

    Answer: B. Why: most AI disappointment comes from mistaking a capable demo for reliable production behavior.

Apply: Module 20 WORKSHOP — run the vendor claim filter on a real claim. One Question: Is this reliable in production conditions like mine, or have I only seen it be capable in a demo?


MODULE 21 — Emerging Patterns (18-Month Horizon)

Essence: The model is commoditizing; competition moves up to orchestration, knowledge, and evaluation. Build model-agnostic infrastructure now.

Learning objectives — you can: 1. Explain the model-commodity shift and its implications 2. Decide long-context vs. RAG vs. hybrid 3. Sort 18-month patterns into plan / prototype / monitor 4. Design model-agnostic infrastructure

Knowledge checks:

  1. The claim "long context will kill RAG" is:
  2. A) True
  3. B) Misleading — they coevolve; hybrid (full-load small/stable + retrieve large/dynamic) outperforms either alone
  4. C) Obviously correct
  5. D) Irrelevant

    Answer: B. Why: long context and RAG are complementary; the production pattern is hybrid, not replacement.

  6. As the model commoditizes, an organization's durable AI advantage shifts to:

  7. A) Having the newest model
  8. B) Proprietary data, deep workflow integration, and evaluation infrastructure
  9. C) The biggest GPU cluster
  10. D) The largest prompt library

    Answer: B. Why: when everyone has good models, the moat is the orchestration/knowledge/eval layer above the model.

Apply: Module 21 WORKSHOP — build an 18-month architecture roadmap. One Question: When the model stops being a differentiator in 18 months, what is my organization's actual AI advantage?


Continued in Part 3 — Practical Layer, Cloud, Responsible AI, Advanced Architecture (Modules 22-37) + Assessment + Capstone.