Skip to content

THE AI ARCHITECT CERTIFICATION — COURSE HANDBOOK

Part 3 of 3 — Practical Layer, Cloud, Responsible AI, Advanced Architecture (Modules 22-37) + Assessment + Capstone

---

THE PRACTICAL LAYER

What an architect actually does day-to-day. This is where most courses stop and this one keeps going.


MODULE 22 — The Practicum

Essence: Knowing the architecture isn't the skill — building it is. Eval suites, prompt iteration, RAG debugging, red-team interpretation, agent state machines, MCP servers.

Objectives: build a 20-case eval suite + judge; iterate a prompt one-failure-at-a-time; diagnose RAG failures in sequence; build a LangGraph state machine and an MCP server.

Knowledge checks: 1. When iterating a prompt, you should change: A) as much as possible at once B) one thing per iteration so you know what worked C) only the temperature D) nothing. > Why: changing one variable at a time isolates cause; batched changes hide which fix helped or broke things. 2. The first step in diagnosing a RAG failure is: A) tune the prompt B) check whether the right content even exists in the KB C) change the model D) re-embed everything. > Why: the four-step sequence starts at "does the content exist," before retrieval/generation.

Apply: LABS 01-04. One Question: Could I sit down right now and actually build this, or do I only understand it well enough to talk about it?


MODULE 23 — The Practice of AI Architecture

Essence: The role is design reviews, PoC scoping, timeline estimation, and the first 90 days — not just system design.

Objectives: run a 90-minute design review and classify findings; scope a PoC (one question + decision gate); estimate including invisible phases; execute a first-90-days plan.

Knowledge checks: 1. A PoC's defining feature is: A) it becomes the production system B) it answers one technical question and is then decommissioned C) it has no deadline D) it includes full production infrastructure. > Why: a PoC validates one question; letting it become production by default is the classic trap. 2. In a design review, you should ask about failure modes: A) before the happy path B) only if time permits C) never D) after deployment. > Why: "what happens when it's wrong" surfaces more risk than walking the happy path.

Apply: Module 23 design-review simulation. One Question: Does my design review actually change designs, or does everyone nod and ship what they already built?


MODULE 24 — AI Production Operations

Essence: AI incidents return HTTP 200 while being wrong. You need quality detection and AI-specific rollback before you need them.

Objectives: run the 30-minute incident playbook; classify by severity; introduce AI via shadow + canary; execute prompt/model/KB rollback.

Knowledge checks: 1. The first diagnostic question in an AI incident is: A) which model? B) what changed in the last 24 hours? C) who's on call? D) what's the budget? > Why: most incidents trace to a recent change (prompt, model, KB, traffic); checking that first is fastest. 2. AI routing should be controlled by a feature flag so that: A) it's cheaper B) rollback is a flag change in seconds, not a deployment C) it's easier to code D) users can choose. > Why: flag-controlled routing makes rollback near-instant and reversible.

Apply: Module 24 incident tabletop + build a runbook. One Question: If the AI started giving wrong answers right now, how would I know, and how fast could I roll back?


Essence: Your job isn't to be the lawyer — it's to tell the lawyer which AI-specific clauses matter. Most AI contracts don't protect the buyer.

Objectives: brief legal on the eight clauses; run a weighted vendor scorecard; conduct a claim-testing demo; complete due diligence.

Knowledge checks: 1. The contract clause most often missing that exposes your data is: A) payment terms B) prohibition on using your data to train the vendor's models C) jurisdiction D) renewal terms. > Why: consumer-style "we may use your content to improve services" language means your data trains their model. 2. In a vendor demo you should: A) watch their prepared scenarios B) supply your own scenarios including failure cases C) trust the accuracy claims D) skip the hard questions. > Why: prepared demos are theater; your scenarios (and failure cases) reveal real behavior.

Apply: Module 25 — build a vendor scorecard + run due diligence. One Question: Does this contract say our data won't train their models, and can we leave with our data in 12 months?


MODULE 26 — AI Requirements Engineering

Essence: "As a user I want X" says nothing about quality, failure, or oversight — the four things that matter most for AI.

Objectives: write an extended AI user story; define AI-specific NFRs; apply the AI definition of done; map a workflow including the failure flow.

Knowledge checks: 1. An AI user story must add, beyond the standard format: A) a budget B) quality acceptance criteria, failure-mode behavior, human-oversight, observability, out-of-scope C) a deadline D) a model name. > Why: these are the AI-specific dimensions that make behavior predictable; standard stories omit them. 2. The AI definition of done adds, beyond standard DoD: A) nothing B) eval suite passing, model pinned, observability + alerts configured, failure path tested C) just more documentation D) a demo. > Why: "works in a demo" isn't done; the AI-specific gates are what make it production-ready.

Apply: Module 26 — rewrite a real user story in the extended format. One Question: Does this requirement say what "good enough" means as a number, and what happens when the AI can't meet it?


MODULE 27 — Organizational Patterns

Essence: Most AI initiatives fail for organizational reasons, not technical ones. Fix the org pattern or the architecture won't matter.

Objectives: identify the four success conditions; recognize and intervene in the ten anti-patterns; design a CoE; assign accountable ownership.

Knowledge checks: 1. The single strongest predictor of AI initiative success is: A) the best model B) a named business owner accountable for the outcome metric C) the biggest budget D) the newest tools. > Why: ownership of the outcome (not IT delivery) is the recurring differentiator between the 5% and the 95%. 2. A separate "innovation lab" producing demos nobody adopts is: A) best practice B) the Innovation Lab anti-pattern C) required D) low-risk. > Why: isolation from production and real workflows is why lab output rarely gets adopted.

Apply: Module 27 — anti-pattern audit of past initiatives. One Question: Who loses their bonus if this doesn't move its business metric — and if nobody, why are we starting?


MODULE 28 — AI Technical Debt

Essence: AI systems accumulate six kinds of invisible debt that stay hidden until a failure, audit, or deprecation makes them suddenly expensive.

Objectives: identify the six debt categories; run a quarterly assessment; prioritize by risk × effort; communicate debt to leadership as risk.

Knowledge checks: 1. The AI debt category with an external deadline you don't control is: A) prompt debt B) model debt (provider deprecation) C) eval debt D) data-pipeline debt. > Why: the provider sets the deprecation date; unplanned, it forces emergency migration at 3x cost. 2. AI technical debt should be communicated to leadership as: A) cleanup work B) risk (with probability and cost) C) a nice-to-have D) an engineering preference. > Why: leadership funds risk reduction, not "refactoring"; framing as risk + cost enables a decision.

Apply: Module 28 — run the debt assessment on a real system. One Question: Which of my systems would fail an audit, a deprecation, or a quality review today — and am I tracking it as risk?


MODULE 29 — AI Architecture Communication

Essence: A decision you can't communicate won't be implemented consistently. AI diagrams must show human touchpoints, data crossings, and failure paths — not an "AI box."

Objectives: diagram with the AI-extended C4 model; diagram an agent's phase-tool boundary; write for four audiences; avoid the diagram anti-patterns.

Knowledge checks: 1. The "AI box" anti-pattern is: A) using color B) representing the AI as a single opaque box, hiding retrieval, gates, escalation, and data flows C) too much detail D) using C4. > Why: an opaque box communicates nothing about how the AI works or fails. 2. The same AI system should be documented for executives as: A) the full ADR set B) a one-pager: what it does, what can go wrong, what it costs, who's responsible C) the source code D) a compliance narrative. > Why: executives need decisions and accountability, not architecture diagrams.

Apply: Module 29 — diagram one system + write for three audiences. One Question: Does my diagram show what happens when the AI fails and where the PII goes, or just the happy path through a magic box?


MODULE 30 — Career & Professional Development

Essence: The role rewards judgment under pressure, maintained competence, and a portfolio of documented decisions — not certifications alone. Character, Competence, Consistency, Commitment, Courage.

Objectives: assess yourself against the skill map; build a decision portfolio; answer interview patterns; apply the 5C frame.

Knowledge checks: 1. For an AI architect role, the most credible evidence of ability is: A) a certificate alone B) a portfolio of documented architectural decisions and trade-offs C) years listed on a resume D) a model leaderboard score. > Why: in a field with no dominant cert, documented real decisions demonstrate judgment best. 2. In a system-design interview, a strong candidate starts with: A) the model choice B) the business problem and success metrics C) the framework D) the database. > Why: problem-first signals architectural maturity; technology-first is the common weak pattern.

Apply: Module 30 — skills gap assessment + mock interview. One Question: In the hard moment, am I the architect who holds the standard or the one who caves?

---

CLOUD & PRODUCTION REALITY


MODULE 31 — Cloud Provider AI Platforms

Essence: Match the AI platform to your existing ecosystem, not to a feature comparison. Native infra, portable logic.

Objectives: select a platform by ecosystem fit; map a pattern to AWS/Azure/GCP; decide native vs. vendor-neutral per component; keep logic portable.

Knowledge checks: 1. The primary criterion for choosing a cloud AI platform is: A) the longest feature list B) ecosystem fit (your cloud, identity, productivity stack) C) the cheapest D) the newest. > Why: switching clouds for AI creates fragmentation costlier than any feature gap. 2. To limit lock-in while using a provider's managed agent service, you should: A) avoid managed services B) keep business logic in portable code so it deploys elsewhere with config changes C) use two providers always D) write everything in the provider's syntax. > Why: native infra is fine; the protection is keeping the logic layer portable.

Apply: Module 31 — map one system to AWS/Azure/GCP + the AgentCore-vs-DIY workshop. One Question: If I had to migrate providers in 12 months, would my agent logic move with a config change, or need a rewrite?


MODULE 32 — AI Agent Implementation Learnings

Essence: 78% have agent pilots; 14% scaled them. Most failures are architecture, integration, and context-engineering failures — not model failures.

Objectives: explain why failures are architectural; design tool schemas that drive correct selection; apply context engineering; apply the production anti-pattern checklist.

Knowledge checks: 1. Adding more conversation history to an agent's context often: A) always helps B) makes reasoning worse, as older tokens compete for attention with the relevant ones C) has no effect D) reduces cost. > Why: context is an attention budget; irrelevant history crowds out the critical constraint. 2. Poorly written tool schemas cause: A) faster responses B) wrong tool selection, wasted context, higher latency and cost C) better accuracy D) nothing. > Why: the tool description is the agent's only selection signal; vague schemas produce wrong calls. 3. The only agent-scaling strategy that reliably works is: A) start broad with all tools B) start narrow (few tools, one task), prove it, expand incrementally C) use the biggest model D) add more agents. > Why: broad scope is undebuggable; narrow scope is testable and compounds reliability.

Apply: Module 32 — run the production anti-pattern checklist on a real agent. One Question: Is this agent narrow enough to debug, lean enough to reason, and are its dangerous actions blocked in code?

---

RESPONSIBLE AI


MODULE 33 — Responsible AI in Practice

Essence: Compliance is the floor; responsible AI is the harder territory above it, where the architect's judgment is the only safeguard. Most AI harms are not illegal.

Objectives: apply the "should we build this" screen; identify the four harm types and choose a fairness definition; design transparency/autonomy/environmental responsibility in; conduct a responsible AI review.

Knowledge checks: 1. You cannot simultaneously satisfy all mathematical definitions of fairness, which means fairness is: A) impossible B) a deliberate, documented choice of which definition fits the context C) a calculation D) the vendor's job. > Why: the impossibility result forces an explicit, accountable choice — not a default. 2. A system optimized purely for engagement, with no wellbeing constraint, will tend to: A) stay neutral B) discover manipulation as an effective strategy C) become more honest D) cost less. > Why: an unbounded objective function will exploit cognitive biases; the architecture must bound it. 3. The most important responsible-AI question is asked: A) after launch B) before design — "should we build this at all?" C) only if regulated D) by the legal team. > Why: some systems cause harm by working as intended; the build decision is the first safeguard.

Apply: Module 33 — run a full responsible AI review on a real system. One Question: Who is harmed if this works exactly as intended — and is the benefit worth it?

---

ADVANCED ARCHITECTURE

Added in 2026. Four deeper modules for practitioners. Each one is about twice the length of an average module, so budget about 2–2.5 hours of core study for each (about 10 hours for the four).


MODULE 34 — Fine-Tuning & Model Customization

Essence: Fine-tuning changes how a model behaves, not what it knows. Exhaust prompting and RAG first. Fine-tune only when you have a dataset, a baseline eval, and an exit path.

Objectives: place a performance gap on the customization spectrum (knowledge → RAG; behavior → fine-tuning); choose the technique (LoRA / QLoRA / full) and the training signal (SFT / DPO / RFT) for a stated constraint; specify a dataset and a baseline eval that also checks for regression and catastrophic forgetting; decide between distillation and an efficient-tier API; govern fine-tuned models (inventory, data provenance, registry, rollback, closed-platform lock-in).

Knowledge checks: 1. A team wants to fine-tune a model on internal policy documents so it "knows" company policy. The right response is: A) full fine-tuning B) LoRA on the policy text C) RAG — this is a knowledge gap, not a behavior gap D) a larger model. > Why: knowledge trained into weights goes stale, cannot be cited, and cannot be updated without retraining; RAG solves the knowledge problem. 2. Before any fine-tuning run you must have: A) a GPU cluster B) a baseline eval of the current system, so improvement and regression are measurable C) a signed vendor contract D) one million examples. > Why: without a baseline you cannot tell whether fine-tuning helped, or what it broke. 3. Reinforcement fine-tuning (RFT) is the right training signal when: A) you need a consistent style B) the task is checkable by a programmatic grader, and the base model gets it right sometimes but not reliably C) the model lacks facts D) no eval exists. > Why: RFT learns from verifiable rewards; the grader must be validated against experts before training, or the model learns to game it. 4. A model fine-tuned on a closed provider's platform should be treated as: A) a portable asset B) deliberate lock-in — you cannot export or retrain it elsewhere if the provider changes course C) free to move D) open-weight. > Why: one major provider began winding down its self-serve fine-tuning platform in 2026; fine-tune open-weight models, or record the dependency in the lock-in ledger (Module 37).

Apply: Module 34 WORKSHOP — build a fine-tuning plan (dataset spec, technique, 20-case baseline eval, cost, deployment and rollback). One Question: Is this gap about what the model knows or how it behaves — and have I proved, with an eval, that prompting and RAG cannot close it?


MODULE 35 — Multimodal Architecture

Essence: Multimodal is the baseline in 2026. Every non-text modality must become text or embeddings before it can be retrieved. The architecture question is when that happens, at what resolution, and at what cost.

Objectives: design a document-intelligence pipeline (parser, table strategy, figure captioning, metadata, citations); choose caption embeddings or image embeddings for vision RAG; choose a voice pipeline or speech-to-speech within a latency budget; choose a frame pipeline or native video based on reuse; estimate multimodal cost from each provider's image and audio token rules; secure multimodal input (image-level PII, visual prompt injection, BAAs, residency).

Knowledge checks: 1. A vision RAG system covers 30,000 figures. The cost-correct pattern is: A) call the vision model on every figure for every query B) caption each figure once at ingestion, retrieve captions, and include only the retrieved images at generation time C) ignore figures D) re-caption on every query for freshness. > Why: processing once and caching by image hash turns millions of vision calls into a one-time ingestion cost. 2. For document Q&A ("what was Q3 revenue in this chart?"), image embeddings (CLIP/SigLIP) are: A) the best choice B) weak — they encode visual appearance, not semantic content; use text captions C) required D) equivalent to captions. > Why: two bar charts about different metrics look alike; semantic retrieval needs a description of what the chart says. 3. The token cost of an image is: A) the same on every provider B) different by provider and model (patch-, tile-, or per-image rules that change between generations), so you measure it with the provider's token counter C) always about 1,000 tokens D) free. > Why: the same image can cost several times more on a newer model; image cost is a model-profile fact, not a constant.

Apply: Module 35 WORKSHOP — multimodal RAG design for 10,000 engineering diagrams (ingestion, index, retrieval, eval, cost). One Question: What information lives outside the text layer of my data, and does my pipeline capture it once, at a cost I have actually measured?


MODULE 36 — GraphRAG & Knowledge Graph Architecture

Essence: Vector similarity cannot traverse relationships. GraphRAG augments vector RAG, rather than replacing it, for multi-hop, entity-network, and corpus-wide questions. It costs 5–20x more to index, and it needs a graph that someone owns.

Objectives: recognize the three query types that defeat vector RAG; explain GraphRAG's pipeline and its local and global retrieval modes; design extraction with entity resolution and quality metrics; choose among graph-augmented RAG, graph-first retrieval, and a hybrid router; apply the GraphRAG-vs-vector decision framework and justify the indexing cost; govern the graph (ontology, freshness, PII and erasure, access control, query audit).

Knowledge checks: 1. "Which customers have contracts with our subsidiary in Singapore?" fails on vector RAG because it is: A) too long B) a multi-hop relationship query that requires traversing entity chains, which similarity search cannot do C) ambiguous D) a keyword query. > Why: Customer → Contract → Entity → Subsidiary → Jurisdiction is a traversal, not a similarity match. 2. When entity resolution is unsure whether two mentions are the same entity, you should: A) merge them B) split them and queue them for review C) drop both D) let the LLM decide at query time. > Why: a false merge creates incorrect relationships, which is worse than a missed link. 3. The trigger that justifies investing in GraphRAG is: A) it is the newest pattern B) a documented, recurring, business-consequential failure of vector RAG on entity-relationship queries C) a large corpus alone D) a vendor recommendation. > Why: the ROI is lower failure rates on a query class that vector RAG cannot answer, not novelty.

Apply: Module 36 WORKSHOP — knowledge graph design (ontology, extraction pipeline, retrieval pattern, database selection, freshness plan, GDPR assessment). One Question: Do my users' failing questions ask about relationships between entities, and who will own the graph after it is built?


MODULE 37 — Model Portability: Engineering the Model as a Replaceable Part

Essence: The model is commoditizing; your integration with it is not. Make the model a replaceable part behind a contract you own: a capability interface, a model profile, a behavioral contract, and a fallback you have actually exercised.

Objectives: diagnose a model swap across the ten incompatibility layers and separate loud breaks from silent ones; design the capability layer, model profiles, and adapter, keeping the gateway as transport and policy; write a behavioral contract (invariants, quality tolerance, SLOs, style) with a correctly sized dataset, judged on cost per completed task; design hot, warm, and cold fallbacks with keep-alive traffic and a degradation ladder; run the swap playbook and score portability; decide when deliberate lock-in is right and record it.

Knowledge checks: 1. Your gateway exposes every provider through one API. This means: A) you are portable B) only transport is normalized — parameters, tool calling, reasoning defaults, tokenizers, and behavior still differ, so you need model profiles and contract tests C) prompts transfer unchanged D) fallbacks are tested. > Why: the gateway is a dependency of portability, not the portability layer. 2. A same-vendor upgrade lowers the default reasoning effort from high to medium. This is: A) a loud break B) a silent break — calls succeed while quality drifts; set reasoning depth explicitly in the profile and catch drift with stratified contract evals C) harmless D) a gateway problem. > Why: silent breaks (defaults, tokenizers, behavior) are the ones that reach customers. 3. A fallback model is "hot" when it: A) is listed in a config file B) continuously serves a small share of real traffic, with provisioned quota, monitoring, and a passing contract C) is the cheapest model D) was tested last year. > Why: a fallback you have never sent traffic to is a hypothesis, not a fallback. 4. To decide a Tier A swap near a "−2 points" tolerance, you need: A) 30 examples B) hundreds of cases (about 500–1,000), paired comparison such as McNemar's test, and stratified results C) a public benchmark D) a vendor's word. > Why: at n = 100 the margin of error is about ±7 points, which cannot resolve a 2-point tolerance.

Apply: Module 37 WORKSHOP — build the portability kit for one capability (interface, two model profiles, base prompt + overlays, contract, fallback design, scorecard). One Question: If my primary provider gave ten days' notice, which capabilities would move with a configuration change, and which would become a project?

---

ASSESSMENT

Level 1 Exam (Foundation)

50 questions across Modules 1, 2, 3, 4, 6, 19, 20, 33. 70% to pass. ~60% scenario, ~30% application, ~20% recall. Plus one short case study: given a business problem, propose a high-level architecture and justify three decisions.

Level 2 Exam (Practitioner)

100 questions across all 37 modules + the capstone below. The total stays at 100. Modules 34–37 contribute about 12 items (about 3 each), and the remaining ~88 are spread across Modules 1–33 (about 2–3 per module), with the same scenario / application / recall mix.

Using the Knowledge Checks

The checks in this handbook are your formative practice. The exams draw on the same question types at higher density and difficulty. If you can answer every module's checks and its One Question with specifics, you are ready.


THE CAPSTONE PROJECT

The credibility anchor. You produce a complete architecture package for a provided scenario (e.g., a financial-services customer-facing AI assistant). Deliverables:

  1. Architecture design — context + container diagrams (Module 29), ≥4 ADRs (Appendix D), RAG/agent design with confidence gating and human oversight
  2. Governance & risk — model risk tiering (Module 11), security threat model + red-team plan (Module 9), responsible AI review (Module 33)
  3. Evaluation & operations — eval strategy and gates (Module 12), operational runbook (Module 24), rollback procedures (Module 24)
  4. Business case — five-element case (Module 19), cost model (Module 13)
  5. Delivery plan — phased timeline (Module 23), team capability assessment (Module 22), PoC scope and gate (Module 23)

Scored on a 100-point rubric (see the Course Framework doc). Pass: 70. Distinction: 85. The rubric rewards designing for failure, explicit trade-off reasoning, and realistic delivery thinking.


COMPLETION

To earn the credential you must: - ☐ Complete all module knowledge checks - ☐ Complete the hands-on labs (01-05) - ☐ Pass the level exam (50 for L1, 100 for L2) - ☐ Pass the capstone (L2) - ☐ For L3: original peer-reviewed architecture submission + oral defense


THE JOURNEY IN ONE VIEW

FOUNDATIONS (1-3)        → how to think, the models, prompts as code
KNOWLEDGE (4-5)          → RAG and the data underneath it
AGENTIC (6-8)            → systems that act
SECURITY & GOV (9-11)    → safe, compliant, accountable
PRODUCTION (12-17)       → measure, cost, host, integrate, platform
LANDSCAPE (18-21)        → market, value, hype, horizon
PRACTICE (22-30)         → the actual job: reviews, ops, contracts, org, career
CLOUD (31)               → AWS/Azure/GCP mapping
AGENT REALITY (32)       → what production teaches
RESPONSIBLE AI (33)      → the judgment above the floor
ADVANCED (34-37)         → fine-tuning, multimodal, GraphRAG, model portability
        ↓
LABS (do) + KNOWLEDGE CHECKS (test) + CAPSTONE (prove)
        ↓
CERTIFIED AI ARCHITECT

This handbook is the course. The modules are the depth; the labs are the practice; the capstone is the proof. Work the loop, module by module, and you don't just learn AI architecture — you can do it.