THE AI ARCHITECT CERTIFICATION — COURSE HANDBOOK¶
Part 1 of 3 — Orientation & Acts I-III (Modules 1-8)¶
This handbook is the course. It sequences the learning journey and, for each module, gives you: the learning objectives, the one-line essence, knowledge-check questions with answers, and the hands-on work that proves you can do it. The full module text lives in the module files; this handbook is how you move through them.
HOW THIS COURSE WORKS¶
The Learning Loop (repeat for every module)¶
1. READ the field card (the essence — 1 page) → orient
2. READ the full module → learn
3. ANSWER the knowledge checks → self-test
4. DO the lab or workshop → apply
5. REVISIT the module's One Question → synthesize
You don't truly know a module until you can answer its One Question with specifics and complete its hands-on work. Reading alone is necessary but not sufficient.
The Three Levels¶
- Foundation (L1): Modules 1, 2, 3, 4, 6, 19, 20, 33 — 10-12 hours — for those entering AI architecture
- Practitioner (L2): all 37 modules — 45-55 hours — for working architects
- Expert (L3): live workshops + peer-reviewed submission — for principals and leads
How You're Assessed¶
- Knowledge checks (in this handbook): self-test as you go
- Level exams: 50 questions (L1) / 100 questions (L2), 70% to pass
- Capstone: the full architecture package (see Part 3) — the credibility anchor
- Labs: hands-on, runnable, with worked solutions
Knowledge-Check Answer Key Convention¶
Each question shows the answer and a one-line why immediately after, in this format:
Answer: B. Why: ...
Cover the answer line, commit to a choice, then check. Guessing right for the wrong reason still counts as not knowing it.
---¶
ACT I — FOUNDATIONS¶
The mindset and the building blocks. Everything else assumes these.
MODULE 1 — The AI Architect's Operating Model¶
Essence: Your job is not to know the most about AI — it's to make sound decisions under uncertainty and be the person who asks what happens when it's wrong.
Learning objectives — you can: 1. Distinguish the three altitudes (strategy / system / component) and place a decision at the right one 2. Decide what to own versus delegate 3. Recognize the three traps as they happen 4. Communicate an AI decision to a non-technical stakeholder
Knowledge checks:
- A product team asks you to decide which vector database to use. Before answering, what should you establish first?
- A) The fastest vector database on benchmarks
- B) Whether this decision needs you at all, or whether the team can make it with a decision framework you provide
- C) The cheapest option
-
D) What the last project used
Answer: B. Why: the architect's first move is choosing the right altitude — many component decisions should be delegated with a framework, not made by the architect.
-
An executive says "our competitors are using AI agents, we need them too." This is a decision at which altitude?
-
A) Component B) System C) Strategy D) Implementation
Answer: C. Why: whether/why to adopt a capability is a strategy-altitude question; jumping to "which agent framework" would be drifting to the wrong altitude.
-
Which is the clearest sign an architect has fallen into a trap?
- A) They ask about failure modes
- B) They approve every design that comes to them
- C) They request a business metric
- D) They delegate a component choice
Answer: B. Why: approving everything means the review adds no value — nobody trusts the architect who never says "not yet."
Apply: WORKSHOP from Module 1 — map three current decisions in your org to their correct altitude. One Question: Is this a decision I need to own, or one I need to enable someone else to make well?
MODULE 2 — The Model Ecosystem¶
Essence: The model is no longer the differentiator — it's a commodity input. Design for substitutability, not for a model.
Learning objectives — you can: 1. Categorize any model (frontier / open-weight / reasoning / SLM) and state its sweet spot 2. Design a routing strategy matching task to model tier 3. Assess and mitigate vendor lock-in 4. Apply the selection framework to a constrained use case
Knowledge checks:
- A workflow does high-volume document classification on PII-sensitive internal data. Best model choice?
- A) The top frontier model via public API
- B) A small or open-weight model, likely self-hosted
- C) Whatever has the best benchmark score
-
D) A reasoning model
Answer: B. Why: high volume + sensitive data + a simple task favors a small/open-weight model, self-hosted for residency and cost; frontier via public API is overkill and a data-residency risk.
-
Why pin a model version in production rather than use a floating alias like "latest"?
- A) Pinned versions are cheaper
- B) A silent provider update can change behavior and break your evals without warning
- C) Aliases are deprecated
-
D) Pinning improves latency
Answer: B. Why: model behavior can shift on a silent update; pinning makes behavior reproducible and changes intentional.
-
The strongest hedge against vendor lock-in is:
- A) Signing a long contract B) Coding against capabilities (not models) behind a gateway, with a model profile per model and a fallback you have tested C) Using only one model D) Avoiding open-weight models
Answer: B. Why: a gateway alone normalizes only the wire call; parameters, tool calling, and behavior still differ between models. The capability layer, profiles, and a tested fallback are what make a switch a configuration change (Module 37).
Apply: WORKSHOP — design a routing table for three real task types in your org. One Question: If my primary model doubled in price or got deprecated next month, how much would it cost me to switch?
MODULE 3 — Prompt Engineering as a Discipline¶
Essence: Prompts are production code. Version, test, review, and govern them — or they become ungovernable liabilities.
Learning objectives — you can: 1. Treat prompts as version-controlled, tested code 2. Construct a layered production prompt 3. Design injection-resistant prompts via structural separation 4. Build a prompt regression test and change process
Knowledge checks:
- The single most important reason to put prompts in version control is:
- A) To save disk space
- B) So you can roll back a bad change fast and know exactly what changed
- C) Because Git is free
-
D) To make prompts shorter
Answer: B. Why: production prompts change; without versioning you can't roll back or attribute a regression.
-
The best defense against a user embedding "ignore your instructions" in their input is:
- A) A longer system prompt
- B) Structurally separating user content from instructions (e.g., typed fields / delimiters) so user text is never interpreted as instructions
- C) Asking the model politely to ignore such attempts
-
D) Lowering temperature
Answer: B. Why: structural separation is architectural; relying on the prompt to resist injection is fragile.
-
A prompt change should be promoted to production only after:
- A) The author is confident B) It passes the eval suite C) It's shorter than before D) A week has passed
Answer: B. Why: prompt changes are code changes — the eval suite is the gate, not someone's confidence.
Apply: Module 3 WORKSHOP + LAB 01 (eval suite gates a prompt change). One Question: Can I roll back a prompt change in under five minutes, and would a test catch it if it broke something?
---¶
ACT II — KNOWLEDGE ARCHITECTURE¶
How AI systems access and reason over your information.
MODULE 4 — RAG Architecture¶
Essence: RAG quality is determined by retrieval, not the model. If the right content isn't retrieved, no model can save the answer.
Learning objectives — you can: 1. Identify which of the seven naive-RAG failure modes is occurring 2. Design a production pipeline (hybrid search, re-ranking, confidence gate) 3. Diagnose a RAG failure using the four-step sequence 4. Evaluate RAG quality with faithfulness, relevance, precision
Knowledge checks:
- A RAG system gives a wrong answer. The retrieved chunks don't contain the right info, but the source document does. Root cause?
-
A) Generation failure B) Retrieval failure C) Knowledge base gap D) Prompt failure
Answer: B. Why: the content exists in the KB but wasn't retrieved — that's a retrieval failure (chunking/embedding/filter), not generation or a KB gap.
-
Why is pure vector search insufficient in production?
- A) It's too slow
- B) It can miss exact-match terms (codes, names, IDs) that keyword search catches
- C) It's too expensive
-
D) It requires GPUs
Answer: B. Why: dense search captures semantics but misses exact lexical matches; hybrid (dense + BM25) covers both.
-
The confidence gate's job is to:
- A) Speed up retrieval
- B) Decide when retrieval is too weak to answer, routing to escalation instead of guessing
- C) Reduce cost
- D) Rank chunks
Answer: B. Why: the gate prevents the system from answering on weak retrieval — the safety mechanism most teams skip.
Apply: LAB 02 (debug a broken RAG system) + Module 4 WORKSHOP. One Question: When the system is wrong, do I check whether the right content was retrieved before I touch the prompt?
MODULE 5 — AI Data Architecture¶
Essence: AI runs on unstructured data you've never had to make machine-readable before. The data work is ~80% of the project and always underestimated.
Learning objectives — you can: 1. Distinguish the AI data stack from the traditional stack 2. Design an ingestion pipeline with versioning, lineage, PII handling 3. Assess data readiness before committing to a timeline 4. Prevent PII exposure across the pipeline
Knowledge checks:
- The most common reason AI project timelines slip is:
- A) The model is too slow
- B) Data readiness work (parsing, quality, integration, PII) was underestimated
- C) The prompt needs tuning
-
D) The team is too small
Answer: B. Why: data work is typically 3-5x the estimate and is the usual hidden critical path.
-
User-submitted content should never flow directly into the vector store because:
- A) It's too large
- B) It could carry poisoned/injected content that later influences retrieval (data poisoning)
- C) It's slow to embed
-
D) It's always low quality
Answer: B. Why: unfiltered user content in the KB is a data-poisoning vector (OWASP LLM04).
-
"Data readiness" assessment is done:
- A) After the model is chosen B) Before committing to AI timelines C) During production D) Only if there's a problem
Answer: B. Why: readiness gates the realistic timeline; assessing it late guarantees slippage.
Apply: Module 5 WORKSHOP — run a data readiness assessment. One Question: Is the data this AI needs actually ready — profiled, governed, integrated — or am I assuming it is?
---¶
ACT III — AGENTIC SYSTEMS¶
Systems that take actions, not just produce text.
MODULE 6 — Agentic Systems¶
Essence: An agent is a loop with tools and a stopping condition. Most agent failures are missing stopping conditions and guardrails — architecture, not intelligence.
Learning objectives — you can: 1. Define an agent architecturally (loop + tools + controls + stop) 2. Apply the "should this be an agent?" test 3. Design code-enforced controls (iteration limit, cost ceiling, tool scope, approval) 4. Diagnose and fix an agent stuck in a loop
Knowledge checks:
- The right place to enforce "never delete files outside the project directory" is:
- A) A strong instruction in the system prompt
- B) Code that blocks the action regardless of what the model decides
- C) A polite reminder in the user prompt
-
D) A lower temperature
Answer: B. Why: the model can reason around any prompt instruction; irreversibility must be code-enforced.
-
An agent calls the same tool repeatedly with the same result and never stops. The fix is:
- A) A better model
- B) A code-enforced iteration limit / escalation after N attempts
- C) A longer prompt
-
D) More tools
Answer: B. Why: loops are a missing-stopping-condition problem; the limit must be in code.
-
Before building an agent, the first question is:
- A) Which framework? B) Should this even be an agent, or would a deterministic workflow do? C) Which model? D) How many tools?
Answer: B. Why: many "agent" problems are better solved deterministically; agents add cost and failure surface.
Apply: LAB 03 (build an agent state machine) + Module 6 PONDER. One Question: What's the worst thing this agent can do before a human sees it — and is that limit in code or just in the prompt?
MODULE 7 — Multi-Agent, MCP & A2A¶
Essence: Most "multi-agent" problems are better solved by one well-designed agent. When you do need many, the hard part is communication, state, and emergent failure.
Learning objectives — you can: 1. Decide whether a problem needs multiple agents or one 2. Explain MCP and A2A and when each applies 3. Design a multi-agent system with a deterministic orchestrator 4. Govern the MCP servers agents can reach
Knowledge checks:
- In a multi-agent system, routing and sequencing decisions should be made by:
- A) The most capable LLM
- B) Deterministic code (the orchestrator)
- C) A vote among agents
-
D) The user
Answer: B. Why: non-deterministic orchestration is unpredictable and hard to debug; routing is code.
-
MCP is best described as:
-
A) A model B) A standard for connecting agents to tools/data C) A vector database D) A prompt format
Answer: B. Why: MCP standardizes agent-to-tool connectivity (the "USB-C of AI tools"); A2A is agent-to-agent.
-
The realistic size of a production multi-agent system today is:
- A) 100+ autonomous agents B) 3-7 tightly scoped, supervised agents C) Exactly 2 D) As many as possible
Answer: B. Why: large autonomous agent networks produce undebuggable emergent failure; production systems stay small and supervised.
Apply: Module 7 WORKSHOP — justify (or eliminate) the second agent in a design. One Question: Could one well-designed agent do this — and if not, what specifically requires the second one?
MODULE 8 — Copilot Ecosystem & Skills¶
Essence: M365 Copilot's biggest risk isn't the AI — it's that it surfaces everything the user already had over-broad access to. Oversharing is the threat.
Learning objectives — you can: 1. Explain how Copilot inherits permissions and why oversharing is the core risk 2. Run a pre-deployment oversharing assessment 3. Govern Copilot Studio agents as production apps 4. Build a Copilot threat model
Knowledge checks:
- The primary risk when rolling out M365 Copilot broadly is:
- A) The AI hallucinates
- B) It makes latent over-permissioned access instantly actionable (oversharing)
- C) It's too slow
-
D) It costs too much
Answer: B. Why: Copilot inherits the user's permissions; years of access drift become suddenly exploitable via search/synthesis.
-
The right work to do before broad Copilot rollout is:
-
A) Tune prompts B) Audit and remediate access permissions C) Buy more licenses D) Train users on prompting
Answer: B. Why: the pre-deployment work is permissions remediation, not AI configuration.
-
Citizen-built Copilot Studio agents should be:
- A) Unrestricted B) Governed like production applications C) Banned D) Ignored
Answer: B. Why: they access real data and act in real systems; they need production-grade governance.
Apply: Module 8 — run the oversharing pre-deployment checklist. One Question: Before we turn this on for everyone, do we know what each user's Copilot can reach, and is that what we intend?
Continued in Part 2 — Acts IV-VI (Modules 9-21).