Skip to content

THE AI ARCHITECT CERTIFICATION — COURSE HANDBOOK

Part 1 of 3 — Orientation & Acts I-III (Modules 1-8)

This handbook is the course. It sequences the learning journey and, for each module, gives you: the learning objectives, the one-line essence, knowledge-check questions with answers, and the hands-on work that proves you can do it. The full module text lives in the module files; this handbook is how you move through them.


HOW THIS COURSE WORKS

The Learning Loop (repeat for every module)

1. READ the field card (the essence — 1 page)        → orient
2. READ the full module                              → learn
3. ANSWER the knowledge checks                       → self-test
4. DO the lab or workshop                            → apply
5. REVISIT the module's One Question                 → synthesize

You don't truly know a module until you can answer its One Question with specifics and complete its hands-on work. Reading alone is necessary but not sufficient.

The Three Levels

  • Foundation (L1): Modules 1, 2, 3, 4, 6, 19, 20, 33 — 10-12 hours — for those entering AI architecture
  • Practitioner (L2): all 37 modules — 45-55 hours — for working architects
  • Expert (L3): live workshops + peer-reviewed submission — for principals and leads

How You're Assessed

  • Knowledge checks (in this handbook): self-test as you go
  • Level exams: 50 questions (L1) / 100 questions (L2), 70% to pass
  • Capstone: the full architecture package (see Part 3) — the credibility anchor
  • Labs: hands-on, runnable, with worked solutions

Knowledge-Check Answer Key Convention

Each question shows the answer and a one-line why immediately after, in this format:

Answer: B. Why: ...

Cover the answer line, commit to a choice, then check. Guessing right for the wrong reason still counts as not knowing it.

---

ACT I — FOUNDATIONS

The mindset and the building blocks. Everything else assumes these.


MODULE 1 — The AI Architect's Operating Model

Essence: Your job is not to know the most about AI — it's to make sound decisions under uncertainty and be the person who asks what happens when it's wrong.

Learning objectives — you can: 1. Distinguish the three altitudes (strategy / system / component) and place a decision at the right one 2. Decide what to own versus delegate 3. Recognize the three traps as they happen 4. Communicate an AI decision to a non-technical stakeholder

Knowledge checks:

  1. A product team asks you to decide which vector database to use. Before answering, what should you establish first?
  2. A) The fastest vector database on benchmarks
  3. B) Whether this decision needs you at all, or whether the team can make it with a decision framework you provide
  4. C) The cheapest option
  5. D) What the last project used

    Answer: B. Why: the architect's first move is choosing the right altitude — many component decisions should be delegated with a framework, not made by the architect.

  6. An executive says "our competitors are using AI agents, we need them too." This is a decision at which altitude?

  7. A) Component B) System C) Strategy D) Implementation

    Answer: C. Why: whether/why to adopt a capability is a strategy-altitude question; jumping to "which agent framework" would be drifting to the wrong altitude.

  8. Which is the clearest sign an architect has fallen into a trap?

  9. A) They ask about failure modes
  10. B) They approve every design that comes to them
  11. C) They request a business metric
  12. D) They delegate a component choice

    Answer: B. Why: approving everything means the review adds no value — nobody trusts the architect who never says "not yet."

Apply: WORKSHOP from Module 1 — map three current decisions in your org to their correct altitude. One Question: Is this a decision I need to own, or one I need to enable someone else to make well?


MODULE 2 — The Model Ecosystem

Essence: The model is no longer the differentiator — it's a commodity input. Design for substitutability, not for a model.

Learning objectives — you can: 1. Categorize any model (frontier / open-weight / reasoning / SLM) and state its sweet spot 2. Design a routing strategy matching task to model tier 3. Assess and mitigate vendor lock-in 4. Apply the selection framework to a constrained use case

Knowledge checks:

  1. A workflow does high-volume document classification on PII-sensitive internal data. Best model choice?
  2. A) The top frontier model via public API
  3. B) A small or open-weight model, likely self-hosted
  4. C) Whatever has the best benchmark score
  5. D) A reasoning model

    Answer: B. Why: high volume + sensitive data + a simple task favors a small/open-weight model, self-hosted for residency and cost; frontier via public API is overkill and a data-residency risk.

  6. Why pin a model version in production rather than use a floating alias like "latest"?

  7. A) Pinned versions are cheaper
  8. B) A silent provider update can change behavior and break your evals without warning
  9. C) Aliases are deprecated
  10. D) Pinning improves latency

    Answer: B. Why: model behavior can shift on a silent update; pinning makes behavior reproducible and changes intentional.

  11. The strongest hedge against vendor lock-in is:

  12. A) Signing a long contract B) Coding against capabilities (not models) behind a gateway, with a model profile per model and a fallback you have tested C) Using only one model D) Avoiding open-weight models

    Answer: B. Why: a gateway alone normalizes only the wire call; parameters, tool calling, and behavior still differ between models. The capability layer, profiles, and a tested fallback are what make a switch a configuration change (Module 37).

Apply: WORKSHOP — design a routing table for three real task types in your org. One Question: If my primary model doubled in price or got deprecated next month, how much would it cost me to switch?


MODULE 3 — Prompt Engineering as a Discipline

Essence: Prompts are production code. Version, test, review, and govern them — or they become ungovernable liabilities.

Learning objectives — you can: 1. Treat prompts as version-controlled, tested code 2. Construct a layered production prompt 3. Design injection-resistant prompts via structural separation 4. Build a prompt regression test and change process

Knowledge checks:

  1. The single most important reason to put prompts in version control is:
  2. A) To save disk space
  3. B) So you can roll back a bad change fast and know exactly what changed
  4. C) Because Git is free
  5. D) To make prompts shorter

    Answer: B. Why: production prompts change; without versioning you can't roll back or attribute a regression.

  6. The best defense against a user embedding "ignore your instructions" in their input is:

  7. A) A longer system prompt
  8. B) Structurally separating user content from instructions (e.g., typed fields / delimiters) so user text is never interpreted as instructions
  9. C) Asking the model politely to ignore such attempts
  10. D) Lowering temperature

    Answer: B. Why: structural separation is architectural; relying on the prompt to resist injection is fragile.

  11. A prompt change should be promoted to production only after:

  12. A) The author is confident B) It passes the eval suite C) It's shorter than before D) A week has passed

    Answer: B. Why: prompt changes are code changes — the eval suite is the gate, not someone's confidence.

Apply: Module 3 WORKSHOP + LAB 01 (eval suite gates a prompt change). One Question: Can I roll back a prompt change in under five minutes, and would a test catch it if it broke something?

---

ACT II — KNOWLEDGE ARCHITECTURE

How AI systems access and reason over your information.


MODULE 4 — RAG Architecture

Essence: RAG quality is determined by retrieval, not the model. If the right content isn't retrieved, no model can save the answer.

Learning objectives — you can: 1. Identify which of the seven naive-RAG failure modes is occurring 2. Design a production pipeline (hybrid search, re-ranking, confidence gate) 3. Diagnose a RAG failure using the four-step sequence 4. Evaluate RAG quality with faithfulness, relevance, precision

Knowledge checks:

  1. A RAG system gives a wrong answer. The retrieved chunks don't contain the right info, but the source document does. Root cause?
  2. A) Generation failure B) Retrieval failure C) Knowledge base gap D) Prompt failure

    Answer: B. Why: the content exists in the KB but wasn't retrieved — that's a retrieval failure (chunking/embedding/filter), not generation or a KB gap.

  3. Why is pure vector search insufficient in production?

  4. A) It's too slow
  5. B) It can miss exact-match terms (codes, names, IDs) that keyword search catches
  6. C) It's too expensive
  7. D) It requires GPUs

    Answer: B. Why: dense search captures semantics but misses exact lexical matches; hybrid (dense + BM25) covers both.

  8. The confidence gate's job is to:

  9. A) Speed up retrieval
  10. B) Decide when retrieval is too weak to answer, routing to escalation instead of guessing
  11. C) Reduce cost
  12. D) Rank chunks

    Answer: B. Why: the gate prevents the system from answering on weak retrieval — the safety mechanism most teams skip.

Apply: LAB 02 (debug a broken RAG system) + Module 4 WORKSHOP. One Question: When the system is wrong, do I check whether the right content was retrieved before I touch the prompt?


MODULE 5 — AI Data Architecture

Essence: AI runs on unstructured data you've never had to make machine-readable before. The data work is ~80% of the project and always underestimated.

Learning objectives — you can: 1. Distinguish the AI data stack from the traditional stack 2. Design an ingestion pipeline with versioning, lineage, PII handling 3. Assess data readiness before committing to a timeline 4. Prevent PII exposure across the pipeline

Knowledge checks:

  1. The most common reason AI project timelines slip is:
  2. A) The model is too slow
  3. B) Data readiness work (parsing, quality, integration, PII) was underestimated
  4. C) The prompt needs tuning
  5. D) The team is too small

    Answer: B. Why: data work is typically 3-5x the estimate and is the usual hidden critical path.

  6. User-submitted content should never flow directly into the vector store because:

  7. A) It's too large
  8. B) It could carry poisoned/injected content that later influences retrieval (data poisoning)
  9. C) It's slow to embed
  10. D) It's always low quality

    Answer: B. Why: unfiltered user content in the KB is a data-poisoning vector (OWASP LLM04).

  11. "Data readiness" assessment is done:

  12. A) After the model is chosen B) Before committing to AI timelines C) During production D) Only if there's a problem

    Answer: B. Why: readiness gates the realistic timeline; assessing it late guarantees slippage.

Apply: Module 5 WORKSHOP — run a data readiness assessment. One Question: Is the data this AI needs actually ready — profiled, governed, integrated — or am I assuming it is?

---

ACT III — AGENTIC SYSTEMS

Systems that take actions, not just produce text.


MODULE 6 — Agentic Systems

Essence: An agent is a loop with tools and a stopping condition. Most agent failures are missing stopping conditions and guardrails — architecture, not intelligence.

Learning objectives — you can: 1. Define an agent architecturally (loop + tools + controls + stop) 2. Apply the "should this be an agent?" test 3. Design code-enforced controls (iteration limit, cost ceiling, tool scope, approval) 4. Diagnose and fix an agent stuck in a loop

Knowledge checks:

  1. The right place to enforce "never delete files outside the project directory" is:
  2. A) A strong instruction in the system prompt
  3. B) Code that blocks the action regardless of what the model decides
  4. C) A polite reminder in the user prompt
  5. D) A lower temperature

    Answer: B. Why: the model can reason around any prompt instruction; irreversibility must be code-enforced.

  6. An agent calls the same tool repeatedly with the same result and never stops. The fix is:

  7. A) A better model
  8. B) A code-enforced iteration limit / escalation after N attempts
  9. C) A longer prompt
  10. D) More tools

    Answer: B. Why: loops are a missing-stopping-condition problem; the limit must be in code.

  11. Before building an agent, the first question is:

  12. A) Which framework? B) Should this even be an agent, or would a deterministic workflow do? C) Which model? D) How many tools?

    Answer: B. Why: many "agent" problems are better solved deterministically; agents add cost and failure surface.

Apply: LAB 03 (build an agent state machine) + Module 6 PONDER. One Question: What's the worst thing this agent can do before a human sees it — and is that limit in code or just in the prompt?


MODULE 7 — Multi-Agent, MCP & A2A

Essence: Most "multi-agent" problems are better solved by one well-designed agent. When you do need many, the hard part is communication, state, and emergent failure.

Learning objectives — you can: 1. Decide whether a problem needs multiple agents or one 2. Explain MCP and A2A and when each applies 3. Design a multi-agent system with a deterministic orchestrator 4. Govern the MCP servers agents can reach

Knowledge checks:

  1. In a multi-agent system, routing and sequencing decisions should be made by:
  2. A) The most capable LLM
  3. B) Deterministic code (the orchestrator)
  4. C) A vote among agents
  5. D) The user

    Answer: B. Why: non-deterministic orchestration is unpredictable and hard to debug; routing is code.

  6. MCP is best described as:

  7. A) A model B) A standard for connecting agents to tools/data C) A vector database D) A prompt format

    Answer: B. Why: MCP standardizes agent-to-tool connectivity (the "USB-C of AI tools"); A2A is agent-to-agent.

  8. The realistic size of a production multi-agent system today is:

  9. A) 100+ autonomous agents B) 3-7 tightly scoped, supervised agents C) Exactly 2 D) As many as possible

    Answer: B. Why: large autonomous agent networks produce undebuggable emergent failure; production systems stay small and supervised.

Apply: Module 7 WORKSHOP — justify (or eliminate) the second agent in a design. One Question: Could one well-designed agent do this — and if not, what specifically requires the second one?


MODULE 8 — Copilot Ecosystem & Skills

Essence: M365 Copilot's biggest risk isn't the AI — it's that it surfaces everything the user already had over-broad access to. Oversharing is the threat.

Learning objectives — you can: 1. Explain how Copilot inherits permissions and why oversharing is the core risk 2. Run a pre-deployment oversharing assessment 3. Govern Copilot Studio agents as production apps 4. Build a Copilot threat model

Knowledge checks:

  1. The primary risk when rolling out M365 Copilot broadly is:
  2. A) The AI hallucinates
  3. B) It makes latent over-permissioned access instantly actionable (oversharing)
  4. C) It's too slow
  5. D) It costs too much

    Answer: B. Why: Copilot inherits the user's permissions; years of access drift become suddenly exploitable via search/synthesis.

  6. The right work to do before broad Copilot rollout is:

  7. A) Tune prompts B) Audit and remediate access permissions C) Buy more licenses D) Train users on prompting

    Answer: B. Why: the pre-deployment work is permissions remediation, not AI configuration.

  8. Citizen-built Copilot Studio agents should be:

  9. A) Unrestricted B) Governed like production applications C) Banned D) Ignored

    Answer: B. Why: they access real data and act in real systems; they need production-grade governance.

Apply: Module 8 — run the oversharing pre-deployment checklist. One Question: Before we turn this on for everyone, do we know what each user's Copilot can reach, and is that what we intend?


Continued in Part 2 — Acts IV-VI (Modules 9-21).