THE AI ARCHITECT CERTIFICATION¶
Course Design & Certification Framework¶
This document turns the 37-module reference into a structured, teachable, certifiable course. It defines the certification levels, the learning objectives for every module, the assessment model, and the capstone. The modules are the content; this is the course.
PART 1 — THE CERTIFICATION STRUCTURE¶
Three Levels¶
LEVEL 1 — AI ARCHITECTURE FOUNDATION
For: engineers, tech leads, and managers entering AI architecture
Modules: 1, 2, 3, 4, 6, 9, 19, 20, 33
Time: 12-15 hours
Assessment: 50-question exam + 1 case study
Credential: Certified AI Architecture Foundation
LEVEL 2 — AI ARCHITECTURE PRACTITIONER
For: working architects building and governing AI systems
Modules: all 37 (Modules 34-37 are the Advanced Architecture act)
Time: 45-55 hours core study; 75-95 hours with all exercises,
workshops, and labs
Assessment: 100-question exam + capstone project
Credential: Certified AI Architecture Practitioner
Prerequisite: Level 1 (or equivalent experience + placement test)
LEVEL 3 — AI ARCHITECTURE EXPERT
For: principal architects, AI platform leads, AI risk leaders
Prerequisite: Level 2 certification + minimum 2 years hands-on
AI architecture experience in production systems
Format: 3-day live intensive workshop + 90-day submission window
Assessment:
- Original architecture submission: a real or anonymised
production AI system (not a hypothetical), reviewed against
the L3 rubric by 2 certified Expert-holders independently;
discordant scores go to a third reviewer
- Oral defense: 45-minute live session — 20-min presentation
of key design decisions, 25-min Q&A from panel
Pass: 80/100 from peer review + pass the oral defense
Distinction: 90/100 + panel commendation
Credential: Certified AI Architecture Expert
Recertification: required every 2 years (see Part 7)
Entry Requirements¶
LEVEL 1:
Minimum: familiarity with software development concepts —
can read code, understands APIs, knows what a database is.
No prior AI experience required. The course starts from
first principles on all AI topics.
Unsure if L1 or L2 is right? Take the placement test
(30-question diagnostic, results in under 20 minutes).
LEVEL 2:
Level 1 certification, or equivalent experience verified
via the placement test (90-minute exam covering all 9
Foundation module topics; 70% pass mark).
LEVEL 3:
Level 2 certification + minimum 2 years of hands-on
production AI architecture experience. Verified by
submitting a portfolio of 3 real architectural decisions
made in production (ADR format, anonymised if required).
The Module-to-Level Map¶
L1 L2 L3
ACT I — FOUNDATIONS
1 Operating Model ● ● ●
2 Model Ecosystem ● ● ●
3 Prompt Engineering Discipline ● ● ●
ACT II — KNOWLEDGE
4 RAG Architecture ● ● ●
5 Data Architecture ● ●
ACT III — AGENTIC
6 Agentic Systems ● ● ●
7 Multi-Agent / MCP / A2A ● ●
8 Copilot Ecosystem ●
ACT IV — SECURITY & GOVERNANCE
9 AI Security ● ● ●
10 Shadow AI & Governance ● ●
11 Compliance & Model Risk ● ●
ACT V — PRODUCTION & PLATFORMS
12 Observability & Evals ● ●
13 Cost Engineering ● ●
14 Infrastructure ● ●
15 Coding Assistants ●
16 Integration & AI-Native ● ●
17 Platform Engineering ● ●
ACT VI — LANDSCAPE, VALUE & DIRECTION
18 Startup Landscape ●
19 Business Value ● ● ●
20 Hype vs. Reality ● ● ●
21 Emerging Patterns ● ●
PRACTICAL LAYER
22 The Practicum ● ●
23 Practice of Architecture ● ●
24 Production Operations ● ●
25 Commercial & Legal ● ●
26 Requirements Engineering ● ●
27 Organizational Patterns ● ●
28 Technical Debt ● ●
29 Communication ● ●
30 Career & Development ● ●
CLOUD & PRODUCTION REALITY
31 Cloud Platforms ● ●
32 Agent Implementation Learnings ● ●
RESPONSIBLE AI
33 Responsible AI in Practice ● ● ●
ADVANCED ARCHITECTURE
34 Fine-Tuning & Customization ●
35 Multimodal Architecture ●
36 GraphRAG & Knowledge Graphs ●
37 Model Portability ● ●
Extension note (Modules 34-37). The four added modules extend Level 2 only. Level 1 is unchanged (same 9 modules, same hours). At Level 3, Module 37 is included because it operationalizes Module 21's commodity thesis, Module 28's model debt, and Module 24's rollback practice, which Expert candidates are expected to apply to production systems. Modules 34-36 are Practitioner depth and are not separately assessed at L3. L3 candidates are expected to have passed the full L2 content.
PART 2 — LEARNING OBJECTIVES (ALL 37 MODULES)¶
Format: "By the end of this module, the learner can…" — written as assessable capabilities, not topics. These map directly to exam items and the capstone rubric.
Module 1 — Operating Model¶
- Distinguish the three altitudes of AI architecture decisions and identify which altitude a given decision belongs to.
- Determine which decisions an architect must own versus delegate.
- Identify the three common traps and recognize when they are occurring.
- Communicate an AI architecture decision to a non-technical stakeholder.
Module 2 — Model Ecosystem¶
- Categorize any model into frontier, open-weight, reasoning, or SLM and state its architectural sweet spot.
- Design a model-routing strategy that matches task complexity to model tier.
- Assess vendor lock-in risk and design for model substitutability.
- Apply the model selection framework to a specific use case with stated constraints.
Module 3 — Prompt Engineering Discipline¶
- Treat prompts as version-controlled, tested production code.
- Construct a layered production prompt with role, constraints, context, and output format.
- Design injection-resistant prompts using structural separation.
- Build a prompt regression test and a change-control process.
Module 4 — RAG Architecture¶
- Identify which of the seven naive-RAG failure modes is occurring in a given system.
- Design a production RAG pipeline with hybrid search, re-ranking, and a confidence gate.
- Diagnose a RAG failure using the four-step sequence.
- Evaluate RAG quality with faithfulness, relevance, and precision metrics.
Module 5 — Data Architecture¶
- Distinguish the AI data stack from the traditional data stack.
- Design a document ingestion pipeline with versioning, lineage, and PII handling.
- Assess data readiness before committing to an AI timeline.
- Prevent PII exposure across the AI data pipeline.
Module 6 — Agentic Systems¶
- Define an agent architecturally as a loop with tools, controls, and a stopping condition.
- Apply the "should this be an agent?" test.
- Design code-enforced controls: iteration limit, cost ceiling, tool scope, human approval.
- Diagnose and fix an agent stuck in a loop.
Module 7 — Multi-Agent / MCP / A2A¶
- Decide whether a problem requires multiple agents or one well-designed agent.
- Explain MCP and A2A and when each applies.
- Design a multi-agent system with a deterministic orchestrator.
- Govern the MCP servers agents can reach.
Module 8 — Copilot Ecosystem¶
- Explain how M365 Copilot inherits user permissions and why oversharing is the core risk.
- Run a pre-deployment oversharing assessment.
- Govern Copilot Studio agents as production applications.
- Build a Copilot threat model.
Module 9 — AI Security¶
- Map the AI attack surface and the OWASP LLM Top 10 to architectural responses.
- Design defense-in-depth that assumes any single layer is bypassed.
- Run and interpret a red-team exercise using Garak.
- Convert security findings into regression eval cases.
Module 10 — Shadow AI & Governance¶
- Detect shadow AI usage in an organization.
- Design an enable-and-govern posture rather than a block posture.
- Implement policy-as-code guardrails.
- Assess organizational AI governance maturity.
Module 11 — Compliance & Model Risk¶
- Apply enduring model-risk principles even where genAI guidance is still forming.
- Build a model inventory and tier models by materiality.
- Design explainability for adverse decisions.
- Establish ongoing fairness and performance monitoring with a named owner.
Module 12 — Observability & Evals¶
- Instrument the three AI observability signals using OpenTelemetry.
- Build an offline eval suite that gates deployment.
- Design and calibrate an LLM-as-judge.
- Detect drift and attribute cost in production.
Module 13 — Cost Engineering¶
- Build a full AI cost model across its eight components.
- Apply routing, caching, and context management to reduce cost.
- Calculate the build-vs-host break-even for a workload.
- Enforce per-task and monthly cost ceilings in code.
Module 14 — Infrastructure¶
- Choose between API, cloud-VPC, and on-prem deployment by data sensitivity and volume.
- Configure vLLM for production self-hosted serving.
- Select a quantization approach for given hardware.
- Design autoscaling for variable inference load.
Module 15 — Coding Assistants¶
- Quantify the productivity gain and the security risk of AI coding tools.
- Design a governance framework with mandatory security review of AI code.
- Match coding-tool configuration to enterprise IP and data requirements.
- Assess the agentic coding tools' specific risks.
Module 16 — Integration & AI-Native Design¶
- Distinguish AI-augmented from AI-native architecture and choose the right one.
- Apply the strangler fig and sidecar patterns to add AI to legacy systems.
- Design confidence and provenance into AI service responses.
- Identify and avoid the AI integration anti-patterns.
Module 17 — Platform Engineering¶
- Define the scope of an AI platform team.
- Design the LLM gateway as the enforcement point for cost, PII, and audit.
- Build a golden path that makes the governed way the easy way.
- Measure the platform as a product (adoption, developer experience).
Module 18 — Startup Landscape¶
- Identify the four moat types and apply them to a build-vs-buy decision.
- Map an AI capability to its startup category and leading vendors.
- Run a vendor competitive-intelligence assessment.
- Recognize where the SaaSpocalypse changes the build-vs-buy calculus.
Module 19 — Business Value¶
- Construct the five-element AI business case.
- Identify which of the six pilot-failure patterns threatens a given initiative.
- Design a pilot that proves economic, not just technical, feasibility.
- Translate AI metrics into CFO, COO, and board language.
Module 20 — Hype vs. Reality¶
- Distinguish capability from reliability for any AI claim.
- Place a capability in the reliable / overhyped / research tier.
- Apply the vendor claim filter.
- Apply the earned-autonomy model to an agent deployment.
Module 21 — Emerging Patterns¶
- Explain the model-commodity shift and its architectural implications.
- Decide between long-context, RAG, and hybrid for a use case.
- Identify which 18-month patterns to plan for, prototype, or monitor.
- Design model-agnostic infrastructure.
Module 22 — The Practicum¶
- Build a 20-case eval dataset and an LLM-as-judge from scratch.
- Iterate a prompt one-failure-at-a-time to a quality bar.
- Diagnose a RAG failure in the fixed sequence.
- Translate a workflow into a LangGraph state machine and build an MCP server.
Module 23 — Practice of Architecture¶
- Run a 90-minute AI design review and classify findings.
- Scope a PoC with one question, an out-of-scope list, and a decision gate.
- Estimate an AI project timeline including the invisible phases.
- Execute a structured first-90-days plan in a new AI architect role.
Module 24 — Production Operations¶
- Execute the 30-minute AI incident playbook.
- Classify an incident by the AI severity matrix.
- Introduce AI safely via shadow mode and canary deployment.
- Execute prompt, model, and knowledge-base rollback procedures.
Module 25 — Commercial & Legal¶
- Brief a legal team on the eight AI-specific contract clauses.
- Run a vendor evaluation with a weighted scorecard.
- Conduct a vendor demo that stress-tests claims.
- Complete AI vendor due diligence.
Module 26 — Requirements Engineering¶
- Write an extended AI user story with quality, failure, oversight, and observability criteria.
- Define AI-specific non-functional requirements.
- Apply the AI definition of done.
- Map an AI workflow including the failure flow.
Module 27 — Organizational Patterns¶
- Identify the four conditions that predict AI initiative success.
- Recognize and intervene in the ten organizational anti-patterns.
- Design an AI Center of Excellence with engagement tiers.
- Assign accountable business ownership for an initiative.
Module 28 — Technical Debt¶
- Identify the six categories of AI technical debt.
- Run a quarterly AI debt assessment and score it.
- Prioritize debt remediation by risk and effort.
- Communicate AI debt to leadership as risk.
Module 29 — Communication¶
- Diagram an AI system using the AI-extended C4 model with human touchpoints and data classification.
- Diagram an agent's phase-tool security boundary.
- Write for four audiences: ADR, brief, executive one-pager, compliance narrative.
- Avoid the diagram anti-patterns.
Module 30 — Career & Development¶
- Assess one's own skills against the AI architect skill map.
- Build an architecture portfolio of documented decisions.
- Answer the common AI architect interview patterns effectively.
- Apply the 5C framework to professional practice.
Module 31 — Cloud Platforms¶
- Select a cloud AI platform by ecosystem fit.
- Map a course pattern to its AWS, Azure, and GCP implementations.
- Decide provider-native versus vendor-neutral for each component.
- Keep business logic portable across providers.
Module 32 — Agent Implementation Learnings¶
- Explain why most agent failures are architecture failures.
- Design tool schemas that drive correct tool selection.
- Apply context engineering to maintain reasoning quality.
- Apply the production agent anti-pattern checklist.
Module 33 — Responsible AI¶
- Apply the "should we build this" screen.
- Identify the four harm types and choose a fairness definition.
- Design transparency, autonomy, and environmental responsibility into a system.
- Conduct a responsible AI review.
Module 34 — Fine-Tuning & Model Customization¶
- Place a performance gap on the customization spectrum (prompting, few-shot, RAG, fine-tuning, full fine-tuning) and determine whether it is a knowledge gap or a behavior gap.
- Select the technique (LoRA, QLoRA, full fine-tuning) and training signal (SFT, DPO, RFT with a validated grader) for a stated task, dataset, and hardware constraint.
- Specify a fine-tuning dataset and a baseline evaluation that covers task performance, regression, catastrophic forgetting, and safety.
- Decide between fine-tuning, distillation, and an efficient-tier API, and govern the resulting model (inventory, data provenance, registry, rollback, closed-platform lock-in).
Module 35 — Multimodal Architecture¶
- Design a document-intelligence pipeline with element-level parsing, table and figure strategies, metadata, and citations.
- Design a vision RAG system, choosing between caption embeddings and image embeddings, and process each asset once at ingestion.
- Choose between a voice pipeline and a speech-to-speech model within a latency, control, and cost budget, and between a frame pipeline and native video based on reuse.
- Estimate multimodal cost from each provider's image, audio, and video token rules, and secure multimodal input (image-level PII, visual prompt injection, BAAs, residency).
Module 36 — GraphRAG & Knowledge Graph Architecture¶
- Recognize the query types (multi-hop, entity-network, corpus-wide) that vector RAG structurally cannot answer.
- Design a knowledge-graph extraction pipeline with entity resolution and quality metrics.
- Choose among graph-augmented RAG, graph-first retrieval, and a hybrid router, and apply the GraphRAG-versus-vector decision framework with a cost justification.
- Govern a knowledge graph as a production data asset (ontology, freshness, PII and erasure, access control, query audit).
Module 37 — Model Portability¶
- Diagnose a model swap across the ten incompatibility layers and distinguish loud breaks from silent ones.
- Design a capability layer, model profiles, and an adapter, and keep the gateway as transport and policy rather than the portability layer.
- Write a behavioral contract with a correctly sized dataset and judge a swap on cost per completed task.
- Design hot, warm, and cold fallbacks with a degradation ladder, run the model swap playbook, and decide when deliberate lock-in is the right call.
PART 3 — ASSESSMENT MODEL¶
Knowledge Checks (per module)¶
Each module has 5-10 knowledge-check questions. The question types:
QUESTION TYPE MIX (per module)
SCENARIO (50%): A situation is described; learner selects the best
architectural response. Tests judgment, not recall.
Example: "A RAG system returns a wrong answer. The retrieved chunks
do not contain the correct information, but the source document does.
What is the root cause?" → (knowledge base gap | retrieval failure |
generation failure | prompt failure)
APPLICATION (30%): Learner applies a framework to a new case.
Example: "Given this workflow, which steps require an LLM and which
are deterministic?"
RECALL (20%): Foundational facts that must be known.
Example: "What is the relationship between input and output token cost?"
Each module's "One Question" (from the Field Cards) is the anchor
synthesis question — the capstone of that module's knowledge check.
The Level 1 Exam¶
50 questions across the 9 Foundation modules (1, 2, 3, 4, 6, 9, 19, 20, 33). 70% to pass. Scenario-weighted. Plus one short case study: given a described business problem, propose a high-level AI architecture and justify three key decisions.
The Level 2 Exam + Capstone¶
100-question exam across all 37 modules. Plus the capstone (below).
Question allocation. The total stays at 100, because the exam length is a constraint on candidate time, not on module count. The new modules are weighted at the same density as the others:
L2 EXAM ALLOCATION (100 items)
Modules 1-33: about 88 items (about 2-3 per module; the Foundation
modules, Security, and Model Risk sit at the upper end)
Modules 34-37: about 12 items (about 3 per module)
Mix unchanged: ~50% scenario, ~30% application, ~20% recall
Allocation is an item-bank design rule, not a per-sitting guarantee. A sitting samples across modules without requiring every module to appear.
PART 4 — THE CAPSTONE PROJECT¶
The capstone is the credibility anchor of the certification. It demonstrates that the learner can do the work, not just answer questions about it.
The Brief¶
Time expectation: 20–30 hours for a submission-grade package. This is not a 2-hour exercise — it is a professional deliverable that should reflect the same rigour as a real client engagement. Budget accordingly: 2–4 hours for diagrams, 4–6 hours for ADRs, 3–4 hours for the evaluation strategy, 2–3 hours for the business case, 2–3 hours for the delivery plan, remainder for integration and review.
CAPSTONE PROJECT BRIEF
You are the AI architect for [scenario provided — e.g., a mid-size
financial services firm wanting to deploy an AI assistant for
customer-facing fee and product questions].
Produce a complete architecture package:
1. ARCHITECTURE DESIGN
- Context and container diagrams (AI-extended C4, Module 29)
- Key design decisions as ADRs (minimum 4, Appendix D)
- The RAG/agent design with confidence gating and human oversight
2. GOVERNANCE & RISK
- Model risk tiering and the governance plan (Module 11)
- Security: threat model + planned red-team approach (Module 9)
- Responsible AI review (Module 33)
3. EVALUATION & OPERATIONS
- The eval strategy: what you measure and the gates (Module 12)
- The operational runbook outline (Module 24)
- The rollback procedures (Module 24)
4. BUSINESS CASE
- The five-element business case (Module 19)
- Cost model and projection (Module 13)
5. DELIVERY PLAN
- Realistic timeline with phases (Module 23)
- Team capability requirements (Module 22)
- The PoC scope and decision gate (Module 23)
The Rubric¶
CAPSTONE SCORING RUBRIC (100 points)
ARCHITECTURE QUALITY (25 pts)
├── Design addresses the actual business problem (5)
├── Failure modes designed for, not just happy path (5)
├── Confidence gating and human oversight present (5)
├── ADRs show reasoning and trade-offs, not just decisions (5)
└── Diagrams show human touchpoints and data classification (5)
GOVERNANCE & RISK (20 pts)
├── Model risk tiering is appropriate and justified (5)
├── Security threat model is specific to this system (5)
├── Responsible AI review surfaces real concerns (5)
└── Compliance considerations are correct for the domain (5)
EVALUATION & OPERATIONS (20 pts)
├── Eval strategy is measurable and gates deployment (7)
├── Operational runbook is usable by an on-call engineer (7)
└── Rollback procedures are specific and fast (6)
BUSINESS CASE (15 pts)
├── Cost baseline and ROI are specific and defensible (8)
└── Named business owner and success metric defined (7)
DELIVERY REALISM (10 pts)
├── Timeline includes the invisible phases (data, eval, compliance) (5)
└── Team capability gaps identified with mitigation (5)
JUDGMENT & COMMUNICATION (10 pts)
├── Trade-offs are explicit and well-reasoned (5)
└── The package communicates clearly to its intended audiences (5)
PASS: 70 points. DISTINCTION: 85 points.
The rubric deliberately rewards: designing for failure, explicit
trade-off reasoning, and realistic delivery thinking — the things
that distinguish a real architect from someone who memorized patterns.
PART 5 — DELIVERY FORMATS¶
THREE DELIVERY MODELS (build in this sequence)
1. CORPORATE TRAINING (fastest to revenue — start here)
Format: 2-day intensive workshop, customized to the client's stack
Content: Modules 1, 6, 7, 9, 11, 19, 20, 33 + capstone exercise
Material: exists today — this is deliverable now
Price band: high per-engagement
2. SELF-PACED LEVEL 1 (builds audience)
Format: recorded modules + knowledge checks + exam
Content: the 8 Foundation modules
Build need: video recording, quiz platform, exam
Price band: accessible per-learner
3. LIVE COHORT LEVEL 2 (premium positioning)
Format: 10-week cohort, weekly live sessions + capstone
Content: full 37 modules (Modules 34-37 as a closing block;
extend the cohort by one week if needed)
Build need: cohort facilitation, capstone review process
Price band: premium per-learner
PART 6 — WHAT STILL NEEDS BUILDING¶
To go fully live, in priority order:
REMAINING BUILD CHECKLIST
CONTENT (mostly done):
✓ 37 modules
✓ Appendices A-G
✓ Field Cards (distillation)
✓ The 32 Questions
✓ Learning objectives (this document)
✓ Capstone brief + rubric (this document)
☐ Knowledge-check questions written out (5-10 per module)
☐ Level 1 and Level 2 exam item banks
☐ Case study scenarios for the exams
PRODUCTION:
☐ Professional diagrams (replacing ASCII)
☐ Slide decks (for live/corporate delivery)
☐ Video recordings (for self-paced)
PLATFORM & OPERATIONS:
☐ Platform selection (Teachable / Thinkific / Kajabi / own on 5cway.com)
☐ Credential issuance (Credly badge or own certificate)
☐ Capstone review process and reviewer training
This framework converts the 37-module reference into a structured certification. The content is the hard part and it is done. What remains — knowledge-check items, exams, diagrams, platform — is production work, not intellectual work. The course exists; it now needs to be produced.
PART 7 — RECERTIFICATION¶
AI architecture moves fast. A certification earned today against 2026 content may be materially outdated in 18 months. All levels expire and must be renewed to remain valid.
EXPIRY PERIODS
Level 1 — Foundation: valid 2 years from issue date
Level 2 — Practitioner: valid 2 years from issue date
Level 3 — Expert: valid 2 years from issue date
RENEWAL OPTIONS (choose one)
Option A — Recertification Exam:
A shortened exam covering modules updated since your
certification date.
Level 1: 30 questions, 70% to pass
Level 2: 60 questions, 70% to pass
Available from your certification anniversary date.
Option B — Continuing Education Units (CEUs):
Earn 8 CEUs per 2-year period. CEU-eligible activities:
Attending an AI Architect workshop or live session 1 CEU
Completing a new or updated course module 1 CEU each
Contributing a verified case study to the course 2 CEUs
Speaking at a qualifying AI architecture event 2 CEUs
Earning a higher certification level resets clock
Option C — Level Upgrade:
Passing a higher level exam resets the expiry for all
lower-level credentials automatically.
URGENT UPDATES:
Significant regulatory changes (new EU AI Act enforcement
actions, material NIST AI RMF revisions) may trigger a
mandatory targeted module update outside the 2-year window.
Credential holders will be notified by email with a 90-day
grace period to complete the update. Failure to complete
within the grace period suspends the credential.