Skip to content

MODULE 23 — The Practice of AI Architecture

23.1 What This Module Covers

Modules 1-22 teach the knowledge and skills of AI architecture. This module teaches the practice — the actual work an AI architect does day-to-day. There is a significant gap between understanding AI architecture intellectually and being effective in the role. This module closes that gap.


23.2 Running an AI Design Review

The Purpose

A design review is not a checkpoint. It is a conversation that surfaces risks the team hasn't seen, validates that the design's assumptions are sound, and produces a documented record of what was agreed. Done well, it prevents expensive mistakes. Done poorly, it is theater that delays delivery without adding value.

Who Should Be in the Room

AI DESIGN REVIEW PARTICIPANTS

Required:
  ├── Lead architect (reviewer — you)
  ├── Technical lead of the team being reviewed
  ├── One senior engineer from the team
  └── Product owner (for the first 30 minutes — business context only)

Invite based on the system:
  ├── Security architect (if the system handles sensitive data)
  ├── Data engineer or data lead (if complex data pipelines)
  ├── Compliance/model risk (if regulated use case)
  └── Platform team representative (if using new platform components)

NOT required:
  ├── The entire development team (intimidates honest conversation)
  ├── Executives (changes the conversation to performance)
  └── Vendors (no external parties in internal reviews)

The 90-Minute Agenda

AI DESIGN REVIEW AGENDA

0:00-0:15 — Context Setting (Product Owner present)
  Team presents: What problem are we solving? Who are the users?
  What does success look like in 90 days?
  Architect asks: Why is AI the right approach for this?
  [Product Owner leaves at 0:15]

0:15-0:40 — Architecture Walkthrough
  Team walks through: the system design, data flows, AI components,
  integration points, deployment model.
  Architect LISTENS. No questions yet. Take notes.
  One rule: the team presents their design, not a revised design they
  think the architect wants to see.

0:40-1:00 — Architectural Questions (High Altitude)
  Architect asks the questions that surface the highest-impact risks.
  Use the review checklist (Appendix C) as a mental guide, not a script.
  Priority order: data governance → security → failure modes →
                  evaluation → cost → operational concerns

  KEY: Ask about failure modes before asking about the happy path.
  "Walk me through what happens when the model returns a wrong answer."
  reveals more than "walk me through a successful interaction."

1:00-1:20 — Open Discussion
  "What are you most uncertain about in this design?"
  "What would you do differently if you had more time?"
  "What is the thing most likely to go wrong in the first 30 days?"
  These questions surface the team's own concerns — which are usually
  the most accurate predictors of production problems.

1:20-1:30 — Findings and Next Steps
  Architect summarizes: required changes, recommendations, observations.
  See classification below.
  No surprises. If a finding is significant, the team should hear it
  framed as "I want to flag something significant" before the meeting ends.

Classifying Findings

FINDING CLASSIFICATION

REQUIRED (blocks deployment):
  The design has a material gap that must be addressed before production.
  Examples: no confidence gate on a customer-facing RAG system,
            no audit trail on a regulated AI decision,
            PII being sent to external LLM without policy approval.

  Document as: "Before this system can deploy to production, [specific change]
               must be made. Owner: [name]. Deadline: [date]."

RECOMMENDED (should address before deployment):
  Not a blocker, but a meaningful quality or risk improvement.
  Examples: no eval suite, no model version pinning, weak fallback design.

  Document as: "The team should address [specific item] before production.
               If not addressed, the risk is [specific consequence]."

OBSERVATION (for future consideration):
  Something worth noting for future iterations.
  Examples: the chosen approach will hit scaling limits at X volume,
            this component is a candidate for the platform team to own.

  Document as: "For future consideration: [observation]."

COMMENDATION (yes, document the good too):
  Specific good decisions that should be recognized.
  This is not flattery — it teaches others what good looks like.
  Document as: "The team's decision to [X] is a good pattern for
               other teams to follow."

The Written Findings Document

Send within 24 hours of the review. Never more than one page.

AI DESIGN REVIEW FINDINGS

System: [Name]
Review date: [Date]
Attendees: [List]
Reviewer: [Your name]

REQUIRED CHANGES (must complete before production deployment):
  1. [Specific change, specific reason, owner, deadline]

RECOMMENDATIONS (address before production):
  1. [Specific item, specific risk if not addressed]

OBSERVATIONS (future consideration):
  1. [Observation and why it matters longer-term]

COMMENDATIONS:
  1. [Specific decision and why it's a good pattern]

APPROVED TO PROCEED TO: [Staging / Limited Production / Full Production]
with the condition that Required changes are completed by [date].

23.3 How to Scope a Proof of Concept

The PoC's Only Job

A PoC has one job: answer a specific technical question that cannot be answered any other way. Not "demonstrate the value of AI." Not "build a prototype we can show the board." Those are different artifacts with different scopes.

If you cannot write the technical question in a single sentence, you do not have a PoC. You have a project disguised as a PoC.

GOOD POC QUESTIONS (single technical question):
  "Can we achieve >85% faithfulness on our policy document corpus
   using a RAG approach with standard chunking and hybrid search?"

  "Does the Claude Sonnet model produce medically appropriate triage
   suggestions for our specific patient intake scenarios at acceptable
   accuracy?"

  "Can a LangGraph state machine handle the exception patterns in our
   claims workflow without requiring human intervention more than 20%
   of the time?"

BAD POC QUESTIONS (vague or multiple):
  "Can AI improve our customer experience?" — Not a technical question.
  "Should we use OpenAI or Anthropic?" — Vendor selection, not a PoC.
  "Will AI save us money and improve quality?" — Business case, not a PoC.

The PoC Scope Constraint

Everything that is not needed to answer the question is out of scope. This sounds obvious. It is not, in practice.

POC SCOPE TEMPLATE

Question: Can we achieve >85% faithfulness on our policy document corpus?

IN SCOPE:
  ├── Representative sample of production documents (50-100 docs)
  ├── Basic RAG pipeline (chunking, embedding, retrieval)
  ├── Faithfulness evaluation against 30 human-authored test cases
  └── Comparison of 2-3 chunking strategies

OUT OF SCOPE (explicitly excluded):
  ├── Production infrastructure (use local env or single cloud instance)
  ├── Authentication and authorization
  ├── User interface
  ├── Document lifecycle management
  ├── Integration with production systems
  ├── Monitoring and alerting
  ├── Error handling beyond basic try/catch
  └── Performance optimization

Time-box: 2 weeks, no extensions.

The out-of-scope list is as important as the in-scope list.
Without it, engineers add "just one more thing" until the PoC
is a half-built production system that runs for 6 months.

PoC Decision Gate

Define the pass/fail criteria before the PoC starts. Not after you see the results.

POC DECISION GATE

Pass criteria: faithfulness > 0.85 on 30 test cases with production documents
Fail criteria: faithfulness ≤ 0.75

If pass: Proceed to design a production architecture. Budget approved.
If fail: Investigate root cause. Is it chunking? Embedding model?
         Document quality? Is the question answerable with a
         different approach?
If inconclusive (0.75-0.85): Run targeted experiments to understand
         what's limiting quality. 2 additional weeks maximum.

What this PoC does NOT decide:
  - Which vendor to use (that's a separate decision)
  - What the production architecture looks like (PoC is NOT a prototype)
  - Whether to scale this to the full document corpus (requires production design)

The PoC-to-Production Trap

The most expensive PoC failure is when the PoC "succeeds" and becomes the production system by default. This happens when: - The PoC was scoped as a prototype, not a technical validation - There was no explicit decision gate, so the PoC just kept running - The team made architectural decisions during the PoC under the assumption they'd be replaced later — they weren't

The rule: A PoC that passes its decision gate is decommissioned. A new production design begins. Code from the PoC may inform the design but is not promoted to production directly.


23.4 Estimating AI Project Timelines

Why AI Timelines Are Different

Traditional software estimation is driven by code complexity. AI project estimation is driven by a different set of variables — most of which engineers underestimate because they don't know they exist.

The AI Project Phase Model

AI PROJECT PHASE MODEL

Phase 0: Data Readiness (often invisible in plans — always runs long)
  Activities: Data audit, quality remediation, pipeline design,
              integration with source systems, PII handling design

  Typical estimate: 2 weeks
  Realistic range: 3-12 weeks
  Multiplier: 3-5x what engineers expect

  Why it runs long: Source systems are messier than documented.
  Legal/compliance review of data usage takes longer than IT work.
  The data team has other priorities.
  What was described as "a clean database" has 8 years of schema drift.

Phase 1: Foundation (often missing from plans)
  Activities: Eval infrastructure setup, baseline measurement,
              gateway and observability integration

  Typical estimate: 0 (teams skip this)
  Realistic range: 2-4 weeks

  Why it runs long: Teams don't budget for eval infrastructure.
  They treat it as "something we'll add later." Without it, there is
  no way to know if Phase 2 is working.

Phase 2: Core AI Build
  Activities: Model selection, prompt development, RAG pipeline,
              agent design, integration coding

  Typical estimate: This is the only phase teams plan for.
  Realistic range: 4-8 weeks for moderate complexity

  Estimation rule: Double whatever the engineering team estimates for
  the first iteration. Add 50% for each integration with a legacy
  system.

Phase 3: Evaluation and Iteration (always underestimated)
  Activities: Running evals, identifying quality gaps,
              iterating on prompts and retrieval, re-evaluation

  Typical estimate: 1 week
  Realistic range: 2-6 weeks

  Why it runs long: The first eval run almost always fails to meet
  the quality bar. Each iteration of "fix the problem, re-evaluate"
  takes 3-5 days minimum. Teams budget for 1 cycle. Reality is 3-6.

Phase 4: Security and Governance Review
  Activities: Red team exercise, compliance review, model risk
              documentation, policy approval

  For regulated industries:
  Typical estimate: 1 week (teams think it's a rubber stamp)
  Realistic range: 4-12 weeks

  Note: Compliance review gates cannot be parallelized with development
  in regulated industries. Plan for them on the critical path.

Phase 5: Production Preparation
  Activities: Infrastructure setup, monitoring configuration,
              runbook creation, team training, change management

  Typical estimate: 1-2 weeks
  Realistic range: 2-4 weeks

TOTAL PROJECT TIMELINE EXAMPLE:
  Engineering team estimate for "moderate complexity AI feature":
  6-8 weeks total

  Realistic estimate with all phases:
  Phase 0: 4-6 weeks (data)
  Phase 1: 2-3 weeks (foundation)
  Phase 2: 4-6 weeks (build)
  Phase 3: 3-5 weeks (eval iteration)
  Phase 4: 4-8 weeks (review — for regulated)
  Phase 5: 2-3 weeks (production prep)

  Realistic total: 19-31 weeks (5-8 months)
  vs. team estimate of 6-8 weeks

  This is not an exaggeration. This is the consistent pattern.

How to Communicate Timeline Estimates

When you give an estimate that is 3x longer than the team expected, you will face pushback. These framings help:

"What can we cut from scope to hit the shorter timeline?" This is the right answer to timeline pressure. If the business needs this in 8 weeks, which phases can be compressed? Data readiness cannot be compressed (it takes what it takes). The compliance review cannot be compressed. What CAN be compressed is Phase 3 (ship with lower quality bar, accept more risk) or Phase 2 (narrower scope). Make the trade-off explicit, not implicit.

"The 6-week estimate assumes this data is ready. Is it?" Most timeline overruns trace back to Phase 0 being wrong. Surface this early. Ask: has the data been profiled? Has legal/compliance approved this data for AI use? The answers to these questions will immediately clarify whether the short timeline is achievable.


23.5 The First 90 Days as an AI Architect

What This Framework Is For

Whether you are newly hired into an AI architect role, transitioning from a different architecture specialty, or being asked to own AI architecture for the first time in your current organization — the first 90 days establish your credibility, your understanding, and your initial direction. This framework prevents the two most common failure modes: moving too fast (committing to a direction before understanding the landscape) and moving too slowly (spending 90 days learning without producing anything).

Days 1-30: Listen and Inventory

The goal of the first 30 days is not to change anything. It is to understand what exists.

DAYS 1-30 ACTIVITIES

Week 1: Stakeholder mapping
  ├── Who are the 10 people most affected by AI architecture decisions?
  ├── Schedule 30-minute conversations with each.
  ├── Question to ask each: "What's the most frustrating AI-related
  │   problem you're dealing with right now?"
  │   DO NOT ask: "What do you think we should build?"
  └── Listen for: pain points, political dynamics, existing commitments

Week 2: AI system inventory
  ├── Audit every AI system currently in production or development.
  ├── For each: purpose, owner, status, data handling, governance status.
  ├── Use the AI inventory schema from Module 11.
  ├── Start from zero — do not assume the inventory is accurate.
  └── Source of truth: interview the team leads, not the documentation.

Week 3: Data and platform assessment
  ├── What data infrastructure exists? (Data warehouse, data lake, CDC)
  ├── What AI platform infrastructure exists? (Gateway, eval, observability)
  ├── Where is the biggest data readiness gap?
  └── Is there a model risk management process? If yes, is it followed?

Week 4: Gap analysis and prioritization
  ├── What are the 3 highest-risk AI systems currently in production?
  ├── What are the 3 most urgent missing governance/platform components?
  ├── What is the most immediate threat (security, compliance, cost)?
  └── Draft your initial findings — show them to one trusted peer before
      sharing more broadly.

Days 31-60: Analyse and Validate

The goal of days 31-60 is to validate your initial findings and produce a credible assessment.

DAYS 31-60 ACTIVITIES

Week 5-6: Deep-dives on high-risk systems
  ├── Run informal design reviews on the 3 highest-risk systems.
  ├── Use the review process from Section 23.2.
  ├── Not formal yet — position as "I'm trying to understand the system."
  ├── Document findings but do not publish yet.
  └── Look for patterns across systems, not just individual problems.

Week 7-8: Platform and tooling assessment
  ├── Evaluate current AI tooling against Module 17's maturity model.
  ├── Where is the organization on the platform maturity scale?
  ├── What is the single highest-leverage platform investment?
  └── Identify the champions — the 2-3 engineers who want better
      infrastructure and will help build it if you create the space.

Checkpoint: Present your preliminary findings to your manager only.
  Not to leadership yet. Validate that your read of the situation
  is accurate. Surface anything you may have misread.

Days 61-90: Produce and Commit

The goal of days 61-90 is to deliver something tangible and commit to a direction.

DAYS 61-90 ACTIVITIES

Deliverable 1: The AI Architecture Assessment (10-15 pages)
  ├── Inventory of current AI systems with risk ratings
  ├── Top 3 governance/security gaps with evidence
  ├── Platform maturity assessment with gap analysis
  ├── 3 recommended immediate actions
  └── 12-month architectural direction (not a roadmap yet — a direction)

Deliverable 2: One quick win
  ├── Identify a gap that can be closed in 2-4 weeks with minimal
  │   engineering effort but meaningful governance value.
  ├── Examples: implement model version pinning organization-wide,
  │   add cost attribution tags to the highest-cost AI service,
  │   run a one-day red team exercise on the highest-risk system.
  └── Complete this before the 90-day mark. It establishes credibility.

Deliverable 3: The 12-month plan (high level)
  ├── Present to leadership by day 90.
  ├── 3 columns: Now (next 90 days), Next (90-180 days), Later (180+ days)
  ├── Each initiative: what, why, estimated effort, who owns it
  └── Prioritization rationale visible — not just a list, a reasoned sequence.

What NOT to do in the first 90 days:
  ├── Do not commit to a specific technology stack before week 6
  ├── Do not run formal design reviews before building relationships
  ├── Do not produce a 12-month roadmap in week 2 (you don't know enough)
  ├── Do not publicly criticize existing systems before understanding
  │   the constraints that produced them
  └── Do not be the person who arrives with all the answers.
      Be the person who arrives with the right questions.

23.6 Running Architecture Office Hours

One of the most effective practices an AI architect can establish, rarely discussed: regular open office hours for engineering teams. Not formal reviews. Not scheduled consultations. 60-90 minutes per week where any team can drop in with AI architecture questions.

Why this matters: - It makes governance feel like help, not oversight - It surfaces problems early, when they are cheap to fix - It builds the relationships that make formal reviews more effective - It teaches you what is actually happening in the organization, not what is documented

How to run them effectively:

AI ARCHITECTURE OFFICE HOURS

Format: 60-90 minutes, recurring weekly
Booking: First-come first-served, 15-minute slots
Platform: Slack channel + video link

Ground rules:
  - Bring a specific question, not a general "can you review this?"
  - Bring the relevant context (design doc, code snippet, error)
  - No decisions made in office hours — decisions require proper review
  - What is said in office hours informs formal reviews but does not
    replace them

What to do:
  - Answer the specific question asked
  - Note patterns you're seeing across teams ("this is the third team
    this week asking about embedding model selection")
  - Connect teams who are working on similar problems
  - Leave with a clear next step: "bring a design doc to the review,
    read this module, spike this experiment"

What NOT to do:
  - Do not solve the team's design problem for them
  - Do not commit to approvals
  - Do not create dependencies on you being present for every decision

23.7 The Architecture Decision Log

An architecture decision log is not the same as an ADR. The ADR is a document about a specific decision. The decision log is the running record of all architectural decisions made across the organization, who made them, and when.

Most organizations have neither. The architecture decision log is the most underbuilt governance artifact in most enterprise AI programs.

ARCHITECTURE DECISION LOG ENTRY

Decision ID: AD-2024-047
Date: 2024-03-15
Decision maker(s): [names]
Affected systems: customer-support-ai, product-recommendation-ai
Status: Active (can be superseded)

Decision: All production AI systems must use the centralized LLM gateway.
Direct API calls to model providers are prohibited in production.

Rationale: Cost visibility, PII scanning, consistent audit trail,
           rate limiting. Without a gateway, these cannot be enforced
           consistently.

Alternatives rejected:
  - Per-service gateway: rejected because it replicates the
    enforcement problem.

Constraints this creates:
  - Teams must onboard to the gateway before deploying AI features.
  - Gateway downtime affects all AI features.

Review trigger: If gateway adds > 50ms latency for > 5% of requests,
                or if a use case cannot be served by the gateway,
                this decision should be revisited.

Related decisions: AD-2024-031 (cost attribution policy),
                   AD-2024-038 (model version pinning)

The value of the log over individual ADRs is the relationships between decisions — when AD-2024-047 was superseded by a later decision, the teams affected by the original decision can be identified and notified.


EXERCISE — Design Review Simulation: Pair with a colleague. One person presents an AI system design (real or hypothetical). The other runs a 45-minute design review using the agenda from Section 23.2. Produce the findings document within 24 hours. Swap roles and repeat. The debrief question: what did the reviewer ask that the designer hadn't considered? What did the designer present that surprised the reviewer?

PONDER — Your First 90 Days: If you started in a new AI architect role tomorrow, what is the single most important thing to understand in the first 30 days? What would the quick win be? What is the thing you would most need to resist doing too early?

WORKSHOP — PoC Scoping: Take an AI initiative your organization is considering. Write the single technical question the PoC should answer. Define what is explicitly in scope and explicitly out of scope. Define the decision gate (pass/fail criteria, timeline). Identify the PoC-to-production trap risk: what is the most likely way this PoC becomes the production system by default, and how do you prevent it?


Next: Module 24 — AI Production Operations