THE AI ARCHITECT'S FIELD CARDS — PART 3¶
Cards 23-33 (The Practical, Cloud & Responsible AI Layer)¶
Part 3 of 4. Covers the Practical Layer (Cards 23-30), Cloud Provider Platforms (Card 31), Agent Implementation Learnings (Card 32), and Responsible AI (Card 33). Part 1 covers Acts I-IV (Cards 1-11). Part 2 covers Acts V-VI (Cards 12-22). Part 4 covers the Advanced Topics (Cards 34-37). Each card follows the same six-section structure: Essence → Core Insight → Key Framework → Decision Rule → Red Flags → The One Question.
---¶
CARD 23 — The Practice of AI Architecture¶
THE ESSENCE The role is design reviews, PoC scoping, timeline estimation, and the first 90 days — not just system design. The practice is what makes the knowledge effective.
THE CORE INSIGHT A design review is a conversation that surfaces hidden risk, not a checkbox. A PoC answers one technical question, then is decommissioned. AI timelines are 3-5x longer than teams estimate because of invisible phases (data, evals, compliance). The first 90 days are for listening and one quick win — not for arriving with all the answers.
THE KEY FRAMEWORK — The Core Practices
Design review: context → walkthrough → failure-mode questions → classified findings
PoC: ONE technical question → explicit out-of-scope → decision gate → decommission
Timeline: data (3-5x) + eval infra + build + eval iteration + compliance + prod prep
First 90 days: listen → analyze → produce one assessment + one quick win
THE DECISION RULE - Ask about failure modes before the happy path in every review - A PoC has one question and a kill date; it is not a prototype - Surface the data-readiness phase explicitly in timelines - In a new role: listen first, deliver one quick win, don't over-commit early
RED FLAGS - Design reviews with no authority to require changes (theater) - A PoC that quietly became the production system - Timelines that only count the build phase
THE ONE QUESTION "Does my design review actually change designs, or does everyone nod and ship what they already built?"
---¶
CARD 24 — AI Production Operations¶
THE ESSENCE AI incidents return HTTP 200 while being wrong. Monitoring shows green while users are harmed. You need quality detection and AI-specific rollback before you need them.
THE CORE INSIGHT The failure mode is content-level, not system-level, so traditional monitoring misses it. The 30-minute incident playbook: triage what's wrong → check what changed in 24h → run evals against production → decide mitigation (rollback / route-to-human / disable). AI has more rollback targets than software: prompt, model, knowledge base — each with its own procedure.
THE KEY FRAMEWORK — Incident + Rollback
Detect (eval alert > complaint) → triage → "what changed in 24h?" →
run evals → mitigate (rollback prompt | model | KB / route to human / disable)
Shadow mode → canary (5→20→50→100%) for safe introduction
THE DECISION RULE - Detect quality drops by monitoring, not customer complaints - Know your rollback time for prompt, model, and KB (target < 30 min) - Control AI routing with feature flags (rollback = flag change, not deploy) - Introduce AI via shadow mode then canary, never big-bang
RED FLAGS - You learn about AI quality problems from angry users - Rollback requires a deployment (too slow) - No shadow/canary process — AI goes straight to 100% of traffic
THE ONE QUESTION "If the AI started giving wrong answers right now, how would I know, and how fast could I roll back?"
---¶
CARD 25 — Commercial & Legal¶
THE ESSENCE Your job isn't to be the lawyer — it's to tell the lawyer what AI-specific clauses matter. Most AI contracts don't protect the buyer because nobody knew what to ask for.
THE CORE INSIGHT The eight clauses that matter: training-data use, retention, model-version notification, accuracy/process SLA, security-incident notification, audit rights, exit/portability, and EU AI Act compliance. Vendor demos are theater — supply your own scenarios, include failure cases, and watch how they react to the hard procurement questions.
THE KEY FRAMEWORK — The Eight Clauses
Training-data use | retention | model-version notice | accuracy/process SLA |
incident notification (72h) | audit rights | exit/portability | EU AI Act
Plus: demo with YOUR scenarios incl. failures; ask "show me when it's wrong."
THE DECISION RULE - Brief legal on the eight AI-specific clauses before contract review - Run demos with your own scenarios, including failure cases - Demand training-data-use prohibition for enterprise data - Verify exit/portability before you're dependent - Check provider termination-notice and change-of-control terms: access has been withdrawn for business reasons with days to weeks of notice (Card 37)
RED FLAGS - Signing a contract silent on training-data use (your data may train their model) - "Enterprise-ready" with no DPA or audit rights - A vendor who gets evasive at standard procurement questions
THE ONE QUESTION "Does this contract say our data won't train their models, and can we leave with our data in 12 months?"
---¶
CARD 26 — AI Requirements Engineering¶
THE ESSENCE "As a user I want X" says nothing about quality, failure, or human oversight — the four things that matter most for AI. Standard user stories ship unpredictable AI.
THE CORE INSIGHT AI requirements need four extensions: measurable quality acceptance criteria, defined failure-mode behavior, human-oversight requirements, and observability requirements — plus an explicit out-of-scope declaration. Without these, engineers build to implicit standards and the gaps surface in production.
THE KEY FRAMEWORK — The Extended AI Story
As a [user] I want [capability] so that [value]
+ QUALITY AC (measurable: "≥85% faithfulness")
+ FAILURE MODE ("when confidence < X, escalate")
+ HUMAN OVERSIGHT ("who reviews what, when")
+ OBSERVABILITY ("log these fields")
+ OUT OF SCOPE ("must NOT do Y")
THE DECISION RULE - Every AI story has quality AC with a number - Define failure-mode behavior explicitly (what happens at low confidence) - Specify the human-oversight model in the requirement - Declare out-of-scope behaviors up front
RED FLAGS - AI feature with no measurable quality target ("done" = doesn't crash) - No defined behavior for low-confidence or out-of-scope inputs - Observability bolted on after launch
THE ONE QUESTION "Does this requirement say what 'good enough' means as a number, and what happens when the AI can't meet it?"
---¶
CARD 27 — Organizational Patterns¶
THE ESSENCE Most AI initiatives fail for organizational reasons, not technical ones: no business owner, data not ready first, governance bolted on, scope too broad. Fix the org pattern or the architecture won't matter.
THE CORE INSIGHT Four conditions predict success: a named business owner accountable for the outcome, data foundation built first, governance designed in, and narrow scope proven before expansion. Ten anti-patterns predict failure — the AI Champion (single point of failure), Innovation Lab (isolated from production), Permanent Pilot, Review Without Authority, Build Everything.
THE KEY FRAMEWORK — Success Conditions vs. Anti-Patterns
SUCCEED: named business owner + data-first + governance-in + start-narrow
FAIL: AI Champion / Innovation Lab / Permanent Pilot / Data Standoff /
Vendor Dependency / "AI figures it out" / activity-not-outcomes /
headcount-fear comms / review-without-authority / build-everything
THE DECISION RULE - Name a business owner accountable for the outcome before starting - Build the data foundation first; it's the gating dependency - Start narrow, prove it, then expand on the same infrastructure - Watch for the ten anti-patterns and intervene early
RED FLAGS - AI initiative owned by an enthusiastic individual, not the business - A separate "AI lab" producing demos nobody adopts - A pilot that's been "almost ready to scale" for a year
THE ONE QUESTION "Who loses their bonus if this AI initiative doesn't move its business metric — and if the answer is 'nobody,' why are we starting?"
---¶
CARD 28 — AI Technical Debt¶
THE ESSENCE AI systems accumulate six kinds of invisible debt — prompt, eval, model, knowledge-base, governance, data-pipeline — that stay hidden until a failure, audit, or deprecation makes them suddenly expensive.
THE CORE INSIGHT Unlike code debt, AI debt is invisible because the system keeps running with quietly declining quality. The most dangerous is model debt (a deprecation deadline you didn't plan for) and governance debt (audit trails that don't exist until an examiner asks). Measure it quarterly; communicate it to leadership as risk, not as cleanup work.
THE KEY FRAMEWORK — The Six Debt Types
PROMPT (bloated, untested) | EVAL (coverage gaps) | MODEL (deprecation, no migration) |
KNOWLEDGE-BASE (stale docs) | GOVERNANCE (missing audit/inventory) |
DATA-PIPELINE (fragile, unmonitored)
Score quarterly; prioritize by risk × remediation effort.
THE DECISION RULE - Run a quarterly AI debt assessment across all six categories - Check model deprecation timelines every quarter (the one with a deadline) - Communicate debt as risk + cost, not as "we need to refactor" - Add an eval case for every production incident (pays down eval debt) - Score model portability per capability; low scores are model debt in waiting (Card 37)
RED FLAGS - A pinned model approaching deprecation with no migration plan - Can't list your AI systems and owners (governance debt) - Knowledge base full of stale docs nobody reviews
THE ONE QUESTION "Which of my AI systems would fail an audit, a model deprecation, or a quality review today — and am I tracking that as risk?"
---¶
CARD 29 — AI Architecture Communication¶
THE ESSENCE A decision you can't communicate won't be implemented consistently. AI systems need diagrams that show the human touchpoints, the data crossings, and the failure paths — not just an "AI box."
THE CORE INSIGHT The "AI box" diagram (Service A → AI → Service B) communicates nothing. Open the box: show retrieval, confidence gate, escalation, the data classification on every cross-boundary flow, and where humans review. Then write for the audience — ADR for engineers, brief for tech leaders, one-pager for executives, compliance narrative for auditors.
THE KEY FRAMEWORK — Diagram + Audience
C4 for AI: show human touchpoints, data-classification on flows, failure paths
Four documents: ADR (engineers) | brief (leaders) | one-pager (execs) |
compliance narrative (auditors) — same system, four translations
THE DECISION RULE - Open the "AI box" — diagram retrieval, gates, escalation, data crossings - Annotate every cross-boundary data flow with its classification - Always draw the failure paths, not just the happy path - Write the right document for each audience
RED FLAGS - An architecture diagram with a single opaque "AI" box - No data-classification annotation (can't tell if PII crosses the boundary) - The same technical doc handed to engineers and the board
THE ONE QUESTION "Does my diagram show what happens when the AI fails and where the PII goes — or just the happy path through a magic box?"
---¶
CARD 30 — Career & Professional Development¶
THE ESSENCE The role rewards judgment shown under pressure, maintained competence in a fast field, and a portfolio of documented decisions — not certifications alone. Character, Competence, Consistency, Commitment, Courage.
THE CORE INSIGHT AI architecture has no dominant certification yet, so a portfolio of documented architectural decisions beats a credential. Maintain competence with a deliberate learning cadence (the field's knowledge half-life is short). Build experience where you are. The 5C frame — Character, Competence, Consistency, Commitment, Courage — is how the practice compounds over a career.
THE KEY FRAMEWORK — Skill Map + 5C
SKILLS: systems + AI/ML + governance + business communication
PORTFOLIO: documented decisions & trade-offs > certifications
5C: Character (hold standards) · Competence (maintain it) · Consistency
(same bar for everyone) · Commitment (the unglamorous maintenance) ·
Courage (ask "what happens when it's wrong?")
THE DECISION RULE - Build a portfolio of documented decisions, not just certs - Maintain a deliberate learning cadence (daily/weekly/monthly/quarterly) - Build AI experience in your current role — title follows the work - Apply existing domain expertise; it's harder to teach than AI patterns
RED FLAGS - Collecting certifications without shipping or documenting real systems - Knowledge frozen at a point in time in a field that moves every 6 weeks - Confidence about domains you've never actually worked in
THE ONE QUESTION "In the hard moment — deadline pressure, an over-promising vendor, a risky shortcut — am I the architect who holds the standard or the one who caves?"
---¶
CARD 31 — Cloud Provider AI Platforms¶
Last reviewed: October 2026. Cloud platform features and product names change frequently. Verify provider-specific details quarterly (Appendix G, G.6).
THE ESSENCE Match the AI platform to your existing ecosystem, not to a feature comparison. Switching clouds for AI creates fragmentation that costs more than any feature gap.
THE CORE INSIGHT AWS Bedrock/AgentCore, Microsoft Foundry (formerly Azure AI Foundry), and Google's Gemini Enterprise Agent Platform (formerly Vertex AI) are converging on capability; they differ on ecosystem fit. Model access is converging too: since 2026 AWS hosts both Claude and OpenAI's frontier models, and Foundry hosts OpenAI and Claude, so the model no longer dictates the cloud. AWS suits AWS-native orgs; its agent runtime instances now run up to 14 days, but Google and Microsoft also offer multi-day runtimes, so run length is no longer a differentiator. Azure suits M365 orgs, with Entra-ID agent identity and the broadest catalog. GCP suits data-heavy GCP orgs that want Google Search grounding. Go provider-native for infrastructure, vendor-neutral for business logic.
THE KEY FRAMEWORK — Fit + Portability
AWS → AWS-native, Claude + OpenAI models, IAM agent identity, AgentCore
(long-running agents: 14 days on instances; Google 7, Foundry 30-day sessions)
AZURE → Microsoft Foundry: M365 integration, Entra-ID "digital employee"
agents, 11,000+ model catalog
GCP → Gemini Enterprise Agent Platform: data-heavy GCP, Search grounding,
A2A-native
Rule: native infra, portable logic — agent code deploys anywhere w/ config change
THE DECISION RULE - Select by ecosystem fit first, feature comparison second - Keep business logic portable (LangGraph/code), even on native infra - Use provider-managed services for the undifferentiated heavy lifting - Define requirements before selecting the platform
RED FLAGS - Choosing a platform on feature checklist while ignoring ecosystem fit - Choosing a cloud because it is the only route to a model (no longer true for the major frontier models) - Architecture docs still naming retired products (Vertex AI, Azure AI Foundry) - Business logic locked into provider-specific orchestration syntax - "Managed service" treated as a substitute for governance
THE ONE QUESTION "If I had to migrate providers in 12 months, would my agent logic move with a config change — or need a rewrite?"
---¶
CARD 32 — AI Agent Implementation Learnings¶
THE ESSENCE 78% of enterprises have agent pilots; 14% scaled them (Capgemini Research Institute, 2025). Most failures are architecture, integration, and context-engineering failures — not model failures.
THE CORE INSIGHT What production teaches: integration is the real bottleneck (you don't control the APIs), tool-schema quality drives tool-selection accuracy, longer context makes things worse (attention competition), evaluation must measure tool selection and multi-step coherence, irreversibility must be code-enforced, and starting narrow is the only strategy that works.
THE KEY FRAMEWORK — The Production Learnings
Architecture > model | integration is the bottleneck | tool-schema quality matters |
context engineering > prompt engineering (more context can hurt) |
observability before production | irreversibility in code | start narrow
THE DECISION RULE - Treat tool-schema design as architecture (it drives tool selection) - Manage context deliberately; don't dump full history into every call - Enforce irreversibility limits in code, never in the prompt - Start with 3 tools done well, expand incrementally with evals
RED FLAGS - Full conversation history passed to every call (attention competition) - Irreversibility "handled" by a careful-sounding prompt instruction - An agent with 15+ tools shipped before any narrow version proved out
THE ONE QUESTION "Is this agent's scope narrow enough to debug, its context lean enough to reason, and its dangerous actions blocked in code — or am I shipping the 86% that fail?"
---¶
RESPONSIBLE AI¶
CARD 33 — Responsible AI in Practice¶
THE ESSENCE "Can we build it?" and "Should we build it?" are different questions. Responsible AI is the discipline of asking the second one — and having the courage to act on the answer.
THE CORE INSIGHT Responsible AI fails when it becomes compliance theater — a checklist signed off before shipping, never revisited. The discipline is four things in practice: a structured "should we build this" screen before any design begins; a deliberate choice of fairness definition (made explicitly, not by default); architectural design-in of transparency and human autonomy; and a responsible AI review conducted with enough authority to stop or change a project. The hard cases — the compliant-but-harmful system, the beneficial-but-biased one, the thing you could build but shouldn't — don't yield to a policy. They yield to a practice and the courage to hold it under deadline pressure.
THE KEY FRAMEWORK — Screen + Four Harm Types + Four Fairness Definitions
SHOULD-WE-BUILD SCREEN (run before design):
Who benefits? Who could be harmed — when it works as designed?
Does benefit accrue to the same people bearing the risk?
Is there a less risky design? Who is accountable if it causes harm?
FOUR HARM TYPES:
Direct (system causes harm) | Facilitated (enables human harm) |
Representational (degrades/misrepresents groups) |
Allocative (distributes opportunities/resources unfairly)
FOUR FAIRNESS DEFINITIONS (choose one explicitly, document why):
Demographic parity | Equal opportunity |
Calibration | Individual fairness
THE DECISION RULE - Run the should-we-build screen before any technical design begins - Choose a fairness definition deliberately and document the choice — a default is a hidden decision someone else made - Design transparency and human autonomy in from the start, not bolted on for compliance - Conduct a responsible AI review before any consequential system ships — and give it actual authority
RED FLAGS - "It's legal" treated as sufficient ("legal" is the floor, not the ceiling) - Aggregate accuracy hiding concentrated harm on a minority group - Responsible AI review skipped under deadline pressure — the moment it's most needed - A "should we build this" conversation that has never actually happened
THE ONE QUESTION "Who could be harmed by this system working exactly as designed — and does that person have a voice in the decision about whether to build it?"
---¶
Continued in Part 4 (Cards 34-37: Fine-Tuning, Multimodal, GraphRAG, Model Portability).