THE AI ARCHITECT'S FIELD CARDS — PART 2¶
Cards 12-22¶
Part 2 of 4. Covers Act V (Production & Platforms, Cards 12-17) and Act VI (Landscape, Value & Direction, Cards 18-22). Part 1 covers Acts I-IV (Cards 1-11). Part 3 covers the Practical Layer, Cloud, and Responsible AI (Cards 23-33). Part 4 covers the Advanced Topics (Cards 34-37). Each card follows the same six-section structure: Essence → Core Insight → Key Framework → Decision Rule → Red Flags → The One Question.
---¶
ACT V — PRODUCTION & PLATFORMS¶
CARD 12 — Observability & Evals¶
THE ESSENCE If you can't measure AI quality with a number, you can't improve it, can't catch regressions, and can't tell a CFO it works. Evals are the test suite for probabilistic systems.
THE CORE INSIGHT Traditional APM tells you the service is up. It doesn't tell you the AI is giving wrong answers with HTTP 200. AI observability adds quality as a first-class signal: offline evals that gate deployment, online evals that sample production, and the feedback loop between them. Without evals, every change is a gamble.
THE KEY FRAMEWORK — Three Signals + Eval Loop
Operational (is it up?) + Quality (is it right?) + Cost (what's it spending?)
Offline evals gate deploys → online evals sample production → failures become
new eval cases → loop. LLM-as-judge must be calibrated before trusted.
THE DECISION RULE - Build the eval suite before production, not after the first incident - Evals gate deployment — failing evals block promotion - Calibrate LLM-as-judge against human judgment before relying on it - Every production incident becomes a new eval case
RED FLAGS - "Quality looks good" with no number behind it - No eval suite; changes shipped on hope - LLM-as-judge scores trusted without calibration
THE ONE QUESTION "What's my faithfulness number, and would my eval suite catch it if this change made the system worse?"
---¶
CARD 13 — Cost Engineering¶
THE ESSENCE AI cost is an architecture problem, not a procurement problem. The biggest savings come from routing, caching, and context discipline — not from negotiating rates.
THE CORE INSIGHT Output tokens cost roughly 5-6x input on closed frontier models, cached input costs about a tenth of uncached input, agents can loop into runaway spend, and the same task can cost 50x more on the wrong model. Cost is designed in: route by complexity, cache aggressively (prompt + semantic), manage context length, and enforce ceilings in code. 2026 added new traps: introductory prices with step-up dates, long-context surcharges on the whole request, and tokenizer changes between versions. The pricing table changes monthly (Appendix G); the engineering principles don't.
THE KEY FRAMEWORK — The Cost Levers
Model routing (cheapest model that meets the bar) + prompt caching +
semantic caching + context management + batch API + per-task cost ceiling
Build-vs-host break-even = f(volume, GPU cost, ops overhead)
THE DECISION RULE - Route by task; don't pay frontier prices for classification - Cache prompt prefixes and semantically-similar queries - Enforce per-task and monthly cost ceilings in code, not policy - Tag every call for cost attribution (team, feature) - Measure cost per completed task (retries and repairs included), not price per token
RED FLAGS - No cost ceiling — an agent loop can run up an unbounded bill - Frontier model used for every task regardless of complexity - You can't attribute cost to a team or feature - Budget built on an introductory price with no note of when it steps up
THE ONE QUESTION "What's the most this system could cost if every safeguard failed, and is that bounded in code?"
---¶
CARD 14 — AI Infrastructure¶
THE ESSENCE The build-vs-host decision is driven by data sensitivity and volume, not by preference. Above a volume threshold with sensitive data, self-hosting wins on both cost and compliance.
THE CORE INSIGHT Self-hosting open-weight models on vLLM gives you data residency, cost control at scale, and no deprecation risk — at the price of operational complexity and GPU spend. The decision is quantitative: model the break-even between API calls and GPU infrastructure for your actual volume.
THE KEY FRAMEWORK — The Deployment Spectrum
API (managed) → low ops, data leaves perimeter, deprecation risk
Cloud VPC host → data in your cloud, moderate ops
On-prem/air-gap → full control, highest ops, sovereignty/compliance
vLLM is the production serving standard for self-hosting.
THE DECISION RULE - Decide by data sensitivity + volume, then model the break-even - Use vLLM for production self-hosted serving (PagedAttention, batching) - Quantize (AWQ/GPTQ/FP8) to fit hardware and cut cost - Match deployment tier to the data classification
RED FLAGS - Self-hosting at low volume (API would be cheaper and simpler) - Sending sovereignty-restricted data to an external API - No autoscaling plan for variable inference load
THE ONE QUESTION "At my actual volume and data sensitivity, does the break-even math favor API or self-hosting — or am I deciding by habit?"
---¶
CARD 15 — AI Coding Assistants¶
THE ESSENCE AI coding tools deliver real 30-55% productivity gains and ship code with a ~29% vulnerability rate. Both are true. Governance is what lets you keep the gain without the breach.
THE CORE INSIGHT The productivity is real and the security risk is real. AI-generated code has high vulnerability rates and meaningful secret-leakage risk. The architectural response is mandatory human security review of AI-generated code, IP-aware tool configuration, and treating these tools as powerful but supervised.
THE KEY FRAMEWORK — Gain vs. Risk
GAIN: faster boilerplate, tests, refactoring, code comprehension
RISK: ~29% vuln rate, secret leakage, IP exposure, skill atrophy
CONTROL: mandatory security review of AI code + enterprise data agreements
THE DECISION RULE - AI-generated code goes through the same (or stricter) security review - Configure enterprise tools with proper data/IP agreements - Match the tool to the engineer (juniors need more guardrails) - Never auto-merge AI-generated code without review
RED FLAGS - AI-generated code skipping security review because "the AI wrote it" - Consumer coding tools with no enterprise data agreement - No policy on what the AI tool can access (secrets, proprietary code)
THE ONE QUESTION "Does AI-generated code get reviewed at least as carefully as human-written code, or are we trusting it more because it looks confident?"
---¶
CARD 16 — Integration Patterns & AI-Native Design¶
THE ESSENCE AI-augmented adds AI to an existing workflow; AI-native rebuilds the workflow around AI. Know which you're doing — the architectures are fundamentally different.
THE CORE INSIGHT Bolting AI onto a legacy system (augmented) and designing a system that assumes AI (native) require different patterns. Augmentation uses the strangler fig and sidecar patterns to add AI safely; native design assumes probabilistic outputs, builds in confidence/provenance from the start, and uses feature flags and content versioning as first-class concerns.
THE KEY FRAMEWORK — Augmented vs. Native
AUGMENTED: strangler fig, AI sidecar, progressive enhancement,
AI fails → fall back to existing system
NATIVE: confidence + provenance in every response, streaming,
feature-flagged AI, stored-output versioning, no "old way" fallback
THE DECISION RULE - Decide augmented vs. native explicitly and document it - Augmented systems need a working non-AI fallback - Native systems need confidence/provenance designed in from day one - Use feature flags so AI can be dialed up/down without deployment - Put typed-decision models (calibrated probabilities over a fixed answer set) at bounded decision points such as routers, confidence gates and agent step guards, with a three-zone cascade (act / LLM / human) and thresholds set on your own data
RED FLAGS - Treating an AI-native build like an augmentation (no fallback narrative) - AI responses with no confidence or provenance signal - AI behavior changeable only via deployment, not feature flag - Trusting a typed-decision model's probabilities without measuring calibration on your own traffic, per segment
THE ONE QUESTION "Am I adding AI to this workflow or rebuilding the workflow around AI — and does my architecture match that answer?"
---¶
CARD 17 — Platform Engineering for AI¶
THE ESSENCE Every team solving LLM gateway, prompt management, and evals from scratch is waste. A platform turns that shared infrastructure into a paved road — and teams without one feel it acutely.
THE CORE INSIGHT The AI platform team provides the gateway, prompt management, eval infrastructure, and a golden path so product teams focus on their use case, not on plumbing. Platform-as-product: the platform succeeds when teams adopt it because it's easier than DIY, not because it's mandated.
THE KEY FRAMEWORK — The AI Platform Stack
LLM Gateway (routing, cost, PII, audit) + Prompt Management +
Eval Infrastructure + Golden Path Template + Developer Portal
Measured by adoption and developer experience, not by features shipped.
THE DECISION RULE - Build the gateway first — it's the enforcement point for cost, PII, audit - Provide a golden path so the easy way is the governed way - Treat the platform as a product with internal customers - Measure adoption and DX, not feature count
RED FLAGS - Beautiful platform nobody uses (the Platform-First/Features-Never trap) - Every team building their own gateway and eval setup - Platform mandated but harder than DIY
THE ONE QUESTION "Do product teams use the platform because it's easier than building their own — or only because they're told to?"
---¶
ACT VI — LANDSCAPE, VALUE & DIRECTION¶
CARD 18 — AI Startup Landscape¶
THE ESSENCE The model isn't the moat. Every winning AI company is defended by proprietary data, deep workflow embedding, distribution, or switching costs — and those are also your build-vs-buy signals.
THE CORE INSIGHT The startups winning in 2026 don't build models — they build vertical applications on top of them, defended by domain data and workflow integration that take years to replicate. For an architect, a vendor's moat is exactly what makes buying better than building: if they have years of domain data and deep integrations, buying gives you their moat.
THE KEY FRAMEWORK — The Four Moats
Proprietary data (from usage) + workflow embedding (habit + integration) +
distribution/brand + switching costs (user state)
These four = the reasons a startup wins AND the reasons to buy vs. build.
THE DECISION RULE - Assess any AI capability against the four moats before building it - If a vendor has a real domain-data moat, buying gives you their head start - If the only "moat" is model access, you can build it — so maybe do - Track the vertical-AI leaders in your industry as competitive intelligence
RED FLAGS - Building a capability a funded startup will sell to your competitors cheaper - Buying a thin model-wrapper at a premium you could replicate in weeks - No view of the startup landscape in your domain
THE ONE QUESTION "For what I'm about to build — what's the moat, and could a funded startup out-build me and sell it to everyone?"
---¶
CARD 19 — Business Value Framework¶
THE ESSENCE 95% of AI pilots deliver no measurable P&L impact — almost always because nobody defined the business metric, the data wasn't ready, or no one owned the outcome. Not because the AI failed.
THE CORE INSIGHT The gap between "AI is productive" and "AI improved the P&L" is the central problem. Pilots that scale define the business metric before starting, plan the data integration as real work, assign a business owner accountable for the outcome, and measure against a baseline. The highest ROI is in boring back-office automation, not glamorous customer-facing AI.
THE KEY FRAMEWORK — The Five-Element Business Case
Cost baseline (specific) + falsifiable impact hypothesis +
total cost (incl. data + change mgmt) + ROI calc (3 scenarios) +
named business owner. Missing any one = incomplete case.
THE DECISION RULE - Define the business metric before the pilot, not after - Total cost includes data engineering, change management, governance - Name a business owner accountable for the outcome metric - Look for ROI in back-office automation before customer-facing flash
RED FLAGS - Pilot measured by "users like it," not a business number - No baseline, so you can't prove the AI changed anything - "Owned by the AI team" — nobody accountable for the business outcome
THE ONE QUESTION "What's the cost baseline, the target metric, and the name of the person accountable for moving it?"
---¶
CARD 20 — Hype vs. Reality¶
THE ESSENCE You're the only person in the room whose job is to be right, not interesting. The discipline: separate capability (can do once) from reliability (does consistently in production).
THE CORE INSIGHT Demos show capability; production requires reliability. Agents work in narrow, governed, human-supervised domains — not as autonomous general workers (58% need override 20-40% of the time). The architect's value is placing each capability correctly: reliable today, overhyped-but-coming, or still research — and not letting an exciting demo skip that step.
THE KEY FRAMEWORK — The Credibility Tiers
RELIABLE TODAY: document extraction, classification, RAG, code assist
OVERHYPED-BUT-COMING: autonomous agents, multi-agent at scale, AI strategy
STILL RESEARCH: AGI, self-improvement, causal reasoning
Capability ≠ reliability. Demos prove the first, not the second.
THE DECISION RULE - Ask for production override rates, not demo accuracy - Place each capability in the right tier before committing - Build near-term on proven tier; treat the rest as bets, labeled - Be the calm voice that asks "what happens when it's wrong?"
RED FLAGS - Planning production on a capability you've only seen demoed - Accepting "autonomous" without an override-rate number - A vendor who can't describe their failure modes
THE ONE QUESTION "Is this capability reliable in production conditions like mine — or have I only seen it be capable in a demo?"
---¶
CARD 21 — Emerging Patterns (18-Month Horizon)¶
THE ESSENCE The model is commoditizing; competition is moving up to orchestration, knowledge, and evaluation. Build model-agnostic infrastructure now so you're positioned as the layer shifts.
THE CORE INSIGHT Over the next 18 months: model capability becomes interchangeable, long context and RAG coevolve (not replace), edge AI matures, MCP becomes as standard as REST, and memory governance becomes a real concern. The architect's move is to invest in the durable layer (orchestration, knowledge, evals) rather than model-specific optimizations. One caveat: capability is commoditizing, but integration is not. API surfaces, behavior, and commercial access still differ, so interchangeability has to be engineered (Card 37).
THE KEY FRAMEWORK — The Horizon Tiers
PLAN NOW: model commodity, MCP standardization, platform investment,
bounded human-gated computer use
PROTOTYPE: hybrid context, reasoning pipelines, voice, agentic commerce controls
MONITOR: unattended computer use for consequential work, agent-commerce
liability rules, physical AI, sovereign AI, A2A inter-org, agentic SRE
THE DECISION RULE - Invest in model-agnostic infrastructure, not model-specific tricks - Standardize tool integration on MCP before the ecosystem expands - Design memory governance alongside persistent-memory features - Pre-identify use cases for capabilities maturing in 12-18 months - Prefer API → MCP → RPA → computer use, in that order; screen automation is the last resort, not the first - Before agents spend money, require signed user mandates and spend limits (agent payment protocols such as AP2 and ACP)
RED FLAGS - Your AI moat is "access to a good model" (that's commoditizing) - Model-specific optimizations that break on a model switch - Assuming "commodity" means "swappable today" without ever having tested a swap - No plan for the shift from model-choice to orchestration competition
THE ONE QUESTION "When the model stops being a differentiator in 18 months, what is my organization's actual AI advantage?"
---¶
CARD 22 — The Practicum (How, Not Just What)¶
THE ESSENCE Knowing the architecture isn't the skill — building it is. Eval suites, prompt iteration, RAG debugging, red-team interpretation, agent state machines, MCP servers: muscle memory, not theory.
THE CORE INSIGHT The gap between understanding RAG and debugging a retrieval failure at midnight is practice. This is the hands-on layer: write the LLM-as-judge prompt, run RAGAS and interpret the score, iterate a prompt one-failure-at-a-time, diagnose RAG failures in sequence, interpret a Garak scan, translate a workflow into a state machine, build an MCP server.
THE KEY FRAMEWORK — The Practitioner Loops
Eval: define-good → 20 cases → judge prompt → run → interpret → gate
Prompt: minimal → test → fix ONE failure → retest → repeat
RAG debug: content exists? → retrieved? → used? → ignored? (in order)
THE DECISION RULE - Define "good" before running anything (avoid confirmation bias) - Fix one failure per prompt iteration (so you know what worked) - Diagnose RAG failures in the fixed sequence, not by guessing - Every security finding becomes a regression eval case
RED FLAGS - You can describe evals but have never written one - Changing multiple prompt things at once, can't tell what helped - "Run Garak" with no idea how to interpret the output
THE ONE QUESTION "Could I sit down right now and actually build this — or do I only understand it well enough to talk about it?"