MODULE 37 — Model Portability: Engineering the Model as a Replaceable Part¶
⚠️ Currency note: The model names, API behaviors, and case-study dates in this module are accurate as of October 2026. They illustrate the patterns and are not recommendations. Model names and prices change quarterly, so verify them against Appendix G — Current Landscape and the provider links in Module 2 §2.1. The architecture in this module is built to outlast those names.
37.1 The Commodity Thesis and Where It Breaks¶
Module 21 (§21.2) makes the strategic argument: the model is becoming a commodity, and the competitive advantage moves to the systems around it — knowledge, workflow, evaluation, and orchestration. Module 2 (§2.5) lists the lock-in vectors. Module 17 (§17.3) makes the LLM gateway the platform's core infrastructure.
This module is about the gap between those three ideas and production reality. The model is commoditizing. Your integration with it is not.
Two different things get called "the model is a commodity":
| What is commoditizing | What is not commoditizing |
|---|---|
| Capability at the "good enough for most enterprise tasks" tier — several providers clear it | The API surface — parameters, tool calling, structured output, reasoning controls |
| Price per unit of capability — falling every quarter | Behavior — verbosity, refusals, instruction-following strictness, formatting |
| Access — multiple clouds now host the same models (e.g., OpenAI and Anthropic models are both on AWS Bedrock as of 2026) | Operational envelope — rate limits, latency, regions, retention requirements, caching |
| The idea of switching — gateways make the wire call look the same | Commercial access — a provider can and does withdraw access for business reasons |
A commodity in the economic sense is fungible: one barrel of oil is the same as another. Models are not fungible in that sense. Electricity is the closer analogy. It is a commodity, but your building still needs a transformer, a breaker panel, and a tested generator before you can switch supply without the lights going out.
The architect's framing:
The model is a replaceable part — but only behind a contract you own. The contract is: a capability interface your application codes against, a model profile that maps each model onto that interface, a behavioral test suite that defines "equivalent," and a fallback you have actually exercised. With those four things, switching models is a configuration change you have already tested. Without them, every model change is a project — and an unplanned one is an incident.
The rest of this module builds those four things.
37.2 Why Model Swaps Fail: The Ten Incompatibility Layers¶
Teams that "switched models in an afternoon" usually swapped the transport and found the other problems in production weeks later. A swap can break at ten distinct layers. A gateway such as LiteLLM, Portkey, or OpenRouter handles only the first one well.
| # | Layer | What breaks on a swap | Real example (as of Oct 2026) | Mitigation |
|---|---|---|---|---|
| 1 | Transport / wire format | Endpoint, auth, request envelope | Every provider's SDK differs | Gateway with an OpenAI-compatible or normalized schema |
| 2 | Request parameters | Parameters accepted by one model are rejected (HTTP 400) by another — including newer models from the same vendor | Claude Opus 5.5 rejects thinking: {type: "disabled"} and budget_tokens with a 400; Claude Opus 4.7+ rejects non-default sampling parameters (temperature, top_p) |
Model profile declares supported parameters; the adapter translates or drops them |
| 3 | Tool calling | Forced tool selection, parallel calls, argument escaping | Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1 return 400 on forced tool_choice (any / specific tool); newer Claude models may escape JSON in tool arguments differently |
Never string-match tool arguments; parse JSON; use auto plus strict schemas; adapter handles forcing semantics |
| 4 | Structured output | Native JSON-schema mode vs. tool-call workaround vs. prompt-only | Assistant-message prefill (a common trick for forcing JSON) returns 400 on Claude 4.6+ models | Use each provider's native structured-output feature through the adapter; validate every response against the schema anyway |
| 5 | Reasoning controls | Token budgets vs. effort levels vs. thinking levels; different defaults | Claude: effort from low to max, and Opus 5.5 defaults to medium where Opus 5 defaulted to high. Gemini 3.8 Flash: thinking levels low/medium/high. OpenAI: reasoning-effort settings |
Set reasoning depth explicitly per capability in the profile; never rely on a default |
| 6 | Context and output limits | A prompt that fits one model overflows another; long answers get truncated | Claude Opus 5.5: 1M context / 128K output. Gemini 3.8 Flash: 1M context / 64K output. GPT-6 bills requests over 272K input tokens at a surcharge | Profile declares limits; the request builder enforces your design ceiling, not the model's maximum |
| 7 | Tokenizer and cost accounting | The same text is a different number of tokens, so cost, truncation, and budgets shift | Anthropic documents that the tokenizer introduced with Opus 4.7 can use ~1.0–1.35× as many tokens as earlier models for the same text | Re-baseline token counts per model; budget in cost per completed task, not tokens |
| 8 | Behavior | Verbosity, tone, refusal patterns, how literally instructions are followed | Anthropic's migration guidance notes that prompts written for prior models are often "too prescriptive" for newer ones and reduce output quality; newer models add new safety-refusal categories | Behavioral contract tests (§37.5); per-model prompt overlays (§37.4) |
| 9 | Caching and state | Prompt caches are model-scoped; reasoning state may be model-bound | Switching models forfeits cache hits. Newer Claude models bind thinking blocks to the producing model, so other models drop them on replay | Plan cache-cold cost into failover; switch at conversation boundaries (§37.7) |
| 10 | Operational, legal, and commercial | Rate limits, regions, data-retention terms, and access itself | Claude Fable 5.1 requires 30-day data retention and is unavailable under zero-data-retention unless Anthropic expressly authorizes it, so it can be a compliance blocker. OpenAI is ending Cursor's model access after SpaceX's acquisition (§37.11) | Eligibility rules in the profile; contract clauses; a pre-provisioned hot fallback |
Three lessons from this table:
- Same-vendor upgrades are swaps too. Layers 2, 3, 5 and 8 broke on routine upgrades within one vendor's flagship line in 2026. "We're staying with our provider, so we don't need portability" is false. Model versioning is model swapping.
- Silent breaks are worse than loud ones. A 400 error is loud: you find it in the first test. A default effort level that drops from
hightomediumis silent. The calls succeed, quality degrades a little, and nobody notices until a customer complaint arrives three weeks later. Layers 5, 7 and 8 are where silent breaks live. - Most of the work is above the gateway. A gateway normalizes transport. Portability requires normalizing semantics, and that is the adapter layer in §37.3, which you own.
37.3 The Portability Architecture¶
The Reference Architecture¶
PORTABILITY REFERENCE ARCHITECTURE
┌──────────────────────────────────────────────────────────────────────┐
│ APPLICATION CODE │
│ calls capabilities, never models: │
│ triage_ticket(ticket) -> TicketDecision │
│ draft_reply(ticket, context) -> Draft │
│ summarize_contract(doc) -> ContractSummary │
└───────────────────────────────┬──────────────────────────────────────┘
│ typed request / typed response
┌───────────────────────────────▼──────────────────────────────────────┐
│ CAPABILITY LAYER (you own this — the portability contract) │
│ ├── Capability registry: name, input/output schema, tier (A/B/C) │
│ ├── Routing policy: primary + ordered fallbacks per capability │
│ ├── Behavioral contract: the tests that define "equivalent" │
│ └── Prompt registry: base prompt + per-model-family overlays │
└───────────────────────────────┬──────────────────────────────────────┘
│ neutral request (messages, tools,
│ schema, reasoning_depth, limits)
┌───────────────────────────────▼──────────────────────────────────────┐
│ MODEL ADAPTER LAYER (you own this — one profile per model) │
│ ├── Parameter translation (reasoning_depth → effort / level / ...) │
│ ├── Feature negotiation (forced tool choice? prefill? native JSON?)│
│ ├── Limit enforcement (design ceilings, not model maximums) │
│ ├── Response normalization (tool calls, refusals, usage, errors) │
│ └── Eligibility checks (data class, region, retention terms) │
└───────────────────────────────┬──────────────────────────────────────┘
│ provider-specific request
┌───────────────────────────────▼──────────────────────────────────────┐
│ LLM GATEWAY (buy or run — Module 17 §17.3) │
│ auth, rate limits, PII scanning, audit logging, retries, │
│ transport-level failover, cost metering │
└──────┬──────────────┬──────────────┬──────────────┬──────────────────┘
▼ ▼ ▼ ▼
Provider A Provider B Cloud-hosted Self-hosted
(direct API) (direct API) (Bedrock / (vLLM / SGLang
Foundry / open-weight)
Agent Platform)
The Principle: Code Against Capabilities, Not Models¶
Most AI codebases look like this:
# ANTI-PATTERN: application code knows about the model
response = llm.chat(
model="provider-x-flagship-2026-05",
messages=[{"role": "system", "content": TRIAGE_PROMPT}, {"role": "user", "content": ticket.text}],
temperature=0,
tool_choice={"type": "tool", "name": "classify"}, # 400 on some newer models
)
category = json.loads(response.tool_calls[0].arguments)["category"]
Every line is coupled to one model: the name, the sampling parameter, the forced tool choice, and the response shape. A swap touches every call site.
The portable version:
# PORTABLE: application code knows about the capability only
decision: TicketDecision = capabilities.triage_ticket(ticket)
That buys four things:
- Swap blast radius goes to near zero. A model change edits one profile and one routing entry. No application call site changes.
- The capability is the unit of evaluation. Tests attach to triage_ticket, not to a prompt or a model, so they survive both kinds of change.
- The capability is the unit of routing. Different capabilities can sit on different providers, which spreads concentration risk without forcing a single "which vendor?" decision.
- The capability is the unit of cost attribution. Module 13's per-feature cost reporting falls out naturally.
A capability is not a thin wrapper around chat(). It owns its input schema, output schema, prompt, validation, retry policy, and routing. Most systems need 5–30 capabilities, not hundreds. If you count more, you have probably mixed up capabilities with prompts.
The Model Profile¶
The model profile is the single file that knows how one specific model behaves. It is the most useful portability artifact you can create, because it turns the ten layers in §37.2 into data instead of tribal knowledge.
# model-profiles/anthropic-claude-opus-5-5.yaml (illustrative, Oct 2026)
id: anthropic/claude-opus-5-5
family: claude-5 # prompt overlays are keyed by family
provider_routes: # same model, several ways to reach it
- anthropic-direct
- aws-bedrock # partner-operated; pricing differs
- gcp-agent-platform
limits:
context_tokens: 1000000
max_output_tokens: 128000
design_ceiling_input: 200000 # OUR ceiling — keeps failover targets viable
features:
native_structured_output: true
strict_tool_schemas: true
forced_tool_choice: false # 'any'/'tool' → 400; adapter uses auto + strict + instruction
assistant_prefill: false # 400 on 4.6+ family
sampling_params: false # temperature/top_p rejected
parallel_tool_calls: true
prompt_caching: { supported: true, scope: model, min_prefix_tokens: model_dependent }
reasoning:
control: effort # low | medium | high | xhigh | max
can_disable: false # thinking cannot be turned off on this model
default: medium # NOTE: differs from claude-opus-5 (high) — always set explicitly
reasoning_state_portable: false # thinking blocks bound to this model
translation: # neutral reasoning_depth → this model's control
reasoning_depth:
minimal: { effort: low }
standard: { effort: medium }
deep: { effort: high }
max: { effort: max }
refusals:
signal: stop_reason == "refusal" # HTTP 200, not an error
categories_exposed: true # stop_details.category
server_side_fallback: available # provider feature; record whether we use it
eligibility:
data_classes_allowed: [public, internal, confidential] # per your policy
zero_data_retention: per_contract
regions: [us, eu] # verify per route
tokenizer_note: "Opus 4.7+ tokenizer; re-baseline counts vs. pre-4.7 models"
pricing_ref: appendix-g#g3 # never hardcode prices here
last_verified: 2026-10-03
owner: ai-platform-team
The design_ceiling_input field is the most underrated line in this profile. If your primary model accepts 1M tokens and your application starts sending 600K-token prompts, every fallback target with a 128K–262K window becomes unusable without notice. Set the ceiling at what your weakest acceptable fallback can handle, and raise it only deliberately, with an ADR.
The Adapter: Feature Negotiation in Practice¶
The adapter turns the neutral request into a provider request by consulting the profile. A sketch of the logic, provider-neutral pseudocode:
def build_request(neutral: NeutralRequest, profile: ModelProfile) -> ProviderRequest:
req = ProviderRequest(model=profile.id)
# 1. Enforce OUR ceiling, not the model's
if count_tokens(neutral, profile) > profile.limits.design_ceiling_input:
raise CapabilityInputTooLarge(neutral.capability) # route to chunking path
# 2. Reasoning depth: always explicit, translated per model
req.reasoning = profile.translation.reasoning_depth[neutral.reasoning_depth]
# 3. Structured output: best available mechanism
if neutral.output_schema:
if profile.features.native_structured_output:
req.output_format = neutral.output_schema
else:
req.tools = [schema_as_tool(neutral.output_schema)]
req.instructions += "Respond only by calling the provided tool."
# 4. Forced tool selection: emulate where unsupported
if neutral.require_tool:
if profile.features.forced_tool_choice:
req.tool_choice = {"type": "tool", "name": neutral.require_tool}
else:
req.tool_choice = {"type": "auto"}
mark_strict(req.tools)
req.instructions += f"You must call the `{neutral.require_tool}` tool."
# 5. Drop parameters the model rejects (log it — silent drops hide intent)
if neutral.temperature is not None and not profile.features.sampling_params:
log.info("dropping temperature for %s", profile.id)
return req
def normalize_response(raw, profile) -> NeutralResponse:
# Refusals are not errors on some providers (HTTP 200) — normalize them
# Tool-call arguments: always json.loads, never string-match
# Usage: map provider token fields to (input, cached_input, output, reasoning)
...
Notice what the adapter does not do: it does not hide failures. If a capability needs forced tool selection and the target cannot emulate it reliably, the adapter should refuse to route there. That shows up in the eval harness as a contract failure, not as a mysterious quality drop.
Gateway vs. Adapter: Who Does What¶
| Concern | Gateway | Adapter / capability layer |
|---|---|---|
| Auth, keys, quotas | ✓ | |
| PII scanning, audit logs | ✓ | |
| Transport retries, 5xx failover | ✓ | |
| Cost metering | ✓ | Cost per capability attribution |
| Parameter semantics, feature negotiation | ✓ | |
| Prompt selection per model family | ✓ | |
| Output validation against schema | ✓ | |
| Quality-triggered failover | ✓ (needs the contract) | |
| "Is this model eligible for this data?" | Partial (policy) | ✓ (per capability) |
The gateway is not your portability layer. It is a dependency of it. Treat the gateway as replaceable too: gateway vendors consolidated sharply in 2026. Palo Alto Networks acquired Portkey (closed May 2026), and other gateway and observability startups were reportedly absorbed, frozen, or pivoted. Keep your capability and adapter layer independent of any one gateway product's proprietary configuration format.
37.4 Prompt Portability¶
Prompts are where portability quietly fails. Every prompt is tuned, deliberately or not, to the model it was developed against.
Three Strategies¶
| Strategy | How it works | Pros | Cons | Use when |
|---|---|---|---|---|
| Lowest common denominator | One prompt, written to work acceptably on every candidate | Simplest; one artifact to govern | Leaves quality on the table on every model; the "common" set shrinks as models diverge | Tier C capabilities; very simple tasks |
| Base + overlays (recommended default) | An invariant base prompt plus a small per-model-family overlay | Policy stays in one place; tuning is isolated and reviewable | Overlays must be evaluated per family | Most Tier A/B capabilities |
| Per-model prompts | Fully separate prompts per model | Maximum quality per model | Policy drift between copies; N× governance cost | Very high-value capabilities with a dedicated owner |
The Invariant / Tuned Split¶
Split every production prompt into what must stay the same on any model and what may vary:
PROMPT LAYERING FOR PORTABILITY
INVARIANT BASE (identical across models; changes go through policy review)
├── Role and scope boundaries ("You triage support tickets for ...")
├── Business rules and policy ("Never promise refunds; escalate if ...")
├── Output contract (schema, required fields, enumerations)
├── Safety and injection-resistance rules (Module 3 §3.5–3.6)
└── Domain knowledge and definitions
MODEL-FAMILY OVERLAY (small; tuned per family; changes go through eval)
├── Emphasis calibration (newer models over-apply ALL-CAPS / "MUST")
├── Verbosity guidance ("be concise" means different things per family)
├── Reasoning guidance (none for reasoning-native models; brief for others)
├── Formatting quirks (markdown vs. plain text tendencies)
└── Known failure workarounds (documented, with the eval case that proves them)
Rules for overlays: - An overlay never contains policy. If an overlay says "don't discuss pricing", that rule belongs in the base, or one model family will quietly lack it. - Every overlay line cites a reason. Either the eval case it fixes or the migration note it implements. An overlay line without a reason gets deleted at the next review. - Keep overlays small. If an overlay grows past ~15–20% of the base, you have drifted into per-model prompts by accident. Decide whether that was intended.
Model-Specific Constructs to Remove from Base Prompts¶
These constructs tie a prompt to one model or generation. Grep for them before a swap:
| Construct | Why it is not portable |
|---|---|
Provider-specific tags (<thinking>, <scratchpad>, special tokens) |
Meaningless or harmful elsewhere; some models leak or mimic them |
| "Think step by step" on reasoning-native models | Redundant with built-in reasoning; can lengthen output without improving quality |
| Assistant-message prefill to force a format | Rejected outright by some newer models (400) |
| Heavy emphasis written for weaker instruction-followers ("CRITICAL!!! YOU MUST ALWAYS...") | Newer models follow instructions more literally and over-trigger |
| References to a model's own name or version | Wrong after the swap |
| Few-shot examples copied from one model's own outputs | Teaches the next model to imitate the previous model's quirks |
| Instructions relying on a specific default (temperature, reasoning depth) | Defaults differ and change between versions |
The Prompt Registry, Keyed for Portability¶
A Module 3 prompt registry entry should be keyed by (capability, prompt_version, model_family). Each (prompt_version, model_family) pair needs an eval result before it can serve production traffic. The registry should refuse to route a capability to a model family that has no evaluated overlay. This one rule prevents the most common portability incident: the fallback model ran a prompt nobody had ever tested on it.
37.5 Behavioral Contracts: Defining "Equivalent"¶
You cannot swap a part without a specification of what the part must do. Public benchmarks are not that specification. They measure general capability on someone else's tasks. A behavioral contract defines equivalence for your capability.
The Four Tiers of a Contract¶
| Tier | What it covers | Threshold style | Example (ticket triage) |
|---|---|---|---|
| 1. Hard invariants | Things that must never break | Absolute — any failure blocks | 100% schema-valid output; 0 PII echoed; 100% of injection cases stay in scope |
| 2. Quality metrics | Task accuracy and faithfulness | Relative to baseline, with tolerance | Category accuracy ≥ baseline − 2 pts; "needs human" recall ≥ 95% |
| 3. Operational | Latency, cost, reliability | Absolute SLOs | P95 latency ≤ 2.5 s; cost per completed triage ≤ $0.004; error + refusal rate ≤ 0.5% |
| 4. Style | Tone, length, formatting | Soft — reviewed, not gating | Median rationale length within ±30% of baseline |
The asymmetry is deliberate. Tier 1 is binary and blocks a swap. Tier 2 compares against the current model, so a swap is judged on whether it is equivalent or better, not on whether it is perfect. Tier 4 is reviewed by a human and does not gate the decision. A contract that gates on style will block every swap.
A Contract File¶
# contracts/triage_ticket.yaml
capability: triage_ticket
tier: A # business-critical: hot fallback required
dataset:
path: evals/triage/v7.jsonl
size: 640 # see sample-size guidance below
composition: { production_sample: 480, edge_cases: 100, adversarial: 60 }
provenance: "prod sample 2026-08, PII-scrubbed; edge cases from incident log"
hard_invariants:
- schema_valid: 1.00
- pii_echo_rate: 0.00
- adversarial_in_scope: 1.00
quality: # relative to current primary
- metric: category_accuracy
tolerance: -0.02
- metric: needs_human_recall
min_absolute: 0.95 # absolute floor regardless of baseline
- metric: urgency_mae
tolerance: +0.15
operational:
- p95_latency_ms: 2500
- cost_per_completed_task_usd: 0.004 # includes retries and repair calls
- error_plus_refusal_rate: 0.005
style: # reported, not gating
- rationale_length_median_delta: 0.30
grader:
deterministic: [schema_valid, category_accuracy, urgency_mae, pii_echo_rate]
llm_judge:
metrics: [adversarial_in_scope]
judge_model: pinned # see §37.6 — never the candidate itself
Sample Size: How Many Cases Is Enough?¶
Teams routinely decide a swap on 30 examples. That is not enough to detect the differences that matter.
- Rule of thumb for one accuracy number: the 95% margin of error is about ±1.96 × √(p(1−p)/n). At an accuracy near 85% with n = 100, the margin is about ±7 points, which is useless for a "−2 points" tolerance. At n = 600 it is about ±2.9 points.
- Use paired comparison. Both models answer the same cases, so compare them case by case and not as two separate averages. Count the discordant cases: b = the baseline was right and the candidate wrong; c = the candidate was right and the baseline wrong. McNemar's test on b and c is far more sensitive than comparing two accuracies, because the cases both models get right or both get wrong carry no information about the difference. Statistic: (|b − c| − 1)² ⁄ (b + c), compared against a χ² distribution with 1 degree of freedom (3.84 at 95%).
- Practical sizing:
- Smoke test (does it work at all?): 30–50 cases.
- Tier C swap decision: 150–300.
- Tier A swap decision near a tolerance boundary: 500–1,000, plus shadow traffic (§37.8).
- Stratify. Report results per segment (ticket category, language, customer tier). An aggregate that holds steady can hide a 15-point regression on one language that is 8% of traffic.
Shared APIs Do Not Mean Equivalent Models¶
Typed-decision models show the point clearly. As of October 2026, Cloudflare's Clef is reported to be API-compatible with TypeSafe's Jev, which makes switching easy to wire. Early comparisons nevertheless report different strengths (for example, better at routing but weaker at judgment) (reported), and each model has its own calibration profile. For these models the contract must include calibration metrics (per-segment calibration error, coverage at the target precision) as well as accuracy. See Module 16 §16.9.
Measure Cost per Completed Task, Not per Token¶
A cheaper model that needs more repair calls, retries, or agent turns to finish the job is not cheaper. Define completed per capability: schema-valid, passed validation, no human repair needed. Then compute:
cost_per_completed_task = (Σ cost of all calls incl. retries, repairs, judge checks)
÷ (number of tasks completed without human repair)
This is the operational-tier cost metric in every contract. It routinely reverses "cheaper model" decisions, in both directions.
37.6 The Cross-Model Eval Harness¶
The contract defines equivalence. The harness measures it, continuously, against candidates you are not yet using.
Harness Components¶
CROSS-MODEL EVAL HARNESS
┌────────────────┐ ┌───────────────────┐ ┌──────────────────────┐
│ Frozen datasets │──►│ Runner │──►│ Graders │
│ per capability │ │ capability × │ │ 1. deterministic │
│ (versioned, │ │ model profile × │ │ (schema, exact, │
│ provenance, │ │ prompt overlay │ │ regex, numeric) │
│ PII-scrubbed) │ │ via the SAME │ │ 2. LLM judge │
└────────────────┘ │ adapter as prod │ │ (pinned model, │
└───────────────────┘ │ different family │
│ from candidate) │
└──────────┬───────────┘
▼
┌──────────────────────────────────┐
│ Portability report per capability │
│ contract pass/fail · paired stats │
│ · stratified results · cost/task │
└──────────────────────────────────┘
Non-negotiables: - Run through the production adapter. An eval that calls the model directly, bypassing the adapter, tests a different system from the one you will deploy. Most of the swap bugs from §37.2 live in the adapter. - Pin the judge, and never let a model judge itself. LLM judges show self-preference bias: a model tends to rate outputs that resemble its own more highly. Use a judge from a different model family from the candidate. Pin its version. Re-calibrate it against human labels whenever the judge model changes (Module 12 §12.5). - Version datasets like code. Every dataset has a version, a provenance note, and a change log. When a production incident occurs, its case goes into the dataset (Module 3 §3.7). - Prefer replayed production traffic. Synthetic cases miss the real distribution. Replay a PII-scrubbed production sample monthly, and keep the hand-built edge and adversarial cases on top.
The Quarterly Candidate Run¶
This makes Module 28's model-debt guidance operational. Once a quarter, run every Tier A and Tier B capability against: 1. the current primary (the baseline), 2. the configured fallbacks, 3. the next version of the primary model (to catch same-vendor breaks early), 4. one or two new candidates from the Appendix G refresh.
PORTABILITY REPORT — triage_ticket — 2026-Q4 run (illustrative)
Model / overlay Contract Accuracy Δ Recall(needs_human) P95 Cost/task Notes
──────────────────────────────────────────────────────────────────────────────────────────────
primary (vendor A flagship) baseline — 0.962 1.9s $0.0031
fallback1 (vendor B mid-tier) PASS −1.1 pts 0.958 1.4s $0.0019
fallback2 (self-hosted MoE) PASS* −1.8 pts 0.951 2.3s $0.0012 *near tolerance
next-ver (vendor A next) FAIL +0.6 pts 0.966 2.1s $0.0034 T1: 3 schema fails —
forced tool_choice
rejected; adapter
emulation needed
candidate (new open-weight) FAIL −4.9 pts 0.931 1.1s $0.0007 Fails on non-English
segment (−14 pts)
That "next-ver FAIL" row is exactly the result you want in a quarterly run, months before the deprecation deadline, and not during an emergency migration.
Harness Cost Is a Line Item¶
A full candidate run costs real money: capabilities × models × cases × (call + judge). For example, 20 capabilities × 5 models × 600 cases × ~$0.004 ≈ $240 per quarterly run, before judge costs. Budget for it explicitly in Module 13's FinOps model. It is the cheapest insurance in the AI budget.
37.7 Fallback Design: Hot, Warm, and Cold¶
A fallback you have never sent traffic to is a hypothesis, not a fallback.
The Three Readiness Levels¶
| Level | Definition | What exists | Time to take full traffic | Required for |
|---|---|---|---|---|
| Hot | Continuously serving a small share of real traffic | Profile, evaluated overlay, contract PASS, provisioned quota, monitoring, live traffic share | Minutes (config flip) | Tier A capabilities |
| Warm | Ready but idle | Profile, evaluated overlay, contract PASS within the last quarter, account and quota in place | Hours to a day (verify quota, shadow briefly, ramp) | Tier B capabilities |
| Cold | Identified, not prepared | A named candidate; maybe a profile | Days to weeks (full playbook §37.8) | Tier C capabilities |
Keep-Alive Traffic¶
Send a small, steady share of production traffic to the hot fallback, typically 1–5%, and evaluate it with the same online monitors as the primary. It costs very little, and it gives you four things: - proof that the credentials, quota, network path, and adapter actually work today; - a live quality comparison on real traffic, not only on eval sets; - an existing commercial relationship and rate-limit history with the fallback provider, which matters when you suddenly need 20× the quota; - organizational familiarity: on-call engineers have seen the fallback's dashboards before the incident.
The Cursor case (§37.11) shows the value of traffic diversity: when one provider withdrew access, it carried only about 5% of traffic.
Failover Triggers¶
Define triggers per capability, and make each one an explicit, logged decision:
| Trigger | Signal | Typical action |
|---|---|---|
| Hard availability | 5xx rate > X% over 2 min; connection failures | Automatic failover (gateway) |
| Sustained throttling | 429 rate > X% over 5 min after retries | Automatic partial shift (overflow to fallback) |
| Latency SLO breach | P95 > SLO for 10 min | Automatic or on-call decision, per tier |
| Refusal spike | Refusal rate > 3× baseline | On-call decision. Investigate first: a refusal spike is often caused by your input, not by the model |
| Quality monitor drop | Online eval score below the contract floor | On-call decision; never automatic on a single signal |
| Commercial / policy event | Access-termination notice, ToS change, retention change, deprecation | Planned migration via the playbook |
The Degradation Ladder¶
When the primary fails, step down a defined ladder and do not improvise:
DEGRADATION LADDER (per capability)
L0 Primary model normal operation
L1 Hot fallback, same quality tier transparent to users
L2 Lower tier model response flagged "reduced quality";
stricter validation; more human review
L3 Self-hosted open-weight reduced capability, data stays inside
L4 Non-AI fallback rules engine, cached answers, search,
or route to human queue (Module 16 Anti-Pattern 2)
Each rung needs an owner, a tested runbook entry (Module 24 §24.2), and a defined user experience (Module 16 Principle 5, "Design the Failure Narrative").
Three Failover Pitfalls Teams Discover Too Late¶
1. The cache-cold cost spike. Prompt caches are scoped to a model. Consider a capability that sends a 50,000-token cached system prompt and knowledge prefix on every request, at 1M requests a month. Using illustrative prices of $2.00 per million input tokens uncached and $0.20 cached, the prefix costs about $10,000 a month while cached and $100,000 a month uncached. On failover the fallback starts cold. If it takes a large share of traffic, the input bill for that period can jump by up to 10×, on top of cache-write premiums. Mitigations: keep-alive traffic keeps the fallback's cache warm; pre-warm on failover; budget a failover cost reserve; keep cacheable prefixes lean.
2. The cascade. The primary fails, all traffic moves to the fallback, and the fallback's quota, sized for 5% keep-alive, throttles within minutes. Now two providers are failing. Mitigations: provision fallback quota for your realistic failover share (contracted or reserved capacity for Tier A); shift traffic progressively; shed load on Tier C capabilities first.
3. The mid-conversation switch. Long conversations and agent loops carry provider-specific state: reasoning blocks bound to the producing model, tool-call ID formats, cached prefixes, and provider-specific content block types. Switching mid-session can drop reasoning context or produce malformed history. Mitigations: - prefer switching at conversation or task boundaries: new sessions go to the fallback while in-flight sessions drain on the primary where possible; - store conversation history in a provider-neutral canonical form (your own message schema) and render it per provider, never storing the provider's raw format as the source of truth; - on a forced mid-session switch, re-render history from the canonical form, drop provider-specific blocks, and add a short neutral summary of prior reasoning if the task depends on it; - for agents, checkpoint task state at step boundaries (Module 6) so a different model can resume from state, not from a transcript.
37.8 The Model Swap Playbook¶
The same playbook applies whether the swap is planned (cost, quality, deprecation) or forced (access revocation, outage, compliance change). Only the time pressure differs. Architecture done well makes the emergency path a compressed version of the planned path, not a different one.
Phase 0 — Trigger and Classify¶
| Trigger class | Example | Typical deadline | Urgency |
|---|---|---|---|
| Deprecation | Pinned version retirement notice | 60–180 days | Planned |
| Cost / quality opportunity | New model passes contract at 40% lower cost per task | None | Planned |
| Same-vendor upgrade | The next flagship version, with breaking API changes | Self-imposed | Planned |
| Compliance / terms change | Provider changes data-retention terms; new regulation | 30–90 days | Elevated |
| Access revocation | Provider ends access (Windsurf 2025: under a week; Cursor 2026: ~10 weeks) | Days to weeks | Emergency |
| Sustained outage / quality collapse | Multi-day degradation | Hours | Emergency |
Phases 1–8¶
MODEL SWAP PLAYBOOK
PHASE 1 INVENTORY planned: 1–2 days · emergency: hours
Source: gateway logs + capability registry (never "ask the teams")
Output: every capability on the outgoing model, traffic share, tier,
provider-specific features used, prompt overlays, data classes
Exit: complete list signed off by platform owner
PHASE 2 FEATURE-GAP ANALYSIS planned: 1–3 days · emergency: hours
Diff outgoing profile vs. candidate profile across the 10 layers (§37.2)
Output: gap list — each gap marked: adapter can emulate / needs code /
blocks the swap
Exit: no unmitigated blocking gap for Tier A capabilities
PHASE 3 ADAPTER + OVERLAY WORK planned: 3–10 days · emergency: 1–2 days
Candidate profile, translations, overlays per family; unit tests per gap
Exit: all capabilities run end-to-end on candidate in staging
PHASE 4 OFFLINE CONTRACT EVAL planned: 1–3 days · emergency: hours
Full contract per capability (§37.5), paired stats, stratified
Exit: Tier 1 invariants 100% · Tier 2 within tolerance · Tier 3 SLOs met
PHASE 5 SHADOW planned: 3–7 days · emergency: skip or <24h
Module 24 §24.3: candidate receives mirrored traffic, outputs logged only
Exit: ≥1,000 shadow requests per Tier A capability, monitors green
PHASE 6 CANARY RAMP planned: 1–2 weeks · emergency: 1–2 days
1% → 5% → 25% → 50% → 100%, with automatic rollback triggers:
contract floor breach · error+refusal > threshold · P95 > SLO ·
cost/task > ceiling · human-escalation rate > baseline + X
Exit: 100% for ≥72h with monitors green
PHASE 7 PROMOTE + RETAIN planned and emergency
New model becomes primary; OUTGOING model becomes warm/hot fallback until
its deprecation date (do not delete the profile on day one)
Update: routing config, ADR, runbooks, Appendix-G-style internal registry
PHASE 8 RETRO + SCORE within 2 weeks
Record actual time-to-switch per capability; update portability score (§37.9);
every surprise becomes an eval case or a profile field
Planned swap, well-architected system: about 3–6 weeks of elapsed time, mostly spent waiting on shadow and canary data, not on engineering. Emergency swap, well-architected system: 2–5 business days for Tier A capabilities that already had a warm or hot fallback. Emergency swap, poorly architected system: weeks, plus a production incident, plus quality regressions found by customers.
Worked Example A — A Same-Vendor Upgrade That Breaks (as of Oct 2026)¶
Scenario: ticket triage runs on Claude Opus 5. The team plans to move to Claude Opus 5.5, which is cheaper per token. Everyone treats it as "just a version bump."
Phase 2 finds three gaps:
| Gap | Outgoing (Opus 5) | Candidate (Opus 5.5) | Class | Resolution |
|---|---|---|---|---|
| Thinking disabled for latency | Accepted at high effort or below |
{type: "disabled"} → 400 at every effort |
Loud | Remove; set effort: low explicitly for latency-sensitive capabilities |
| Forced tool call for schema | tool_choice: {type: "tool"} accepted |
400 on forced any/tool |
Loud | Adapter emulation: auto + strict schema + instruction, or native structured output |
| Default reasoning depth | high |
medium |
Silent | Profile sets effort explicitly; the contract eval confirms accuracy at the chosen effort |
The two loud gaps would have failed in the first staging run. The silent one would not. Calls succeed, and on a hard ticket segment accuracy may drift down by a few points. Only the explicit-effort rule in the profile and the stratified contract eval catch it. The adapter change for all three is a profile diff:
# diff: model-profiles/anthropic-claude-opus-5.yaml → anthropic-claude-opus-5-5.yaml
features:
- forced_tool_choice: true
+ forced_tool_choice: false
reasoning:
- can_disable: true # only at effort <= high
- default: high
+ can_disable: false
+ default: medium # we never rely on it — translation sets effort explicitly
No application code changed. That is what portability looks like in practice.
Worked Example B — An Emergency Cross-Vendor Swap (Composite)¶
Scenario (composite, for illustration): a provider sends a notice ending access in 10 days. The organization runs 14 capabilities: 4 Tier A, 6 Tier B, 4 Tier C.
| Day | Tier A (4) — hot fallback existed | Tier B (6) — warm fallback | Tier C (4) — cold |
|---|---|---|---|
| 0 | Inventory from gateway logs (2h). Raise fallback share 5% → 25% | Inventory; confirm quotas with fallback provider | Inventory; pick candidates |
| 1 | Contract eval re-run on current fallback (all PASS). Ramp to 50% | Offline contract eval (5 PASS, 1 FAIL on a long-document capability: input exceeds fallback context) | Write profiles, start overlays |
| 2 | Ramp to 100%. Outgoing provider kept as fallback until cutoff | Failing capability: enforce design ceiling + chunking path; re-eval PASS. Begin canary | Offline eval |
| 3–5 | Done. Monitor cache-cold cost (as predicted, ~3 days elevated) | Canary 5% → 100% | Canary, with "reduced quality" flags allowed for Tier C |
| 6–10 | Retro | Retro | One capability moved to L4 (rules + human queue) for 2 weeks while its prompt was reworked |
Why it took days instead of weeks: Tier A already had hot fallbacks with keep-alive traffic. Every capability had a contract. All history was stored in canonical form. The only real engineering surprise was the context-window gap on one capability, and a profile-enforced design ceiling would have prevented even that.
37.9 Measuring Portability: The Portability Scorecard¶
What you do not measure decays. Portability decays fast: every new provider-specific feature a team adopts, and every prompt tuned on one model, adds switching cost without anyone deciding to.
Core Metrics¶
| Metric | Definition | Target (Tier A) |
|---|---|---|
| Time-to-switch (TTS) | Measured elapsed time to move a capability's full traffic to its fallback while meeting the contract. Measure it in drills, don't estimate it | ≤ 1 day |
| Fallback readiness | % of Tier A/B traffic whose capability has a hot/warm fallback with a contract PASS in the last quarter | 100% (A), ≥ 90% (B) |
| Provider concentration | % of total AI traffic (or spend) on the single largest provider | Policy-defined; many organizations cap Tier A at ~70–80% on one provider |
| Proprietary-feature count | Number of provider-specific features in use without a documented exit path (lock-in ledger, §37.10) | 0 undocumented |
| Overlay coverage | % of (capability × fallback family) pairs with an evaluated overlay | 100% (A) |
| Contract coverage | % of capabilities with a contract file and a dataset of the required size | 100% (A/B) |
| Profile freshness | Days since each profile's last_verified |
≤ 90 |
The Scorecard¶
Score each Tier A/B capability from 0 to 3 on each dimension, and review it quarterly with the architecture board:
| Dimension | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Interface | App code calls a model directly | Wrapper, but model-specific params leak | Capability interface; some leakage | Pure capability interface |
| Profile | None | Informal notes | Profile exists, stale | Profile current (≤ 90 days) |
| Prompt | Single model-tuned prompt | LCD prompt, untested on fallback | Base + overlay, partially evaluated | Base + evaluated overlay per fallback family |
| Contract | None | Ad hoc examples | Contract, undersized dataset | Contract + adequate dataset + stratification |
| Fallback | None | Cold | Warm | Hot with keep-alive |
| State | Provider-raw history | Partially canonical | Canonical history | Canonical + checkpointed agent state |
| Drill | Never | Tabletop only | Staging drill | Production drill (canary-scale) within 6 months |
A capability scoring under 14 of 21 is a portability risk and belongs on the Module 28 debt backlog. A Tier A capability with a 0 anywhere is a finding, regardless of its total.
The Provider Game Day¶
Run a quarterly game day for Tier A: 1. Pick a capability and announce the drill window. 2. At canary scale (say 10% of traffic), block the primary provider at the gateway. 3. Measure: time to detect, time to fail over, contract metrics during the drill, cost delta, and user-visible errors. 4. Restore the primary, then record actual TTS on the scorecard. 5. Every surprise becomes a runbook fix, a profile field, or an eval case.
The first game day nearly always finds something: an expired fallback API key, a quota sized for keep-alive only, or a capability that turns out to call the primary directly and bypass the gateway.
37.10 When Not to Be Portable: Deliberate Lock-In¶
Portability has costs. Pretending otherwise produces over-abstracted systems that use only the lowest common denominator of every model and adopt nothing new.
What Portability Costs¶
- Feature lag. New provider capabilities, such as server-side agent runtimes, computer-use toolsets, advanced caching modes, or server-side refusal fallbacks, reach a portable system later, because each one needs a gap analysis and an exit path.
- Engineering and eval spend. Profiles, overlays, contracts, keep-alive traffic, quarterly runs, and game days are real work and real money.
- Abstraction drag. Heavy frameworks that hide the provider make debugging harder. When a call misbehaves, you need to see the actual request that went out.
When Lock-In Is the Right Call¶
Accept a provider-specific dependency deliberately when all of these hold: 1. The feature delivers a step change, not an incremental gain, for the capability. 2. The exit cost is bounded and known. You have written down what it would take to replicate or replace the feature. 3. The capability's tier tolerates it: Tier C and B usually yes; Tier A only with a documented degraded mode. 4. There is a revisit trigger, such as a date, a price threshold, or a competitor shipping an equivalent.
The Lock-In Ledger¶
Record every deliberate dependency as a short ADR in a single ledger:
LOCK-IN LEDGER ENTRY (ADR-LK-014)
Dependency: Provider X managed agent runtime (hosted sandbox + memory)
Capabilities: research_assistant (Tier B), report_builder (Tier C)
Why accepted: Replaces ~6 weeks of sandbox/infra build; step change in time-to-value
Exit path: Self-hosted harness + container sandbox; est. 4–6 engineer-weeks;
canonical task state already stored outside the runtime
Degraded mode: research_assistant falls to single-call RAG (L2) during exit
Revisit trigger: 2027-Q2 review, or provider price change > 25%, or access/terms change
Owner: platform-architecture
The ledger turns "we're locked in and didn't notice" into "we chose this, we know what leaving costs, and we review it." The proprietary-feature metric in §37.9 counts ledger entries, and undocumented dependencies are the ones that score against you.
Tiering Sets How Much Portability Each Capability Gets¶
| Capability tier | Definition | Portability investment |
|---|---|---|
| A — Business-critical | Revenue- or safety-impacting; customer-facing at scale; regulated | Hot fallback, full contract, base + overlay, game days, canonical state |
| B — Important | Material productivity impact; internal or limited external | Warm fallback, contract, overlays for one fallback family |
| C — Convenience | Low impact; easily paused | Cold fallback acceptable; LCD prompt; deliberate lock-in welcome |
The architect's job is not to make everything portable. It is to make sure the portability of each capability matches its tier, and that the mismatch is visible when it isn't.
37.11 Case Studies¶
Case 1 — Windsurf and Anthropic (June 2025): Access as a Strategic Decision¶
Shortly after reports that OpenAI would acquire the AI coding tool Windsurf, Anthropic told Windsurf it was cutting nearly all of its direct, first-party capacity for Claude 3.x models, with less than a week's notice. Windsurf had to find capacity through other providers and push users toward bring-your-own-key options. Anthropic's leadership framed the decision as strategic: it would be odd to sell Claude to a competitor's subsidiary. The reported OpenAI acquisition later fell through. The access cut came from the expectation of a change of control, not from a completed one.
Lessons: - Access is a business decision, not only a technical SLA. No uptime clause protects you from a provider deciding it no longer wants you as a customer. - Change of control on your side can trigger it. An acquisition, investment, or partnership can turn you into a competitor's asset overnight. - A week is not enough time to build portability. It is only enough time to use portability you already have.
Case 2 — Anthropic and OpenAI (August 2025): Terms-of-Service Enforcement¶
Anthropic revoked OpenAI's API access to Claude models, citing violations of its terms of service (reportedly heavy use of Claude Code by OpenAI technical staff).
Lesson: Terms of service are part of your dependency surface. Usage patterns that breach a provider's terms, such as benchmarking for competitive development or prohibited use cases, can end access abruptly. Your legal review of provider terms (Module 25) is part of your availability architecture.
Case 3 — Cursor and OpenAI (August–November 2026): Diversification Makes It Survivable¶
SpaceX completed its acquisition of Anysphere, the maker of Cursor, on August 14, 2026. On August 29, OpenAI announced it would end developer access to its models through Cursor, with a proposed shutoff date of November 12, 2026. It said it could not be confident its technology would be used within its terms. Cursor's CEO said OpenAI models served about 5% of Cursor's user traffic, and Cursor already supported models from several other providers, including Anthropic, Google, and SpaceX's own AI unit.
Lessons: - Diversification turned an existential event into a migration. Compare the Windsurf case: concentration on one provider made a similar event acute. - The notice period was ~10 weeks, not 90+ days. Plan your emergency playbook for weeks, not quarters, whatever your contract says (§37.12). - Multi-provider support is itself a product feature for platforms that resell model access. Their customers' portability depends on yours.
Case 4 — The Same-Vendor Breaking Upgrade (2026)¶
Across Anthropic's 2026 releases, several request patterns that had been valid became HTTP 400 errors on newer models: assistant prefill, non-default sampling parameters, fixed thinking budgets, forced tool selection, and disabling thinking. At the same time, defaults changed silently, for example reasoning effort on Claude Opus 5.5 versus Opus 5. Any vendor's release history would show similar patterns. This is how API evolution works, not a criticism of one provider.
Lesson: Staying with one vendor does not avoid swap work. The profile-and-contract discipline in this module pays off just as much on routine version upgrades as on cross-vendor moves, and upgrades happen far more often.
Case 5 — The Gateway Layer Consolidates (2026)¶
The LLM gateway and observability market consolidated quickly in 2026. Palo Alto Networks acquired Portkey and is integrating it into its AI security platform. Other gateway and observability products were reportedly absorbed or frozen after acquisition, or pivoted away from general routing.
Lesson: The tool you use to avoid lock-in can become lock-in. Keep the portability contract (capabilities, profiles, contracts, canonical history) in your own repository and formats. Use the gateway for transport and policy, and make sure you could replace it within a quarter.
37.12 Contracts and Governance for Portability¶
Module 25 covers AI contract clauses in depth. Portability adds these items to negotiate or verify:
| Clause / check | Why it matters for portability | What good looks like |
|---|---|---|
| Termination-for-convenience notice (provider side) | Cases 1 and 3: access can end for business reasons | Minimum notice for Tier A use (ask for 180 days; plan for far less) |
| Change-of-control provisions | Your own M&A can trigger termination | Explicit treatment of acquisitions; transition period |
| Model deprecation notice (Module 25 Clause 3) | Planned swaps need runway | ≥ 90 days; right to stay on the prior version during transition |
| Material behavior-change notification | Silent breaks (§37.2 layers 5, 7, 8) | Notice of default changes, tokenizer changes, new refusal categories |
| Data-retention terms per model | Some models require retention settings that conflict with your policy | Retention terms listed per model, not per account |
| Capacity reservation for failover | Avoiding the cascade (§37.7) | Reserved or committed quota on the fallback provider sized for failover share |
| Data and artifact export (Module 25 Clause 7) | Fine-tunes, stored prompts, eval data, and conversation logs must leave with you | Export in documented formats; fine-tuned weights or a reproducible tuning recipe |
Governance cadence: - Monthly: concentration and fallback-readiness metrics on the AI health dashboard (Module 12 §12.8). - Quarterly: candidate eval run (§37.6), scorecard review (§37.9), game day for one Tier A capability, profile re-verification, and lock-in ledger review. - On every provider event (new model, deprecation, terms change, access news): profile update and a gap check across affected capabilities.
37.13 The Model Portability Checklist¶
Architecture - [ ] Application code calls capabilities, never models or provider SDKs directly? - [ ] Capability registry exists with tier (A/B/C), schemas, and routing per capability? - [ ] A model profile exists for every model in routing, verified within 90 days? - [ ] Design ceiling for input size set to what the weakest acceptable fallback can handle? - [ ] Adapter performs feature negotiation and refuses to route where it cannot emulate? - [ ] Gateway used for transport and policy, and replaceable within a quarter?
Prompts - [ ] Prompts split into an invariant base (policy) and per-family overlays (tuning)? - [ ] No policy content in any overlay? - [ ] Base prompts free of model-specific constructs (tags, prefill, version names)? - [ ] Registry refuses to route to a model family without an evaluated overlay?
Contracts and evals - [ ] Behavioral contract per Tier A/B capability: invariants, quality tolerance, SLOs, style? - [ ] Datasets sized for the decision (500+ for Tier A near tolerance), stratified, versioned? - [ ] Paired comparison (McNemar or equivalent) used for swap decisions? - [ ] Cost per completed task is the cost metric? - [ ] Eval runs through the production adapter, with a pinned judge from a different model family? - [ ] Quarterly candidate run includes the next version of the current primary?
Fallback and state - [ ] Tier A capabilities have a hot fallback with keep-alive traffic? - [ ] Fallback quota sized for realistic failover share, not keep-alive? - [ ] Degradation ladder (L0–L4) defined, owned, and in the runbook? - [ ] Conversation history stored in a provider-neutral canonical form? - [ ] Failover cost reserve accounts for cache-cold periods?
Governance - [ ] Portability scorecard reviewed quarterly; scores under 14/21 on the debt backlog? - [ ] Provider game day run for at least one Tier A capability per quarter? - [ ] Lock-in ledger records every deliberate provider-specific dependency? - [ ] Provider contracts reviewed for termination notice, change of control, and retention terms? - [ ] Provider concentration metric tracked against a policy cap?
EXERCISE — Map the Ten Layers: Pick one production AI capability in your organization and its most plausible fallback model. Walk through the ten incompatibility layers in §37.2 and, for each, write down: what is different between the two models, whether the difference is loud (it would error) or silent (it would degrade), and whether anything you have today would catch it. Count your silent, uncaught differences. That number is your real switching risk, and it is usually larger than the team's estimate.
PONDER — The Ten-Day Notice: Your largest model provider tells you on Friday that access ends in ten days. Leave aside what you would do. List what you would discover: which capabilities call the provider directly, which prompts have never run on another model, which conversation histories are stored in the provider's format, which fallback accounts have no quota, and who owns the decision to send traffic to a lower-quality tier. Then ask: what would it cost to discover those things this quarter instead, on your own schedule?
WORKSHOP — Build the Portability Kit for One Capability: Choose a Tier A candidate capability (for example ticket triage, document extraction, or a RAG answer step) and produce: (1) a capability interface with typed input and output schemas; (2) model profiles for the current primary and one fallback from a different provider, covering all ten layers; (3) a base prompt plus one overlay per model family, with each overlay line justified; (4) a behavioral contract file with all four tiers and a justified dataset size; (5) a fallback design stating its readiness level, keep-alive share, failover triggers, degradation ladder, and an estimate of the cache-cold cost; (6) a scorecard self-assessment with the three lowest-scoring dimensions and a 90-day plan to raise them. Present it as if to an architecture review board that must approve the capability for production.
This is the final module of this reference (37 modules as of October 2026). Related modules: §2.5 (lock-in vectors), §21.2 (the commodity thesis), §17.3 (the LLM gateway), §24.3–24.4 (shadow, canary, rollback), §25.3 (contract clauses), §28.2 (model debt). Volatile facts live in Appendix G.