Skip to content

MODULE 37 — Model Portability: Engineering the Model as a Replaceable Part

⚠️ Currency note: The model names, API behaviors, and case-study dates in this module are accurate as of October 2026. They illustrate the patterns and are not recommendations. Model names and prices change quarterly, so verify them against Appendix G — Current Landscape and the provider links in Module 2 §2.1. The architecture in this module is built to outlast those names.

37.1 The Commodity Thesis and Where It Breaks

Module 21 (§21.2) makes the strategic argument: the model is becoming a commodity, and the competitive advantage moves to the systems around it — knowledge, workflow, evaluation, and orchestration. Module 2 (§2.5) lists the lock-in vectors. Module 17 (§17.3) makes the LLM gateway the platform's core infrastructure.

This module is about the gap between those three ideas and production reality. The model is commoditizing. Your integration with it is not.

Two different things get called "the model is a commodity":

What is commoditizing What is not commoditizing
Capability at the "good enough for most enterprise tasks" tier — several providers clear it The API surface — parameters, tool calling, structured output, reasoning controls
Price per unit of capability — falling every quarter Behavior — verbosity, refusals, instruction-following strictness, formatting
Access — multiple clouds now host the same models (e.g., OpenAI and Anthropic models are both on AWS Bedrock as of 2026) Operational envelope — rate limits, latency, regions, retention requirements, caching
The idea of switching — gateways make the wire call look the same Commercial access — a provider can and does withdraw access for business reasons

A commodity in the economic sense is fungible: one barrel of oil is the same as another. Models are not fungible in that sense. Electricity is the closer analogy. It is a commodity, but your building still needs a transformer, a breaker panel, and a tested generator before you can switch supply without the lights going out.

The architect's framing:

The model is a replaceable part — but only behind a contract you own. The contract is: a capability interface your application codes against, a model profile that maps each model onto that interface, a behavioral test suite that defines "equivalent," and a fallback you have actually exercised. With those four things, switching models is a configuration change you have already tested. Without them, every model change is a project — and an unplanned one is an incident.

The rest of this module builds those four things.


37.2 Why Model Swaps Fail: The Ten Incompatibility Layers

Teams that "switched models in an afternoon" usually swapped the transport and found the other problems in production weeks later. A swap can break at ten distinct layers. A gateway such as LiteLLM, Portkey, or OpenRouter handles only the first one well.

# Layer What breaks on a swap Real example (as of Oct 2026) Mitigation
1 Transport / wire format Endpoint, auth, request envelope Every provider's SDK differs Gateway with an OpenAI-compatible or normalized schema
2 Request parameters Parameters accepted by one model are rejected (HTTP 400) by another — including newer models from the same vendor Claude Opus 5.5 rejects thinking: {type: "disabled"} and budget_tokens with a 400; Claude Opus 4.7+ rejects non-default sampling parameters (temperature, top_p) Model profile declares supported parameters; the adapter translates or drops them
3 Tool calling Forced tool selection, parallel calls, argument escaping Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1 return 400 on forced tool_choice (any / specific tool); newer Claude models may escape JSON in tool arguments differently Never string-match tool arguments; parse JSON; use auto plus strict schemas; adapter handles forcing semantics
4 Structured output Native JSON-schema mode vs. tool-call workaround vs. prompt-only Assistant-message prefill (a common trick for forcing JSON) returns 400 on Claude 4.6+ models Use each provider's native structured-output feature through the adapter; validate every response against the schema anyway
5 Reasoning controls Token budgets vs. effort levels vs. thinking levels; different defaults Claude: effort from low to max, and Opus 5.5 defaults to medium where Opus 5 defaulted to high. Gemini 3.8 Flash: thinking levels low/medium/high. OpenAI: reasoning-effort settings Set reasoning depth explicitly per capability in the profile; never rely on a default
6 Context and output limits A prompt that fits one model overflows another; long answers get truncated Claude Opus 5.5: 1M context / 128K output. Gemini 3.8 Flash: 1M context / 64K output. GPT-6 bills requests over 272K input tokens at a surcharge Profile declares limits; the request builder enforces your design ceiling, not the model's maximum
7 Tokenizer and cost accounting The same text is a different number of tokens, so cost, truncation, and budgets shift Anthropic documents that the tokenizer introduced with Opus 4.7 can use ~1.0–1.35× as many tokens as earlier models for the same text Re-baseline token counts per model; budget in cost per completed task, not tokens
8 Behavior Verbosity, tone, refusal patterns, how literally instructions are followed Anthropic's migration guidance notes that prompts written for prior models are often "too prescriptive" for newer ones and reduce output quality; newer models add new safety-refusal categories Behavioral contract tests (§37.5); per-model prompt overlays (§37.4)
9 Caching and state Prompt caches are model-scoped; reasoning state may be model-bound Switching models forfeits cache hits. Newer Claude models bind thinking blocks to the producing model, so other models drop them on replay Plan cache-cold cost into failover; switch at conversation boundaries (§37.7)
10 Operational, legal, and commercial Rate limits, regions, data-retention terms, and access itself Claude Fable 5.1 requires 30-day data retention and is unavailable under zero-data-retention unless Anthropic expressly authorizes it, so it can be a compliance blocker. OpenAI is ending Cursor's model access after SpaceX's acquisition (§37.11) Eligibility rules in the profile; contract clauses; a pre-provisioned hot fallback

Three lessons from this table:

  1. Same-vendor upgrades are swaps too. Layers 2, 3, 5 and 8 broke on routine upgrades within one vendor's flagship line in 2026. "We're staying with our provider, so we don't need portability" is false. Model versioning is model swapping.
  2. Silent breaks are worse than loud ones. A 400 error is loud: you find it in the first test. A default effort level that drops from high to medium is silent. The calls succeed, quality degrades a little, and nobody notices until a customer complaint arrives three weeks later. Layers 5, 7 and 8 are where silent breaks live.
  3. Most of the work is above the gateway. A gateway normalizes transport. Portability requires normalizing semantics, and that is the adapter layer in §37.3, which you own.

37.3 The Portability Architecture

The Reference Architecture

PORTABILITY REFERENCE ARCHITECTURE

┌──────────────────────────────────────────────────────────────────────┐
│ APPLICATION CODE                                                     │
│   calls capabilities, never models:                                  │
│   triage_ticket(ticket) -> TicketDecision                            │
│   draft_reply(ticket, context) -> Draft                              │
│   summarize_contract(doc) -> ContractSummary                         │
└───────────────────────────────┬──────────────────────────────────────┘
                                │  typed request / typed response
┌───────────────────────────────▼──────────────────────────────────────┐
│ CAPABILITY LAYER  (you own this — the portability contract)          │
│   ├── Capability registry: name, input/output schema, tier (A/B/C)   │
│   ├── Routing policy: primary + ordered fallbacks per capability     │
│   ├── Behavioral contract: the tests that define "equivalent"        │
│   └── Prompt registry: base prompt + per-model-family overlays       │
└───────────────────────────────┬──────────────────────────────────────┘
                                │  neutral request (messages, tools,
                                │  schema, reasoning_depth, limits)
┌───────────────────────────────▼──────────────────────────────────────┐
│ MODEL ADAPTER LAYER  (you own this — one profile per model)          │
│   ├── Parameter translation (reasoning_depth → effort / level / ...) │
│   ├── Feature negotiation (forced tool choice? prefill? native JSON?)│
│   ├── Limit enforcement (design ceilings, not model maximums)        │
│   ├── Response normalization (tool calls, refusals, usage, errors)   │
│   └── Eligibility checks (data class, region, retention terms)       │
└───────────────────────────────┬──────────────────────────────────────┘
                                │  provider-specific request
┌───────────────────────────────▼──────────────────────────────────────┐
│ LLM GATEWAY  (buy or run — Module 17 §17.3)                          │
│   auth, rate limits, PII scanning, audit logging, retries,           │
│   transport-level failover, cost metering                            │
└──────┬──────────────┬──────────────┬──────────────┬──────────────────┘
       ▼              ▼              ▼              ▼
   Provider A     Provider B     Cloud-hosted    Self-hosted
   (direct API)   (direct API)   (Bedrock /      (vLLM / SGLang
                                  Foundry /       open-weight)
                                  Agent Platform)

The Principle: Code Against Capabilities, Not Models

Most AI codebases look like this:

# ANTI-PATTERN: application code knows about the model
response = llm.chat(
    model="provider-x-flagship-2026-05",
    messages=[{"role": "system", "content": TRIAGE_PROMPT}, {"role": "user", "content": ticket.text}],
    temperature=0,
    tool_choice={"type": "tool", "name": "classify"},   # 400 on some newer models
)
category = json.loads(response.tool_calls[0].arguments)["category"]

Every line is coupled to one model: the name, the sampling parameter, the forced tool choice, and the response shape. A swap touches every call site.

The portable version:

# PORTABLE: application code knows about the capability only
decision: TicketDecision = capabilities.triage_ticket(ticket)

That buys four things: - Swap blast radius goes to near zero. A model change edits one profile and one routing entry. No application call site changes. - The capability is the unit of evaluation. Tests attach to triage_ticket, not to a prompt or a model, so they survive both kinds of change. - The capability is the unit of routing. Different capabilities can sit on different providers, which spreads concentration risk without forcing a single "which vendor?" decision. - The capability is the unit of cost attribution. Module 13's per-feature cost reporting falls out naturally.

A capability is not a thin wrapper around chat(). It owns its input schema, output schema, prompt, validation, retry policy, and routing. Most systems need 5–30 capabilities, not hundreds. If you count more, you have probably mixed up capabilities with prompts.

The Model Profile

The model profile is the single file that knows how one specific model behaves. It is the most useful portability artifact you can create, because it turns the ten layers in §37.2 into data instead of tribal knowledge.

# model-profiles/anthropic-claude-opus-5-5.yaml   (illustrative, Oct 2026)
id: anthropic/claude-opus-5-5
family: claude-5            # prompt overlays are keyed by family
provider_routes:            # same model, several ways to reach it
  - anthropic-direct
  - aws-bedrock             # partner-operated; pricing differs
  - gcp-agent-platform

limits:
  context_tokens: 1000000
  max_output_tokens: 128000
  design_ceiling_input: 200000   # OUR ceiling — keeps failover targets viable

features:
  native_structured_output: true
  strict_tool_schemas: true
  forced_tool_choice: false      # 'any'/'tool' → 400; adapter uses auto + strict + instruction
  assistant_prefill: false       # 400 on 4.6+ family
  sampling_params: false         # temperature/top_p rejected
  parallel_tool_calls: true
  prompt_caching: { supported: true, scope: model, min_prefix_tokens: model_dependent }
  reasoning:
    control: effort              # low | medium | high | xhigh | max
    can_disable: false           # thinking cannot be turned off on this model
    default: medium              # NOTE: differs from claude-opus-5 (high) — always set explicitly
    reasoning_state_portable: false   # thinking blocks bound to this model

translation:                     # neutral reasoning_depth → this model's control
  reasoning_depth:
    minimal: { effort: low }
    standard: { effort: medium }
    deep: { effort: high }
    max: { effort: max }

refusals:
  signal: stop_reason == "refusal"      # HTTP 200, not an error
  categories_exposed: true              # stop_details.category
  server_side_fallback: available       # provider feature; record whether we use it

eligibility:
  data_classes_allowed: [public, internal, confidential]   # per your policy
  zero_data_retention: per_contract
  regions: [us, eu]                     # verify per route
tokenizer_note: "Opus 4.7+ tokenizer; re-baseline counts vs. pre-4.7 models"
pricing_ref: appendix-g#g3            # never hardcode prices here
last_verified: 2026-10-03
owner: ai-platform-team

The design_ceiling_input field is the most underrated line in this profile. If your primary model accepts 1M tokens and your application starts sending 600K-token prompts, every fallback target with a 128K–262K window becomes unusable without notice. Set the ceiling at what your weakest acceptable fallback can handle, and raise it only deliberately, with an ADR.

The Adapter: Feature Negotiation in Practice

The adapter turns the neutral request into a provider request by consulting the profile. A sketch of the logic, provider-neutral pseudocode:

def build_request(neutral: NeutralRequest, profile: ModelProfile) -> ProviderRequest:
    req = ProviderRequest(model=profile.id)

    # 1. Enforce OUR ceiling, not the model's
    if count_tokens(neutral, profile) > profile.limits.design_ceiling_input:
        raise CapabilityInputTooLarge(neutral.capability)   # route to chunking path

    # 2. Reasoning depth: always explicit, translated per model
    req.reasoning = profile.translation.reasoning_depth[neutral.reasoning_depth]

    # 3. Structured output: best available mechanism
    if neutral.output_schema:
        if profile.features.native_structured_output:
            req.output_format = neutral.output_schema
        else:
            req.tools = [schema_as_tool(neutral.output_schema)]
            req.instructions += "Respond only by calling the provided tool."

    # 4. Forced tool selection: emulate where unsupported
    if neutral.require_tool:
        if profile.features.forced_tool_choice:
            req.tool_choice = {"type": "tool", "name": neutral.require_tool}
        else:
            req.tool_choice = {"type": "auto"}
            mark_strict(req.tools)
            req.instructions += f"You must call the `{neutral.require_tool}` tool."

    # 5. Drop parameters the model rejects (log it — silent drops hide intent)
    if neutral.temperature is not None and not profile.features.sampling_params:
        log.info("dropping temperature for %s", profile.id)

    return req


def normalize_response(raw, profile) -> NeutralResponse:
    # Refusals are not errors on some providers (HTTP 200) — normalize them
    # Tool-call arguments: always json.loads, never string-match
    # Usage: map provider token fields to (input, cached_input, output, reasoning)
    ...

Notice what the adapter does not do: it does not hide failures. If a capability needs forced tool selection and the target cannot emulate it reliably, the adapter should refuse to route there. That shows up in the eval harness as a contract failure, not as a mysterious quality drop.

Gateway vs. Adapter: Who Does What

Concern Gateway Adapter / capability layer
Auth, keys, quotas ✓
PII scanning, audit logs ✓
Transport retries, 5xx failover ✓
Cost metering ✓ Cost per capability attribution
Parameter semantics, feature negotiation ✓
Prompt selection per model family ✓
Output validation against schema ✓
Quality-triggered failover ✓ (needs the contract)
"Is this model eligible for this data?" Partial (policy) ✓ (per capability)

The gateway is not your portability layer. It is a dependency of it. Treat the gateway as replaceable too: gateway vendors consolidated sharply in 2026. Palo Alto Networks acquired Portkey (closed May 2026), and other gateway and observability startups were reportedly absorbed, frozen, or pivoted. Keep your capability and adapter layer independent of any one gateway product's proprietary configuration format.


37.4 Prompt Portability

Prompts are where portability quietly fails. Every prompt is tuned, deliberately or not, to the model it was developed against.

Three Strategies

Strategy How it works Pros Cons Use when
Lowest common denominator One prompt, written to work acceptably on every candidate Simplest; one artifact to govern Leaves quality on the table on every model; the "common" set shrinks as models diverge Tier C capabilities; very simple tasks
Base + overlays (recommended default) An invariant base prompt plus a small per-model-family overlay Policy stays in one place; tuning is isolated and reviewable Overlays must be evaluated per family Most Tier A/B capabilities
Per-model prompts Fully separate prompts per model Maximum quality per model Policy drift between copies; N× governance cost Very high-value capabilities with a dedicated owner

The Invariant / Tuned Split

Split every production prompt into what must stay the same on any model and what may vary:

PROMPT LAYERING FOR PORTABILITY

INVARIANT BASE  (identical across models; changes go through policy review)
  ├── Role and scope boundaries ("You triage support tickets for ...")
  ├── Business rules and policy ("Never promise refunds; escalate if ...")
  ├── Output contract (schema, required fields, enumerations)
  ├── Safety and injection-resistance rules (Module 3 §3.5–3.6)
  └── Domain knowledge and definitions

MODEL-FAMILY OVERLAY  (small; tuned per family; changes go through eval)
  ├── Emphasis calibration (newer models over-apply ALL-CAPS / "MUST")
  ├── Verbosity guidance ("be concise" means different things per family)
  ├── Reasoning guidance (none for reasoning-native models; brief for others)
  ├── Formatting quirks (markdown vs. plain text tendencies)
  └── Known failure workarounds (documented, with the eval case that proves them)

Rules for overlays: - An overlay never contains policy. If an overlay says "don't discuss pricing", that rule belongs in the base, or one model family will quietly lack it. - Every overlay line cites a reason. Either the eval case it fixes or the migration note it implements. An overlay line without a reason gets deleted at the next review. - Keep overlays small. If an overlay grows past ~15–20% of the base, you have drifted into per-model prompts by accident. Decide whether that was intended.

Model-Specific Constructs to Remove from Base Prompts

These constructs tie a prompt to one model or generation. Grep for them before a swap:

Construct Why it is not portable
Provider-specific tags (<thinking>, <scratchpad>, special tokens) Meaningless or harmful elsewhere; some models leak or mimic them
"Think step by step" on reasoning-native models Redundant with built-in reasoning; can lengthen output without improving quality
Assistant-message prefill to force a format Rejected outright by some newer models (400)
Heavy emphasis written for weaker instruction-followers ("CRITICAL!!! YOU MUST ALWAYS...") Newer models follow instructions more literally and over-trigger
References to a model's own name or version Wrong after the swap
Few-shot examples copied from one model's own outputs Teaches the next model to imitate the previous model's quirks
Instructions relying on a specific default (temperature, reasoning depth) Defaults differ and change between versions

The Prompt Registry, Keyed for Portability

A Module 3 prompt registry entry should be keyed by (capability, prompt_version, model_family). Each (prompt_version, model_family) pair needs an eval result before it can serve production traffic. The registry should refuse to route a capability to a model family that has no evaluated overlay. This one rule prevents the most common portability incident: the fallback model ran a prompt nobody had ever tested on it.


37.5 Behavioral Contracts: Defining "Equivalent"

You cannot swap a part without a specification of what the part must do. Public benchmarks are not that specification. They measure general capability on someone else's tasks. A behavioral contract defines equivalence for your capability.

The Four Tiers of a Contract

Tier What it covers Threshold style Example (ticket triage)
1. Hard invariants Things that must never break Absolute — any failure blocks 100% schema-valid output; 0 PII echoed; 100% of injection cases stay in scope
2. Quality metrics Task accuracy and faithfulness Relative to baseline, with tolerance Category accuracy ≥ baseline − 2 pts; "needs human" recall ≥ 95%
3. Operational Latency, cost, reliability Absolute SLOs P95 latency ≤ 2.5 s; cost per completed triage ≤ $0.004; error + refusal rate ≤ 0.5%
4. Style Tone, length, formatting Soft — reviewed, not gating Median rationale length within ±30% of baseline

The asymmetry is deliberate. Tier 1 is binary and blocks a swap. Tier 2 compares against the current model, so a swap is judged on whether it is equivalent or better, not on whether it is perfect. Tier 4 is reviewed by a human and does not gate the decision. A contract that gates on style will block every swap.

A Contract File

# contracts/triage_ticket.yaml
capability: triage_ticket
tier: A                       # business-critical: hot fallback required
dataset:
  path: evals/triage/v7.jsonl
  size: 640                   # see sample-size guidance below
  composition: { production_sample: 480, edge_cases: 100, adversarial: 60 }
  provenance: "prod sample 2026-08, PII-scrubbed; edge cases from incident log"

hard_invariants:
  - schema_valid: 1.00
  - pii_echo_rate: 0.00
  - adversarial_in_scope: 1.00

quality:                      # relative to current primary
  - metric: category_accuracy
    tolerance: -0.02
  - metric: needs_human_recall
    min_absolute: 0.95        # absolute floor regardless of baseline
  - metric: urgency_mae
    tolerance: +0.15

operational:
  - p95_latency_ms: 2500
  - cost_per_completed_task_usd: 0.004   # includes retries and repair calls
  - error_plus_refusal_rate: 0.005

style:                        # reported, not gating
  - rationale_length_median_delta: 0.30

grader:
  deterministic: [schema_valid, category_accuracy, urgency_mae, pii_echo_rate]
  llm_judge:
    metrics: [adversarial_in_scope]
    judge_model: pinned          # see §37.6 — never the candidate itself

Sample Size: How Many Cases Is Enough?

Teams routinely decide a swap on 30 examples. That is not enough to detect the differences that matter.

  • Rule of thumb for one accuracy number: the 95% margin of error is about ±1.96 × √(p(1−p)/n). At an accuracy near 85% with n = 100, the margin is about ±7 points, which is useless for a "−2 points" tolerance. At n = 600 it is about ±2.9 points.
  • Use paired comparison. Both models answer the same cases, so compare them case by case and not as two separate averages. Count the discordant cases: b = the baseline was right and the candidate wrong; c = the candidate was right and the baseline wrong. McNemar's test on b and c is far more sensitive than comparing two accuracies, because the cases both models get right or both get wrong carry no information about the difference. Statistic: (|b − c| − 1)² ⁄ (b + c), compared against a χ² distribution with 1 degree of freedom (3.84 at 95%).
  • Practical sizing:
  • Smoke test (does it work at all?): 30–50 cases.
  • Tier C swap decision: 150–300.
  • Tier A swap decision near a tolerance boundary: 500–1,000, plus shadow traffic (§37.8).
  • Stratify. Report results per segment (ticket category, language, customer tier). An aggregate that holds steady can hide a 15-point regression on one language that is 8% of traffic.

Shared APIs Do Not Mean Equivalent Models

Typed-decision models show the point clearly. As of October 2026, Cloudflare's Clef is reported to be API-compatible with TypeSafe's Jev, which makes switching easy to wire. Early comparisons nevertheless report different strengths (for example, better at routing but weaker at judgment) (reported), and each model has its own calibration profile. For these models the contract must include calibration metrics (per-segment calibration error, coverage at the target precision) as well as accuracy. See Module 16 §16.9.

Measure Cost per Completed Task, Not per Token

A cheaper model that needs more repair calls, retries, or agent turns to finish the job is not cheaper. Define completed per capability: schema-valid, passed validation, no human repair needed. Then compute:

cost_per_completed_task = (Σ cost of all calls incl. retries, repairs, judge checks)
                          ÷ (number of tasks completed without human repair)

This is the operational-tier cost metric in every contract. It routinely reverses "cheaper model" decisions, in both directions.


37.6 The Cross-Model Eval Harness

The contract defines equivalence. The harness measures it, continuously, against candidates you are not yet using.

Harness Components

CROSS-MODEL EVAL HARNESS

  ┌────────────────┐   ┌───────────────────┐   ┌──────────────────────┐
  │ Frozen datasets │──►│ Runner            │──►│ Graders              │
  │ per capability  │   │ capability ×      │   │ 1. deterministic     │
  │ (versioned,     │   │ model profile ×   │   │    (schema, exact,   │
  │  provenance,    │   │ prompt overlay    │   │     regex, numeric)  │
  │  PII-scrubbed)  │   │ via the SAME      │   │ 2. LLM judge         │
  └────────────────┘   │ adapter as prod   │   │    (pinned model,    │
                       └───────────────────┘   │     different family │
                                               │     from candidate)  │
                                               └──────────┬───────────┘
                                                          ▼
                                     ┌──────────────────────────────────┐
                                     │ Portability report per capability │
                                     │ contract pass/fail · paired stats │
                                     │ · stratified results · cost/task  │
                                     └──────────────────────────────────┘

Non-negotiables: - Run through the production adapter. An eval that calls the model directly, bypassing the adapter, tests a different system from the one you will deploy. Most of the swap bugs from §37.2 live in the adapter. - Pin the judge, and never let a model judge itself. LLM judges show self-preference bias: a model tends to rate outputs that resemble its own more highly. Use a judge from a different model family from the candidate. Pin its version. Re-calibrate it against human labels whenever the judge model changes (Module 12 §12.5). - Version datasets like code. Every dataset has a version, a provenance note, and a change log. When a production incident occurs, its case goes into the dataset (Module 3 §3.7). - Prefer replayed production traffic. Synthetic cases miss the real distribution. Replay a PII-scrubbed production sample monthly, and keep the hand-built edge and adversarial cases on top.

The Quarterly Candidate Run

This makes Module 28's model-debt guidance operational. Once a quarter, run every Tier A and Tier B capability against: 1. the current primary (the baseline), 2. the configured fallbacks, 3. the next version of the primary model (to catch same-vendor breaks early), 4. one or two new candidates from the Appendix G refresh.

PORTABILITY REPORT — triage_ticket — 2026-Q4 run (illustrative)

Model / overlay              Contract  Accuracy Δ  Recall(needs_human)  P95    Cost/task  Notes
──────────────────────────────────────────────────────────────────────────────────────────────
primary   (vendor A flagship) baseline  —          0.962               1.9s   $0.0031
fallback1 (vendor B mid-tier) PASS      −1.1 pts   0.958               1.4s   $0.0019
fallback2 (self-hosted MoE)   PASS*     −1.8 pts   0.951               2.3s   $0.0012    *near tolerance
next-ver  (vendor A next)     FAIL      +0.6 pts   0.966               2.1s   $0.0034    T1: 3 schema fails —
                                                                                         forced tool_choice
                                                                                         rejected; adapter
                                                                                         emulation needed
candidate (new open-weight)   FAIL      −4.9 pts   0.931               1.1s   $0.0007    Fails on non-English
                                                                                         segment (−14 pts)

That "next-ver FAIL" row is exactly the result you want in a quarterly run, months before the deprecation deadline, and not during an emergency migration.

Harness Cost Is a Line Item

A full candidate run costs real money: capabilities × models × cases × (call + judge). For example, 20 capabilities × 5 models × 600 cases × ~$0.004 ≈ $240 per quarterly run, before judge costs. Budget for it explicitly in Module 13's FinOps model. It is the cheapest insurance in the AI budget.


37.7 Fallback Design: Hot, Warm, and Cold

A fallback you have never sent traffic to is a hypothesis, not a fallback.

The Three Readiness Levels

Level Definition What exists Time to take full traffic Required for
Hot Continuously serving a small share of real traffic Profile, evaluated overlay, contract PASS, provisioned quota, monitoring, live traffic share Minutes (config flip) Tier A capabilities
Warm Ready but idle Profile, evaluated overlay, contract PASS within the last quarter, account and quota in place Hours to a day (verify quota, shadow briefly, ramp) Tier B capabilities
Cold Identified, not prepared A named candidate; maybe a profile Days to weeks (full playbook §37.8) Tier C capabilities

Keep-Alive Traffic

Send a small, steady share of production traffic to the hot fallback, typically 1–5%, and evaluate it with the same online monitors as the primary. It costs very little, and it gives you four things: - proof that the credentials, quota, network path, and adapter actually work today; - a live quality comparison on real traffic, not only on eval sets; - an existing commercial relationship and rate-limit history with the fallback provider, which matters when you suddenly need 20× the quota; - organizational familiarity: on-call engineers have seen the fallback's dashboards before the incident.

The Cursor case (§37.11) shows the value of traffic diversity: when one provider withdrew access, it carried only about 5% of traffic.

Failover Triggers

Define triggers per capability, and make each one an explicit, logged decision:

Trigger Signal Typical action
Hard availability 5xx rate > X% over 2 min; connection failures Automatic failover (gateway)
Sustained throttling 429 rate > X% over 5 min after retries Automatic partial shift (overflow to fallback)
Latency SLO breach P95 > SLO for 10 min Automatic or on-call decision, per tier
Refusal spike Refusal rate > 3× baseline On-call decision. Investigate first: a refusal spike is often caused by your input, not by the model
Quality monitor drop Online eval score below the contract floor On-call decision; never automatic on a single signal
Commercial / policy event Access-termination notice, ToS change, retention change, deprecation Planned migration via the playbook

The Degradation Ladder

When the primary fails, step down a defined ladder and do not improvise:

DEGRADATION LADDER (per capability)

  L0  Primary model                         normal operation
  L1  Hot fallback, same quality tier       transparent to users
  L2  Lower tier model                      response flagged "reduced quality";
                                            stricter validation; more human review
  L3  Self-hosted open-weight               reduced capability, data stays inside
  L4  Non-AI fallback                       rules engine, cached answers, search,
                                            or route to human queue (Module 16 Anti-Pattern 2)

Each rung needs an owner, a tested runbook entry (Module 24 §24.2), and a defined user experience (Module 16 Principle 5, "Design the Failure Narrative").

Three Failover Pitfalls Teams Discover Too Late

1. The cache-cold cost spike. Prompt caches are scoped to a model. Consider a capability that sends a 50,000-token cached system prompt and knowledge prefix on every request, at 1M requests a month. Using illustrative prices of $2.00 per million input tokens uncached and $0.20 cached, the prefix costs about $10,000 a month while cached and $100,000 a month uncached. On failover the fallback starts cold. If it takes a large share of traffic, the input bill for that period can jump by up to 10×, on top of cache-write premiums. Mitigations: keep-alive traffic keeps the fallback's cache warm; pre-warm on failover; budget a failover cost reserve; keep cacheable prefixes lean.

2. The cascade. The primary fails, all traffic moves to the fallback, and the fallback's quota, sized for 5% keep-alive, throttles within minutes. Now two providers are failing. Mitigations: provision fallback quota for your realistic failover share (contracted or reserved capacity for Tier A); shift traffic progressively; shed load on Tier C capabilities first.

3. The mid-conversation switch. Long conversations and agent loops carry provider-specific state: reasoning blocks bound to the producing model, tool-call ID formats, cached prefixes, and provider-specific content block types. Switching mid-session can drop reasoning context or produce malformed history. Mitigations: - prefer switching at conversation or task boundaries: new sessions go to the fallback while in-flight sessions drain on the primary where possible; - store conversation history in a provider-neutral canonical form (your own message schema) and render it per provider, never storing the provider's raw format as the source of truth; - on a forced mid-session switch, re-render history from the canonical form, drop provider-specific blocks, and add a short neutral summary of prior reasoning if the task depends on it; - for agents, checkpoint task state at step boundaries (Module 6) so a different model can resume from state, not from a transcript.


37.8 The Model Swap Playbook

The same playbook applies whether the swap is planned (cost, quality, deprecation) or forced (access revocation, outage, compliance change). Only the time pressure differs. Architecture done well makes the emergency path a compressed version of the planned path, not a different one.

Phase 0 — Trigger and Classify

Trigger class Example Typical deadline Urgency
Deprecation Pinned version retirement notice 60–180 days Planned
Cost / quality opportunity New model passes contract at 40% lower cost per task None Planned
Same-vendor upgrade The next flagship version, with breaking API changes Self-imposed Planned
Compliance / terms change Provider changes data-retention terms; new regulation 30–90 days Elevated
Access revocation Provider ends access (Windsurf 2025: under a week; Cursor 2026: ~10 weeks) Days to weeks Emergency
Sustained outage / quality collapse Multi-day degradation Hours Emergency

Phases 1–8

MODEL SWAP PLAYBOOK

PHASE 1  INVENTORY                                   planned: 1–2 days · emergency: hours
  Source: gateway logs + capability registry (never "ask the teams")
  Output: every capability on the outgoing model, traffic share, tier,
          provider-specific features used, prompt overlays, data classes
  Exit:   complete list signed off by platform owner

PHASE 2  FEATURE-GAP ANALYSIS                       planned: 1–3 days · emergency: hours
  Diff outgoing profile vs. candidate profile across the 10 layers (§37.2)
  Output: gap list — each gap marked: adapter can emulate / needs code /
          blocks the swap
  Exit:   no unmitigated blocking gap for Tier A capabilities

PHASE 3  ADAPTER + OVERLAY WORK                     planned: 3–10 days · emergency: 1–2 days
  Candidate profile, translations, overlays per family; unit tests per gap
  Exit:   all capabilities run end-to-end on candidate in staging

PHASE 4  OFFLINE CONTRACT EVAL                      planned: 1–3 days · emergency: hours
  Full contract per capability (§37.5), paired stats, stratified
  Exit:   Tier 1 invariants 100% · Tier 2 within tolerance · Tier 3 SLOs met

PHASE 5  SHADOW                                     planned: 3–7 days · emergency: skip or <24h
  Module 24 §24.3: candidate receives mirrored traffic, outputs logged only
  Exit:   ≥1,000 shadow requests per Tier A capability, monitors green

PHASE 6  CANARY RAMP                                planned: 1–2 weeks · emergency: 1–2 days
  1% → 5% → 25% → 50% → 100%, with automatic rollback triggers:
    contract floor breach · error+refusal > threshold · P95 > SLO ·
    cost/task > ceiling · human-escalation rate > baseline + X
  Exit:   100% for ≥72h with monitors green

PHASE 7  PROMOTE + RETAIN                           planned and emergency
  New model becomes primary; OUTGOING model becomes warm/hot fallback until
  its deprecation date (do not delete the profile on day one)
  Update: routing config, ADR, runbooks, Appendix-G-style internal registry

PHASE 8  RETRO + SCORE                              within 2 weeks
  Record actual time-to-switch per capability; update portability score (§37.9);
  every surprise becomes an eval case or a profile field

Planned swap, well-architected system: about 3–6 weeks of elapsed time, mostly spent waiting on shadow and canary data, not on engineering. Emergency swap, well-architected system: 2–5 business days for Tier A capabilities that already had a warm or hot fallback. Emergency swap, poorly architected system: weeks, plus a production incident, plus quality regressions found by customers.

Worked Example A — A Same-Vendor Upgrade That Breaks (as of Oct 2026)

Scenario: ticket triage runs on Claude Opus 5. The team plans to move to Claude Opus 5.5, which is cheaper per token. Everyone treats it as "just a version bump."

Phase 2 finds three gaps:

Gap Outgoing (Opus 5) Candidate (Opus 5.5) Class Resolution
Thinking disabled for latency Accepted at high effort or below {type: "disabled"} → 400 at every effort Loud Remove; set effort: low explicitly for latency-sensitive capabilities
Forced tool call for schema tool_choice: {type: "tool"} accepted 400 on forced any/tool Loud Adapter emulation: auto + strict schema + instruction, or native structured output
Default reasoning depth high medium Silent Profile sets effort explicitly; the contract eval confirms accuracy at the chosen effort

The two loud gaps would have failed in the first staging run. The silent one would not. Calls succeed, and on a hard ticket segment accuracy may drift down by a few points. Only the explicit-effort rule in the profile and the stratified contract eval catch it. The adapter change for all three is a profile diff:

# diff: model-profiles/anthropic-claude-opus-5.yaml → anthropic-claude-opus-5-5.yaml
 features:
-  forced_tool_choice: true
+  forced_tool_choice: false
   reasoning:
-    can_disable: true          # only at effort <= high
-    default: high
+    can_disable: false
+    default: medium            # we never rely on it — translation sets effort explicitly

No application code changed. That is what portability looks like in practice.

Worked Example B — An Emergency Cross-Vendor Swap (Composite)

Scenario (composite, for illustration): a provider sends a notice ending access in 10 days. The organization runs 14 capabilities: 4 Tier A, 6 Tier B, 4 Tier C.

Day Tier A (4) — hot fallback existed Tier B (6) — warm fallback Tier C (4) — cold
0 Inventory from gateway logs (2h). Raise fallback share 5% → 25% Inventory; confirm quotas with fallback provider Inventory; pick candidates
1 Contract eval re-run on current fallback (all PASS). Ramp to 50% Offline contract eval (5 PASS, 1 FAIL on a long-document capability: input exceeds fallback context) Write profiles, start overlays
2 Ramp to 100%. Outgoing provider kept as fallback until cutoff Failing capability: enforce design ceiling + chunking path; re-eval PASS. Begin canary Offline eval
3–5 Done. Monitor cache-cold cost (as predicted, ~3 days elevated) Canary 5% → 100% Canary, with "reduced quality" flags allowed for Tier C
6–10 Retro Retro One capability moved to L4 (rules + human queue) for 2 weeks while its prompt was reworked

Why it took days instead of weeks: Tier A already had hot fallbacks with keep-alive traffic. Every capability had a contract. All history was stored in canonical form. The only real engineering surprise was the context-window gap on one capability, and a profile-enforced design ceiling would have prevented even that.


37.9 Measuring Portability: The Portability Scorecard

What you do not measure decays. Portability decays fast: every new provider-specific feature a team adopts, and every prompt tuned on one model, adds switching cost without anyone deciding to.

Core Metrics

Metric Definition Target (Tier A)
Time-to-switch (TTS) Measured elapsed time to move a capability's full traffic to its fallback while meeting the contract. Measure it in drills, don't estimate it ≤ 1 day
Fallback readiness % of Tier A/B traffic whose capability has a hot/warm fallback with a contract PASS in the last quarter 100% (A), ≥ 90% (B)
Provider concentration % of total AI traffic (or spend) on the single largest provider Policy-defined; many organizations cap Tier A at ~70–80% on one provider
Proprietary-feature count Number of provider-specific features in use without a documented exit path (lock-in ledger, §37.10) 0 undocumented
Overlay coverage % of (capability × fallback family) pairs with an evaluated overlay 100% (A)
Contract coverage % of capabilities with a contract file and a dataset of the required size 100% (A/B)
Profile freshness Days since each profile's last_verified ≤ 90

The Scorecard

Score each Tier A/B capability from 0 to 3 on each dimension, and review it quarterly with the architecture board:

Dimension 0 1 2 3
Interface App code calls a model directly Wrapper, but model-specific params leak Capability interface; some leakage Pure capability interface
Profile None Informal notes Profile exists, stale Profile current (≤ 90 days)
Prompt Single model-tuned prompt LCD prompt, untested on fallback Base + overlay, partially evaluated Base + evaluated overlay per fallback family
Contract None Ad hoc examples Contract, undersized dataset Contract + adequate dataset + stratification
Fallback None Cold Warm Hot with keep-alive
State Provider-raw history Partially canonical Canonical history Canonical + checkpointed agent state
Drill Never Tabletop only Staging drill Production drill (canary-scale) within 6 months

A capability scoring under 14 of 21 is a portability risk and belongs on the Module 28 debt backlog. A Tier A capability with a 0 anywhere is a finding, regardless of its total.

The Provider Game Day

Run a quarterly game day for Tier A: 1. Pick a capability and announce the drill window. 2. At canary scale (say 10% of traffic), block the primary provider at the gateway. 3. Measure: time to detect, time to fail over, contract metrics during the drill, cost delta, and user-visible errors. 4. Restore the primary, then record actual TTS on the scorecard. 5. Every surprise becomes a runbook fix, a profile field, or an eval case.

The first game day nearly always finds something: an expired fallback API key, a quota sized for keep-alive only, or a capability that turns out to call the primary directly and bypass the gateway.


37.10 When Not to Be Portable: Deliberate Lock-In

Portability has costs. Pretending otherwise produces over-abstracted systems that use only the lowest common denominator of every model and adopt nothing new.

What Portability Costs

  • Feature lag. New provider capabilities, such as server-side agent runtimes, computer-use toolsets, advanced caching modes, or server-side refusal fallbacks, reach a portable system later, because each one needs a gap analysis and an exit path.
  • Engineering and eval spend. Profiles, overlays, contracts, keep-alive traffic, quarterly runs, and game days are real work and real money.
  • Abstraction drag. Heavy frameworks that hide the provider make debugging harder. When a call misbehaves, you need to see the actual request that went out.

When Lock-In Is the Right Call

Accept a provider-specific dependency deliberately when all of these hold: 1. The feature delivers a step change, not an incremental gain, for the capability. 2. The exit cost is bounded and known. You have written down what it would take to replicate or replace the feature. 3. The capability's tier tolerates it: Tier C and B usually yes; Tier A only with a documented degraded mode. 4. There is a revisit trigger, such as a date, a price threshold, or a competitor shipping an equivalent.

The Lock-In Ledger

Record every deliberate dependency as a short ADR in a single ledger:

LOCK-IN LEDGER ENTRY (ADR-LK-014)

Dependency:      Provider X managed agent runtime (hosted sandbox + memory)
Capabilities:    research_assistant (Tier B), report_builder (Tier C)
Why accepted:    Replaces ~6 weeks of sandbox/infra build; step change in time-to-value
Exit path:       Self-hosted harness + container sandbox; est. 4–6 engineer-weeks;
                 canonical task state already stored outside the runtime
Degraded mode:   research_assistant falls to single-call RAG (L2) during exit
Revisit trigger: 2027-Q2 review, or provider price change > 25%, or access/terms change
Owner:           platform-architecture

The ledger turns "we're locked in and didn't notice" into "we chose this, we know what leaving costs, and we review it." The proprietary-feature metric in §37.9 counts ledger entries, and undocumented dependencies are the ones that score against you.

Tiering Sets How Much Portability Each Capability Gets

Capability tier Definition Portability investment
A — Business-critical Revenue- or safety-impacting; customer-facing at scale; regulated Hot fallback, full contract, base + overlay, game days, canonical state
B — Important Material productivity impact; internal or limited external Warm fallback, contract, overlays for one fallback family
C — Convenience Low impact; easily paused Cold fallback acceptable; LCD prompt; deliberate lock-in welcome

The architect's job is not to make everything portable. It is to make sure the portability of each capability matches its tier, and that the mismatch is visible when it isn't.


37.11 Case Studies

Case 1 — Windsurf and Anthropic (June 2025): Access as a Strategic Decision

Shortly after reports that OpenAI would acquire the AI coding tool Windsurf, Anthropic told Windsurf it was cutting nearly all of its direct, first-party capacity for Claude 3.x models, with less than a week's notice. Windsurf had to find capacity through other providers and push users toward bring-your-own-key options. Anthropic's leadership framed the decision as strategic: it would be odd to sell Claude to a competitor's subsidiary. The reported OpenAI acquisition later fell through. The access cut came from the expectation of a change of control, not from a completed one.

Lessons: - Access is a business decision, not only a technical SLA. No uptime clause protects you from a provider deciding it no longer wants you as a customer. - Change of control on your side can trigger it. An acquisition, investment, or partnership can turn you into a competitor's asset overnight. - A week is not enough time to build portability. It is only enough time to use portability you already have.

Case 2 — Anthropic and OpenAI (August 2025): Terms-of-Service Enforcement

Anthropic revoked OpenAI's API access to Claude models, citing violations of its terms of service (reportedly heavy use of Claude Code by OpenAI technical staff).

Lesson: Terms of service are part of your dependency surface. Usage patterns that breach a provider's terms, such as benchmarking for competitive development or prohibited use cases, can end access abruptly. Your legal review of provider terms (Module 25) is part of your availability architecture.

Case 3 — Cursor and OpenAI (August–November 2026): Diversification Makes It Survivable

SpaceX completed its acquisition of Anysphere, the maker of Cursor, on August 14, 2026. On August 29, OpenAI announced it would end developer access to its models through Cursor, with a proposed shutoff date of November 12, 2026. It said it could not be confident its technology would be used within its terms. Cursor's CEO said OpenAI models served about 5% of Cursor's user traffic, and Cursor already supported models from several other providers, including Anthropic, Google, and SpaceX's own AI unit.

Lessons: - Diversification turned an existential event into a migration. Compare the Windsurf case: concentration on one provider made a similar event acute. - The notice period was ~10 weeks, not 90+ days. Plan your emergency playbook for weeks, not quarters, whatever your contract says (§37.12). - Multi-provider support is itself a product feature for platforms that resell model access. Their customers' portability depends on yours.

Case 4 — The Same-Vendor Breaking Upgrade (2026)

Across Anthropic's 2026 releases, several request patterns that had been valid became HTTP 400 errors on newer models: assistant prefill, non-default sampling parameters, fixed thinking budgets, forced tool selection, and disabling thinking. At the same time, defaults changed silently, for example reasoning effort on Claude Opus 5.5 versus Opus 5. Any vendor's release history would show similar patterns. This is how API evolution works, not a criticism of one provider.

Lesson: Staying with one vendor does not avoid swap work. The profile-and-contract discipline in this module pays off just as much on routine version upgrades as on cross-vendor moves, and upgrades happen far more often.

Case 5 — The Gateway Layer Consolidates (2026)

The LLM gateway and observability market consolidated quickly in 2026. Palo Alto Networks acquired Portkey and is integrating it into its AI security platform. Other gateway and observability products were reportedly absorbed or frozen after acquisition, or pivoted away from general routing.

Lesson: The tool you use to avoid lock-in can become lock-in. Keep the portability contract (capabilities, profiles, contracts, canonical history) in your own repository and formats. Use the gateway for transport and policy, and make sure you could replace it within a quarter.


37.12 Contracts and Governance for Portability

Module 25 covers AI contract clauses in depth. Portability adds these items to negotiate or verify:

Clause / check Why it matters for portability What good looks like
Termination-for-convenience notice (provider side) Cases 1 and 3: access can end for business reasons Minimum notice for Tier A use (ask for 180 days; plan for far less)
Change-of-control provisions Your own M&A can trigger termination Explicit treatment of acquisitions; transition period
Model deprecation notice (Module 25 Clause 3) Planned swaps need runway ≥ 90 days; right to stay on the prior version during transition
Material behavior-change notification Silent breaks (§37.2 layers 5, 7, 8) Notice of default changes, tokenizer changes, new refusal categories
Data-retention terms per model Some models require retention settings that conflict with your policy Retention terms listed per model, not per account
Capacity reservation for failover Avoiding the cascade (§37.7) Reserved or committed quota on the fallback provider sized for failover share
Data and artifact export (Module 25 Clause 7) Fine-tunes, stored prompts, eval data, and conversation logs must leave with you Export in documented formats; fine-tuned weights or a reproducible tuning recipe

Governance cadence: - Monthly: concentration and fallback-readiness metrics on the AI health dashboard (Module 12 §12.8). - Quarterly: candidate eval run (§37.6), scorecard review (§37.9), game day for one Tier A capability, profile re-verification, and lock-in ledger review. - On every provider event (new model, deprecation, terms change, access news): profile update and a gap check across affected capabilities.


37.13 The Model Portability Checklist

Architecture - [ ] Application code calls capabilities, never models or provider SDKs directly? - [ ] Capability registry exists with tier (A/B/C), schemas, and routing per capability? - [ ] A model profile exists for every model in routing, verified within 90 days? - [ ] Design ceiling for input size set to what the weakest acceptable fallback can handle? - [ ] Adapter performs feature negotiation and refuses to route where it cannot emulate? - [ ] Gateway used for transport and policy, and replaceable within a quarter?

Prompts - [ ] Prompts split into an invariant base (policy) and per-family overlays (tuning)? - [ ] No policy content in any overlay? - [ ] Base prompts free of model-specific constructs (tags, prefill, version names)? - [ ] Registry refuses to route to a model family without an evaluated overlay?

Contracts and evals - [ ] Behavioral contract per Tier A/B capability: invariants, quality tolerance, SLOs, style? - [ ] Datasets sized for the decision (500+ for Tier A near tolerance), stratified, versioned? - [ ] Paired comparison (McNemar or equivalent) used for swap decisions? - [ ] Cost per completed task is the cost metric? - [ ] Eval runs through the production adapter, with a pinned judge from a different model family? - [ ] Quarterly candidate run includes the next version of the current primary?

Fallback and state - [ ] Tier A capabilities have a hot fallback with keep-alive traffic? - [ ] Fallback quota sized for realistic failover share, not keep-alive? - [ ] Degradation ladder (L0–L4) defined, owned, and in the runbook? - [ ] Conversation history stored in a provider-neutral canonical form? - [ ] Failover cost reserve accounts for cache-cold periods?

Governance - [ ] Portability scorecard reviewed quarterly; scores under 14/21 on the debt backlog? - [ ] Provider game day run for at least one Tier A capability per quarter? - [ ] Lock-in ledger records every deliberate provider-specific dependency? - [ ] Provider contracts reviewed for termination notice, change of control, and retention terms? - [ ] Provider concentration metric tracked against a policy cap?


EXERCISE — Map the Ten Layers: Pick one production AI capability in your organization and its most plausible fallback model. Walk through the ten incompatibility layers in §37.2 and, for each, write down: what is different between the two models, whether the difference is loud (it would error) or silent (it would degrade), and whether anything you have today would catch it. Count your silent, uncaught differences. That number is your real switching risk, and it is usually larger than the team's estimate.

PONDER — The Ten-Day Notice: Your largest model provider tells you on Friday that access ends in ten days. Leave aside what you would do. List what you would discover: which capabilities call the provider directly, which prompts have never run on another model, which conversation histories are stored in the provider's format, which fallback accounts have no quota, and who owns the decision to send traffic to a lower-quality tier. Then ask: what would it cost to discover those things this quarter instead, on your own schedule?

WORKSHOP — Build the Portability Kit for One Capability: Choose a Tier A candidate capability (for example ticket triage, document extraction, or a RAG answer step) and produce: (1) a capability interface with typed input and output schemas; (2) model profiles for the current primary and one fallback from a different provider, covering all ten layers; (3) a base prompt plus one overlay per model family, with each overlay line justified; (4) a behavioral contract file with all four tiers and a justified dataset size; (5) a fallback design stating its readiness level, keep-alive share, failover triggers, degradation ladder, and an estimate of the cache-cold cost; (6) a scorecard self-assessment with the three lowest-scoring dimensions and a 90-day plan to raise them. Present it as if to an architecture review board that must approve the capability for production.


This is the final module of this reference (37 modules as of October 2026). Related modules: §2.5 (lock-in vectors), §21.2 (the commodity thesis), §17.3 (the LLM gateway), §24.3–24.4 (shadow, canary, rollback), §25.3 (contract clauses), §28.2 (model debt). Volatile facts live in Appendix G.