MODULE 2 — The Model Ecosystem: Closed, Open, Reasoning, Small¶
⚠️ Currency disclaimer — model names, versions, prices and benchmark figures in this module are accurate as of October 2026. New models ship every few weeks, rankings shift, and older versions get deprecated. Before you make a decision or quote a figure, check the current state with the sources in §2.1 — Where to Check the Latest Models. The course keeps the full dated reference (models, pricing, platforms, regulation) in Appendix G — Current Landscape. The architectural patterns (routing, fallback, data residency, version pinning) do not expire. The specific model names do.
2.1 The Landscape Has Permanently Fractured¶
For most of 2023 and 2024, the model question was simple: use GPT-4 unless you have a reason not to. That era is over. The model landscape in 2025–2026 has fractured into competing tiers with no single winner across all categories.
LANDSCAPE SNAPSHOT — October 2026 (this box dates fastest; re-verify before quoting it)
- Frontier flagships: GPT-6 Astra (OpenAI, Sep 2026); Claude Fable 5.1 and Claude Opus 5.5 (Anthropic); Gemini 3.1 Pro (Google, preview). Rankings depend on which benchmark or index you read
- Balanced / everyday tier: GPT-6 Sol (OpenAI), Claude Sonnet 5.5 (Anthropic), Gemini 3.8 Flash (Google, GA)
- Common default for production agentic and coding workflows: Claude Opus 5.x family (Anthropic)
- Efficient tier: GPT-6 Luna (OpenAI), Claude Haiku 4.5 (Anthropic), Gemini Flash-Lite (Google)
- Best cost-efficiency near frontier: DeepSeek V4 (Pro / Flash) — MIT license, 1M-token context, a fraction of flagship pricing
- Open-weight leaders: DeepSeek V4, Kimi K3 (Moonshot), GLM-5.2 (Zhipu), Qwen 3.6 (Alibaba) — the open frontier is now led by Chinese labs, while Meta has shifted its newest flagship (Muse Spark) to closed weights
- Best for edge/on-device: Gemma 4 small variants (Google), Phi-4-mini (Microsoft), Qwen small variants (Alibaba)
Treat every name above as perishable. The categories and the reasoning in the rest of this module are what to retain. Pricing and the full model reference are in Appendix G (G.2, G.3).
Where to Check the Latest Models¶
The models named in this module are a snapshot as of October 2026 and will change. Use these sources to get the current picture:
Official model catalogs and deprecation schedules (authoritative)
| Provider | Current models | Deprecations / retirements |
|---|---|---|
| OpenAI | platform.openai.com/docs/models | platform.openai.com/docs/deprecations |
| Anthropic (Claude) | docs.anthropic.com — Models overview | docs.anthropic.com — Model deprecations |
| Google (Gemini) | ai.google.dev/gemini-api/docs/models | ai.google.dev/gemini-api/docs/changelog |
| Mistral | docs.mistral.ai — Models | Listed on the same page |
| DeepSeek | api-docs.deepseek.com | "News" section of the same API docs |
| Managed private deployments | Microsoft Foundry (Azure OpenAI) models · AWS Bedrock models · Google Cloud Model Garden | Provider docs |
Independent leaderboards and benchmarks (comparative; read several, not one)
- LMArena — rankings from crowd-sourced human preference votes, including coding and vision arenas
- Artificial Analysis — intelligence index, price, speed and latency compared across providers
- OpenRouter Rankings — what developers actually use, measured by token volume
- SWE-bench — real-world software engineering benchmark
- Epoch AI — Benchmarking hub — independent benchmark tracking and trends
Open-weight and small models
- Hugging Face Models (sorted by trending) — new open-weight releases, model cards, licenses
- Hugging Face Open LLM Leaderboard — standardized open-model evals
- Ollama Library — what runs locally right now, by size
Rule of thumb: use the provider's own docs for facts (names, context windows, prices, deprecation dates), leaderboards for relative comparisons, and your own eval set (§2.4, Step 4) for the actual decision. Benchmarks measure general capability, not fit to your task.
The practical implication for architects: "Which model should we use?" is now the wrong question. The right question is "Which model family for which task class, at what cost tier, within what data residency constraint?" Every production AI system of any scale should be running more than one model, routing tasks to the appropriate tier.
As one benchmark tracking service noted: the AI industry logged 255 model releases from major organizations in Q1 2026 alone. An architect who is not following the model landscape is making model selection decisions based on 6-month-old assumptions in a field where the competitive picture changes every 6 weeks.
2.2 The Five Model Categories Every Architect Must Understand¶
Category 1: Frontier Closed Models (API-only)¶
What they are: The highest-capability models from major labs — OpenAI (GPT-6 / GPT-5.x), Anthropic (Claude Fable 5.1, Opus / Sonnet 5.x), Google (Gemini 3.x), SpaceXAI (Grok, formerly xAI), and now Meta (Muse Spark, closed-weight). Accessible only via API. Weights are not released. Data is processed on the provider's infrastructure.
Architectural implications:
Data residency. Customer data, internal documents, PII — everything sent to these models crosses your network perimeter to a third-party cloud. This is the single biggest architectural concern for regulated industries. Before selecting any frontier closed model, the question is not "is the provider trustworthy?" but "does our data classification policy permit this data to leave our infrastructure?" Many organizations answer this question informally and inconsistently. It must be answered formally, in writing, by legal and compliance, before any production deployment.
Vendor dependency and deprecation risk. These models are deprecated on the provider's schedule. GPT-4-0613 was deprecated. Claude 2 and Claude 3 Opus were deprecated. Every model you deploy against has a deprecation date. Architecturally, this means: (1) never hardcode a model name in application code — always go through a configuration layer or gateway that can be updated without a deployment; (2) run regression evals when the model version changes, because behavior changes subtly with each version update; (3) have a multi-vendor fallback — if your primary provider has an outage or deprecates a key model, what is the fallback? If the answer is "we wait," that is not an architecture.
Silent behavior changes. Model providers update model behavior with silent updates (safety filter changes, response style changes, capability changes) that are not announced as version bumps. A model you tested in production can behave differently 6 weeks later without any notification. This is why prompt regression testing (Module 3) and eval pipelines (Module 7 of the integration patterns reference) are not optional.
Pricing volatility. Frontier model pricing has dropped dramatically — costs that would have been $30/M tokens in 2023 are now $3–6/M for equivalent capability. But pricing is set unilaterally by the provider. A production system that is economically viable today may not be viable if the provider changes pricing. The architecture mitigation: abstract the model call through a gateway, maintain cost models at the task level not the provider level, and design for model substitutability.
When to use: Tasks where the quality ceiling matters — complex reasoning, nuanced generation, agentic workflows requiring sophisticated multi-step planning. The cost is justified when the quality delta between a frontier model and a cheaper alternative is meaningful for the use case.
When not to use: High-volume classification, entity extraction, summarization of structured data, simple FAQ retrieval. Using a frontier model for tasks a smaller model handles adequately is a cost architecture failure.
The middle path: Managed Private Deployments
Between "send data to OpenAI's public API" and "self-host Llama on your own GPU cluster" is a frequently overlooked architectural option: managed private deployments of frontier models.
-
Microsoft Foundry (formerly Azure AI Foundry / Azure OpenAI Service) — Microsoft deploys OpenAI models (GPT-5.x / GPT-6) inside Azure's infrastructure, and the Foundry catalog also includes Claude and thousands of other models. Your data stays within your Azure tenant. Microsoft contractually commits that your prompts are not used to train models. For organizations that need frontier model quality with stronger data residency than the public API provides, this is the default enterprise path for Microsoft-aligned organizations.
-
AWS Bedrock — Hosts Claude (Anthropic), Llama (Meta), Mistral, DeepSeek, and others inside AWS VPC. Since mid-2026 it also hosts OpenAI's frontier models (GPT-5.5 onward, including the GPT-6 family), so "Azure is the only enterprise route to OpenAI" no longer holds. One cloud can now host your primary and your cross-vendor fallback (see Module 37). Requests stay within your AWS environment. Supports private endpoints, removing public internet exposure entirely.
-
Google Gemini Enterprise Agent Platform (formerly Vertex AI, renamed April 2026) — Hosts Gemini (and Claude) models inside Google Cloud with similar data residency and access control capabilities.
These managed services solve the primary data residency concern of frontier closed models without requiring you to operate GPU infrastructure. The trade-off: you are still dependent on the cloud provider's model availability and deprecation schedule, and the models available are typically a version behind the public API release. For most regulated enterprises, this is the right default before self-hosting is considered.
Category 2: Open-Weight Models¶
What they are: Models whose weights are publicly available — DeepSeek V4, Moonshot Kimi K3, Zhipu GLM-5.2, Alibaba Qwen 3.6, Mistral Large 3, Meta Llama 4 (Scout, Maverick), Google Gemma 4, NVIDIA Nemotron. (Llama 4 Behemoth was announced but never released as open weights, and Meta's newest flagship, Muse Spark, is closed-weight. A vendor's open-weight commitment is itself a risk to track.) Can be self-hosted, fine-tuned, and run inside your infrastructure perimeter.
Why open-weight changes the architecture fundamentally:
Data never leaves your perimeter. For financial services, healthcare, government, and any organization handling data that cannot be sent to external APIs — open-weight models running on internal infrastructure are the only viable path. This is not a preference. It is a compliance requirement for many data classifications.
DeepSeek's geopolitical consideration. DeepSeek (Chinese origin, MIT license, outstanding quality-per-cost) is the cost efficiency leader in the open-weight category. However, for organizations with data sovereignty requirements vis-à-vis China's Data Security Law — or for any data that should not be processed on servers subject to Chinese jurisdiction — using the DeepSeek API is architecturally inappropriate. Self-hosting DeepSeek weights on your own infrastructure resolves the data residency concern while retaining the cost advantage. This is a real architectural decision that teams are getting wrong by defaulting to the DeepSeek API without assessing data residency implications.
Infrastructure cost and complexity. Self-hosting a 70B parameter model requires meaningful GPU infrastructure. DeepSeek V4 Pro at full size (1.6T parameters, MoE architecture with ~49B active per inference) requires a multi-node GPU cluster. The calculus: at sufficient volume, the GPU infrastructure cost is lower than the API call cost. The break-even point depends on usage volume, GPU instance pricing, and operational overhead. For low-to-medium volume, the API is cheaper. For high volume with sensitive data requirements, self-hosting is cheaper and compliant.
Operational burden. You now own the inference infrastructure. Scaling, failover, patching, model version management — these are your operational responsibilities. A team that selects self-hosted open models without accounting for this operational burden will end up with a system that has better data governance and worse reliability than the API alternative.
Fine-tuning capability. Open-weight models can be fine-tuned on your domain data. This allows quality improvement for domain-specific tasks at lower parameter counts than a general frontier model requires. A fine-tuned 7B model can outperform a general-purpose 70B model on a narrow, well-defined task. The architectural trade-off: fine-tuning creates a versioning and retraining lifecycle that you must own and operate.
Key open-weight models and their architectural sweet spots (as of October 2026):
| Model | Parameters | Primary Strength | Self-Host Requirement |
|---|---|---|---|
| DeepSeek V4 Pro | 1.6T / ~49B active | Near-frontier general, reasoning, code; 1M context; MIT | Multi-node GPU cluster |
| DeepSeek V4 Flash | 284B / ~13B active | Cost-efficient general; 1M context; MIT | Single 8-GPU high-end node |
| Kimi K3 (Moonshot) | 2.8T / ~50B active | Agentic, long context (1M) | Multi-node GPU cluster |
| Mistral Large 3 | 675B MoE | Instruction following, multilingual; Apache 2.0; EU vendor | Multi-node GPU cluster |
| Llama 4 Maverick | 400B / 17B active | General purpose, 1M context, image input | Single 8-GPU node (quantized) / multi-node |
| Llama 4 Scout | ~109B MoE | Cost-efficient, long context | Single high-end node |
| Qwen 3.6 (35B-A3B) | 35B / 3B active | Very cheap inference, 262K context; Apache 2.0 | Single mid-tier GPU |
| Gemma 4 (E2B/E4B · 26B MoE · 31B) | ~2–4B effective (edge) to 31B dense | Apache 2.0; edge models take image/video/audio input (128K ctx); 26B/31B for local servers (256K ctx) | Phone/laptop (E2B/E4B) · single GPU (26B/31B) |
| Phi-4-mini | 3.8B | On-device, edge, strong reasoning per param | CPU or single consumer GPU |
Jurisdiction note: four of the strongest open-weight families (DeepSeek, Kimi, GLM, Qwen) come from Chinese labs. Self-hosting the weights addresses data residency, as described above. It does not address supply-chain or procurement policies that restrict models by origin, so check those separately.
Category 3: Reasoning Models — A Structurally Different Architecture Decision¶
Reasoning models represent a qualitatively different approach to model capability and have distinct architectural implications. The category began as separate model lines (OpenAI's o-series, DeepSeek R1, Claude 3.7 extended thinking). By 2026, reasoning is mostly a mode of the flagship models rather than a separate family: GPT-5.x/6 with a reasoning-effort setting, Claude Opus/Sonnet 5.x with extended/adaptive thinking, Gemini thinking modes, and DeepSeek V4. The architectural question has moved from "which reasoning model?" to "how much reasoning, on which calls?"
What makes them different:
Standard LLMs generate responses using a single forward pass through the network. They "think" and "answer" in one step. Reasoning models use test-time compute — they perform an extended internal reasoning process (often visible as a "thinking" chain) before producing a final response. The quality of the answer improves as more compute is allocated to the reasoning phase.
The architectural implication: reasoning models trade latency and cost for accuracy on hard problems. This is a spectrum, not a binary.
REASONING MODEL COST/LATENCY/QUALITY TRADE-OFF
Simple Query (email draft, FAQ answer, classification)
─────────────────────────────────────────────────────
Use: Standard model (fast, cheap)
Reasoning model here = paying 10x for no quality gain
Medium Complexity (document summarization, code review, data extraction)
────────────────────────────────────────────────────────────────────────
Use: Standard frontier or efficient open-weight
Reasoning adds modest quality improvement at significant cost
High Complexity (multi-step analysis, legal reasoning, mathematical derivation,
complex code generation, compliance policy interpretation)
───────────────────────────────────────────────────────────────────────────────
Use: Reasoning model with calibrated thinking budget
Quality improvement is material and justifies cost
Critical/Irreversible Decisions (fraud investigation, medical diagnosis support,
contract risk analysis, security vulnerability assessment)
──────────────────────────────────────────────────────────────────────────────────────────
Use: Reasoning model + human review
The combination of deep reasoning + human judgment is the architecture
The thinking budget concept:
Reasoning depth is a developer-controlled setting. Early APIs exposed it as an explicit token budget (e.g., Claude's original extended thinking budget_tokens: 1K thinking tokens for a medium task, 32K for a critical analysis). Current APIs mostly expose an effort level (low / medium / high) or adaptive thinking, where the model decides how long to think within the bounds you set. Either way, the cost-quality trade-off is a tunable parameter, not a fixed property of the model. Set it per task class in your routing configuration (§2.3), not ad hoc in prompts. This is an architectural tuning lever, not a prompt engineering trick.
Reasoning models in agentic systems — important caveat:
Reasoning models are significantly slower than standard models. A single high-effort reasoning call may take 30–90 seconds for a complex problem. In an agentic loop that makes 10 LLM calls to complete a task, using a reasoning model for every call produces a workflow that takes 5–15 minutes. Most agentic tasks do not require reasoning model quality for every step — only for the planning and decision steps. The pattern: use a reasoning model for the planning phase, use a standard model for the execution and tool-call interpretation steps. This preserves quality where it matters while keeping latency manageable.
Cost reality check:
Reasoning tokens are typically billed at 3–5x the rate of standard output tokens. A reasoning model call that uses 20,000 thinking tokens costs roughly equivalent to 60,000–100,000 standard output tokens. At scale, routing everything through a reasoning model is financially indefensible unless the task genuinely requires it. The routing decision — which tasks go to reasoning, which to standard — is an architectural decision with direct cost implications.
Category 4: Small Language Models (SLMs) — The Edge Revolution¶
What they are: Models under ~10B parameters designed for efficiency — Gemma 4 small variants (e.g., E4B), Phi-4-mini (3.8B), Qwen 3.5 small variants (0.8B–9B), Llama 3.2 (1B, 3B), SmolLM3. Designed to run on consumer hardware, edge devices, and air-gapped environments.
Why this category matters architecturally:
The quality improvement in SLMs between 2023 and 2026 has been extraordinary. Microsoft's Phi-4-mini (3.8B parameters) now outperforms the 70B-class models of 2023 on structured reasoning and instruction following. A 3.8B model that fits comfortably on a single GPU — or on a CPU — and delivers what GPT-4-level quality of 18 months ago is not a compromise. It is a different deployment topology.
Deployment tiers:
TIER 1: Mobile and IoT (sub-3B parameters, <4GB RAM)
Models: Gemma 4 E4B, Phi-4-mini, Llama 3.2 1B/3B, SmolLM3
Hardware: Smartphone NPU, Raspberry Pi, IoT sensors
Latency: <100ms on-device
Use cases: On-device classification, offline assistant,
real-time sensor data interpretation
Data: Never leaves the device
TIER 2: Edge Server / Desktop (3B–9B parameters, 8–16GB RAM)
Models: Phi-4, Qwen 3.5 9B, Gemma 3 12B
Hardware: NVIDIA Jetson Orin, standard workstation, NUC
Latency: 80–200ms
Use cases: Air-gapped enterprise workflows, manufacturing
quality control, healthcare record summarization
where data cannot leave the facility
Data: Stays within facility network
TIER 3: Private Cloud (9B–70B parameters, 1–4 GPUs)
Models: Qwen 3.6 35B-A3B (MoE), Gemma 3 27B,
Qwen 3.5 72B, Llama 3.3 70B
Hardware: Single high-end GPU node (A100, H100)
Latency: 100–500ms
Use cases: High-volume API tasks with data residency
requirements, fine-tuned domain models
Data: Within your cloud VPC
When SLMs are the right architectural choice:
Data sovereignty is non-negotiable. Healthcare patient records, classified government data, financial transaction data under strict jurisdiction requirements — if data cannot leave the building, SLMs are the only path.
Latency requirements below API feasibility. API round-trips add 200–500ms of network latency. On-device inference adds 20–100ms. For real-time applications — manufacturing quality control at 30fps, point-of-sale interactions, vehicle sensor processing — API latency is architecturally infeasible. Local SLMs are the only viable approach.
High volume with simple task profiles. If a task is high-volume and well-defined (extract these 5 fields from this document template, classify this support ticket into one of 12 categories), a fine-tuned SLM may outperform a frontier model on that specific task at 1/100th the cost. The "more parameters = better for every task" assumption has been empirically disproven.
Air-gapped environments. Factories, military, financial infrastructure with air-gap requirements — the model must run where the network cannot reach. SLMs are the architecture, full stop.
What SLMs cannot replace:
Complex reasoning, nuanced creative generation, open-ended multi-domain questions, tasks requiring broad world knowledge. The right architecture often uses SLMs for classification and extraction at the edges and frontier models for complex reasoning at the center.
Category 5: Typed-Decision Models (Non-Generative) — The Newest Tier¶
What they are: Models that do not generate text at all. You define the answer shape in advance — a fixed set of categories, a bounded numeric score, a boolean, an enum of up to a few hundred values — and send the model a piece of state plus that schema. It returns the schema filled in, with a calibrated probability on each field, sampled in parallel rather than decoded token by token. The first commercial example (TypeSafe AI's "Jev," released in early access on September 15, 2026) kicked off a small category — "System One models," named after Kahneman's fast/intuitive System 1 vs. slow/deliberate System 2 — and several providers followed with their own versions within weeks.
Why the probability can be trusted — the training objective is different, not just the architecture: Standard RLHF/DPO (Module 34) trains a model to produce outputs a human (or a preference model standing in for one) prefers. These models are trained with a different target called RLCD — Reinforcement Learning for Calibrated Decisions (unrelated to an older, same-acronym 2023 paper on contrastive preference distillation — a pure naming collision, not the same technique). RLCD rewards the model for calibration: on a static dataset of known-correct answers, the reward is based on how closely the model's stated probability distribution matches the actual outcome rate — not on whether a human liked the answer. Concretely: if the model says 80% confidence across 10,000 decisions, roughly 8,000 of those should actually be correct. That is a fundamentally different thing to optimize for than "pick the right answer" or "write what a rater prefers," and it's the reason the probability attached to each field is meant to be read literally rather than as a vague confidence gesture — though as with any vendor-reported training claim, verify calibration on your own data before trusting it in a gate (Module 20's vendor-claim filter applies here too).
Why this is architecturally different, not just a smaller/cheaper LLM:
Every category above — frontier, open-weight, reasoning, SLM — still generates text autoregressively and still needs the downstream apparatus this whole course spends dozens of pages on: output parsing, schema validation, a faithfulness check, a confidence gate, retry-on-malformed-output. A typed-decision model skips that apparatus for the specific step it's used on, because the output space was enumerated before the call was made. It cannot return a value outside the schema, and it cannot drift into prose. The trade is total: you give up open-ended generation completely in exchange for a decision step that is fast, cheap, and structurally incapable of hallucinating the kind of output your code wasn't expecting.
GENERATIVE MODEL (any of Categories 1-4) TYPED-DECISION MODEL (Category 5)
Prompt: "Classify this ticket and Schema: { category: enum[12 values],
explain your reasoning" urgency: int[1-5],
│ needs_human: bool }
▼ State: {ticket text, account tier, history}
[Autoregressive decode, token by token] │
│ ▼
▼ [Parallel sampling against the
Free-text output fixed schema — no decode loop]
│ │
▼ ▼
[Parser] → [Schema validator] → { category: "billing", urgency: 4,
[Retry if malformed] → [Confidence gate] needs_human: false, p: 0.92 }
│ │
▼ ▼
Usable value (maybe) Usable value (always, by construction)
When this is the right architectural choice:
Any step you were already trying to make deterministic. This course repeatedly argues for keeping classification, routing, and gating decisions out of free-text LLM calls precisely because of the parse/validate/retry overhead and residual hallucination risk (Module 6's agent controls, Module 12's LLM-as-judge calibration, the routing classifier in §2.3 below). A typed-decision model is a new, model-backed way to implement exactly those steps — it belongs at the classifier/gate layer, not as a chat replacement.
High-volume, narrow, well-specified decisions. Ticket triage, fraud-score-and-flag, content moderation categories, approve/escalate gates — anywhere the answer space was already a fixed list before you wrote the prompt. Reported cost/latency gains over calling a frontier model for the same decision are large (vendor-reported, verify independently): Jev, for example, is reported at 40–200x faster than frontier LLMs and priced at $0.042 per million input tokens, because there's no decode loop.
Anywhere "the model made up a category we didn't define" is a production incident. Since the output space is the schema, that failure mode is structurally impossible rather than something you catch downstream.
What it cannot replace:
Anything open-ended — drafting, summarizing in the user's own words, multi-turn conversation, reasoning chains you want to inspect, or any task where you don't already know the shape of a correct answer. It is not a cheaper GPT-5.x; it is a different kind of component that sits next to your LLM calls, not in place of all of them. Treat vendor benchmark numbers (speed/cost multipliers, "cannot hallucinate" claims) with the same skepticism Module 20 teaches for any vendor claim — the mechanism (enumerated output space, no decoding) is sound and durable; the specific multiples and the specific vendor landscape (currently a handful of competing providers, a few months old) will move fast and are not something to anchor on by name.
The category is already multi-vendor (October 2026, reported). Besides Jev, there is Laya (Convai Innovations; open source, a ~421M-parameter encoder) and Cloudflare's Clef and Clef-flash (released October 1, 2026; open weights under Apache 2.0, post-trained from Qwen models). Clef is reported to be API-compatible with Jev, and not every model in the category is built the same way: Laya is an encoder, while Clef starts from a decoder language model. TypeSafe's RLCD method itself has not been published, so what is public is the behaviour (typed questions in, calibrated probabilities out), not the recipe. For where to place these models in a system, how to use their probabilities, and how to evaluate and adopt one, see Module 16 §16.9.
2.3 The Model Routing Architecture¶
The maturity marker of an enterprise AI architecture is not which model it uses — it is whether it routes intelligently across models. Model routing is the practice of directing different tasks to different models based on complexity, cost, latency, and data sensitivity.
INCOMING REQUEST
│
▼
[REQUEST CLASSIFIER] (lightweight model or rule-based)
│
├── Task type: classification/extraction
│ └── Route: SLM or efficient open-weight (~$0.001/call)
│
├── Task type: standard generation/summarization
│ └── Route: Efficient frontier (DeepSeek V4 Flash, Gemini Flash, Claude Sonnet-class, etc.) (~$0.01/call)
│
├── Task type: complex reasoning/analysis
│ └── Route: Frontier model or reasoning model (~$0.10/call)
│
├── Task type: critical/high-stakes decision support
│ └── Route: Reasoning model + human review gate (~$0.50–$2.00/call)
│
└── Data sensitivity: regulated/confidential
└── Override: Self-hosted model regardless of task complexity
The classifier itself should be a lightweight, fast model (or rule-based for clear categories). The overhead of the routing decision should be negligible — 5–10ms — not a second LLM call that costs more than the task itself. A typed-decision model (§2.2, Category 5) is a natural fit for this classifier step specifically, since routing is itself a bounded, enumerable decision.
Enterprise model selection should not chase a single strongest model but adopt a hybrid routing strategy — routing simple tasks to low-cost models and complex reasoning tasks to premium models — which can reduce API costs by 60–80% while maintaining over 95% quality.
The routing decision record. The routing rules are architectural decisions, not configuration details. They should be documented: "Tasks of type X go to model Y because of constraint Z." When the routing rules change — because a new cheaper model is available, or because quality requirements have changed — the change should go through the same review process as any architectural change.
2.4 The Model Selection Framework¶
When a team asks "which model should we use?" — this is the framework for answering it.
Step 1: Define the task profile
- What is the input? (structured data, natural language, documents, code, images)
- What is the output? (classification, generation, extraction, reasoning, code)
- What is the quality bar? (human-reviewed outputs, customer-facing, internal tooling)
- What is the latency SLA? (real-time <500ms, interactive <3s, batch >3s)
- What is the volume? (requests/day, requests/month)
Step 2: Apply the constraints
- Data residency: can this data be sent to external APIs?
- Compliance: does any regulation restrict which providers handle this data?
- Geography: which jurisdictions does data cross?
- Vendor policy: does your organization have approved vendor list requirements?
Step 3: Calculate unit economics
- Estimate input tokens per request (typical and P95)
- Estimate output tokens per request
- Multiply by candidate model pricing
- Calculate monthly cost at current volume and at 3x volume
- Compare against the cost of a simpler model that meets the quality bar
Step 4: Evaluate quality fit
- Run the candidate model against 20–50 representative examples from your actual data
- Score against your quality bar (human evaluation or automated scoring)
- Compare quality vs. cost across 2–3 candidate models
- The question is not "which model scores higher on benchmarks" — benchmark scores measure general capability, not fit to your specific task profile
Step 5: Plan for change
- What is the model's deprecation risk? (newer model, more stable track record = lower risk)
- What is the fallback if this model is unavailable or deprecated?
- How is the model name/version abstracted in the codebase? (it must be configurable without deployment)
- What regression eval will you run when you switch models?
2.5 Vendor Lock-In Strategy¶
Every frontier model provider wants to be your primary — and ideally only — model vendor. The features, SDKs, tooling integrations, and pricing incentives are all designed to make switching harder. This is a genuine architectural risk that most teams underestimate.
Lock-in vectors:
API format differences. OpenAI's API format (chat completions) has become a de facto standard that many providers implement. But tool calling schemas, structured output formats, streaming protocols, and advanced features differ enough that switching providers requires code changes. Mitigation: abstract all LLM calls through a gateway (LiteLLM, Portkey) that normalizes API formats.
Context window dependencies. 1M-token context windows are now common at the frontier (Claude, DeepSeek V4, Kimi K3, GLM-5.2), but many cheaper, self-hosted, and edge models still run at 32K–262K. If your application is designed around stuffing a 1M-token context and you need to fail over to a 128K model, the architecture changes significantly. Design for the lowest common denominator context window you might need to target.
Feature dependencies. Relying on provider-specific features (OpenAI's Assistants API, Anthropic's extended thinking, Google's multimodal file handling) creates lock-in at the feature level. For each provider-specific feature you adopt, evaluate: what does this cost in switching friction if we need to change providers?
Embedding model lock-in. If your vector store is populated with embeddings from OpenAI's text-embedding-3-large, switching to a different embedding model requires re-embedding every vector. At millions of documents, this is a non-trivial migration. The mitigation: use an open-source embedding model (e.g., nomic-embed-text, bge-m3) that you can self-host, so the embedding layer is not tied to a single vendor.
The multi-vendor architecture:
LLM GATEWAY
├── Primary: Claude Opus 5.x (complex reasoning, agentic)
├── Secondary: GPT-5.x / GPT-6 (fallback, specialized tasks)
├── Cost tier: DeepSeek V4 Flash / Gemini Flash (high-volume, lower-complexity)
└── Private: Self-hosted DeepSeek V4 / Qwen 3.6 / Llama 4 (sensitive data)
EMBEDDING LAYER
└── Open-source embedding model (self-hosted, no vendor dependency)
FALLBACK CHAIN
└── Primary unavailable → Secondary
Secondary unavailable → Cost tier (with quality flag in response)
All API unavailable → Self-hosted (reduced capability, but available)
This is not about using every model everywhere. It is about ensuring that no single vendor failure or deprecation decision takes your entire AI capability offline.
Going deeper: A gateway with several vendors is necessary but not sufficient. A gateway normalizes the API call, not the model's behavior. Module 37 — Model Portability covers the rest: the ten layers where model swaps break, capability interfaces, model profiles, prompt overlays, behavioral contracts, hot/warm/cold fallbacks, a step-by-step swap playbook, and real cases of providers withdrawing access.
2.6 Model Versioning and Deprecation Management¶
Model version pinning. Never use a model name without a version. gpt-4 is not a stable target — it changes. gpt-4-0613 is pinned. claude-3-opus-20240229 was pinned (and has since been retired, which makes the point). When the default alias changes behavior, pinned versions maintain consistent behavior until you explicitly migrate.
Deprecation timeline awareness. Every pinned model version has a deprecation date. Track these. Set a calendar reminder 90 days before deprecation to begin migration testing. A model that is deprecated with 30 days notice and no tested alternative is a production incident.
Migration process:
When a model is deprecated or you choose to upgrade: 1. Identify all services using the model (this requires the gateway's usage logs) 2. Run the eval suite against the candidate replacement model 3. Compare outputs for a sample of production inputs 4. Deploy the new model in shadow mode (both models run, responses compared, old model serves users) 5. If eval scores and shadow comparison are acceptable, promote the new model 6. Retire the old model from the configuration
Without this process, model migrations become emergency incidents driven by deprecation deadlines.
2.7 Multimodal Models — Architectural Considerations¶
The frontier models are all multimodal: text, images, audio, documents, video (to varying degrees). For most production AI applications today, multimodality is an optional enhancement. For some use cases, it changes the architecture fundamentally.
When multimodality changes the architecture:
Document understanding. Processing PDFs, invoices, contracts, medical records with complex layouts, tables, charts, and mixed text/image content. A text-only pipeline requires a separate OCR/parsing step (Textract, DocumentAI, Unstructured) before the LLM sees the content. A multimodal model can process the document as-is. Trade-off: multimodal calls are more expensive than text-only; the parsing approach gives you more control over chunking and metadata extraction.
Real-time voice AI. OpenAI's gpt-realtime-2.x (GPT-5-class reasoning, 128K context, parallel tool calls) and Google's Gemini Live API enable low-latency voice-to-voice AI interactions. The two differ on cost, session length, and video support, so evaluate both. The architecture is different: audio streams in, audio streams out, the traditional text I/O pipeline does not apply. If you are building voice-first interfaces, this is a specialized architecture domain requiring specific latency, streaming, and audio processing patterns.
Visual QA on structured data. Charts, dashboards, CAD drawings, medical images — these require visual understanding that text models cannot provide. The architectural decision: preprocess visual content into text (chart-to-table extraction, image captioning) and use a text model, or use a multimodal model directly. The former is more transparent and auditable; the latter is simpler and often more accurate.
What multimodality does not solve:
Multimodal does not mean all-format. Video understanding is nascent and expensive. Real-time audio processing at production scale has different infrastructure requirements than text. Adding multimodal capability to a system adds cost, latency, and complexity. The question is always: does this use case require multimodality, or can the task be solved by pre-processing the non-text content into text?
2.8 The Questions That Surface Model Architecture Gaps¶
When a team presents a model selection decision, these are the questions that reveal whether the decision was thought through:
"If your primary model provider has a 4-hour outage tonight, what happens to your users?" This surfaces fallback architecture gaps. The answer should be: "Users experience [specific degradation mode] because [fallback configuration]." Not: "It would be a significant incident."
"When was the last time this model's behavior changed silently, and how did you find out?" This surfaces eval and monitoring gaps. Teams that don't monitor model behavior in production cannot answer this question.
"How much did your AI features cost last month, broken down by feature?" This surfaces cost attribution gaps. The answer should be a dashboard, not an estimate.
"What data do we send to this model provider, and is that consistent with our data classification policy?" This surfaces data governance gaps. The answer should be a documented data flow, not "I think it's fine."
"If we needed to switch this model to an open-weight self-hosted alternative tomorrow, what would break and how long would it take?" This surfaces lock-in gaps. A well-architected system can answer this precisely. A locked-in system cannot.
EXERCISE — Model Routing Design: You have a customer support system that handles 500,000 requests per month. Request types: 40% are simple FAQ lookups, 35% are account status queries with structured data, 20% are complaint analysis and response drafting, 5% are complex escalation analysis requiring reasoning. Design the model routing architecture. For each tier: which model, estimated cost per request, monthly total, and the quality bar that determines whether the cheaper model is acceptable.
PONDER — The Deprecation Question: List every AI model your organization is currently using in production. For each: do you have the version pinned or are you using an alias? When is the pinned version deprecated? Do you have a tested fallback? If you are using an alias (unpinned), what changed the last time the alias was silently updated? If you don't know — that is the architectural gap.
WORKSHOP — Build vs. Buy vs. Host: For a specific use case in your organization (choose one with real data sensitivity requirements): compare three paths — (1) OpenAI/Anthropic API, (2) DeepSeek API, (3) Self-hosted open-weight model (DeepSeek V4, Qwen 3.6, or Llama 4) on internal GPU cluster. For each path: data residency compliance, monthly cost at current volume, operational burden, quality estimate, vendor lock-in exposure. Present the recommendation with explicit trade-off acknowledgment.
Next: Module 3 — Prompt Engineering as a Software Engineering Discipline