Skip to content

MODULE 13 — Cost Engineering & Token Economics

⚠️ Currency note: This module teaches the structure of AI cost: asymmetries, multipliers, and the decisions that move them. Those do not expire. Prices do. Every dollar figure in the worked examples is illustrative, based on October 2026 list prices. Current numbers live in Appendix G §G.3 — Pricing Reference. Re-run any example with current prices before you put it in front of a client or a CFO.

13.1 The Cost Problem Is Architecture, Not Procurement

Most organizations discover their AI cost problem when the monthly bill arrives. The CFO asks for a breakdown by product. Nobody can provide one. The DevOps team starts talking to the LLM provider about discounts. The engineering team starts looking at cheaper models.

None of this is cost engineering. Cost engineering is designing AI systems from the ground up with cost as a first-class architectural concern — before the system is built, not after the bill arrives.

The architectural decisions that have the largest impact on AI cost are made early and are expensive to change: - Which model for which task (affects cost by 100x) - How much context is included in every request (affects cost by 3-10x) - Whether semantic caching is implemented (can reduce API calls by 30-60%) - Whether model routing is implemented (often reduces cost by 50-80%) - Whether the async batch API is used for non-real-time workloads (50% discount)

These decisions, made casually in the first sprint, determine 80% of the system's cost profile at scale. Making them intentionally is cost engineering.


13.2 The Token Pricing Landscape: Structure, Not Snapshots

Token prices move every month. The shape of the price list moves much more slowly, and that shape is what an architect designs against. This section teaches the shape. The current numbers are in Appendix G §G.3, which is refreshed every quarter. Do not copy prices from this module into a cost model.

The Stable Structure of Token Pricing

FIVE STRUCTURAL FACTS (hold across providers and refreshes; numbers in G.3)

1. OUTPUT COSTS A MULTIPLE OF INPUT
   Closed frontier and mid-tier models: output ≈ 5–6x input per token.
   DeepSeek-style open-weight APIs: ≈ 2–3x.
   → Favor input-heavy, output-light designs: retrieve more, generate less.

2. CACHED INPUT COSTS ~1/10 OF UNCACHED INPUT (or less)
   True across the major providers as of October 2026. Cache writes may carry
   a premium, and caches expire (TTL) and are scoped to one model.
   → Caching is the largest "free" lever (Decision 2 below), and a failover
     to a model with a cold cache is a cost spike (Module 37 §37.7).

3. PRICE SPANS ~100x BETWEEN TIERS
   Efficient-tier models cost about 1/40 to 1/100 of frontier models per token.
   → Routing (Decision 1) is worth more than any negotiated discount.

4. BATCH ≈ 50% OFF
   Asynchronous batch APIs are priced at about half the synchronous rate at the
   major providers.
   → Anything that does not need an answer in seconds should be a batch job.

5. REASONING IS BILLED AS OUTPUT
   Thinking tokens are charged at the output rate, and the reasoning-depth
   control (effort / thinking level) sets how many you buy.
   → Set reasoning depth explicitly per task (Module 37 §37.2, layer 5).

The 2026 Traps: Pricing Patterns That Break Naive Cost Models

Four patterns became common in 2026. Each can make a cost model that multiplies tokens by a single price per token wrong by a large factor:

Trap What happens Example (October 2026, see G.3) Defense
Introductory prices with a step-up date A model launches at a discount that ends on a fixed date. The same traffic costs more the next day with no change on your side Gemini 3.8 Flash list price doubles on 2027-01-01 Record the step-up date in the model profile; model costs at the post-step-up price; set a calendar alert
Long-context surcharges on the whole request Above an input-size threshold, the entire request is billed at a higher rate, not only the tokens over the threshold GPT-6: above 272K input tokens, 2x input and 1.5x output for the whole request Know each model's threshold; enforce a design ceiling below it (Module 37 §37.3)
Tokenizer changes between versions The same text becomes a different number of tokens after a version upgrade, so cost, truncation, and budgets shift Anthropic documents that the tokenizer introduced with Opus 4.7 uses ~1.0–1.35x as many tokens as earlier models for the same text Re-baseline token counts on every model change, including same-vendor upgrades
Per-token price drops on newer flagships A newer version is cheaper per token but may use more tokens, retries, or reasoning to finish a task Claude Opus 5.5 is cheaper per token than Opus 5 Compare cost per completed task, not price per token (Module 37 §37.5)

Illustrative Tier Table (October 2026)

The table below shows the tiers and their typical price ratios. It is a teaching aid for the worked examples in this module, not a price list. It is a subset of Appendix G §G.3, which has the full table, cached-input prices, cloud-partner notes, and the verification links.

ILLUSTRATIVE TIER PRICING (October 2026 list prices, USD per 1M tokens, see G.3)

Tier                 Example models                       Input     Output   Out/In
──────────────────────────────────────────────────────────────────────────────────
Premium frontier     Claude Fable 5.1, GPT-6 Astra        $10.00    $50.00   5x
Frontier             Claude Opus 5.5                      $4.00     $20.00   5x
Mid-tier             Claude Sonnet 5.5, GPT-6 Sol         $2.00     $10.00   5x
Fast mid-tier        Gemini 3.8 Flash (introductory       $0.75     $3.75    5x
                       through 2026-12-31; $1.50 / $7.50
                       from 2027-01-01)
Small / fast         Claude Haiku 4.5                     $1.00     $5.00    5x
Efficient            GPT-6 Luna                           $0.10     $0.50    5x
Open-weight API      DeepSeek V4 Flash (reported range)   ~$0.10–0.14  ~$0.20–0.28  ~2x

Self-hosted open-weight: ~$0 per token, but you pay for GPUs whether or not
they are busy (see §13.5 for the break-even math).

Worked examples in this module use these illustrative prices. Wherever an example says "Claude Sonnet 5.5 ($2/$10)", substitute the current G.3 price for whichever model you actually route to.

The Output Token Cost Asymmetry

At the illustrative prices above, a closed mid-tier model charges $2.00 per million input tokens and $10.00 per million output tokens, a 5x multiplier. Closed frontier models sit at roughly 5–6x. This asymmetry has direct architectural consequences:

Verbose outputs are expensive. A system prompt that says "provide a comprehensive analysis" will cost significantly more than one that says "provide a 3-sentence summary." The difference in output length is a direct cost driver. This is not a quality trade-off — shorter, more precise outputs are often higher quality than verbose ones.

Reasoning tokens compound the asymmetry. Reasoning is now a mode of flagship models rather than a separate model family: effort levels on Claude, reasoning-effort settings on GPT, thinking levels on Gemini (Appendix G §G.2). Thinking tokens are charged at the output rate. A request that generates 20,000 thinking tokens before it responds pays for those 20,000 tokens at the output rate. At an illustrative frontier output price of $20/M, that is $0.40 in thinking alone, before the actual answer. Default reasoning depth also changes between versions (for example, Opus 5.5 defaults to a lower effort than Opus 5). Route reasoning selectively and set its depth explicitly. Do not rely on defaults.

Long context shifts cost to input volume. Per token, output is the expensive side. Per request, volume usually decides. Reading a 1M-token context at an illustrative $4/M input costs $4.00. A 500-token answer at $20/M output costs $0.01. The long read costs about 400x more than the answer. On a model with a long-context surcharge, crossing the threshold also reprices the whole request. In long-context designs, input volume and caching (§13.9) are the cost levers that matter.


13.3 The Full AI Cost Model: What You're Actually Paying For

Most architects model LLM API cost and miss 40-60% of the actual AI system cost. The full cost model has eight components:

FULL AI SYSTEM COST MODEL

1. LLM INFERENCE COST (usually the largest component)
   Input tokens × input_price + Output tokens × output_price
   For agents: multiply by average iteration count
   For reasoning models: add thinking tokens × output_price

2. EMBEDDING COST (often overlooked)
   Embedding API calls for ingestion: doc_chunks × embedding_price
   Embedding API calls for queries: queries × embedding_price

   Example (illustrative; verify current embedding prices):
   OpenAI text-embedding-3-large = $0.13/M tokens
   100K documents × avg 500 tokens/chunk = 50M tokens
   Ingestion embedding cost: 50M × $0.13/M = $6.50 (one-time)
   Daily query embedding: 10K queries × 200 tokens = 2M tokens = $0.26/day

3. VECTOR STORAGE COST
   Weaviate Cloud / Pinecone / Qdrant Cloud: charged by vectors stored,
   and increasingly by reads/writes on serverless tiers
   Example (illustrative): 10M vectors ≈ ~$70/month — price your actual
   index size, dimensions, and query volume on the vendor's calculator
   Self-hosted pgvector: infrastructure cost only

4. RE-RANKING COST (if using cross-encoder re-ranking)
   Hosted rerankers bill per search (one query + up to N candidates).
   Example: Cohere Rerank 4 Fast ≈ $2.00 per 1,000 searches (reported;
   verify on Cohere's pricing page) = $0.002/query at 20 candidates.
   Long candidates may be split into chunks and billed as extra searches.
   At 100K queries/day: $200/day = $6,000/month — material cost

5. PROMPT CACHING SAVINGS (negative cost)
   Providers offer cache discounts for repeated prompt prefixes
   Major providers (Anthropic, OpenAI, Google, DeepSeek): cached input
     ≈ 1/10 of the uncached price or less (G.3, October 2026)
   Watch for: cache-write premiums, cache TTLs, minimum prefix sizes,
     and caches that are scoped to one model (a failover starts cold —
     Module 37 §37.7)
   For systems with stable system prompts: large savings at scale

6. ORCHESTRATION OVERHEAD (often ignored)
   LangChain/LlamaIndex API calls for routing decisions
   Lightweight classifier calls for model routing
   Quality evaluation calls (LLM-as-judge on sample traffic)
   Typically 5-15% of primary LLM cost

7. SEMANTIC CACHE INFRASTRUCTURE
   Redis instance for semantic caching: $50-200/month (managed)
   Vector similarity computation: typically negligible
   Cache miss rate × primary LLM cost = net savings

8. MONITORING AND OBSERVABILITY
   LangSmith / Arize / gateway-native observability: usage-based pricing
   OTel collector infrastructure
   Storage for traces and eval results
   Typically 2-5% of total AI spend

Cost Model Template

MONTHLY COST ESTIMATION TEMPLATE
(illustrative, October 2026 prices — see Appendix G §G.3;
 primary model: a mid-tier model at $2.00 input / $10.00 output per 1M)

Input parameters:
  daily_queries = 10,000
  avg_input_tokens = 2,500   (system prompt + context + user message)
  avg_output_tokens = 250
  agent_tasks_daily = 500
  agent_avg_iterations = 3
  agent_avg_tokens_per_call = 3,000

LLM inference cost (simple RAG queries):
  daily_queries × (avg_input × $input_price + avg_output × $output_price)
  = 10,000 × (2,500 × $0.000002 + 250 × $0.00001)
  = 10,000 × ($0.005 + $0.0025)
  = 10,000 × $0.0075
  = $75.00/day → $2,250/month

LLM inference cost (agent tasks):
  (conservative simplification: all agent tokens priced at the output rate)
  agent_tasks × iterations × (avg_tokens × $output_price)
  = 500 × 3 × (3,000 × $0.00001)
  = 500 × 3 × $0.03
  = $45.00/day → $1,350/month

Embedding cost (queries):
  10,000 × 200 tokens × $0.00000013 = $0.26/day → $7.80/month

Re-ranking cost (every query, 20 candidates = 1 search):
  10,000 searches × $0.002 = $20.00/day → $600/month

Vector storage: $70/month (fixed)

Total estimated: ~$4,278/month

At 3x scale (30,000 queries/day, storage held flat): ~$12,693/month
Are unit economics sustainable at scale? (~$14.26 per 1,000 queries,
  blended, including the agent workload)

Before you trust the total, apply the 2026 traps (§13.2):
  - Is any routed model on an introductory price? Re-run at the step-up price.
  - Can any request cross a long-context surcharge threshold?
  - Was the token count measured on the model you will actually run?

13.4 The Most Impactful Cost Engineering Decisions

Decision 1: Model Routing (often 50-80% cost reduction)

Routing different tasks to appropriately-priced models is the highest-leverage cost engineering decision. Tiered routing commonly cuts costs by half or more while maintaining quality on difficult queries. The exact saving depends on your traffic mix and on the price gap between tiers, and that gap moves every time a price changes.

MODEL ROUTING COST IMPACT
(illustrative, October 2026 prices — see Appendix G §G.3)

Without routing: all queries go to Claude Sonnet 5.5 ($2.00/$10.00)
  10,000 queries/day × avg 2,500 input + 250 output
  = $75.00/day

With routing (tiered by complexity):
  40% simple  → GPT-6 Luna ($0.10/$0.50):          cost = $1.50/day
  40% medium  → Gemini 3.8 Flash ($0.75/$3.75,
                introductory):                      cost = $11.25/day
  20% complex → Claude Sonnet 5.5 ($2.00/$10.00):  cost = $15.00/day
  Total: $27.75/day

Cost reduction: 63% while maintaining quality for complex queries

THE STEP-UP CHECK (the 2026 trap, §13.2):
  Gemini 3.8 Flash moves to $1.50/$7.50 on 2027-01-01.
  Medium tier becomes $22.50/day → total $39.00/day
  Cost reduction falls from 63% to 48% with no change in your system.
  → Model every routed tier at its post-step-up price, and re-evaluate
    the routing table when a step-up date approaches.

Each tier in the routing table is a model you depend on. Give each one a model profile and a fallback (Module 37 §37.3, §37.7), so a price change or access change on one tier is a configuration change, not a rewrite.

The routing classifier design:

The classifier that decides which tier a query belongs to must itself be cheap and fast — otherwise it adds cost and latency without proportional value.

ROUTING CLASSIFIER OPTIONS

Option A: Rule-based (zero cost)
  if query contains ["explain", "summarize", "what is"]:
    route_to("cheap_tier")
  elif query contains ["analyze", "compare", "evaluate"]:
    route_to("mid_tier")
  else:
    route_to("expensive_tier")

  Pros: Zero cost, zero latency overhead
  Cons: Brittle, misses nuanced complexity signals

Option B: Small LLM classifier (very low cost)
  Use an efficient-tier model (e.g., GPT-6 Luna or Gemini Flash-Lite) to classify:
    "Classify this query as: simple | medium | complex"

  Cost (illustrative): ~200-token classification prompt × ~$0.10/M input
    ≈ $0.00002/query, plus a few output tokens
  At 10K queries/day: ~$0.20/day (negligible next to the $75/day it optimizes)
  Pros: Better accuracy than rule-based
  Cons: Small latency overhead (~50ms)

Option C: Embedding similarity to labeled examples (low cost)
  Pre-embed 100 examples per tier
  For each query, find nearest examples by cosine similarity
  Route to the tier of the nearest examples

  Cost: embedding cost only (negligible)
  Latency: 10-20ms for similarity lookup

Decision 2: Prompt Caching (40-90% reduction on repeated context)

Prompt caching is the most underutilized cost optimization in enterprise AI. All the major providers (Anthropic, OpenAI, Google, DeepSeek) discount cached prompt content heavily: content that appears in the same position at the beginning of requests repeatedly. Cached input is priced at roughly 1/10 of uncached input or less (Appendix G §G.3).

PROMPT CACHING MECHANICS
(illustrative, October 2026 prices: $2.00/M input, $0.20/M cached input)

Standard request structure:
  [System prompt: 800 tokens] + [Context: 2,000 tokens] + [Query: 200 tokens]
  Total input: 3,000 tokens @ $2.00/M = $0.006/request

With prompt caching:
  [System prompt: 800 tokens] — CACHED after first request
  [Context: 2,000 tokens] — CACHED if same context
  [Query: 200 tokens] — NOT cached (unique per request)

  Subsequent requests:
  Cached tokens: 2,800 × $0.20/M (10% of original price) = $0.00056
  Non-cached tokens: 200 × $2.00/M = $0.0004
  Total: $0.00096/request (vs $0.006 without caching)

  Savings: 84% reduction on requests with the same system prompt + context

WHAT CAN BE CACHED:
  ├── System prompt (almost always the same — high cache hit rate)
  ├── Retrieved RAG context (if queries are similar, same chunks retrieved)
  ├── Few-shot examples (if they appear in every request)
  └── Tool definitions (function calling schemas)

WHAT CANNOT BE CACHED:
  └── The user's unique query (different every time)

Implementation requirement:
  Cached content must appear at the BEGINNING of the prompt, before
  any variable content. Reordering the prompt to put cached content
  first is sometimes necessary but always worth it.

CACHING CAVEATS THAT CHANGE THE MATH:
  ├── Cache writes can cost MORE than normal input on some providers;
  │     savings depend on the hit rate, not only the discount
  ├── Caches expire (TTL); low-traffic capabilities may rarely hit
  ├── Caches are scoped to one model: a version upgrade or a failover
  │     starts cold, and the input bill can jump up to ~10x until the
  │     new cache warms (Module 37 §37.7, "cache-cold cost spike")
  └── Budget a failover cost reserve for heavily cached capabilities

Decision 3: Semantic Caching (30-60% reduction on repeated queries)

Semantic caching stores LLM responses and serves cached responses when a new query is semantically similar to a previously answered query. Unlike HTTP caching (exact match), semantic caching catches questions that mean the same thing even with different wording.

SEMANTIC CACHE ARCHITECTURE

New query: "What is the annual fee for a basic account?"
  └── Embed query → similarity search against cached responses
  └── Find: "How much does a basic checking account cost per year?"
         (stored response: "The annual fee is $0 for basic checking accounts.")
         similarity: 0.94 > threshold (0.85)
  └── Return cached response — NO LLM API call

New query: "What are the fees for premium accounts?"
  └── Embed query → similarity search
  └── Find: "Basic account annual fee" response — similarity: 0.61 < threshold
  └── Cache miss → LLM API call → store response in cache

CACHE HIT RATE BY USE CASE:
  Customer FAQ system: 40-60% (many similar questions)
  Document Q&A: 20-35% (more varied queries)
  Custom analysis: 5-15% (most queries unique)
  Internal helpdesk: 50-70% (highly repetitive questions)

SEMANTIC CACHE INVALIDATION:
  ├── Time-based TTL: cached responses expire after N days
  ├── Knowledge base update: when underlying documents change,
  │     related cached responses must be invalidated
  └── Manual invalidation: for urgent corrections

Tools: GPTCache (open source), Redis with vector search,
       gateway-native semantic caches (e.g., LiteLLM; Portkey, now part of
       Palo Alto Networks — Appendix G §G.5). If the cache lives in your
       gateway, it is one more reason to keep the gateway replaceable
       (Module 37 §37.3).

Decision 4: Context Window Management

The context window is both your most valuable resource and your most expensive one. Every token in context costs money on every request. Context that was added for convenience rather than necessity is pure cost overhead.

CONTEXT WINDOW COST MATH

A 10-turn conversation where each turn averages 200 tokens:
  Turn 1: 200 tokens context
  Turn 5: 1,200 tokens context (5 turns × 200 + 200 new)
  Turn 10: 2,200 tokens context
  Turn 20: 4,200 tokens context

If you include full conversation history on every request:
  Cumulative input tokens across 20 turns:
  200 + 400 + 600 + ... + 4,000 = ~42,000 tokens

If you summarize after every 5 turns:
  Summary after turn 5: replace 1,000 tokens with 100-token summary
  Cumulative tokens for same 20 turns: ~8,000 tokens

Cost reduction: 81% on context overhead alone

CONTEXT WINDOW OPTIMIZATION STRATEGIES:

1. Selective history inclusion
   Include only the last N turns of conversation history.
   For most chatbot use cases, the last 5-10 turns provide
   sufficient context. The conversation from 40 turns ago
   is rarely needed.

2. Progressive summarization
   After every 10 turns: summarize the prior conversation into
   a structured 100-200 token summary.
   Replace prior turns with summary. Drop the raw turns.

3. Structured context injection
   Instead of: [full conversation history]
   Use: {
     user_intent: "compare account fees",
     established_facts: ["user has basic checking account",
                          "last discussed: fee waiver"],
     open_questions: ["user wants to understand premium upgrade cost"]
   }
   This is more informative and less expensive than raw history.

4. RAG context optimization
   Retrieve 3 high-quality chunks rather than 5-10 lower-quality chunks.
   Better retrieval precision (re-ranking) reduces context cost.
   Each additional retrieved chunk adds ~500 tokens × input price.

Decision 5: Async Batch Processing

For workloads that don't need real-time responses, the async batch API is the highest-discount option available.

BATCH API PRICING (structure; verify current terms per provider)

OpenAI Batch API:
  ~50% discount vs. synchronous pricing
  24-hour completion window
  Results retrieved by polling or webhook

Anthropic Message Batches API:
  50% discount on all token types (input, output, cache reads/writes)
  24-hour maximum; most batches finish much sooner
  Up to 100,000 requests or 256 MB per batch
  Separate rate-limit pool from synchronous traffic

Google Gemini API batch mode: also discounted vs. synchronous (verify rate)

Use cases for batch processing:
  ├── Nightly report generation
  ├── Document classification at ingestion
  ├── Embedding generation for new content (use embedding batch API)
  ├── Eval runs on production traffic samples
  ├── Synthetic data generation for eval datasets
  └── Analytics and insights generation on historical data

Workflow change required:
  Instead of: HTTP request → wait → response (synchronous)
  Use: Submit batch → webhook/poll → retrieve results (async)
  Application must be designed for async — not a drop-in replacement

SAVINGS EXAMPLE:
  Nightly document classification of 50,000 documents:
  Average 500 input tokens + 50 output tokens per document

  (illustrative, October 2026: Claude Sonnet 5.5 at $2.00/$10.00 — see G.3)

  Synchronous: 50,000 × ($0.001 + $0.0005) = $75.00/night
  Batch API: $75.00 × 50% = $37.50/night
  Annual savings: ~$13,700

Decision 6: Prompt Compression

For systems where the retrieved context dominates input token cost, prompt compression reduces tokens while preserving semantic content.

LLMLingua (Microsoft Research) selectively removes low-information tokens from prompts — achieving 3-20x compression with minimal quality degradation on many tasks.

PROMPT COMPRESSION ECONOMICS

Retrieved context: 5 chunks × 500 tokens = 2,500 tokens
After LLMLingua compression: 2,500 → 800 tokens (3x reduction)
Cost reduction: 68% on retrieved context tokens

At 100K queries/day:
  Without compression: 250M input tokens/day (context portion)
  With compression: 80M input tokens/day
  At an illustrative $2/M input: savings of $340/day = $10,200/month

WHEN TO USE:
  ├── Retrieved context is consistently long (>1,000 tokens per chunk)
  ├── Quality testing shows acceptable degradation (<5% quality loss)
  └── High-volume RAG systems where context costs dominate

WHEN NOT TO USE:
  ├── Short contexts (not enough compression benefit)
  ├── Complex technical content (compression may remove critical details)
  └── Legal/regulatory content where precision matters more than cost

13.5 The Build vs. Buy vs. Host Cost Model

The decision between API (buy), self-hosted open-weight (host), and custom infrastructure requires honest unit economics modeling.

BREAK-EVEN ANALYSIS: API VS SELF-HOSTED
(illustrative — every input below is an assumption to replace with your own
 measurements and current prices from Appendix G §G.3)

Self-hosting an open-weight model (e.g., an MoE such as Llama 4 Scout or
Qwen 3.6, quantized to fit one 80GB GPU):
  Infrastructure: one A100/H100-class 80GB GPU
  Cloud cost: ~$2,500/month dedicated (carried forward from May 2026;
    verify current H100/B200 pricing)

  Sustained blended (input + output) throughput with vLLM and continuous
  batching: assume ~1,000 tokens/second (ILLUSTRATIVE — it varies by an
  order of magnitude with model, quantization, batch size, and the
  input/output mix; measure it with a load test on your own traffic)
  Monthly capacity at 100% utilization:
    1,000 tokens/s × 3,600 × 24 × 30 = ~2.6B tokens

  Effective self-host cost per million tokens:
    at 100% utilization: $2,500 ÷ 2,592M tokens = ~$0.96/M
    at 50% utilization:                          ~$1.93/M
    (before operations staff, monitoring, and a second GPU for redundancy)

BLENDED API PRICE (10:1 input:output mix, illustrative October 2026 prices):
  Efficient tier (e.g., GPT-6 Luna, DeepSeek V4 Flash):  ~$0.11–0.15/M
  Mid-tier (e.g., Claude Sonnet 5.5 at $2/$10):         ~$2.73/M

BREAK-EVEN CALCULATION:
  Monthly API cost = Monthly self-host cost
  X tokens × blended_price = $2,500

  vs. efficient tier: X = $2,500 ÷ $0.15/M ≈ 16.7B tokens/month (more at $0.11)
    → about 6x the GPU's entire capacity. Self-hosting does not beat
      efficient-tier APIs on cost at this scale.
  vs. mid-tier:       X = $2,500 ÷ $2.73/M ≈ 0.9B tokens/month
    → about 35% sustained utilization. Self-hosting can win on cost,
      IF the open-weight model meets the quality contract the mid-tier
      model meets (Module 37 §37.5). Add a second GPU for availability
      and the break-even doubles.

  Practical implication: self-hosting competes on cost with mid-tier and
  frontier APIs at sustained high utilization, and rarely with efficient-
  tier APIs. The stronger reasons to self-host are usually data control,
  residency, latency, and customization (Module 14, Module 34) — and a
  self-hosted model is a useful L3 fallback rung (Module 37 §37.7).

HYBRID APPROACH (most practical for enterprises):
  ├── Low-to-medium volume: API (no infrastructure overhead)
  ├── High-volume, data-sensitive: Self-hosted (better data control + cost)
  ├── Bursty workloads: API (pay per use; no idle infrastructure)
  └── Continuous high-volume: Self-hosted (amortize infrastructure)

13.6 FinOps for AI: The Organizational Layer

Cost attribution, showback, and chargeback are organizational governance practices, not just technical ones. Without them, AI cost accountability is diffuse and optimization is impossible.

The FinOps AI Maturity Levels

Level 1: Inform — you know what you're spending

  Total monthly AI spend: visible
  Spend by model: visible
  Spend by team: NOT visible (no attribution tags)
  Spend by feature: NOT visible
  Cost per interaction: NOT computable

  Most teams start here.

Level 2: Optimize — you know where to optimize

  Total monthly AI spend: visible
  Spend by model: visible
  Spend by team: visible (attribution tags in place)
  Spend by feature: visible
  Cost per interaction by feature: computable
  Model routing implemented: yes
  Prompt caching implemented: yes

  Cost optimization is possible because attribution exists.

Level 3: Operate — cost is a first-class product metric

  All Level 2 capabilities plus:
  Cost per successful interaction: tracked and trended
  Cost budget per feature: enforced with automated alerts
  Team chargeback: each team receives their AI spend monthly
  Cost efficiency reviews: quarterly per team
  Cost as a deployment gate: features with cost regressions
    require approval before deployment

The Chargeback Architecture

Chargeback means each team's AI spend is attributed to their budget, not a shared central infrastructure budget. This creates the right incentives: teams optimize their own AI usage because it affects their own budget.

AI CHARGEBACK MODEL

Technical layer:
  ├── Every LLM call tagged with: team_id, feature_id, workflow_type
  ├── Cost computed at the gateway: tokens × price = cost, using a
  │     versioned price table (uncached / cached / output / reasoning,
  │     long-context tiers, and step-up dates), refreshed with Appendix G
  ├── Cost events streamed to the billing aggregation system
  └── Monthly report: cost by team, feature, model, workflow

Organizational layer:
  ├── Each team has a monthly AI budget (set in planning)
  ├── Teams receive weekly alerts when they hit 50% and 80% of budget
  ├── Teams that exceed budget require approval for continued use
  │     or must optimize before next billing period
  └── Finance receives monthly AI cost report by team
        for actual financial chargeback or showback

Budget guardrails (technical enforcement):
  ├── Gateway enforces per-team token budgets
  ├── Exceeding budget: requests queued or rate-limited (not hard-blocked
  │     for real-time customer-facing features)
  └── Budget exhaustion on batch features: job suspended, owner notified

13.7 The Hidden Costs Most Teams Miss

Evaluation cost. Running LLM-as-judge on 5% of production traffic costs money. At 10,000 queries/day, 5% = 500 eval calls/day. At $0.005/eval call (using a cheap model as judge) = $2.50/day = $75/month. Negligible. But eval calls that use a frontier model as judge can cost 10-100x more.

Re-embedding cost. When you update your embedding model or when documents change, re-embedding is required. 100K documents × 500 tokens average × $0.13/M = $6.50 one-time. At quarterly document refresh cycles, add this to the annual budget.

Agent iteration variance. Agent tasks have a P99 cost that is much higher than the mean. An agent designed to complete in 3 iterations occasionally runs 15-20 iterations on pathological inputs. The P99 cost can be 5-7x the mean. This variance must be accounted for in cost modeling and controlled with hard budget limits.

Model version upgrade migration. When a model is deprecated, migrating requires engineering time plus eval runs to validate the new model. At 4-6 model migrations per year, the engineering cost is material even if the token cost doesn't change. Module 37 (§37.6, §37.8) turns this into a planned, budgeted activity: a quarterly cross-model candidate run is a line item (on the order of hundreds of dollars per run for a mid-size portfolio, before judge costs), and it is far cheaper than an emergency migration.

Tokenizer differences. Different models use different encoding schemes, so the same text yields different token counts across providers. The 2026 surprise is that this also happens between versions from the same vendor: Anthropic documents that the tokenizer introduced with Opus 4.7 uses about 1.0–1.35x as many tokens as earlier models for the same text (Appendix G §G.3). When you switch models, including a same-vendor upgrade, re-measure token counts on your own prompts and re-validate the cost model. A lower per-token price can be partly or fully offset by a higher token count.

Failover and cache-cold cost. Prompt caches are scoped to one model. When traffic fails over to a fallback, or moves to a new version, the fallback starts with a cold cache. For a heavily cached capability the input bill can jump by up to ~10x until the cache warms, on top of any cache-write premium. Module 37 §37.7 works through an example (a 50K-token cached prefix at 1M requests/month: about $10,000/month cached vs. $100,000/month uncached, at illustrative prices). Keep-alive traffic to the fallback, lean cacheable prefixes, and an explicit failover cost reserve keep this from becoming a surprise.

Price changes you did not make. Introductory prices that step up on a fixed date, and long-context surcharges that reprice the whole request (§13.2), change the bill without any change in your system. Track step-up dates and surcharge thresholds per model, in the model profile (Module 37 §37.3), and alert on them.


13.8 The Cost Engineering Checklist

Pricing and modeling - [ ] Current pricing verified for all models in use against Appendix G §G.3 and provider pages (prices change monthly)? - [ ] Introductory-price step-up dates and long-context surcharge thresholds recorded per model, with alerts? - [ ] Token counts re-measured after every model or version change (tokenizers change between versions)? - [ ] Full cost model includes: LLM, embedding, vector storage, re-ranking, caching infrastructure, monitoring? - [ ] Cost estimated at current volume AND 3x volume? - [ ] Unit economics (cost per successful interaction) computed by feature?

Technical optimizations - [ ] Model routing implemented? (high value, complexity, data sensitivity tiers) - [ ] Prompt caching enabled for stable system prompts and common context? - [ ] Context window management implemented (summarization, selective history)? - [ ] Async batch API used for non-real-time workloads? - [ ] Output length controlled (no "comprehensive analysis" when "3 sentences" suffices)? - [ ] Semantic caching evaluated for use cases with repetitive queries?

FinOps - [ ] Cost attribution tags on every LLM call (team, feature, workflow)? - [ ] Weekly cost report distributed to owning teams? - [ ] Per-feature budget alerts configured (80% and 100% of budget)? - [ ] Agent tasks have hard per-task cost ceilings (code-enforced)? - [ ] Model routing reviewed quarterly — is routing logic still optimal? - [ ] Failover cost reserve budgeted for heavily cached capabilities (cache-cold spike, Module 37 §37.7)? - [ ] Quarterly cross-model eval run budgeted as a line item (Module 37 §37.6)?

Architecture gates - [ ] Cost regression detection: does a prompt change that increases cost per request by >10% require approval? - [ ] New features require cost model before deployment? - [ ] P99 agent cost modeled and bounded, not just mean cost? - [ ] Model changes judged on cost per completed task, not price per token (Module 37 §37.5)?


EXERCISE — Cost Model a RAG System: You are designing a customer support RAG system with these parameters: 50,000 queries/day, average 2,000 input tokens (system prompt + 3 RAG chunks + query), average 200 output tokens, 10% of queries trigger re-ranking (20 candidates per query). Using current tier pricing from Appendix G §G.3, compute: monthly cost at current volume, monthly cost at 3x volume, and the cost impact of model routing (40% simple queries to an efficient-tier model, 40% medium to a fast mid-tier model, 20% complex to a mid-tier or frontier model). Show your math. Then repeat the routing calculation after any introductory price step-up that falls within the next 12 months, and state how the saving changes.

PONDER — The Routing Decision: For your most expensive AI feature, what percentage of queries genuinely require a frontier model versus a cheaper alternative? If you honestly assessed the task complexity distribution, how much of your LLM spend is on frontier-model responses that a mid-tier model would have produced correctly 95% of the time?

WORKSHOP — FinOps Maturity Assessment: Using the three-level FinOps maturity model, assess your organization's current level. For each gap between current and Level 3: what is the technical change required, who owns it, and what is the monthly cost savings if implemented? Build a 90-day FinOps roadmap for AI that gets you from current maturity to Level 2.

WORKSHOP — Self-Host vs. API Decision: Your organization processes 5 billion tokens per month across AI features, all currently via a mid-tier API model (illustrative: Claude Sonnet 5.5 at $2/$10 per million tokens, October 2026). A team proposes self-hosting an open-weight model (for example Llama 4 Scout, Qwen 3.6, or DeepSeek V4 Flash) to reduce costs. Build the complete business case: infrastructure cost (GPU instances, operations, maintenance), break-even analysis, quality risk assessment, data governance implications, and the recommendation. Use actual current pricing from Appendix G §G.3 and your measured GPU throughput in your calculation, and include the cost of keeping the API model as a fallback.


13.9 Long Context vs. RAG: The Emerging Cost Trade-off

With 1M-token context windows now common on frontier and mid-tier models (Claude Fable / Opus / Sonnet 5.x, Gemini 3.x, DeepSeek V4, Llama 4 Maverick — Appendix G §G.2), a question that was previously theoretical is now architectural: for some use cases, is it cheaper to load your entire knowledge base into the context window on every request than to build and maintain a RAG infrastructure?

LONG CONTEXT vs. RAG COST ANALYSIS

Scenario: 500-page policy manual (approx. 250K tokens)
Query volume: 10,000 queries/day
Model: Claude Sonnet 5.5 ($2.00/$10.00 per M tokens; cached input $0.20/M)
  (illustrative, October 2026 prices — see Appendix G §G.3)

OPTION A: Full-context approach (load entire manual per query)
  Input tokens per query: 250,000 (manual) + 200 (query) = 250,200
  Output tokens per query: 250
  Cost per query: (250,200 × $2/M) + (250 × $10/M)
                = $0.5004 + $0.0025 = $0.503
  Daily cost: 10,000 × $0.503 = $5,029
  Monthly cost: ~$150,870

  WITH PROMPT CACHING (manual is stable, cached input at ~1/10 price):
  Cached input: 250,000 × $0.20/M = $0.05
  Non-cached input: 200 × $2/M = $0.0004
  Output: 250 × $10/M = $0.0025
  Cost per query with cache: $0.0529
  Monthly cost: ~$15,870 (excluding cache-write premiums and TTL misses)

OPTION B: RAG approach
  Retrieve 3 relevant chunks (avg 500 tokens each = 1,500 tokens)
  System prompt: 500 tokens
  Input tokens: 2,200; Output tokens: 250
  Cost per query: (2,200 × $2/M) + (250 × $10/M) = $0.0044 + $0.0025 = $0.0069
  Monthly cost: ~$2,070 + RAG infrastructure ($500) = ~$2,570

COMPARISON:
  Full context (no cache):    ~$150,870/month
  Full context (with cache):  ~$15,870/month
  RAG:                        ~$2,570/month

THREE 2026 RISKS SPECIFIC TO THE FULL-CONTEXT OPTION:
  ├── Long-context surcharge: if the manual grows past a model's threshold
  │     (e.g., 272K input tokens on GPT-6), EVERY request is repriced,
  │     not only the excess
  ├── Cache-cold failover: the cached option is ~10x cheaper only while
  │     the cache is warm. A failover or version change sends the cost
  │     toward the no-cache line until the new cache warms (Module 37 §37.7)
  └── Fallback viability: a 250K-token prompt rules out any fallback with
        a smaller context window. Set a design ceiling (Module 37 §37.3)

WHEN FULL CONTEXT BEATS RAG:
  ├── Document is small enough that cached cost is cheap
  ├── Queries frequently need content from across the entire document
  │     (RAG retrieval can't find all relevant chunks)
  ├── Document updates are infrequent (high cache hit rate)
  └── RAG infrastructure build cost would be 6+ months of cost savings

WHEN RAG IS STILL BETTER:
  ├── Knowledge base is large (>100K tokens) — caching doesn't help much
  ├── Knowledge base updates frequently (cache constantly invalidated)
  ├── Query volume is very high (even cached queries add up)
  └── The full document approach retrieves irrelevant context that
        degrades generation quality ("lost in the middle" problem)

PRACTICAL GUIDANCE:
  ├── < 50K token knowledge base, < 1,000 queries/day: consider full context
  ├── 50K-500K tokens, stable content, <5,000 queries/day: evaluate with math
  └── > 500K tokens OR > 5,000 queries/day: RAG almost always wins on cost

Next: Module 14 — AI Infrastructure: Inference, Self-Hosting & Cloud Platforms