Skip to content

MODULE 14 — AI Infrastructure: Inference, Self-Hosting & Cloud Platforms

14.1 The Infrastructure Decision Has Changed

In 2023 and most of 2024, the AI infrastructure decision was simple: use the cloud API. The quality gap between hosted frontier models and anything you could run yourself was too large to overcome, and the engineering cost of self-hosting wasn't justified.

In 2026, that picture has changed fundamentally. Open-source inference engines have closed 90% of the performance gap with proprietary alternatives. Stripe achieved a 73% inference cost reduction via vLLM migration. The models that are competitive with frontier models (DeepSeek, Llama 4, Qwen) are open-weight and can be self-hosted. The inference serving software (vLLM) has matured to production-grade quality with enterprise features.

This does not mean self-hosting is the right answer for most teams. It means the infrastructure decision is now a genuine architectural choice that requires analysis, not a default to the cloud API.

The architect's job: frame the decision correctly, with honest tradeoffs.


14.2 The Inference Serving Landscape

The software that actually runs LLM inference — the inference server — is a distinct layer from the model itself. Choosing the right inference server for your deployment topology is as important as choosing the right model.

vLLM: The Production Standard

vLLM has become the default serving engine for production LLM inference. The current stable release was v0.20.2 as of May 2026; by October 2026 it had already moved to v0.30.0 — vLLM ships new minor releases roughly every two weeks, so treat any version number here as illustrative and check vllm.ai/GitHub releases for the current one. It is the de facto choice for enterprises self-hosting open-weight models at scale.

Why vLLM dominates production deployments:

PagedAttention. vLLM's foundational innovation. Traditional inference servers pre-allocate contiguous GPU memory per request — when a batch of requests has variable lengths, the memory between allocations is wasted (internal fragmentation). PagedAttention manages the KV cache in non-contiguous pages, like virtual memory in an operating system. This eliminates 60-80% of memory waste from KV cache fragmentation, dramatically increasing the number of concurrent requests a single GPU can serve.

Continuous batching. Rather than waiting for a full batch to assemble before starting inference, vLLM processes requests as they arrive, inserting new requests into the batch as slots become available. This reduces average latency significantly compared to static batching.

OpenAI-compatible API. vLLM exposes an OpenAI-compatible REST API out of the box. Switching from the OpenAI API to self-hosted vLLM requires changing the base_url and removing the API key — nothing else in the application changes.

Disaggregated prefill/decode (2025-2026 feature). Separates the compute-intensive prompt processing phase (prefill) from the memory-bandwidth-intensive token generation phase (decode). This allows different hardware scheduling for each phase and improves overall GPU utilization.

Automatic prefix caching. When multiple requests share the same prefix (e.g., all requests start with the same system prompt), vLLM caches the computed KV cache for that prefix and reuses it across requests. This eliminates the redundant computation of processing the system prompt on every request. For production deployments where all requests share a 500-800 token system prompt, prefix caching reduces Time to First Token significantly — the cached prefix computation is skipped entirely. Enable with --enable-prefix-caching in vLLM. This is distinct from semantic caching (which caches full responses) — prefix caching caches intermediate computation, not final outputs.

Speculative decoding. Uses a small, fast "draft" model to predict multiple tokens ahead, then verifies them with the large "target" model in one pass. When the draft model's predictions are correct (common for predictable output patterns), multiple tokens are accepted at once, effectively increasing throughput. The Speculators library integrates with vLLM for this. Most beneficial for workloads with repetitive output patterns (code generation, structured outputs, translation). Less beneficial for open-ended generation where predictions are frequently rejected.

Performance benchmark: vLLM achieved 793 tokens per second compared to Ollama's 41 TPS at equivalent configurations — nearly 20x higher throughput under concurrent load. At 50+ concurrent users, Ollama hit a performance cliff and request success rate degraded; vLLM maintained linear scaling.

Critical limitation: vLLM requires a CUDA-capable NVIDIA GPU. It does not support CPU-only inference and does not support Apple Silicon. Teams on Apple hardware or CPU-only servers cannot use vLLM.

Key configuration decisions for production vLLM:

VLLM PRODUCTION CONFIGURATION CONSIDERATIONS

Model loading:
  --model: path to model weights (local) or HuggingFace model ID
  --tensor-parallel-size: number of GPUs for tensor parallelism
    (2 GPUs = model split across 2 GPUs; required for models
     that don't fit in a single GPU's VRAM)
  --pipeline-parallel-size: number of pipeline stages
    (for very large models across many GPUs)

Performance:
  VLLM_USE_V2_MODEL_RUNNER=1: enables Model Runner V2 (up to 56% throughput gain)
  --max-num-batched-tokens: max tokens processed in one batch
  --max-num-seqs: max concurrent sequences
  --block-size: KV cache page size (increase to 32 for long-context models)

Quantization:
  --quantization awq: AWQ quantization (fast, minimal quality loss)
  --quantization gptq: GPTQ quantization (alternative, slightly slower)
  --dtype bfloat16: use BF16 instead of FP32 (2x memory reduction, minimal quality loss)

Serving:
  --host 0.0.0.0 --port 8000: API endpoint
  --api-key: add authentication
  --max-model-len: maximum context length (reduce if VRAM constrained)

Ollama: Development and Edge

Ollama simplifies local model running for development and edge deployments. It handles model download, format conversion, and serving behind a simple API.

Where Ollama is appropriate: - Developer workstations and laptops (including Apple Silicon via Metal) - Edge deployment on devices without enterprise GPU - Initial prototyping before investing in vLLM infrastructure - Single-user applications where <50 concurrent requests is the ceiling

Where Ollama is NOT appropriate: - Production serving for multiple concurrent users - Any workload requiring >100 concurrent requests - High-throughput batch processing

The 20x throughput gap between Ollama and vLLM under concurrent load is not a configuration issue — it is an architectural difference. Ollama was not designed for concurrent production serving.


TensorRT-LLM: Maximum NVIDIA Performance

NVIDIA's TensorRT-LLM provides the highest peak throughput on NVIDIA hardware by compiling models into optimized CUDA kernels specific to the target GPU architecture.

Where TensorRT-LLM wins: - Maximum throughput is the priority and latency target is tight - You are committed to NVIDIA hardware long-term - The model is stable (not changing frequently) - You have the engineering resources for the compilation pipeline

Why TensorRT-LLM adds friction: - Model-specific compilation: every time you change the model or the target GPU architecture, you recompile. This takes hours and requires GPU resources. - NVIDIA lock-in: the optimized engines only run on NVIDIA GPUs. Switching to AMD or another provider requires recompilation. - Narrower model compatibility than vLLM

The vLLM vs. TensorRT-LLM decision: vLLM for most enterprises — broad model compatibility, simpler deployment, OpenAI-compatible API, active community. TensorRT-LLM for latency-critical production at scale where maximum throughput justifies the operational complexity.


SGLang: Structured Generation Focus

SGLang shares architectural similarities with vLLM (PagedAttention-derived memory management) but is optimized for structured generation — applications that require the model to output JSON, code, or other formatted structures reliably and efficiently.

When SGLang is worth evaluating: - High-volume applications where structured output (JSON mode) is the primary use case - Agentic applications where tool call format reliability matters - Use cases where the latency savings from structured generation justify migration


The Inference Engine Decision Matrix

INFERENCE ENGINE SELECTION

                    vLLM       Ollama      TensorRT-LLM   SGLang
Production serving   ✓✓✓        ✗            ✓✓✓           ✓✓
Development/local    ✓          ✓✓✓          ✗             ✓
Apple Silicon        ✗          ✓✓✓          ✗             ✗
CPU-only             ✗          ✓            ✗             ✗
OpenAI-compatible    ✓✓✓        ✓✓           Partial        ✓✓
Model flexibility    ✓✓✓        ✓✓✓          ✓              ✓✓
Structured output    ✓          ✓            ✓              ✓✓✓
Multi-GPU support    ✓✓✓        ✗            ✓✓✓           ✓✓
Vendor lock-in       None       None         NVIDIA heavy   None
Active development   ✓✓✓        ✓✓           ✓✓             ✓✓

14.3 Hardware: GPU Economics and Selection

The GPU Landscape (2026)

GPU SELECTION FOR LLM INFERENCE

NVIDIA H100 (SXM/PCIe, 80GB VRAM)
  Designed for: Large-scale production inference
  Throughput: Best in class for LLM inference
  Memory: 80GB HBM3 — fits 70B models in FP16 without splitting
  Cloud cost: ~$30-50/hour (on-demand), ~$8-12/hour (spot)
  Dedicated: ~$5,000-8,000/month per card
  Best for: High-throughput production serving at scale

NVIDIA A100 (SXM/PCIe, 80GB VRAM)
  Designed for: Training and inference
  Throughput: ~60-70% of H100 for inference
  Memory: 80GB HBM2e
  Cloud cost: ~$3-4/hour (on-demand)
  Dedicated: ~$2,500-3,500/month per card
  Best for: Production inference where H100 budget is not justified

NVIDIA L40S (48GB VRAM)
  Designed for: Inference-optimized
  Throughput: Strong for medium-sized models
  Memory: 48GB GDDR6 — fits 34B models comfortably
  Cloud cost: ~$2-3/hour (on-demand)
  Best for: Cost-efficient production serving of 7B-34B models

NVIDIA RTX 4090 (24GB VRAM)
  Designed for: Consumer/workstation
  Memory: 24GB GDDR6X — fits 7B-13B models (with quantization)
  Cost: ~$1,800 to purchase (consumer market)
  Best for: Development environments, single-user applications,
           edge deployment with tight budget constraints

AMD Instinct MI300X (192GB HBM3)
  Designed for: Competing with H100 on large model serving
  Memory: 192GB — fits 70B models with significant headroom
  Increasingly competitive: ROCm support in vLLM improved significantly
  Best for: Large model serving where VRAM capacity is the bottleneck

The VRAM Sizing Rule

VRAM is the primary constraint in LLM inference. The model weights must fit in VRAM along with the KV cache for active requests.

VRAM REQUIREMENTS BY MODEL SIZE

Model size          FP16 (half precision)   INT8 (8-bit quant)   INT4 (4-bit quant)
7B parameters       ~14GB                   ~7GB                 ~4GB
13B parameters      ~26GB                   ~13GB                ~7GB
34B parameters      ~68GB                   ~34GB                ~17GB
70B parameters      ~140GB                  ~70GB                ~35GB
Llama 4 Scout(109B) ~218GB                  ~109GB               ~55GB

GPU VRAM available:
  RTX 4090: 24GB → can run 7B FP16 or 13B INT4
  A100 80GB: 80GB → can run 34B FP16 or 70B INT4
  H100 80GB: 80GB → same as A100 but faster
  2× H100 (tensor parallel): 160GB → can run 70B FP16
  4× H100 (tensor parallel): 320GB → can run 70B FP16 with KV cache headroom

Rule of thumb: model_size_GB × 1.2 for the model weights alone.
Add 20-30% for KV cache at target concurrency.

Quantization: Reducing VRAM Requirements

Quantization reduces the numerical precision of model weights, shrinking VRAM requirements at the cost of some quality.

QUANTIZATION FORMATS (2026)

GGUF (used by Ollama, CPU-compatible):
  Q4_K_M: ~4-bit, good quality-to-size balance, CPU-runnable
  Q8_0:   ~8-bit, near-lossless quality, still smaller than FP16
  Pros: CPU support, broad compatibility
  Cons: Lower throughput than GPU-native formats

AWQ (Activation-aware Weight Quantization):
  4-bit quantization designed for GPU inference
  Calibrated on representative data — better quality than naive INT4
  Supported natively by vLLM: --quantization awq
  Pros: Minimal quality loss vs. FP16, half the VRAM requirement
  Cons: Requires AWQ-quantized model weights

GPTQ:
  Alternative 4-bit GPU quantization
  More widely available for older models
  Supported by vLLM: --quantization gptq
  Slightly lower throughput than AWQ in most benchmarks

FP8 (8-bit floating point, H100 native):
  H100 GPUs have native FP8 hardware support
  Better quality retention than INT8 at similar memory reduction
  Supported in vLLM for H100: --dtype fp8
  Recommended for H100 production deployments

The quantization recommendation: For production H100/A100 deployments — use AWQ or FP8 (H100). For edge/CPU deployments — use GGUF Q4_K_M or Q8_0 via Ollama. Never use unquantized FP32 in production — it provides no benefit over FP16 and uses 2x the VRAM.


14.4 The Private Deployment Architectures

Different data residency and security requirements call for different deployment topologies.

Architecture 1: Cloud VPC Deployment

Host the model on cloud GPU instances inside your organization's VPC. Data stays within your cloud account and never reaches external model provider infrastructure.

CLOUD VPC DEPLOYMENT ARCHITECTURE

┌────────────────────────────────────────────────────────────────┐
│                    YOUR CLOUD VPC                               │
│                                                                │
│  ┌─────────────────────────────────────────────────────────┐  │
│  │  Application Layer                                       │  │
│  │  ├── API Services (your existing apps)                   │  │
│  │  └── LLM Gateway (LiteLLM) → routes to internal servers │  │
│  └─────────────────────┬───────────────────────────────────┘  │
│                         │ Internal VPC network only            │
│  ┌──────────────────────▼─────────────────────────────────┐   │
│  │  Inference Layer                                        │   │
│  │  ├── vLLM Server 1 (H100 instance) — primary           │   │
│  │  ├── vLLM Server 2 (H100 instance) — replica/scale-out │   │
│  │  └── Load balancer (internal only)                      │   │
│  └─────────────────────┬───────────────────────────────────┘  │
│                         │                                      │
│  ┌──────────────────────▼─────────────────────────────────┐   │
│  │  Model Storage                                          │   │
│  │  ├── Model weights in S3/GCS/Azure Blob (private)       │   │
│  │  └── Loaded to GPU on startup                           │   │
│  └─────────────────────────────────────────────────────────┘  │
│                                                                │
│  [NO traffic leaves the VPC to external model provider APIs]   │
└────────────────────────────────────────────────────────────────┘

Infrastructure management:
  ├── GPU auto-scaling: KEDA (Kubernetes Event-Driven Autoscaling)
  │     Scale up on queue depth; scale down on idle
  ├── Health monitoring: vLLM metrics endpoint → Prometheus → Grafana
  ├── Model versioning: model weights tagged and versioned in object storage
  └── Failover: multi-region deployment for high-availability

Architecture 2: Air-Gapped On-Premises

For environments with no external internet connectivity — government, defense, highly regulated financial institutions.

AIR-GAPPED DEPLOYMENT ARCHITECTURE

Physical requirements:
  ├── On-premises GPU server (H100 or A100 DGX system)
  ├── Local object storage for model weights (MinIO, NetApp)
  ├── No outbound internet from the inference network
  └── Model weights physically transported and cryptographically verified

Software stack:
  ├── vLLM: inference server (runs entirely offline)
  ├── Harbor/Nexus: private container registry (no Docker Hub access)
  ├── Private pip mirror: all Python dependencies pre-cached
  └── Prometheus + Grafana: local observability (no cloud monitoring)

Model acquisition process:
  ├── Download model weights to an internet-connected staging environment
  ├── Verify SHA256 checksums against published values
  ├── Transfer to air-gapped environment via secure media
  └── Verify checksums again on the air-gapped side before loading

Update process:
  ├── Model updates require physical media transfer
  ├── This creates a natural change control gate
  └── Each model version formally validated before deployment

Architecture 3: Managed Cloud AI (Middle Path)

For organizations that need better data control than consumer APIs without the operational burden of self-hosted infrastructure.

MANAGED CLOUD AI OPTIONS

AWS Bedrock:
  ├── Hosts: Claude (Anthropic), OpenAI models (added 2026), Llama (Meta), Mistral, Amazon's own models
  ├── Data: stays within your AWS account/region
  ├── No training on your data (contractual)
  ├── Private endpoints: via PrivateLink (no public internet)
  ├── Model versions: available approximately 2-4 weeks after provider release
  └── Best for: AWS-first organizations, regulated industries

Azure OpenAI Service:
  ├── Hosts: OpenAI GPT-6 and GPT-5.x families plus image and speech models;
  │     Microsoft Foundry (which now fronts Azure OpenAI) also hosts Claude
  ├── Data: stays within your Azure tenant
  ├── Enterprise data agreement: Microsoft's enterprise terms apply
  ├── Private endpoints: via Azure Private Link
  ├── Regional compliance: EU data residency, government regions
  └── Best for: Microsoft-first organizations, strict data sovereignty

Google Gemini Enterprise Agent Platform (formerly Vertex AI):
  ├── Hosts: Gemini family + partner models (including Claude)
  ├── Data: stays within your Google Cloud project
  ├── Private endpoints: via VPC Service Controls
  └── Best for: Google Cloud-first organizations

DECISION FACTORS FOR MANAGED VS SELF-HOSTED:

  Use Managed Cloud when:
  ├── You need frontier model quality (not available self-hosted)
  ├── Data residency is met by the cloud provider's controls
  ├── Engineering team cannot absorb GPU infrastructure operations
  └── Usage pattern is bursty (pay-per-token, no idle cost)

  Use Self-Hosted when:
  ├── Data cannot leave your infrastructure perimeter (air-gap required)
  ├── Volume is high enough that self-hosting cost is lower than API cost
  ├── Open-weight model quality meets requirements (often the case for 2026)
  └── Fine-tuning on proprietary data is required

14.5 Scaling Self-Hosted Inference

Key Metrics to Monitor for Scaling Decisions

INFERENCE SERVING METRICS (what vLLM exposes)

Throughput:
  avg_prompt_throughput_toks_per_s   — input processing rate
  avg_generation_throughput_toks_per_s — output generation rate

Latency:
  Time to First Token (TTFT)          — latency to first output token
  Time per Output Token (TPOT)        — latency between tokens (affects UX)
  End-to-end request latency

GPU utilization:
  gpu_cache_usage_perc                — KV cache utilization (0-100%)
    > 90%: GPU near capacity, requests will queue
    < 30%: GPU underutilized, scale down or increase batch size
  gpu_prefix_cache_hit_rate           — prompt prefix caching effectiveness

Queue depth:
  num_requests_waiting                — requests queued for GPU slots
  num_requests_running                — requests currently processing

When to scale out:
  ├── gpu_cache_usage_perc > 85% consistently (VRAM bottleneck)
  ├── num_requests_waiting > 10 consistently (throughput bottleneck)
  └── P99 TTFT > SLA threshold consistently (latency bottleneck)

Autoscaling Architecture

KUBERNETES-BASED AUTOSCALING FOR VLLM

KEDA (Kubernetes Event-Driven Autoscaling):
  ScaledObject:
    trigger: Prometheus metric — num_requests_waiting
    minReplicas: 1
    maxReplicas: 8
    threshold: 10  (scale up when queue > 10 requests)

  Scale-down:
    stabilization_window: 300s (wait 5 minutes before scaling down)
    cooldown: 120s (wait 2 minutes between scale events)

  Important: GPU pods take 3-5 minutes to start (model loading time)
  Solution: Keep min_replicas = 1 always on; only the burst capacity scales

Node management:
  ├── GPU node pool: dedicated to inference workloads
  ├── Node auto-provisioning: provision GPU instances on demand
  ├── Preemptible/spot instances: for batch workloads (not real-time)
  └── Node cleanup: decommission GPU instances during off-peak
        (GPU instances cost money even when idle)

14.6 Model Serving Patterns

Beyond the infrastructure layer, the serving pattern — how models are accessed and managed — has significant implications for flexibility, governance, and cost.

Pattern 1: Single Model, Direct Access

Client → vLLM Server (single model)

Simplest. No routing overhead. No flexibility. If you need a different model or the model changes, the client code changes. Appropriate for: single-use-case deployments where model diversity is not needed.


Pattern 2: LLM Gateway with Model Registry

Client → LLM Gateway (LiteLLM / Portkey)
           ├── Route: simple queries → Self-hosted Llama 4 Scout
           ├── Route: complex queries → Claude Sonnet (API)
           ├── Route: sensitive data → Self-hosted Mistral (air-gapped)
           └── Fallback: primary unavailable → secondary provider

The recommended pattern for enterprises with multiple model needs. Discussed extensively in the integration patterns reference (Artifact 2, Pattern 1). The key benefit: any change to model routing is a gateway configuration change, not an application code change.


Pattern 3: Model-per-Domain (Specialized Deployments)

Customer Support Domain → Customer Support vLLM Server (fine-tuned Llama 4)
Financial Analysis Domain → Analysis vLLM Server (Mistral-based, fine-tuned)
Code Generation Domain → Code vLLM Server (code-specialized open-weight model)

Each domain has its own fine-tuned model and serving infrastructure. The specialization improves quality for domain-specific tasks. The cost: significantly more infrastructure to operate and govern.

Appropriate when: a single general model consistently underperforms for specific domains even with good prompting, and the volume justifies the infrastructure overhead.


14.7 Fine-Tuning: When and How (Briefly)

Fine-tuning is the process of adapting an open-weight base model on domain-specific data to improve performance on domain-specific tasks. This module covers only the architectural implications — the full fine-tuning methodology is a data science topic.

When Fine-Tuning Makes Architectural Sense

Fine-tuning is warranted when: - A general model consistently underperforms on the task even with optimized prompting and RAG - The task has a domain-specific vocabulary or format that the base model doesn't handle well - You need consistent output format that prompting alone can't guarantee - You have high-quality labeled examples (minimum 500-1,000) and the infrastructure to fine-tune

Fine-tuning is NOT warranted when: - The base model's failure is a knowledge gap (use RAG, not fine-tuning) - You have fewer than 500 high-quality examples (use few-shot prompting) - The team has no experience validating fine-tuned models (creates ungoverned risk) - You need to deploy quickly (fine-tuning + evaluation takes weeks)

Fine-Tuning Architecture Implications

FINE-TUNING CREATES A MODEL LIFECYCLE

Base model → Fine-tuned adapter (LoRA weights)
               ├── Stored in your model registry
               ├── Version-controlled (v1.0, v1.1, v2.0)
               ├── Requires evaluation before deployment
               │     (fine-tuned model may have unexpected behaviors)
               ├── Requires re-validation when base model is updated
               └── Falls under model risk management (Module 11)
                     if used for consequential decisions

LoRA (Low-Rank Adaptation):
  - Train only a small adapter layer on top of the frozen base model
  - Adapter size: ~100MB-1GB vs. the base model's 40-140GB
  - At inference: merge LoRA weights with base model (small overhead)
  - License note: Apache 2.0 models (Qwen, some Llama variants) and
    MIT models (DeepSeek V4, Phi-4) permit fine-tuned derivatives

Fine-tuning tools:
  Unsloth: fast, memory-efficient fine-tuning (supports LoRA, QLoRA)
  Axolotl: flexible fine-tuning framework, good for experimentation
  LLaMA-Factory: multi-model fine-tuning platform

14.8 The Infrastructure Governance Checklist

Hardware and capacity - [ ] VRAM sizing validated: model + KV cache at target concurrency fits? - [ ] GPU instance type matched to workload (inference-optimized for inference)? - [ ] Quantization evaluated: quality impact validated against task requirements? - [ ] Autoscaling configured with appropriate scale-up/down thresholds? - [ ] Minimum replica kept running to avoid cold-start latency?

Inference server configuration - [ ] vLLM version pinned to current stable (check vllm.ai/GitHub releases — new minors ship every ~2 weeks)? - [ ] Model Runner V2 evaluated (VLLM_USE_V2_MODEL_RUNNER=1)? - [ ] API authentication configured (--api-key or upstream auth)? - [ ] Max model length configured (prevent VRAM overflow on long requests)? - [ ] OpenAI-compatible API verified (clients can switch without code changes)?

Observability - [ ] vLLM Prometheus metrics endpoint active? - [ ] Key metrics in dashboard: gpu_cache_usage, TTFT, queue_depth? - [ ] Autoscaling alerts configured (queue_depth > threshold)? - [ ] Cost per GPU hour tracked (idle GPU is wasted spend)?

Security and governance - [ ] Model weights verified (SHA256 checksums against published values)? - [ ] Model weights stored in access-controlled object storage? - [ ] No direct internet access from inference nodes (egress restricted)? - [ ] Fine-tuned models in model registry with version history? - [ ] Fine-tuned models under model risk management if used for decisions?

Managed cloud (if applicable) - [ ] Private endpoints configured (no public internet for inference traffic)? - [ ] Data processing agreement reviewed? - [ ] Data residency region confirmed and matches regulatory requirements? - [ ] Model version pinned (not using default alias that auto-updates)?


EXERCISE — Capacity Planning: You need to serve Llama 4 Scout (109B parameters) at 100 concurrent requests with P99 latency under 3 seconds for first token. Using the VRAM sizing rules and GPU specifications from this module: (1) How many A100 80GB GPUs do you need with FP16 weights? (2) How many with AWQ 4-bit quantization? (3) What is the monthly infrastructure cost at each configuration? (4) What is the break-even request volume vs. using the Claude Sonnet API at $3/$15 per million tokens?

PONDER — The Self-Host Decision: Your organization currently spends $80,000/month on LLM API calls across 5 applications. The usage pattern is relatively consistent (not bursty). A team proposes self-hosting Llama 4 Maverick to reduce costs. What information do you need to evaluate this proposal? What are the non-cost considerations (operations, quality, governance) that must be assessed alongside the cost comparison?

WORKSHOP — Infrastructure Architecture Design: Design the inference infrastructure for a financial services organization that: processes 500,000 document classification tasks per day (batch, non-real-time), serves a real-time customer chat assistant at peak 200 concurrent users, and has a strict requirement that customer data never leaves the organization's cloud VPC. Design: the model selection for each workload, the inference server configuration, the hardware requirements, the autoscaling rules, and the observability stack. Calculate the monthly infrastructure cost.


Next: Module 15 — AI Coding Assistants: Architecture, Governance & Risk