APPENDIX G — The Current Landscape (Volatile Data Reference)¶
⚠️ THIS IS THE DOCUMENT IN THE COURSE THAT GOES STALE QUICKLY.
All time-sensitive data — model versions, pricing, valuations, market statistics, platform feature status, and regulatory state — is consolidated here. Modules reference this appendix rather than embedding numbers inline, so the course stays current by updating this one document. (Module 2's landscape snapshot and Module 37's examples are also dated; refresh them alongside this appendix.)
Last updated: October 2026 (full refresh of G.2, G.3, G.5–G.8; G.4 carried forward, see note) Refresh cadence: Quarterly (or before each teaching cohort) Next scheduled review: January 2027 (note: Gemini 3.8 Flash introductory pricing ends December 31, 2026, see G.3)
When you update this appendix, update the "Last updated" date and log the change in Section G.9.
G.1 How to Use This Appendix¶
The modules teach reasoning that doesn't expire. This appendix holds the facts that do. When a module says "see the Current Landscape appendix for pricing," the number lives here.
The maintenance principle: Never embed a price, a model version, a valuation, or a market statistic in a module. Put it here and reference it. That keeps all the decay in one maintainable document.
The decay rates to expect: - Pricing: changes monthly; verify before every cohort. New in 2026: introductory prices that step up on a fixed date, and long-context surcharges. - Model versions: new releases every few weeks; verify quarterly. - Platform names: three major rebrands in 12 months (Azure AI Foundry → Microsoft Foundry; Vertex AI → Gemini Enterprise Agent Platform; xAI → SpaceXAI); verify quarterly. - Provider access and terms: episodic but high-impact (see Module 37 §37.11); check on every cohort. - Startup valuations/ARR: change constantly; verify quarterly. - Market statistics: annual studies; verify annually. - Regulatory status: episodic; verify quarterly for pending changes.
Where to verify: Provider model catalogs, deprecation pages, and independent leaderboards are listed in Module 2 §2.1 — Where to Check the Latest Models. Pricing pages are listed in G.3.
Source quality note: Entries marked (reported) come from secondary or press sources that were not confirmed against a primary source during this refresh. Entries marked (verify) had incomplete or conflicting data. Do not quote either in client-facing material without checking.
G.2 Model Reference (October 2026)¶
Maps to: Module 2 (Model Ecosystem), Module 21 (Emerging Patterns), Module 37 (Model Portability) Verify against: provider documentation, lmarena.ai, artificialanalysis.ai
Frontier Models (closed, API-only)¶
FRONTIER TIER (October 2026)
Provider Flagship / tiers Reasoning control Notes
──────────────────────────────────────────────────────────────────────────────────────────────
OpenAI GPT-6 Astra (flagship, Sep 2026) Reasoning-effort setting GPT-6 family: Astra (most capable),
GPT-6.1 Sol / GPT-6 Sol (balanced) Sol (balanced), Luna (cheapest).
GPT-6 Luna (efficient) GPT-5.5 / 5.6 still served.
Also on AWS Bedrock (2026).
Anthropic Claude Fable 5.1 (most capable Adaptive thinking + Opus 5.5 = default flagship;
widely released) effort: low → max Sonnet 5.5 = everyday tier;
Claude Opus 5.5 · Sonnet 5.5 · (fixed token budgets Haiku 4.5 = small/fast.
Haiku 4.5 removed on newest models) 1M context on Fable/Opus/Sonnet.
Google Gemini 3.8 Flash (GA, newest) Thinking levels: Gemini 3.1 Pro listed as preview.
Gemini 3.1 Pro (preview) low / medium / high 3.8 Live (real-time), 3.8 Flash
Gemini 3.5 Flash-Lite Cyber variants. 1M context.
No "Gemini 4" as of Oct 2026.
SpaceXAI Grok 4 / 4.1 Reasoning variants xAI merged into SpaceX (Feb 2026).
(formerly Grok 5 in training, unreleased
xAI) as of Oct 2026 (reported).
Meta Muse Spark (Apr 2026) — Closed-weight, cloud-only. Marks
Meta's shift away from open-weight
for its newest flagship.
Open-Weight Models¶
OPEN-WEIGHT TIER (October 2026)
Model Size (total / active) License Architectural Sweet Spot
─────────────────────────────────────────────────────────────────────────────────────────────
DeepSeek V4 Pro 1.6T / ~49B MoE MIT Near-frontier general + code; 1M ctx
DeepSeek V4 Flash 284B / ~13B MoE MIT Cost-efficient general; 1M ctx
Kimi K3 (Moonshot) 2.8T / ~50B MoE Modified MIT Agentic, long context (1M); Jul 2026
GLM-5.2 (Zhipu) MoE (verify) (verify) General; 1M ctx extension; Jun 2026
Qwen 3.6 (35B-A3B) 35B / 3B MoE Apache 2.0 Very cheap inference; 262K ctx; Apr 2026
Mistral Large 3 675B MoE Apache 2.0 Multilingual; EU vendor; Dec 2025
Llama 4 Maverick 400B / 17B MoE Llama license General; 1M ctx; Meta's latest open flagship
Llama 4 Scout ~109B MoE Llama license Cost-efficient, long context
Google Gemma 4 E2B/E4B · 26B MoE (4B Apache 2.0 Edge (E2B/E4B, 128K ctx) to local servers
active) · 31B dense (26B/31B, 256K ctx); Apr 2, 2026
OpenAI gpt-oss Various Apache 2.0 Open-weight from OpenAI; on Bedrock GovCloud
NVIDIA Nemotron Various NVIDIA open Reasoning, NVIDIA-optimized
Llama 4 Behemoth: announced, never released as open weights; paused (reported).
JURISDICTION NOTE: The open-weight frontier is now led by Chinese labs
(DeepSeek, Kimi, GLM, Qwen). Self-hosting the weights resolves data
residency (Module 2) but NOT procurement/supply-chain policies that
restrict models by origin — check those separately.
Reasoning Models¶
REASONING (October 2026)
Reasoning is now a MODE of flagship models, not a separate family:
OpenAI: reasoning-effort setting on GPT-5.x / GPT-6
Anthropic: adaptive thinking + effort (low/medium/high/xhigh/max);
fixed budget_tokens deprecated, then removed on newest models;
thinking cannot be disabled on Opus 5.5 / Fable 5.1
Google: thinking levels (low/medium/high) on Gemini 3.x
Open: DeepSeek V4 (unified), Kimi K3, Nemotron reasoning variants
Trajectory: the control moved from "token budget" to "effort level" in
~12 months, and defaults now differ between versions (e.g., Opus 5.5
defaults to medium effort; Opus 5 defaulted to high). Always set reasoning
depth explicitly (Module 37 §37.2, layer 5).
Small Language Models (SLMs)¶
SLM TIER (October 2026)
Google Gemma 4 E2B / E4B (Apr 2026, Apache 2.0) — 128K ctx, image/video/
audio input, runs offline; E4B ~5GB RAM at 4-bit
Microsoft Phi-4-mini (3.8B) — strong reasoning per parameter
Alibaba Qwen 3.5 small variants (0.8B–9B); Qwen3 1.7B for multilingual
Meta Llama 3.2 (1B / 3B)
Hugging Face SmolLM3
Sweet spot: edge deployment, on-device, high-volume classification,
cost-sensitive routing tier.
Typed-Decision ("System One") Models¶
TYPED-DECISION TIER (October 2026) — all figures vendor-reported / secondary
Model Vendor / date Form Notes
─────────────────────────────────────────────────────────────────────────────────────
Jev TypeSafe AI, early access Hosted API ~$0.042/M input tokens; 40–200x
Sep 15, 2026 faster than frontier LLMs (claimed)
Laya Convai Innovations Open source, local ~421M-param bidirectional encoder;
~33 ms/call; check licence
Clef / Clef- Cloudflare, Oct 1, 2026 Open weights (Apache 2.0) Post-trained from Qwen3.8-27B /
flash Qwen3.5-9B; Jev-compatible API;
~39 ms median (flash, reported)
All use typed questions (choice / score / yes-no) and return calibrated
probabilities; all are described as trained with RLCD (reward reported to be a
strictly proper scoring rule). The RLCD method is unpublished. Independent
comparisons are early and mixed — benchmark on your own labelled data.
Where to use them: Module 16 §16.9.
Real-Time Voice¶
VOICE (October 2026)
OpenAI gpt-realtime-2.1 / 2.1-mini (Jul 2026) — GPT-5-class reasoning,
128K ctx, parallel tool calls; plus GPT-Realtime-Translate / -Whisper
Google Gemini Live API / Gemini 3.8 Live — lower cost, video input
Market Velocity Statistic¶
255 model releases from major organizations in Q1 2026 (carried forward; Q3 2026 count not refreshed). The competitive picture changes roughly every 6 weeks. (Used in Module 2. Refresh the number each quarter.)
G.3 Pricing Reference (October 2026)¶
Maps to: Module 13 (Cost Engineering), Module 2 (Model Ecosystem), Module 37 (Portability) ⚠️ HIGHEST DECAY RATE IN THE COURSE — verify before every cohort. Verify against: Anthropic · OpenAI · Google Gemini · DeepSeek · AWS Bedrock · Google Cloud · Azure
TOKEN PRICING REFERENCE (October 2026, USD per 1M tokens, first-party API list prices)
Model Input Output Cached input Notes
────────────────────────────────────────────────────────────────────────────────────────
FRONTIER:
Claude Fable 5.1 $10.00 $50.00 $0.25 Requires 30-day retention (no ZDR
unless expressly authorized)
GPT-6 Astra $10.00 $50.00 (verify) >272K input: 2x input / 1.5x output
for the whole request
Claude Opus 5.5 $4.00 $20.00 $0.20 Fast mode: $8 / $40
GPT-5.5 $5.00 $30.00 (verify)
GPT-5.5 Pro $30.00 $180.00 n/a
Gemini 3.1 Pro (preview) $2.00 (verify) (verify) Tiered by prompt length (verify)
MID-TIER:
Claude Sonnet 5.5 $2.00 $10.00 $0.20
GPT-6 Sol $2.00 $10.00 (verify) GPT-6.1 Sol: verify price
Gemini 3.8 Flash $0.75 $3.75 $0.075 INTRODUCTORY through 2026-12-31;
$1.50 / $7.50 / $0.15 from 2027-01-01
Claude Haiku 4.5 $1.00 $5.00 ~$0.10
EFFICIENT:
GPT-6 Luna $0.10 $0.50 (verify)
Gemini 2.5/3.5 Flash-Lite ~$0.10 ~$0.40 (verify)
DeepSeek V4 Flash ~$0.10–0.14 ~$0.20–0.28 ~1/10 input Reported range across snapshots
DeepSeek V4 Pro ~$1.32–1.74 ~$3.48–3.96 ~1/10 input (reported; verify on DeepSeek page)
SELF-HOSTED:
Open-weight models: ~$0 per token, but GPU cost — e.g., ~$2,000–3,000/month per
A100-class GPU (carried forward from May 2026; verify current H100/B200 pricing)
CLOUD-HOSTED NOTE:
Claude on Microsoft Foundry is billed at Anthropic first-party rates.
Claude on Bedrock / Google Cloud and OpenAI models on Bedrock are
partner-priced — check each cloud's page; do not assume parity.
STRUCTURAL FACTS (stable even as numbers change):
1. Output tokens cost ~5–6x input for closed frontier models (Claude 5x,
GPT-6 5x, GPT-5.5 6x, Gemini 3.8 Flash 5x); ~2–3x for DeepSeek V4.
→ favor input-heavy / output-light designs.
2. Cached input is ~1/10 (or less) of uncached input across major providers.
→ caching is the largest free lever; failover to a cold cache is a cost
spike (Module 37 §37.7).
3. NEW 2026 PATTERNS that break naive cost models:
- Introductory pricing with a fixed step-up date (Gemini 3.8 Flash → 2x on 2027-01-01)
- Long-context surcharges applied to the WHOLE request (GPT-6 >272K)
- Tokenizer changes between versions (same text, different token count —
e.g., Anthropic's Opus 4.7+ tokenizer: ~1.0–1.35x vs. earlier models)
- Per-token price DROPS on newer flagships (Opus 5.5 is cheaper than Opus 5)
→ re-run cost-per-completed-task, not just price-per-token
G.4 Market Statistics (carried forward from May 2026)¶
Maps to: Module 19 (Business Value), Module 20 (Hype vs Reality), Module 32 (Agent Learnings), Module 10 (Shadow AI) Verify against: original studies (cited); refresh annually.
October 2026 note: These statistics were not re-verified in this refresh. Most come from annual studies, so they are still the most recent published figures as far as we know. Quote them as "as of the cited study," not "as of October 2026." Check for new MIT / IBM / Gartner / Forrester releases before the next cohort.
AI ADOPTION & ROI STATISTICS (as of studies cited, May 2026 baseline)
PILOT-TO-PRODUCTION:
79% of organizations report productivity gains from AI (IBM)
5% achieve substantial ROI (IBM)
95% of generative AI pilots fail to deliver measurable P&L impact (MIT 2025)
88% of AI pilots never reach production
46% of AI POCs abandoned before production
AI AGENTS SPECIFICALLY (March 2026 survey, 650 enterprise tech leaders):
78% have at least one agent pilot running
14% have scaled an agent to organization-wide operational use
5 gaps account for 89% of scaling failures:
integration complexity, inconsistent output quality at volume,
absent monitoring tooling, unclear ownership, insufficient domain data
AGENT TASK COMPLETION:
30.3% — best AI agent models on real-world office tasks
(Carnegie Mellon TheAgentCompany benchmark — likely superseded
by newer model generations; re-check before quoting)
58% of deployed agents require manual override 20-40% of the time
(Forrester Q1 2026)
73% of "agentic AI" vendors had no published accuracy benchmarks
(Compyl 2026)
67% of GRC buyers say vendor autonomy claims are overstated/misleading
(451 Research 2026)
SHADOW AI (Module 10):
45% of employees are regular AI users
~70% of AI use is outside IT oversight
SECURITY (Module 9):
89.6% roleplay-based prompt injection success rate against frontier
models (2025 study, 1,400+ adversarial prompts)
<17 min average jailbreak time for GPT-4 (historical — GPT-4 era)
86% of prompt injection attacks partially succeeded against web agents
(Meta research)
29.1% security vulnerability rate in AI-generated Python (Module 15)
6.4% secret leakage rate in AI coding tool usage (Module 15)
OWASP 2026 LLM Top 10 methodology drew on 6,639 real incidents (Sep 2026)
BUDGET ALLOCATION (Module 19):
>50% of genAI budgets go to sales/marketing
Highest ROI is in back-office automation (the inverse — MIT finding)
67% buy success rate vs 33% internal build success rate (MIT)
PLATFORM SCALE (new, Oct 2026):
Amazon Bedrock: 125,000+ customers, ~80% of Fortune 100 (Q1 2026, reported;
an earlier draft said 225,000+, which could not be confirmed)
G.5 Startup and Vendor Landscape Snapshot (October 2026)¶
Maps to: Module 18 (Startup Landscape), Module 37 (Portability case studies), Appendix E ⚠️ Valuations and ARR change constantly — verify quarterly. Verify against: company announcements, Crunchbase, press.
VERTICAL AI VALUATIONS/TRACTION (October 2026)
Legal:
Harvey $15.6B valuation ($550M round, 2026); ARR >$400M (reported)
(was $11B / ~$200M ARR in May 2026 snapshot)
Legora $5.55B valuation, €550M Series D (May 2026; not re-verified)
Healthcare:
Abridge $5.3B valuation (May 2026; not re-verified)
Hippocratic AI Hospital partnerships in production
Enterprise Search:
Glean >$100M ARR (May 2026; not re-verified)
Perplexity $9B+ valuation (May 2026; not re-verified)
AI Coding:
Cursor ACQUIRED by SpaceX — $60B all-stock; closed Aug 14, 2026.
~$3B+ ARR before acquisition (reported ~$4B annualized by Jun 2026).
OpenAI ending model access via Cursor; proposed shutoff
Nov 12, 2026 (OpenAI ≈5% of Cursor traffic) — Module 37 Case 3
GitHub Copilot 4.7M paid seats, 90% Fortune 100 (May 2026; not re-verified)
Claude Code (refresh adoption figures before quoting)
Replit $150M ARR annualized (May 2026; not re-verified)
AI Labs (structural changes):
xAI Merged into SpaceX (Feb 2026); operates as SpaceXAI
Meta Meta Superintelligence Labs; Muse Spark closed-weight (Apr 2026)
Infrastructure:
Groq, Cerebras (custom silicon)
ElevenLabs, Deepgram (voice)
Weaviate, Qdrant, Pinecone (vector DB)
Unstructured.io, Scale AI, Databricks (data)
LLM Gateways / AI control plane (Module 17, Module 37):
Portkey ACQUIRED by Palo Alto Networks (closed May 29, 2026);
becoming the AI gateway for Prisma AIRS
LiteLLM Open-source, self-hostable; still the de facto OSS gateway
OpenRouter Hosted aggregator; 400+ models, 70+ providers (company figures)
Helicone Reported absorbed by Mintlify (Mar 2026) and frozen (reported)
MARKET CONTEXT:
AI captured ~80% of Q1 2026 VC funding (May 2026 snapshot)
AI startups raised $100B+ globally in 2025
"SaaSpocalypse" — ~$2T wiped from enterprise SaaS valuations (May 2026 snapshot)
Provider access withdrawals are now a recurring pattern:
Windsurf (Jun 2025), OpenAI→Claude (Aug 2025), Cursor→OpenAI (Aug 2026)
G.6 Cloud Platform Feature Status (October 2026)¶
Maps to: Module 31 (Cloud Provider AI Platforms), Module 2 (Managed Private Deployments) Verify against: AWS What's New · Microsoft Foundry what's new · Gemini Enterprise Agent Platform release notes
AWS BEDROCK / AGENTCORE (October 2026)
Amazon Bedrock AgentCore: GA October 13, 2025
AgentCore Policy: GA March 3, 2026
AgentCore Evaluations: GA March 31, 2026
Components: Runtime, Gateway (MCP + API/Lambda tools), Memory, Identity,
Observability (CloudWatch + OTel), Code Interpreter, Browser Tool
NEW 2026:
OpenAI models GA on Bedrock — GPT-5.5 / 5.4 / Codex (Jun 2026);
GPT-5.6 Sol / Terra / Luna (Jul 2026, 1M ctx + prompt caching, Aug 2026);
GPT-6 Astra, GPT-6 Sol / Luna, GPT-6.1 Sol (Sep 2026)
AgentCore Runtime: microVM sessions up to 8 hours; runtime instances
(dedicated compute) up to 14 days — GA Aug 2026
Strands Agents SDK open-sourced May 2025; first public Strands harness
releases (Python + TypeScript) Sep 21, 2026 (reported)
Bedrock Agents Classic: closed to new customers from Jul 30, 2026 (no new
features; AgentCore is the recommended path)
OpenAI models on Bedrock at launch (Jun 2026): AWS stated pricing matches
OpenAI first-party rates for GPT-5.5/5.4 — verify per model
OpenAI gpt-oss + NVIDIA Nemotron in AWS GovCloud (Apr 2026)
A2A support: yes (since Oct 2025)
MICROSOFT FOUNDRY (formerly Azure AI Foundry) (October 2026)
Renamed at Ignite (Nov 2025); official in Product Terms Jan 2026 — folds
Azure AI Foundry, Azure AI Studio, and Azure AI Services into one resource
Foundry Agent Service: GA March 2026
Hosted Agents (managed runtime, per-session sandbox): GA July 9, 2026
Entra ID for Agents (each agent gets a directory identity)
Microsoft Agent Framework 1.0 (open-source, MCP + A2A) — Apr 3, 2026
Model catalog: 11,000+ models (company figure); includes OpenAI and Claude
Claude on Foundry billed at Anthropic first-party rates
Deep M365 integration (Teams, Outlook, SharePoint, Graph)
GOOGLE — GEMINI ENTERPRISE AGENT PLATFORM (formerly Vertex AI) (October 2026)
Announced Apr 22, 2026 as the replacement/evolution of Vertex AI;
Vertex AI services now delivered through Agent Platform
Gemini family (primary), Google Search grounding, Claude available
Agent Runtime (reported: formerly Agent Engine): agents up to 7 days;
Memory Bank, Agent Identity, Agent Gateway
A2A protocol (Google-originated) — native support
Best for data-heavy GCP workloads
PROVIDER POSITIONING (updated):
AWS — AWS-native orgs; now hosts BOTH Claude and OpenAI frontier models,
so the "Azure is the only enterprise OpenAI path" assumption no
longer holds; long-running agents (14 days on runtime instances),
though Google's Agent Runtime (7 days) and Foundry Hosted Agents
(30-day sessions) now overlap
Azure — Microsoft 365 orgs; Entra ID identity; broadest catalog
GCP — data-heavy GCP orgs; Google Search grounding; Gemini-first
G.7 Regulatory & Standards Status (October 2026)¶
Maps to: Module 9 (Security), Module 11 (Compliance), Module 33 (Responsible AI) Verify against: regulator announcements, OWASP, EU AI Act timeline.
MODEL RISK MANAGEMENT (US financial services):
SR 11-7 / OCC 2011-12 superseded April 17, 2026 by updated interagency
MRM guidance (Federal Reserve SR 26-2; OCC; FDIC)
Principles: risk-based, tailored, commensurate with size/complexity/model use
Generative AI / agentic AI: EXPLICITLY OUT OF SCOPE
("novel and rapidly evolving")
Agencies plan a REQUEST FOR INFORMATION on MRM and AI, incl. genAI/agentic
WATCH FOR: the RFI, then genAI-specific guidance (the key pending change)
SECURITY STANDARDS:
OWASP Top 10 for LLM Applications 2026 (published Aug 2026; formally
launched Sep 1, 2026 alongside the Agent Control Standard):
LLM01 Prompt Injection LLM06 Unbounded Consumption
LLM02 Sensitive Info Disclosure LLM07 Misinformation
LLM03 Excessive Agency (↑ from 6) LLM08 Hidden Context Exposure
LLM04 Supply Chain (expanded from System Prompt Leakage)
LLM05 Data and Model Poisoning LLM09 Vector and Embedding Weaknesses
LLM10 Improper Output Handling
Methodology: practitioner voting (75%) + 6,639 real incidents (25%)
Biggest rank moves: Unbounded Consumption 10→6, Excessive Agency 6→3,
Improper Output Handling 5→10
Agent Control Standard (ACS) v0.1, Sep 1, 2026: runtime hooks at agent
execution points; declarative controls enforced at runtime
OWASP Top 10 for Agentic Applications (ASI01–ASI10): Dec 9, 2025
NIST AI RMF: in active use
EU AI ACT — DIGITAL OMNIBUS ON AI (Regulation (EU) 2026/1744):
Political agreement May 7, 2026; adopted Jun 2026; IN FORCE Jul 27, 2026
Annex III stand-alone high-risk obligations: moved Aug 2, 2026 → Dec 2, 2027
Annex I (product-embedded) high-risk: → Aug 2, 2028
Article 50 transparency obligations (disclose AI interaction): STILL Aug 2, 2026
— narrow watermarking grace for already-deployed systems to Dec 2, 2026
Prohibited practices: enforced (since Feb 2025)
WATCH FOR: harmonised standards and Commission guidelines ahead of Dec 2027
INDUSTRY-SPECIFIC:
HIPAA (healthcare), FDA SaMD guidance, fair lending (ECOA, 4/5 rule)
— stable; see Module 11 for application
G.8 Tool Version Reference (October 2026)¶
Maps to: Module 14 (Infrastructure), Module 17 (Platform), Appendix A (Tool Matrix) Verify against: project release pages (e.g., vLLM releases), refresh quarterly.
INFERENCE & TOOLING VERSIONS (October 2026)
vLLM: v0.28.x (Aug 27, 2026 — sparse attention, faster speculative
decoding) (reported; confirm on the releases page)
Ollama: current stable (local/edge, Apple Silicon)
TensorRT-LLM: current (NVIDIA-optimized)
SGLang: current (structured generation focus)
LangGraph: production standard for agent state machines
LlamaIndex: RAG-first framework
LangChain: broad ecosystem
Strands Agents / Strands harness (AWS, open-source, Apache 2.0)
Microsoft Agent Framework (open-source)
RAGAS: RAG evaluation standard
PromptFoo: prompt regression / CI-CD security testing
Garak: LLM vulnerability scanner (NVIDIA)
PROTOCOLS:
MCP spec 2026-07-28 (stateless core; long-running tasks moved to an
extension; Roots/Sampling/Logging and HTTP+SSE deprecated; 12-month
deprecation window). MCP governed by the Agentic AI Foundation
(Linux Foundation) since Dec 9, 2025
A2A v1.0 (early 2026); donated to the Linux Foundation Jun 2025
GATEWAYS (Module 17, Module 37):
LiteLLM: de facto open-source, self-hostable gateway
Portkey: now part of Palo Alto Networks (Prisma AIRS)
OpenRouter: hosted multi-provider aggregator
→ keep your portability contract independent of any one gateway
Note: version numbers move fast; the tool CHOICES (Appendix A) are
more stable than the version numbers. Verify versions before
recommending specific releases.
G.9 Change Log¶
Log every update here so the teaching team knows what changed.
CHANGE LOG
2026-05 Initial consolidated landscape created.
Baseline: model versions, pricing, market stats, startup
valuations, platform feature status, regulatory state all
current as of May 2026.
2026-10 Full quarterly refresh (Aug review slipped to Oct).
G.2 Models: GPT-6 family (Astra/Sol/Luna); Claude Fable 5.1 / Opus 5.5 /
Sonnet 5.5; Gemini 3.8 Flash + 3.1 Pro (preview) — no "Gemini 4";
xAI → SpaceXAI, Grok 5 unreleased; Meta Muse Spark (closed).
Open-weight: DeepSeek V4, Kimi K3, GLM-5.2, Qwen 3.6, Gemma 4.
Reasoning reframed as a mode; budget → effort transition.
Added real-time voice entry.
G.3 Pricing table rebuilt for Oct 2026. Added structural facts on
introductory pricing, long-context surcharges, tokenizer changes.
Output/input ratio updated (5–6x closed; 2–3x DeepSeek).
G.4 NOT re-verified — carried forward with explicit note.
G.5 Harvey $15.6B; Cursor acquired by SpaceX ($60B, closed Aug 14) and
OpenAI access ending Nov 12; Portkey acquired by Palo Alto Networks;
xAI merged into SpaceX; gateway section added.
G.6 OpenAI models on Bedrock (Jun–Sep 2026); 14-day agents; Azure AI
Foundry → Microsoft Foundry; Vertex AI → Gemini Enterprise Agent
Platform. Provider positioning updated.
G.7 SR 26-2 named; RFI pending. OWASP LLM Top 10 2026 list. EU AI Act
Digital Omnibus (Reg. 2026/1744): Annex III → Dec 2, 2027.
G.2 Typed-decision (System One) models added: Jev, Laya, Clef (Oct 2026).
Integration Patterns reference (docs/reference): reviewed and updated; OWASP table
corrected (was the 2023 list labelled 2025); Patterns 13-15 added (context and
harness, model portability, typed-decision gates); Appendices A-C refreshed.
Tool-status facts used there: Portkey (Palo Alto Networks, May 2026), Helicone
(Mintlify, Mar 2026, reported), Langfuse (ClickHouse, Jan 2026), AutoGen
(maintenance mode, Oct 2025), Llama Guard 4, OTel GenAI conventions (Development).
G.8 vLLM v0.28.x; gateways consolidated; Strands harness (Sep 2026).
Modules touched in same refresh: Module 2 (snapshot, links), new Module 37.
Module 13 §13.2 rewritten (structure, not snapshots). Module 34: OpenAI
fine-tuning platform wind-down (announced May 7, 2026; new jobs end
Jan 6, 2027). Module 11: SR 26-2 details ($30B scope, non-enforceable).
Gemma 4 licence corrected to Apache 2.0; MCP spec 2026-07-28 and A2A v1.0
added to G.8; OWASP LLM Top 10 2026 dates reconciled.
[Next entry: January 2027 scheduled review]
- Gemini 3.8 Flash price step-up (effective 2027-01-01) — update G.3
- Cursor/OpenAI shutoff outcome (proposed 2026-11-12) — update G.5
- Grok 5 release status; any Gemini Pro GA; next Claude / GPT releases
- US banking agencies' AI/MRM RFI — update G.7
- Re-verify all "(reported)" and "(verify)" entries
- Refresh G.4 market statistics if new annual studies are out
G.10 The Quarterly Refresh Checklist¶
When refreshing this appendix (quarterly or before a cohort), work through this list:
QUARTERLY REFRESH CHECKLIST
PRICING (G.3) — highest priority:
□ Check each provider's current pricing page (links in G.3)
□ Update the token pricing table
□ Check for introductory prices with step-up dates (flag the date)
□ Check for long-context / tiered-by-length surcharges
□ Check cloud-partner pricing differences (Bedrock / Foundry / Google Cloud)
□ Verify the input/output asymmetry ratio still holds
MODELS (G.2):
□ Check for new frontier releases (OpenAI, Anthropic, Google, SpaceXAI, Meta)
□ Check for new open-weight releases (DeepSeek, Moonshot, Zhipu, Qwen,
Mistral, Meta, Google, NVIDIA, OpenAI gpt-oss)
□ Check reasoning-control changes (new effort levels, changed defaults)
□ Check deprecation pages for retirements affecting course examples
□ Update the model velocity statistic (releases per quarter)
□ Update SLM leaders if new small models released
PROVIDER ACCESS & TERMS (new — Module 37):
□ Any provider withdrawing access from a customer/competitor?
□ Any data-retention or ToS changes per model?
□ Any acquisitions changing who owns a provider, gateway, or tool?
MARKET STATS (G.4):
□ Check for new MIT/IBM/Gartner/Forrester AI adoption studies
□ Update agent production statistics if new survey published
□ Keep the prior numbers as "as of [date]" if no new study
STARTUPS (G.5):
□ Check for new funding rounds / valuation changes for listed companies
□ Add any major new entrant in a covered category
□ Note any company that failed or was acquired
PLATFORMS (G.6):
□ Check AWS/Microsoft/Google what's-new pages for AI agent features
□ Check for product renames (three in the last 12 months)
□ Update which frontier models each cloud hosts
□ Update GA dates for newly-released capabilities
REGULATORY (G.7):
□ Check for genAI-specific MRM guidance / the agencies' RFI
□ Check OWASP for new Top 10 revisions and the Agent Control Standard
□ Check EU AI Act phase-in and harmonised-standards status
After updating:
□ Update "Last updated" date at the top
□ Update "Next scheduled review" date
□ Add a change log entry (G.9)
□ Update Module 2's landscape snapshot box to match G.2
□ Note any module that references a number that changed materially
This appendix is the maintenance surface of the course. Keep it current, and the 37 modules stay current with it. The intellectual architecture in the modules is built to last. This appendix is where the facts that decay are quarantined and refreshed.