MODULE 17 — Platform Engineering for AI¶
⚠️ Currency note: Model IDs and vendor names in this module are illustrative examples as of October 2026. Current models, platform names, and the gateway vendor landscape are in Appendix G (G.2, G.5, G.8). The platform patterns do not depend on which vendors are named.
17.1 The Platform Engineering Mandate¶
Platform engineering for AI answers a specific organizational failure mode: every team building AI from scratch, making the same mistakes, duplicating the same infrastructure, and introducing the same security and governance gaps.
Without an AI platform function, the typical enterprise ends up with: - 12 different teams calling LLM APIs with 12 different authentication mechanisms and 12 different cost attribution approaches - 7 different prompt management strategies ranging from hardcoded strings to careful version control - No shared evaluation infrastructure — each team builds (or more commonly, skips) its own - Security controls that vary by team and change when engineers leave - No visibility into total AI spend or which products are responsible for it - No consistent approach to human oversight, audit trails, or model risk
The platform engineering team's job is to make good AI architecture the path of least resistance. Not to mandate it, not to police it — to make it easier to do it right than to do it wrong. This distinction matters: a mandate with no tooling creates bureaucracy. A golden path with great tooling creates adoption.
17.2 The AI Platform Team Scope¶
The AI platform team owns the infrastructure and tooling that other teams use to build AI features. It does not own the AI features themselves. The distinction is the same as the relationship between the platform engineering team and product feature teams in any mature engineering organization.
What the AI platform team owns:
AI PLATFORM TEAM SCOPE
Infrastructure:
├── LLM gateway (centralized access, authentication, routing, audit)
├── Evaluation infrastructure (eval pipelines, dataset management,
│ LLM-as-judge configuration)
├── Observability stack (AI-specific dashboards, cost attribution)
├── Prompt management system (versioning, deployment, A/B testing)
└── Self-hosted model infrastructure (vLLM servers, model registry)
if the organization has self-hosting requirements
Standards and golden paths:
├── The approved AI service template (how to build a new AI feature correctly)
├── Security patterns (injection defenses, PII handling, content filtering)
├── Evaluation standards (what evals are required before production)
└── Governance requirements (what must be documented, who must approve)
Enablement:
├── Internal developer documentation
├── Example implementations
├── Office hours / consulting for teams building AI features
└── Onboarding support for first-time AI feature teams
What the AI platform team does NOT own:
├── Individual product AI features (owned by product teams)
├── Domain-specific knowledge bases (owned by domain teams)
├── Feature-level eval cases (owned by the feature team)
└── Business decisions about AI use (owned by product and leadership)
17.3 The LLM Gateway: The Platform's Core Infrastructure¶
The LLM gateway is the most important piece of AI platform infrastructure. Everything else depends on it. It is the single point through which all LLM API calls are made, and it provides centralized enforcement of security, cost, and governance policies.
Covered in depth in the Integration Patterns reference (Artifact 2, Pattern 1). From the platform perspective, what matters is:
The Gateway as the Policy Enforcement Point¶
The gateway is the only place where policies can be enforced consistently across all AI features. This is its most important characteristic.
GATEWAY AS POLICY ENFORCEMENT
Policies enforced at the gateway (not in application code):
├── Authentication: every API call must present a valid service identity
├── Authorization: this service is allowed to use these models
├── Rate limiting: this service's token budget is X per period
├── PII scanning: no SSN, credit card, or sensitive identifiers
│ leave the perimeter without pseudonymization
├── Model version: this service's pinned model version is enforced
│ (no accidental use of unpinned aliases)
├── Content filtering: input and output pass through guardrails
└── Audit logging: every call logged with service identity, model,
token count, cost, prompt template ID
Why application code cannot be the enforcement point:
Application code changes. Developers forget. New services are
created without inheriting the patterns. The gateway enforces
the same policies on day one and day 1,000, for all services.
Gateway Configuration Per Service¶
Each service that uses the gateway must be registered with its own configuration profile:
GATEWAY SERVICE PROFILE
service_id: "customer-support-ai"
team: "customer-experience"
# Model IDs below are illustrative (October 2026). In a mature platform the
# service profile references a MODEL PROFILE (Module 37 §37.3) rather than a raw
# model ID, so a model swap edits one profile, not every service profile.
approved_models:
- model: "claude-sonnet-5-5" # or: model_profile: "anthropic-claude-sonnet-5-5"
max_input_tokens: 4000 # design ceiling, not the model's maximum
max_output_tokens: 500
- model: "claude-haiku-4-5" # for high-volume simple classification
max_input_tokens: 1000
max_output_tokens: 100
fallback_models: # pre-approved, so failover needs no ticket
- model: "<second-provider mid-tier>" # must have an evaluated prompt overlay
token_budgets:
monthly_ceiling_usd: 2000
per_request_input_ceiling: 4000
per_request_output_ceiling: 500
daily_alert_threshold_usd: 150 # alert when daily spend hits this
pii_policy:
scan_input: true
scan_output: true
on_pii_detected: "pseudonymize" # pseudonymize | block | allow_and_log
content_policy:
input_filter: "standard_enterprise"
output_filter: "standard_enterprise"
additional_blocked_categories: [] # can add service-specific
audit:
log_prompt_template_id: true
log_full_content: false # full content logged only in debug mode
retention_days: 365
data_handling:
data_residency: "us"
approved_external_models: true # can send to external providers
This profile is the architectural contract between the service team and the AI platform team. It is version-controlled and reviewed when the service team requests changes.
The Gateway Is Transport and Policy, Not the Portability Layer¶
A gateway normalizes the wire call: auth, keys, quotas, PII scanning, audit, retries, and transport-level failover. It does not make two models behave the same. Parameter semantics, forced tool selection, structured-output mechanisms, reasoning-depth controls and defaults, tokenizer differences, and refusal behavior all differ between models, and between versions from the same vendor. Those are handled above the gateway, in a capability and model-adapter layer that you own, with a behavioral contract that defines "equivalent." Treat the gateway product itself as replaceable too. The gateway market consolidated sharply in 2026: Portkey is now part of Palo Alto Networks, LiteLLM remains the de facto open-source option, and OpenRouter is a hosted aggregator (Appendix G §G.5, §G.8). Keep routing rules, model profiles, and prompts in your own repository and formats, not in one gateway's proprietary configuration. See Module 37 — Model Portability (§37.3, "Gateway vs. Adapter: Who Does What").
17.4 The Prompt Management System¶
Prompts are the most impactful and least managed artifact in most AI systems. The prompt management system provides the infrastructure for treating prompts as first-class software artifacts.
What the Prompt Management System Must Do¶
PROMPT MANAGEMENT SYSTEM CAPABILITIES
Storage and versioning:
├── Store prompt templates in a structured format
├── Version every change (semantic versioning: major.minor.patch)
├── Associate every version with author, timestamp, and change rationale
└── Never delete versions (deprecate and archive instead)
Deployment lifecycle:
├── Environments: dev → staging → production
├── Deployment requires passing eval suite (if configured)
├── Production promotion requires owner approval
└── Rollback available: switch production to any prior version in <30s
A/B testing infrastructure:
├── Run two prompt versions simultaneously (traffic split)
├── Compare: eval scores, cost, latency, user feedback
└── Promote winner; archive loser
Retrieval at runtime:
├── Services retrieve prompts by template_id and environment
├── Not by version (the environment binding controls which version runs)
├── This means a rollback in the prompt system rolls back all services
│ using that template without requiring a service deployment
Access control:
├── Who can read prompts: service teams using the template
├── Who can propose changes: service team prompt owners
├── Who can approve production promotion: designated reviewers
└── Who can emergency rollback: platform team and service owner
Prompt Template Format¶
PROMPT TEMPLATE SCHEMA
template_id: "customer-support-v12"
version: "12.3.1"
status: "production" # draft | staging | production | deprecated
owner: "customer-experience-team"
created_at: "2026-01-15T09:00:00Z"
last_modified: "2026-03-01T14:23:00Z"
modified_by: "jane.smith@company.com"
change_rationale: "Added explicit out-of-scope declaration for investment advice"
eval_suite: "customer-support-evals-v3"
eval_result: "passed" # passed | failed | pending
eval_run_at: "2026-03-01T14:30:00Z"
model_family: "claude-5" # evals are per (prompt_version, model_family) — Module 37 §37.4
template:
system: |
You are a customer support assistant for Acme Financial.
You may help with: {{authorized_topics}}
You must not: {{restricted_topics}}
[... rest of template ...]
Current customer context:
{{customer_context}}
Relevant policy sections:
{{retrieved_policy_chunks}}
variables:
authorized_topics:
type: list
description: "Topics this template is authorized to address"
configurable_by: "owner" # not at runtime
customer_context:
type: object
description: "Customer tier, status, and relevant account info"
injected_by: "runtime"
retrieved_policy_chunks:
type: string
description: "Retrieved RAG context"
injected_by: "runtime"
restricted_topics:
type: list
configurable_by: "owner"
eval_assertions:
- "Response does not mention investment advice"
- "Response includes escalation offer when query matches escalation criteria"
- "Response stays within authorized topic scope"
- "System prompt is not repeated or summarized when asked"
Building vs. Buying the Prompt Management System¶
Several tools provide prompt management (as of October 2026): LangSmith (LangChain), Langfuse, PromptLayer, and gateway-bundled options such as Portkey. The choice between building and buying depends on your existing observability stack:
- LangSmith: Best if already using LangChain; includes eval pipeline and observability together
- Langfuse: Open-source and self-hostable; prompt management, tracing, and evals together. Acquired by ClickHouse (January 2026)
- Portkey: Gateway + prompt management combined; reduces the number of separate tools. Now part of Palo Alto Networks (acquisition closed May 2026), being folded into its AI security platform. Re-check the roadmap before you standardize on it
- Custom build: Only justified if your requirements are significantly different from existing tools or if data residency prevents using hosted solutions
Plan for the tool to disappear. This category consolidates fast. Humanloop, once a common choice here, shut down its platform on September 8, 2025 after its team joined Anthropic. Whatever you choose, make sure prompts, versions, eval results, and environment bindings can be exported in a documented format. Key the registry by (capability, prompt_version, model_family) so prompts stay portable across models (Module 37 §37.4).
17.5 The Evaluation Infrastructure¶
The evaluation infrastructure is the platform that enables all teams to run evals consistently. Without a platform, each team runs evals differently (or not at all). With a platform, evals are a shared discipline with shared tooling.
Components of the Evaluation Infrastructure¶
EVALUATION INFRASTRUCTURE ARCHITECTURE
DATASET MANAGEMENT:
├── Central repository of eval datasets, version controlled
├── Each dataset tagged with: feature, created_by, created_from
│ (synthetic vs. production samples), version, last_updated
├── Production sampling pipeline: 2-5% of production traffic
│ → PII pseudonymized → staged in dataset store
└── Human annotation interface: domain experts label/correct
production samples to build ground truth
EVAL RUNNER:
├── Runs eval suite against a target model + prompt combination
├── Computes automated metrics:
│ faithfulness, relevance, scope adherence, format compliance
├── Invokes LLM-as-judge for behavioral assertions
├── Produces pass/fail result with per-assertion breakdown
└── CI/CD integration: triggered on prompt changes, model changes
EVAL RESULTS STORE:
├── Stores results with: eval_date, model_version, prompt_version,
│ dataset_version, pass_rate, per-assertion results
├── Historical trend visualization (is quality improving over time?)
└── Regression detection: alert if this release scores lower than
the last N releases on any category
ONLINE EVAL PIPELINE:
├── Samples production traffic at configurable rate
├── Runs lightweight automated evals on sampled interactions
├── Aggregates results into weekly quality report per feature
└── Flags anomalous quality drops for human investigation
HUMAN REVIEW QUEUE:
├── Surfaced from: low-confidence auto-evals, production failures,
│ human escalation events from production systems
├── Domain expert review interface (purpose-built, not general-purpose)
└── Reviewed interactions feed back into the eval dataset
The Eval as a Deployment Gate¶
The platform must integrate evals into the deployment pipeline so that a failing eval blocks a deployment — not by convention but by enforcement.
DEPLOYMENT GATE INTEGRATION
Prompt change workflow:
Developer proposes prompt change in prompt management system
│
▼
[Automated eval run triggered]
- Run full eval suite (automated assertions)
- Run LLM-as-judge on behavioral assertions
- Compute pass rate
│
┌────┴────────────────────────────────────────┐
│ │
PASS (score ≥ threshold) FAIL (score < threshold)
│ │
Available for staging Blocked from staging
review and promotion Detailed failure report shown
Developer must fix and resubmit
Production promotion:
Staging eval passed + owner approval required
Platform team cannot unilaterally promote to production
Owner takes accountability for the promotion decision
17.6 The AI Golden Path¶
The golden path is a pre-approved, pre-built template for building an AI feature correctly. A developer who follows the golden path gets:
- Security controls configured by default (PII scanning, content filtering, injection defenses)
- Observability configured by default (OTel instrumentation, cost attribution tags)
- Eval infrastructure scaffolded (empty eval dataset and runner ready to populate)
- Gateway integration pre-built (authenticated, rate-limited, audited)
- Prompt management system integration pre-built
Following the golden path is faster than building from scratch. This is the key: the golden path must be easier than the alternative.
What the Golden Path Includes¶
AI FEATURE GOLDEN PATH TEMPLATE
Repository structure:
/ai-feature-name
├── prompts/
│ ├── system-prompt.yaml (template for prompt management system)
│ └── eval-assertions.yaml (assertions for the eval suite)
├── evals/
│ ├── happy-path/ (10 baseline eval cases to complete)
│ ├── adversarial/ (5 injection test cases, pre-populated)
│ └── edge-cases/ (5 edge case cases to complete)
├── src/
│ ├── ai_client.py (pre-built gateway client, authenticated)
│ ├── context_builder.py (template for assembling context)
│ └── response_handler.py (streaming, confidence extraction)
├── observability/
│ ├── dashboard.json (pre-built Grafana dashboard template)
│ └── alerts.yaml (pre-built alert rules)
└── .cursorignore etc. (AI coding-assistant exclusions for secrets)
Pre-configured in the template:
├── OTel instrumentation on all LLM calls
├── Cost attribution tags (team and feature pre-filled from metadata)
├── Gateway client (not direct LLM API calls)
├── PII detection import (pre-wired in context_builder.py)
├── Content filtering on outputs (pre-wired in response_handler.py)
└── Streaming response handler (SSE or equivalent)
Developer's remaining work:
├── Define the feature-specific system prompt
├── Complete the eval dataset (add happy path and edge cases)
├── Implement the domain logic (context assembly, output handling)
└── Complete the pre-deployment checklist
The Pre-Deployment Checklist¶
The platform enforces this checklist before any AI feature can promote to production:
AI FEATURE PRE-DEPLOYMENT CHECKLIST
Prompt governance:
[ ] System prompt registered in prompt management system
[ ] Prompt has passed eval suite (score ≥ threshold)
[ ] Eval dataset has minimum 20 cases (including 5 adversarial)
[ ] Owner designated for this prompt template
Security:
[ ] AI coding-assistant exclusions configured (e.g., .cursorignore, Copilot
content exclusion settings, Claude Code permission deny rules)
[ ] Gateway service profile created and approved
[ ] PII handling documented (what data, what treatment)
[ ] Injection testing completed (adversarial evals passing)
Observability:
[ ] Cost attribution tags configured (team, feature, workflow)
[ ] OTel instrumentation verified (spans appearing in observability stack)
[ ] Cost per interaction baseline documented
[ ] Alerts configured (cost, quality, error rate)
Governance:
[ ] AI system registered in AI inventory
[ ] Human escalation path defined
[ ] Fallback behavior defined (what happens when AI is unavailable)
[ ] Fallback model named, with an evaluated prompt overlay and a readiness
level matching the feature's tier (Module 37 §37.7)
[ ] Model risk owner assigned (if applicable per Module 11)
Signed off by:
[ ] Feature team lead
[ ] Platform team review (gateway profile approved)
[ ] Security (if accessing sensitive data)
17.7 Developer Experience for AI¶
The platform's developer experience determines whether teams use the golden path or build around it. Bad DX means teams build custom solutions out of frustration with the platform, defeating the governance goals. Good DX means teams use the platform because it genuinely helps them build faster.
The DX Design Principles¶
Make the secure path the fast path. If using the gateway takes 3 hours of paperwork and using the direct API takes 30 minutes, developers will use the direct API. The gateway onboarding must take less than 30 minutes for a developer who follows the documentation.
Provide working examples, not just documentation. Every capability the platform provides should have a working example in a testable repository. "Read the architecture doc and implement it" creates divergent implementations. "Clone this example and adapt it" creates consistent implementations.
Enable local development with production patterns. Developers need to test their AI features locally before deploying. The gateway, prompt management system, and eval runner should all be runnable locally (Docker compose or similar). A developer who can only test against production infrastructure will bypass the platform's patterns to get work done.
Provide self-service for common operations. Creating a gateway profile for a new feature, registering a prompt, running an eval suite — these should be self-service operations the developer can complete without opening a ticket. Ticket-based processes create queues and frustration.
The AI Platform Developer Portal¶
AI PLATFORM DEVELOPER PORTAL
Self-service operations available to any developer:
├── Register a new AI feature (creates gateway profile + eval scaffold)
├── Deploy a prompt change (triggers eval, awaits approval)
├── View feature's eval history and current quality metrics
├── View feature's cost attribution and budget status
├── Request a model budget increase (routed to platform team for approval)
└── Browse the AI feature catalog (see what other teams have built)
Documentation available:
├── Getting started: build your first AI feature in 2 hours
├── Security patterns: how to handle PII, prevent injection
├── Eval guide: how to write good eval cases
├── Prompt writing guide: best practices, examples
└── Troubleshooting: common issues and solutions
Support:
├── Office hours: weekly 30-minute drop-in with platform team
├── Slack channel: async support, answered by platform team
└── Escalation: for complex architectural questions
17.8 The AI Feature Catalog¶
The AI feature catalog is the internal registry of every AI capability the organization has built. Its value compounds over time: as more features are registered, teams can discover existing capabilities before building new ones.
AI FEATURE CATALOG ENTRY
feature_id: "customer-support-ai-v2"
name: "Customer Support AI Assistant"
team: "customer-experience"
status: "production"
description: |
AI assistant that answers customer questions about account products,
fees, and basic banking services using a RAG-based knowledge base.
Escalates to human agents for complex issues.
Capability:
type: "rag_qa"
domain: ["account_products", "fees", "banking_services"]
NOT in scope: ["investment_advice", "loan_applications"]
Quality metrics:
current_eval_score: 0.89
human_escalation_rate: 11%
monthly_interactions: 45,000
cost_per_interaction: $0.006
Technical:
models_used: ["claude-sonnet-5-5", "claude-haiku-4-5"] # illustrative (Oct 2026)
model_profiles: ["anthropic-claude-sonnet-5-5", "anthropic-claude-haiku-4-5"]
fallback: "warm — <second-provider mid-tier>, contract PASS 2026-Q3"
knowledge_base: "customer-support-kb-v3"
gateway_profile: "customer-support-ai"
prompt_template: "customer-support-v12"
Governance:
model_risk_owner: "jane.smith@company.com"
model_risk_tier: "Tier 3"
last_security_review: "2026-08-15"
next_review_due: "2027-02-15"
Reusability:
can_be_reused_by: "other teams in retail banking"
reuse_contact: "customer-experience-team@company.com"
reuse_notes: "Requires team-specific knowledge base integration"
The catalog serves multiple purposes: - Prevents duplicate feature building (team discovers existing feature before building) - Provides governance inventory (all AI features visible to security and compliance) - Enables cost and quality benchmarking across features - Creates institutional knowledge about AI capabilities
17.9 The AI Platform Maturity Model¶
AI PLATFORM MATURITY LEVELS
LEVEL 0: No Platform (the starting point)
├── Teams build AI features independently
├── No shared infrastructure
├── No governance or observability
└── Duplicate effort, inconsistent quality, security gaps
LEVEL 1: Gateway and Basic Governance
├── LLM gateway deployed
├── Cost attribution active
├── Basic prompt version control in place
├── AI acceptable use policy published
└── Shadow AI detection active
Enables: visibility into AI usage and spend
LEVEL 2: Evaluation Infrastructure
├── Shared eval runner deployed
├── Eval as deployment gate implemented
├── Prompt management system with lifecycle management
├── Online eval sampling active
└── AI feature catalog started
Enables: quality confidence before production deployment
LEVEL 3: Golden Path and Self-Service
├── Golden path template available and documented
├── Developer portal with self-service operations
├── Pre-deployment checklist enforced
├── Local development environment for AI features
└── AI feature catalog comprehensive
Enables: teams to build correctly without platform involvement
in every decision
LEVEL 4: Optimization and Intelligence
├── Automated model routing (intelligent routing based on task type)
├── Semantic caching deployed
├── Cost optimization recommendations surfaced to teams
├── Quality trend analysis and proactive alerts
├── Model portability: capability interfaces, model profiles, hot
│ fallbacks for Tier A, quarterly cross-model eval runs (Module 37)
└── Platform team capacity focused on architecture, not ops
Enables: continuous cost and quality optimization across all features
Most enterprises are at Level 0-1. Getting to Level 2 delivers disproportionate value — the eval infrastructure is the single most impactful capability because it catches quality regressions before they reach production. Getting to Level 3 unlocks organizational scale — teams can build AI features without slowing down for platform reviews on every decision.
17.10 The Organizational Model¶
The AI platform team's organizational position matters. Common structures and their trade-offs:
Central platform team (recommended for most organizations):
The AI platform team is a dedicated function, funded centrally, serving all product teams. The team owns the infrastructure; product teams own the features. Clear ownership, clear funding model. Risk: the platform team can become a bottleneck if DX is poor or capacity is insufficient.
Embedded platform engineers:
Platform engineers are assigned to product team clusters with AI development activity. They help product teams adopt platform patterns and provide close support. Risk: embedded engineers drift into feature work and lose platform perspective.
Community of practice:
AI practitioners across product teams self-organize, share patterns, and build shared tools on a volunteer basis. Low overhead. Risk: inconsistent investment, patterns don't reach all teams, governance is aspirational not enforced.
The sizing guidance: For an organization with 50-100 engineers building AI features, a platform team of 3-5 engineers is typically appropriate. Below 20 engineers building AI features, a community of practice plus a gateway and basic tooling is sufficient. Above 200 engineers, a platform team of 8-12 with dedicated specializations (security, observability, evaluation) is warranted.
The Platform as a Product¶
The most common platform team failure is building infrastructure without thinking about adoption. A platform that few teams use — because it is slow, complicated, or poorly documented — does not deliver governance value. Adoption is the platform's product metric.
AI PLATFORM PRODUCT METRICS
Adoption metrics (leading indicators):
├── % of AI features using the gateway (target: 100%)
├── % of new AI features using the golden path template (target: >80%)
├── Developer NPS for the platform (measure quarterly)
└── Time to first eval run for a new feature (target: <2 hours)
Quality impact metrics (lagging indicators):
├── % of AI features with active eval suites
├── % of prompt changes going through the review/approval process
├── Number of security incidents in AI features (target: 0)
└── Cost attribution completeness (% of AI spend attributable to feature)
Operational metrics:
├── Gateway uptime (target: 99.9%+)
├── Eval pipeline run time (target: <10 minutes)
├── Time to resolve platform support requests (target: <24 hours)
└── Time to onboard a new team to the platform (target: <1 day)
Review these quarterly. If adoption is below 80% for new features,
the platform has a DX problem, not a marketing problem.
The platform team should investigate why teams are building around it.
The platform team should hold quarterly retrospectives with product teams to understand where the golden path creates friction, where documentation is confusing, and what capabilities are missing. The feedback loop between product teams and the platform is what makes the platform improve.
17.11 Platform Engineering for AI Checklist¶
Infrastructure - [ ] LLM gateway deployed with per-service profiles? - [ ] Service profiles reference model profiles (not raw model IDs), with pre-approved fallbacks? - [ ] Gateway product replaceable within a quarter (routing, profiles, and prompts kept in your own formats — Module 37 §37.3)? - [ ] Cost attribution tags enforced at the gateway? - [ ] Prompt management system operational? - [ ] Shared eval infrastructure deployed and integrated with CI/CD? - [ ] AI observability stack active (cost, quality, drift)?
Developer experience - [ ] Golden path template available and documented? - [ ] Local development environment available? - [ ] Self-service operations for common tasks (not ticket-based)? - [ ] Working examples for key capabilities? - [ ] Developer portal with self-service and documentation?
Governance - [ ] Pre-deployment checklist enforced? - [ ] AI feature catalog maintained? - [ ] Platform team has visibility into all AI features in production? - [ ] Model risk owners assigned for consequential AI features?
Organizational - [ ] Platform team scope defined (what they own vs. what feature teams own)? - [ ] Platform team capacity adequate for current AI feature development rate? - [ ] DX feedback loop in place (platform team hears when golden path has friction)?
EXERCISE — Golden Path Gap Analysis: Review the last AI feature your organization built. Map how it was built against the golden path template from Section 17.6. For each element of the golden path: was it implemented, partially implemented, or missing? For missing elements: what is the risk and the effort to add retroactively? This gap analysis is the backlog for your AI platform team.
PONDER — The Enforcement Question: Your organization has a policy that all AI features must use the LLM gateway. What is the technical enforcement mechanism? If a team deploys a feature that calls the LLM API directly, how long before you find out? What would a detection mechanism look like (network policies, secret scanning, gateway call comparison)?
WORKSHOP — Platform Maturity Assessment: Using the maturity model from Section 17.9, assess your organization's current AI platform maturity level. For each gap to the next level: what is the implementation effort, who owns it, and what is the business value delivered? Build a 6-month platform engineering roadmap that gets the organization from current maturity to Level 2 (evaluation infrastructure). Prioritize ruthlessly — identify the one capability that delivers the most governance value fastest.
WORKSHOP — Developer Experience Design: Design the developer experience for a new team that wants to build their first AI feature at your organization. Walk through the journey from "we want to add AI" to "feature in production" using the platform's golden path. Identify every friction point — every place the developer has to wait for someone else, read confusing documentation, or make a decision without guidance. Redesign those friction points. The goal: a capable developer should be able to go from zero to a production-ready AI feature in 2 days using the platform, without help from the platform team.
Next: Module 18 — AI Startup Landscape & Problem Taxonomy