Skip to content

MODULE 17 — Platform Engineering for AI

⚠️ Currency note: Model IDs and vendor names in this module are illustrative examples as of October 2026. Current models, platform names, and the gateway vendor landscape are in Appendix G (G.2, G.5, G.8). The platform patterns do not depend on which vendors are named.

17.1 The Platform Engineering Mandate

Platform engineering for AI answers a specific organizational failure mode: every team building AI from scratch, making the same mistakes, duplicating the same infrastructure, and introducing the same security and governance gaps.

Without an AI platform function, the typical enterprise ends up with: - 12 different teams calling LLM APIs with 12 different authentication mechanisms and 12 different cost attribution approaches - 7 different prompt management strategies ranging from hardcoded strings to careful version control - No shared evaluation infrastructure — each team builds (or more commonly, skips) its own - Security controls that vary by team and change when engineers leave - No visibility into total AI spend or which products are responsible for it - No consistent approach to human oversight, audit trails, or model risk

The platform engineering team's job is to make good AI architecture the path of least resistance. Not to mandate it, not to police it — to make it easier to do it right than to do it wrong. This distinction matters: a mandate with no tooling creates bureaucracy. A golden path with great tooling creates adoption.


17.2 The AI Platform Team Scope

The AI platform team owns the infrastructure and tooling that other teams use to build AI features. It does not own the AI features themselves. The distinction is the same as the relationship between the platform engineering team and product feature teams in any mature engineering organization.

What the AI platform team owns:

AI PLATFORM TEAM SCOPE

Infrastructure:
  ├── LLM gateway (centralized access, authentication, routing, audit)
  ├── Evaluation infrastructure (eval pipelines, dataset management,
  │     LLM-as-judge configuration)
  ├── Observability stack (AI-specific dashboards, cost attribution)
  ├── Prompt management system (versioning, deployment, A/B testing)
  └── Self-hosted model infrastructure (vLLM servers, model registry)
      if the organization has self-hosting requirements

Standards and golden paths:
  ├── The approved AI service template (how to build a new AI feature correctly)
  ├── Security patterns (injection defenses, PII handling, content filtering)
  ├── Evaluation standards (what evals are required before production)
  └── Governance requirements (what must be documented, who must approve)

Enablement:
  ├── Internal developer documentation
  ├── Example implementations
  ├── Office hours / consulting for teams building AI features
  └── Onboarding support for first-time AI feature teams

What the AI platform team does NOT own:
  ├── Individual product AI features (owned by product teams)
  ├── Domain-specific knowledge bases (owned by domain teams)
  ├── Feature-level eval cases (owned by the feature team)
  └── Business decisions about AI use (owned by product and leadership)

17.3 The LLM Gateway: The Platform's Core Infrastructure

The LLM gateway is the most important piece of AI platform infrastructure. Everything else depends on it. It is the single point through which all LLM API calls are made, and it provides centralized enforcement of security, cost, and governance policies.

Covered in depth in the Integration Patterns reference (Artifact 2, Pattern 1). From the platform perspective, what matters is:

The Gateway as the Policy Enforcement Point

The gateway is the only place where policies can be enforced consistently across all AI features. This is its most important characteristic.

GATEWAY AS POLICY ENFORCEMENT

Policies enforced at the gateway (not in application code):
  ├── Authentication: every API call must present a valid service identity
  ├── Authorization: this service is allowed to use these models
  ├── Rate limiting: this service's token budget is X per period
  ├── PII scanning: no SSN, credit card, or sensitive identifiers
  │     leave the perimeter without pseudonymization
  ├── Model version: this service's pinned model version is enforced
  │     (no accidental use of unpinned aliases)
  ├── Content filtering: input and output pass through guardrails
  └── Audit logging: every call logged with service identity, model,
        token count, cost, prompt template ID

Why application code cannot be the enforcement point:
  Application code changes. Developers forget. New services are
  created without inheriting the patterns. The gateway enforces
  the same policies on day one and day 1,000, for all services.

Gateway Configuration Per Service

Each service that uses the gateway must be registered with its own configuration profile:

GATEWAY SERVICE PROFILE

service_id: "customer-support-ai"
team: "customer-experience"
# Model IDs below are illustrative (October 2026). In a mature platform the
# service profile references a MODEL PROFILE (Module 37 §37.3) rather than a raw
# model ID, so a model swap edits one profile, not every service profile.
approved_models:
  - model: "claude-sonnet-5-5"       # or: model_profile: "anthropic-claude-sonnet-5-5"
    max_input_tokens: 4000           # design ceiling, not the model's maximum
    max_output_tokens: 500
  - model: "claude-haiku-4-5"        # for high-volume simple classification
    max_input_tokens: 1000
    max_output_tokens: 100
fallback_models:                     # pre-approved, so failover needs no ticket
  - model: "<second-provider mid-tier>"   # must have an evaluated prompt overlay

token_budgets:
  monthly_ceiling_usd: 2000
  per_request_input_ceiling: 4000
  per_request_output_ceiling: 500
  daily_alert_threshold_usd: 150  # alert when daily spend hits this

pii_policy:
  scan_input: true
  scan_output: true
  on_pii_detected: "pseudonymize"  # pseudonymize | block | allow_and_log

content_policy:
  input_filter: "standard_enterprise"
  output_filter: "standard_enterprise"
  additional_blocked_categories: []  # can add service-specific

audit:
  log_prompt_template_id: true
  log_full_content: false  # full content logged only in debug mode
  retention_days: 365

data_handling:
  data_residency: "us"
  approved_external_models: true  # can send to external providers

This profile is the architectural contract between the service team and the AI platform team. It is version-controlled and reviewed when the service team requests changes.

The Gateway Is Transport and Policy, Not the Portability Layer

A gateway normalizes the wire call: auth, keys, quotas, PII scanning, audit, retries, and transport-level failover. It does not make two models behave the same. Parameter semantics, forced tool selection, structured-output mechanisms, reasoning-depth controls and defaults, tokenizer differences, and refusal behavior all differ between models, and between versions from the same vendor. Those are handled above the gateway, in a capability and model-adapter layer that you own, with a behavioral contract that defines "equivalent." Treat the gateway product itself as replaceable too. The gateway market consolidated sharply in 2026: Portkey is now part of Palo Alto Networks, LiteLLM remains the de facto open-source option, and OpenRouter is a hosted aggregator (Appendix G §G.5, §G.8). Keep routing rules, model profiles, and prompts in your own repository and formats, not in one gateway's proprietary configuration. See Module 37 — Model Portability (§37.3, "Gateway vs. Adapter: Who Does What").


17.4 The Prompt Management System

Prompts are the most impactful and least managed artifact in most AI systems. The prompt management system provides the infrastructure for treating prompts as first-class software artifacts.

What the Prompt Management System Must Do

PROMPT MANAGEMENT SYSTEM CAPABILITIES

Storage and versioning:
  ├── Store prompt templates in a structured format
  ├── Version every change (semantic versioning: major.minor.patch)
  ├── Associate every version with author, timestamp, and change rationale
  └── Never delete versions (deprecate and archive instead)

Deployment lifecycle:
  ├── Environments: dev → staging → production
  ├── Deployment requires passing eval suite (if configured)
  ├── Production promotion requires owner approval
  └── Rollback available: switch production to any prior version in <30s

A/B testing infrastructure:
  ├── Run two prompt versions simultaneously (traffic split)
  ├── Compare: eval scores, cost, latency, user feedback
  └── Promote winner; archive loser

Retrieval at runtime:
  ├── Services retrieve prompts by template_id and environment
  ├── Not by version (the environment binding controls which version runs)
  ├── This means a rollback in the prompt system rolls back all services
  │     using that template without requiring a service deployment

Access control:
  ├── Who can read prompts: service teams using the template
  ├── Who can propose changes: service team prompt owners
  ├── Who can approve production promotion: designated reviewers
  └── Who can emergency rollback: platform team and service owner

Prompt Template Format

PROMPT TEMPLATE SCHEMA

template_id: "customer-support-v12"
version: "12.3.1"
status: "production"  # draft | staging | production | deprecated
owner: "customer-experience-team"
created_at: "2026-01-15T09:00:00Z"
last_modified: "2026-03-01T14:23:00Z"
modified_by: "jane.smith@company.com"
change_rationale: "Added explicit out-of-scope declaration for investment advice"
eval_suite: "customer-support-evals-v3"
eval_result: "passed"  # passed | failed | pending
eval_run_at: "2026-03-01T14:30:00Z"
model_family: "claude-5"  # evals are per (prompt_version, model_family) — Module 37 §37.4

template:
  system: |
    You are a customer support assistant for Acme Financial.

    You may help with: {{authorized_topics}}

    You must not: {{restricted_topics}}

    [... rest of template ...]

    Current customer context:
    {{customer_context}}

    Relevant policy sections:
    {{retrieved_policy_chunks}}

variables:
  authorized_topics:
    type: list
    description: "Topics this template is authorized to address"
    configurable_by: "owner"  # not at runtime
  customer_context:
    type: object
    description: "Customer tier, status, and relevant account info"
    injected_by: "runtime"
  retrieved_policy_chunks:
    type: string
    description: "Retrieved RAG context"
    injected_by: "runtime"
  restricted_topics:
    type: list
    configurable_by: "owner"

eval_assertions:
  - "Response does not mention investment advice"
  - "Response includes escalation offer when query matches escalation criteria"
  - "Response stays within authorized topic scope"
  - "System prompt is not repeated or summarized when asked"

Building vs. Buying the Prompt Management System

Several tools provide prompt management (as of October 2026): LangSmith (LangChain), Langfuse, PromptLayer, and gateway-bundled options such as Portkey. The choice between building and buying depends on your existing observability stack:

  • LangSmith: Best if already using LangChain; includes eval pipeline and observability together
  • Langfuse: Open-source and self-hostable; prompt management, tracing, and evals together. Acquired by ClickHouse (January 2026)
  • Portkey: Gateway + prompt management combined; reduces the number of separate tools. Now part of Palo Alto Networks (acquisition closed May 2026), being folded into its AI security platform. Re-check the roadmap before you standardize on it
  • Custom build: Only justified if your requirements are significantly different from existing tools or if data residency prevents using hosted solutions

Plan for the tool to disappear. This category consolidates fast. Humanloop, once a common choice here, shut down its platform on September 8, 2025 after its team joined Anthropic. Whatever you choose, make sure prompts, versions, eval results, and environment bindings can be exported in a documented format. Key the registry by (capability, prompt_version, model_family) so prompts stay portable across models (Module 37 §37.4).


17.5 The Evaluation Infrastructure

The evaluation infrastructure is the platform that enables all teams to run evals consistently. Without a platform, each team runs evals differently (or not at all). With a platform, evals are a shared discipline with shared tooling.

Components of the Evaluation Infrastructure

EVALUATION INFRASTRUCTURE ARCHITECTURE

DATASET MANAGEMENT:
  ├── Central repository of eval datasets, version controlled
  ├── Each dataset tagged with: feature, created_by, created_from
  │     (synthetic vs. production samples), version, last_updated
  ├── Production sampling pipeline: 2-5% of production traffic
  │     → PII pseudonymized → staged in dataset store
  └── Human annotation interface: domain experts label/correct
        production samples to build ground truth

EVAL RUNNER:
  ├── Runs eval suite against a target model + prompt combination
  ├── Computes automated metrics:
  │     faithfulness, relevance, scope adherence, format compliance
  ├── Invokes LLM-as-judge for behavioral assertions
  ├── Produces pass/fail result with per-assertion breakdown
  └── CI/CD integration: triggered on prompt changes, model changes

EVAL RESULTS STORE:
  ├── Stores results with: eval_date, model_version, prompt_version,
  │     dataset_version, pass_rate, per-assertion results
  ├── Historical trend visualization (is quality improving over time?)
  └── Regression detection: alert if this release scores lower than
        the last N releases on any category

ONLINE EVAL PIPELINE:
  ├── Samples production traffic at configurable rate
  ├── Runs lightweight automated evals on sampled interactions
  ├── Aggregates results into weekly quality report per feature
  └── Flags anomalous quality drops for human investigation

HUMAN REVIEW QUEUE:
  ├── Surfaced from: low-confidence auto-evals, production failures,
  │     human escalation events from production systems
  ├── Domain expert review interface (purpose-built, not general-purpose)
  └── Reviewed interactions feed back into the eval dataset

The Eval as a Deployment Gate

The platform must integrate evals into the deployment pipeline so that a failing eval blocks a deployment — not by convention but by enforcement.

DEPLOYMENT GATE INTEGRATION

Prompt change workflow:
  Developer proposes prompt change in prompt management system
         │
         ▼
  [Automated eval run triggered]
  - Run full eval suite (automated assertions)
  - Run LLM-as-judge on behavioral assertions
  - Compute pass rate
         │
    ┌────┴────────────────────────────────────────┐
    │                                             │
  PASS (score ≥ threshold)               FAIL (score < threshold)
    │                                             │
  Available for staging                  Blocked from staging
  review and promotion                   Detailed failure report shown
                                         Developer must fix and resubmit

Production promotion:
  Staging eval passed + owner approval required
  Platform team cannot unilaterally promote to production
  Owner takes accountability for the promotion decision

17.6 The AI Golden Path

The golden path is a pre-approved, pre-built template for building an AI feature correctly. A developer who follows the golden path gets:

  • Security controls configured by default (PII scanning, content filtering, injection defenses)
  • Observability configured by default (OTel instrumentation, cost attribution tags)
  • Eval infrastructure scaffolded (empty eval dataset and runner ready to populate)
  • Gateway integration pre-built (authenticated, rate-limited, audited)
  • Prompt management system integration pre-built

Following the golden path is faster than building from scratch. This is the key: the golden path must be easier than the alternative.

What the Golden Path Includes

AI FEATURE GOLDEN PATH TEMPLATE

Repository structure:
  /ai-feature-name
    ├── prompts/
    │     ├── system-prompt.yaml     (template for prompt management system)
    │     └── eval-assertions.yaml  (assertions for the eval suite)
    ├── evals/
    │     ├── happy-path/           (10 baseline eval cases to complete)
    │     ├── adversarial/          (5 injection test cases, pre-populated)
    │     └── edge-cases/           (5 edge case cases to complete)
    ├── src/
    │     ├── ai_client.py          (pre-built gateway client, authenticated)
    │     ├── context_builder.py    (template for assembling context)
    │     └── response_handler.py   (streaming, confidence extraction)
    ├── observability/
    │     ├── dashboard.json        (pre-built Grafana dashboard template)
    │     └── alerts.yaml           (pre-built alert rules)
    └── .cursorignore etc.          (AI coding-assistant exclusions for secrets)

Pre-configured in the template:
  ├── OTel instrumentation on all LLM calls
  ├── Cost attribution tags (team and feature pre-filled from metadata)
  ├── Gateway client (not direct LLM API calls)
  ├── PII detection import (pre-wired in context_builder.py)
  ├── Content filtering on outputs (pre-wired in response_handler.py)
  └── Streaming response handler (SSE or equivalent)

Developer's remaining work:
  ├── Define the feature-specific system prompt
  ├── Complete the eval dataset (add happy path and edge cases)
  ├── Implement the domain logic (context assembly, output handling)
  └── Complete the pre-deployment checklist

The Pre-Deployment Checklist

The platform enforces this checklist before any AI feature can promote to production:

AI FEATURE PRE-DEPLOYMENT CHECKLIST

Prompt governance:
  [ ] System prompt registered in prompt management system
  [ ] Prompt has passed eval suite (score ≥ threshold)
  [ ] Eval dataset has minimum 20 cases (including 5 adversarial)
  [ ] Owner designated for this prompt template

Security:
  [ ] AI coding-assistant exclusions configured (e.g., .cursorignore, Copilot
        content exclusion settings, Claude Code permission deny rules)
  [ ] Gateway service profile created and approved
  [ ] PII handling documented (what data, what treatment)
  [ ] Injection testing completed (adversarial evals passing)

Observability:
  [ ] Cost attribution tags configured (team, feature, workflow)
  [ ] OTel instrumentation verified (spans appearing in observability stack)
  [ ] Cost per interaction baseline documented
  [ ] Alerts configured (cost, quality, error rate)

Governance:
  [ ] AI system registered in AI inventory
  [ ] Human escalation path defined
  [ ] Fallback behavior defined (what happens when AI is unavailable)
  [ ] Fallback model named, with an evaluated prompt overlay and a readiness
        level matching the feature's tier (Module 37 §37.7)
  [ ] Model risk owner assigned (if applicable per Module 11)

Signed off by:
  [ ] Feature team lead
  [ ] Platform team review (gateway profile approved)
  [ ] Security (if accessing sensitive data)

17.7 Developer Experience for AI

The platform's developer experience determines whether teams use the golden path or build around it. Bad DX means teams build custom solutions out of frustration with the platform, defeating the governance goals. Good DX means teams use the platform because it genuinely helps them build faster.

The DX Design Principles

Make the secure path the fast path. If using the gateway takes 3 hours of paperwork and using the direct API takes 30 minutes, developers will use the direct API. The gateway onboarding must take less than 30 minutes for a developer who follows the documentation.

Provide working examples, not just documentation. Every capability the platform provides should have a working example in a testable repository. "Read the architecture doc and implement it" creates divergent implementations. "Clone this example and adapt it" creates consistent implementations.

Enable local development with production patterns. Developers need to test their AI features locally before deploying. The gateway, prompt management system, and eval runner should all be runnable locally (Docker compose or similar). A developer who can only test against production infrastructure will bypass the platform's patterns to get work done.

Provide self-service for common operations. Creating a gateway profile for a new feature, registering a prompt, running an eval suite — these should be self-service operations the developer can complete without opening a ticket. Ticket-based processes create queues and frustration.

The AI Platform Developer Portal

AI PLATFORM DEVELOPER PORTAL

Self-service operations available to any developer:
  ├── Register a new AI feature (creates gateway profile + eval scaffold)
  ├── Deploy a prompt change (triggers eval, awaits approval)
  ├── View feature's eval history and current quality metrics
  ├── View feature's cost attribution and budget status
  ├── Request a model budget increase (routed to platform team for approval)
  └── Browse the AI feature catalog (see what other teams have built)

Documentation available:
  ├── Getting started: build your first AI feature in 2 hours
  ├── Security patterns: how to handle PII, prevent injection
  ├── Eval guide: how to write good eval cases
  ├── Prompt writing guide: best practices, examples
  └── Troubleshooting: common issues and solutions

Support:
  ├── Office hours: weekly 30-minute drop-in with platform team
  ├── Slack channel: async support, answered by platform team
  └── Escalation: for complex architectural questions

17.8 The AI Feature Catalog

The AI feature catalog is the internal registry of every AI capability the organization has built. Its value compounds over time: as more features are registered, teams can discover existing capabilities before building new ones.

AI FEATURE CATALOG ENTRY

feature_id: "customer-support-ai-v2"
name: "Customer Support AI Assistant"
team: "customer-experience"
status: "production"
description: |
  AI assistant that answers customer questions about account products,
  fees, and basic banking services using a RAG-based knowledge base.
  Escalates to human agents for complex issues.

Capability:
  type: "rag_qa"
  domain: ["account_products", "fees", "banking_services"]
  NOT in scope: ["investment_advice", "loan_applications"]

Quality metrics:
  current_eval_score: 0.89
  human_escalation_rate: 11%
  monthly_interactions: 45,000
  cost_per_interaction: $0.006

Technical:
  models_used: ["claude-sonnet-5-5", "claude-haiku-4-5"]   # illustrative (Oct 2026)
  model_profiles: ["anthropic-claude-sonnet-5-5", "anthropic-claude-haiku-4-5"]
  fallback: "warm — <second-provider mid-tier>, contract PASS 2026-Q3"
  knowledge_base: "customer-support-kb-v3"
  gateway_profile: "customer-support-ai"
  prompt_template: "customer-support-v12"

Governance:
  model_risk_owner: "jane.smith@company.com"
  model_risk_tier: "Tier 3"
  last_security_review: "2026-08-15"
  next_review_due: "2027-02-15"

Reusability:
  can_be_reused_by: "other teams in retail banking"
  reuse_contact: "customer-experience-team@company.com"
  reuse_notes: "Requires team-specific knowledge base integration"

The catalog serves multiple purposes: - Prevents duplicate feature building (team discovers existing feature before building) - Provides governance inventory (all AI features visible to security and compliance) - Enables cost and quality benchmarking across features - Creates institutional knowledge about AI capabilities


17.9 The AI Platform Maturity Model

AI PLATFORM MATURITY LEVELS

LEVEL 0: No Platform (the starting point)
  ├── Teams build AI features independently
  ├── No shared infrastructure
  ├── No governance or observability
  └── Duplicate effort, inconsistent quality, security gaps

LEVEL 1: Gateway and Basic Governance
  ├── LLM gateway deployed
  ├── Cost attribution active
  ├── Basic prompt version control in place
  ├── AI acceptable use policy published
  └── Shadow AI detection active

  Enables: visibility into AI usage and spend

LEVEL 2: Evaluation Infrastructure
  ├── Shared eval runner deployed
  ├── Eval as deployment gate implemented
  ├── Prompt management system with lifecycle management
  ├── Online eval sampling active
  └── AI feature catalog started

  Enables: quality confidence before production deployment

LEVEL 3: Golden Path and Self-Service
  ├── Golden path template available and documented
  ├── Developer portal with self-service operations
  ├── Pre-deployment checklist enforced
  ├── Local development environment for AI features
  └── AI feature catalog comprehensive

  Enables: teams to build correctly without platform involvement
  in every decision

LEVEL 4: Optimization and Intelligence
  ├── Automated model routing (intelligent routing based on task type)
  ├── Semantic caching deployed
  ├── Cost optimization recommendations surfaced to teams
  ├── Quality trend analysis and proactive alerts
  ├── Model portability: capability interfaces, model profiles, hot
  │     fallbacks for Tier A, quarterly cross-model eval runs (Module 37)
  └── Platform team capacity focused on architecture, not ops

  Enables: continuous cost and quality optimization across all features

Most enterprises are at Level 0-1. Getting to Level 2 delivers disproportionate value — the eval infrastructure is the single most impactful capability because it catches quality regressions before they reach production. Getting to Level 3 unlocks organizational scale — teams can build AI features without slowing down for platform reviews on every decision.


17.10 The Organizational Model

The AI platform team's organizational position matters. Common structures and their trade-offs:

Central platform team (recommended for most organizations):

The AI platform team is a dedicated function, funded centrally, serving all product teams. The team owns the infrastructure; product teams own the features. Clear ownership, clear funding model. Risk: the platform team can become a bottleneck if DX is poor or capacity is insufficient.

Embedded platform engineers:

Platform engineers are assigned to product team clusters with AI development activity. They help product teams adopt platform patterns and provide close support. Risk: embedded engineers drift into feature work and lose platform perspective.

Community of practice:

AI practitioners across product teams self-organize, share patterns, and build shared tools on a volunteer basis. Low overhead. Risk: inconsistent investment, patterns don't reach all teams, governance is aspirational not enforced.

The sizing guidance: For an organization with 50-100 engineers building AI features, a platform team of 3-5 engineers is typically appropriate. Below 20 engineers building AI features, a community of practice plus a gateway and basic tooling is sufficient. Above 200 engineers, a platform team of 8-12 with dedicated specializations (security, observability, evaluation) is warranted.

The Platform as a Product

The most common platform team failure is building infrastructure without thinking about adoption. A platform that few teams use — because it is slow, complicated, or poorly documented — does not deliver governance value. Adoption is the platform's product metric.

AI PLATFORM PRODUCT METRICS

Adoption metrics (leading indicators):
  ├── % of AI features using the gateway (target: 100%)
  ├── % of new AI features using the golden path template (target: >80%)
  ├── Developer NPS for the platform (measure quarterly)
  └── Time to first eval run for a new feature (target: <2 hours)

Quality impact metrics (lagging indicators):
  ├── % of AI features with active eval suites
  ├── % of prompt changes going through the review/approval process
  ├── Number of security incidents in AI features (target: 0)
  └── Cost attribution completeness (% of AI spend attributable to feature)

Operational metrics:
  ├── Gateway uptime (target: 99.9%+)
  ├── Eval pipeline run time (target: <10 minutes)
  ├── Time to resolve platform support requests (target: <24 hours)
  └── Time to onboard a new team to the platform (target: <1 day)

Review these quarterly. If adoption is below 80% for new features,
the platform has a DX problem, not a marketing problem.
The platform team should investigate why teams are building around it.

The platform team should hold quarterly retrospectives with product teams to understand where the golden path creates friction, where documentation is confusing, and what capabilities are missing. The feedback loop between product teams and the platform is what makes the platform improve.


17.11 Platform Engineering for AI Checklist

Infrastructure - [ ] LLM gateway deployed with per-service profiles? - [ ] Service profiles reference model profiles (not raw model IDs), with pre-approved fallbacks? - [ ] Gateway product replaceable within a quarter (routing, profiles, and prompts kept in your own formats — Module 37 §37.3)? - [ ] Cost attribution tags enforced at the gateway? - [ ] Prompt management system operational? - [ ] Shared eval infrastructure deployed and integrated with CI/CD? - [ ] AI observability stack active (cost, quality, drift)?

Developer experience - [ ] Golden path template available and documented? - [ ] Local development environment available? - [ ] Self-service operations for common tasks (not ticket-based)? - [ ] Working examples for key capabilities? - [ ] Developer portal with self-service and documentation?

Governance - [ ] Pre-deployment checklist enforced? - [ ] AI feature catalog maintained? - [ ] Platform team has visibility into all AI features in production? - [ ] Model risk owners assigned for consequential AI features?

Organizational - [ ] Platform team scope defined (what they own vs. what feature teams own)? - [ ] Platform team capacity adequate for current AI feature development rate? - [ ] DX feedback loop in place (platform team hears when golden path has friction)?


EXERCISE — Golden Path Gap Analysis: Review the last AI feature your organization built. Map how it was built against the golden path template from Section 17.6. For each element of the golden path: was it implemented, partially implemented, or missing? For missing elements: what is the risk and the effort to add retroactively? This gap analysis is the backlog for your AI platform team.

PONDER — The Enforcement Question: Your organization has a policy that all AI features must use the LLM gateway. What is the technical enforcement mechanism? If a team deploys a feature that calls the LLM API directly, how long before you find out? What would a detection mechanism look like (network policies, secret scanning, gateway call comparison)?

WORKSHOP — Platform Maturity Assessment: Using the maturity model from Section 17.9, assess your organization's current AI platform maturity level. For each gap to the next level: what is the implementation effort, who owns it, and what is the business value delivered? Build a 6-month platform engineering roadmap that gets the organization from current maturity to Level 2 (evaluation infrastructure). Prioritize ruthlessly — identify the one capability that delivers the most governance value fastest.

WORKSHOP — Developer Experience Design: Design the developer experience for a new team that wants to build their first AI feature at your organization. Walk through the journey from "we want to add AI" to "feature in production" using the platform's golden path. Identify every friction point — every place the developer has to wait for someone else, read confusing documentation, or make a decision without guidance. Redesign those friction points. The goal: a capable developer should be able to go from zero to a production-ready AI feature in 2 days using the platform, without help from the platform team.


Next: Module 18 — AI Startup Landscape & Problem Taxonomy