Skip to content

MODULE 24 — AI Production Operations

24.1 Production is Where Architecture Gets Tested

Every architectural decision made in design either holds up in production or it doesn't. Most AI production operations guidance either doesn't exist or is written for ML engineers (model retraining, drift detection) rather than for architects and engineers operating production AI systems.

This module covers three operational areas where most organizations have no documented process: incident response when AI goes wrong in production, safely introducing AI into a running system, and the rollback procedures that work differently for AI than for traditional software.


24.2 AI Incident Response: The Runbook

Why AI Incidents Are Different

A traditional software incident: the service is down or returning errors. Status codes, error logs, and health checks tell you something is wrong.

An AI incident: the service is returning HTTP 200. The responses are confident. They are wrong, harmful, or out-of-scope. The monitoring shows green. Users are experiencing harm before anyone in engineering knows there's a problem.

AI incidents are detected late and diagnosed slowly because the failure mode is content-level, not system-level. Every team needs to have run through this runbook at least once before a real incident — not for the first time during one.

The 30-Minute Incident Playbook

AI INCIDENT RESPONSE: FIRST 30 MINUTES

T+0 — Detection
  Signal sources (in order of reliability):
  1. Automated eval alert: quality metric dropped below threshold
  2. Human escalation rate spike: agents flagging AI outputs as wrong
  3. User complaint pattern: support tickets mentioning AI errors
  4. Internal report: team member noticing anomalous outputs

  Common detection failure: Teams don't have automated quality
  monitoring. They learn about AI incidents from angry customers.
  If this is your detection method, fix the monitoring (Module 12)
  before you need this runbook.

T+0-5 — Triage Questions (answer before doing anything else)
  □ What specifically is wrong? (wrong answers, out-of-scope responses,
    harmful content, complete failures to respond)
  □ How many users are affected? (isolated reports vs. widespread)
  □ Is this getting worse, stable, or intermittent?
  □ When did it start? (correlate with: deployments, model updates,
    document updates, traffic changes)
  □ Is there immediate safety or regulatory risk?
    If YES: escalate to legal/compliance NOW, before further diagnosis

T+5-15 — Initial Diagnosis (the four-question check)
  Q1: Did anything change in the last 24 hours?
    Check: prompt version deployed, model version changed, knowledge
    base updated, traffic volume spike, new user segment
    If YES: the change is the most likely cause. Test by reverting.
    If NO: go to Q2.

  Q2: Is the model returning something different from baseline?
    Check: run your standard eval suite against production right now.
    If evals PASS: the model is behaving normally. Problem is upstream
    (retrieval, data) or downstream (rendering, integration).
    If evals FAIL: model behavior has changed. Go to Q3.

  Q3: Is this a model version change?
    Check the actual model version returned in production API responses.
    Compare to the pinned version in configuration.
    If DIFFERENT: the model was updated without your knowledge.
    Contact the provider. Rollback to pinned version (Section 24.4).
    If SAME: the same model is behaving differently. Possible causes:
    context content changed, prompt changed, provider silent update
    within the version. Escalate to provider with reproduction case.

  Q4: Is this a retrieval/data problem?
    Retrieve the actual chunks being returned for affected queries.
    Is there outdated, incorrect, or injected content in the knowledge base?
    If YES: identify affected document/chunks. Immediate mitigation:
    archive affected content (Section 24.4). Investigate root cause.

T+15-30 — Mitigation Decision
  Option A: Rollback the change that caused the incident.
    (See Section 24.4 for rollback procedures)

  Option B: Route affected traffic to human agents.
    Raise the confidence gate threshold to near-1.0.
    This escalates nearly all queries to humans.
    Service continues; AI is effectively bypassed.

  Option C: Disable the AI feature.
    Return to the baseline (non-AI) experience.
    Only for AI-augmented systems with a working baseline.
    For AI-native systems, coordinate with leadership first.

  Option D: Accept and monitor.
    If the impact is low and the root cause is understood,
    accepting while you implement a proper fix is sometimes right.
    Must be an explicit decision, not a default.

T+30+ — Communication and Resolution
  Internal stakeholder communication within 1 hour of declaration.
  User communication if there is potential for harm from incorrect outputs.
  Root cause investigation (not during the incident — after mitigation).
  Post-mortem within 5 business days (Module 22 template).

The AI Incident Severity Matrix

SEVERITY CLASSIFICATION FOR AI INCIDENTS

SEV 1 (Immediate response, leadership notified within 30 minutes):
  - AI producing outputs that could cause direct user harm
    (medical misinformation, unsafe instructions, discriminatory content)
  - Regulatory exposure: GDPR data breach via AI output,
    ECOA violation via AI decision, SR 11-7 control failure
  - AI agent taking irreversible consequential actions (financial
    transactions, account modifications) without intended authorization
  - Complete AI service failure affecting > 1,000 users

SEV 2 (Response within 2 hours, manager notified):
  - AI producing systematically wrong answers (>20% error rate)
  - Significant quality degradation affecting user experience
  - Security incident: prompt injection succeeded, data exfiltrated
  - AI ignoring scope boundaries at scale
  - Cost spike > 5x normal without explanation

SEV 3 (Response within 24 hours, tracked):
  - Isolated quality failures without clear pattern
  - Intermittent failures with no user-visible impact
  - Eval score regression detected in monitoring (not yet user-visible)
  - Minor scope boundary violations (low frequency)

SEV 4 (Logged, addressed in next sprint):
  - Quality issues only visible in eval data
  - Performance degradation within acceptable SLA
  - Minor prompt inefficiencies detected

24.3 Shadow Mode and Canary Deployment for AI

Why AI Deployments Need Special Care

Deploying a change to a deterministic service is relatively safe: the new version either works or doesn't, and integration tests catch most failures before production. AI introduces uncertainty: even a well-tested prompt change can behave differently on production traffic distributions than on your eval dataset. Shadow mode and canary deployment are the risk management tools that give you production signal before full commitment.

Shadow Mode Deployment

Shadow mode runs the AI alongside the existing system without surfacing its outputs to users. Both the existing system and the AI process the same requests. The AI's outputs are logged and evaluated. Users receive the existing system's response.

SHADOW MODE ARCHITECTURE

User request
     │
     ├──► [Existing System] ──► User Response (shown to user)
     │
     └──► [AI System] ──► AI Response (logged, never shown to user)
                              │
                         [Shadow Evaluation]
                         - Compare AI response to existing response
                         - Run eval assertions on AI response
                         - Log: agreement rate, quality scores,
                                latency, cost per request
                         - Alert if quality below threshold

Duration: Run shadow mode until you have signal on 1,000+ requests
          across the full range of user query types.

What to measure in shadow mode:
  ├── Agreement rate: how often does AI agree with existing system?
  │     (for replacement of a deterministic system)
  ├── Quality score: eval suite scores on production traffic sample
  ├── Latency distribution: P50, P95, P99 of AI response time
  ├── Cost per request: confirms the cost model from design phase
  └── Failure rate: % of requests where AI fails to produce output

What shadow mode tells you:

  • High agreement + high quality → AI is ready for canary
  • High agreement + low quality → AI agrees with a bad existing system. Agreement is not the goal. Quality is.
  • Low agreement + high quality → AI behaves differently from existing system. This may be intentional (AI is better) or a problem (AI is wrong differently). Requires human review of disagreements before canary.
  • Low agreement + low quality → Go back to development.

Canary Deployment

Canary deployment routes a small percentage of real user traffic to the AI system. Users receive the AI response. The canary percentage is increased gradually based on quality signal.

CANARY DEPLOYMENT SEQUENCE

Week 1: 5% canary
  ├── 5% of traffic routes to AI system
  ├── 95% routes to existing system (or no AI)
  ├── Monitor: quality metrics, user feedback, error rate, cost
  └── Decision criteria to advance: quality metrics stable for 3 days,
      no SEV 2+ incidents, no anomalous user complaint spike

Week 2: 20% canary (if week 1 criteria met)
  ├── New scale may reveal distribution shifts not visible at 5%
  └── Same monitoring, same advancement criteria

Week 3: 50% canary (if week 2 criteria met)
  └── At 50%, you have enough signal to confidently assess production behavior

Week 4: Full deployment (if week 3 criteria met)
  └── The "existing system" path becomes the fallback, not the default

ROLLBACK TRIGGER (return to previous percentage immediately):
  ├── Any SEV 2 incident
  ├── Quality metric drops > 15% from canary baseline
  ├── User complaint rate increases > 50% vs. control group
  └── Cost exceeds 2x projected at this traffic level

Feature Flags for AI

The canary deployment percentage should be controlled by a feature flag, not a deployment parameter. This means the percentage can be changed without a deployment:

# Feature flag controls AI routing
# Change in the feature flag system, not in code

def route_request(user_id: str, request: Request) -> Response:
    # Feature flag determines if this user gets AI
    if feature_flags.ai_chat_enabled(
        user_id=user_id,
        rollout_percentage=CANARY_PERCENTAGE  # controlled externally
    ):
        return ai_handler.process(request)
    else:
        return existing_handler.process(request)

This architecture means: - Rollback is a feature flag change, not a deployment (seconds, not minutes) - Specific user segments can be included or excluded from the canary - The canary percentage can be paused mid-way without a rollback


24.4 Rollback Procedures for AI Systems

AI systems have more rollback targets than traditional software. A traditional rollback is: revert the code to the previous version. AI rollback may need to target the prompt, the model, the knowledge base, or the code — or any combination.

Rollback Target 1: Prompt Version Rollback

PROMPT ROLLBACK PROCEDURE

Time to execute: < 5 minutes (if prompt management system is in place)
Time to execute: 30-60 minutes (if prompts are in deployment configs)

Step 1: Identify the prompt version that was active before the incident
  └── Check the prompt management system's deployment history
      or the prompt_version field in recent audit logs.

Step 2: Assess whether rollback is safe
  └── Would rolling back also revert a security or compliance fix?
      If yes: rolling back is not safe. Must fix forward.
      If no: proceed.

Step 3: Execute the rollback
  Prompt management system: 
    → Select the previous version in the UI
    → Click "Promote to production"
    → The gateway's prompt retrieval immediately returns the prior version

  Manual (no prompt management system):
    → Update the configuration value in the deployment config
    → Trigger a deployment to apply the change
    → Verify the old prompt is live by checking a known test query

Step 4: Verify
  └── Run 5 test queries against the known failure scenario.
      Do they produce correct results with the prior prompt?
      If yes: incident is mitigated. Investigation continues.
      If no: prompt is not the root cause. Escalate to model investigation.

Step 5: Document
  └── Log the rollback: what version was active, what version is now active,
      who authorized the rollback, what time it took effect.

Rollback Target 2: Model Version Rollback

MODEL VERSION ROLLBACK PROCEDURE

Prerequisite: the model version was pinned (not using a floating alias)
If not pinned: there is no prior version to roll back to.
              This is an architecture failure. Add to post-mortem.

Time to execute: 5-15 minutes

Step 1: Identify the pinned model version that was working
  └── Check the gateway's model configuration history.
      What was the model version before the most recent change?

Step 2: Update the gateway service profile
  └── In the gateway configuration, change:
        model: "claude-opus-5-5"  (broken version; illustrative ID)
      to:
        model: "claude-opus-5"    (prior working version; illustrative ID)
      Real model IDs come from the provider's model list, never from memory.

Step 3: Verify the gateway is routing to the prior version
  └── Make a test API call. Check the response header or response
      metadata for the actual model version used.

Step 4: Run quick eval validation
  └── Run the 5 most critical eval cases. Do they pass with the prior version?

Step 5: Contact the model provider
  └── File a support ticket documenting the behavioral change.
      Include: what version changed behavior, what the change is,
      reproduction steps. This is your evidence for requesting an
      investigation or a hotfix.

Rollback Target 3: Knowledge Base Rollback

KNOWLEDGE BASE ROLLBACK PROCEDURE

Scenario: A document update introduced incorrect or injected content.
         Users are receiving wrong answers based on the new content.

Immediate mitigation (< 5 minutes):
  Step 1: Identify the problematic document(s)
    └── Search the knowledge base for chunks related to the
        failing queries. Which document are they from?

  Step 2: Archive the problematic chunks
    └── Update the chunk metadata: status = "archived"
        The query filter (status = "active") will immediately
        stop returning these chunks.
        Do NOT delete — preserve for investigation.

  Step 3: Verify mitigation
    └── Run the failing query. Are the problematic chunks still returned?
        If no: mitigation is effective.
        If yes: the filter is not working. Check the query code.

Full rollback (if archiving is insufficient):
  Step 4: Restore prior document version
    └── If using the document registry from Module 5:
        - Find the prior version of the document (status: archived)
        - Restore it to status: active
        - The new version reverts to status: archived
    └── If not using document registry:
        - Re-ingest the prior version of the document
        - Manually archive all chunks from the problematic version

  Step 5: Validate retrieval quality
    └── Run the 10 queries most likely to have been affected.
        Are the results correct with the restored document?

24.5 The AI Operational Runbook

Every production AI system should have an operational runbook — a document that any on-call engineer can use to diagnose and respond to production issues without needing to wake up the AI architect. This is the template.

AI OPERATIONAL RUNBOOK TEMPLATE

System: [Name]
Last updated: [Date]
Contacts: [Primary on-call, escalation, platform team, model provider support]

HEALTH CHECKS
  Normal state indicators:
  ├── Gateway: HTTP 200, P99 latency < Xms
  ├── Eval score (daily run): faithfulness > Y, relevance > Z
  ├── Human escalation rate: < X% per day
  ├── Cost: < $X per day (check daily cost report)
  └── Model version: [expected version string]

  Check these first when investigating an issue.

COMMON ISSUES AND RESPONSES

Issue: Eval scores dropped > 15% overnight
  Likely cause: prompt change, knowledge base update, model update
  Response:
    1. Check what changed in the last 24 hours (deployments, KB updates)
    2. Run the specific failing eval cases to identify the pattern
    3. If prompt change: consider rollback (Section 24.4)
    4. If KB update: check for stale or incorrect content
    5. Escalate to AI architect if root cause not identified in 30 min

Issue: Human escalation rate spike (>2x normal)
  Likely cause: AI quality degraded, scope boundary broken, retrieval failing
  Response:
    1. Sample 10 escalated interactions. What type of failure?
    2. Run eval suite. Identify which assertions are failing.
    3. Check retrieval quality: are the right chunks being returned?
    4. Escalate to AI architect if pattern not identified in 1 hour

Issue: Cost spike (>2x normal)
  Likely cause: prompt grew longer, agent looping, traffic spike, routing failure
  Response:
    1. Check tokens per request trend (is it growing?)
    2. Check agent iteration count P95 (agents looping?)
    3. Check if model routing is working (expensive model for cheap tasks?)
    4. Escalate to platform team if not resolved in 2 hours

Issue: Model returning unexpected responses
  Response:
    1. Check model version in API responses vs. configured version
    2. Run eval suite against production
    3. If version mismatch: rollback to pinned version (Section 24.4)
    4. If version correct but behavior changed: contact model provider

ESCALATION PATH
  L1 On-call engineer:
    Response: Triage, common issues resolution, escalation

  L2 AI Architect:
    When to escalate: root cause not identified in 30 min (SEV 2),
    any SEV 1, any rollback required

  L3 Model Risk Owner (regulated systems):
    When to escalate: SEV 1, any incident with regulatory implications

  L4 Model Provider Support:
    When to escalate: suspected model version behavior change,
    API issues, deprecation notice received

EXERCISE — Incident Simulation: With a colleague, run a tabletop exercise for an AI production incident. Scenario: an AI customer support system has been returning incorrect fee information to customers for 4 hours. Work through the 30-minute playbook: triage questions, initial diagnosis, mitigation decision. Document the decisions made and the rationale. Debrief: which step took longest? Which information was missing?

PONDER — Your Rollback Readiness: For any AI system currently in production at your organization: how long would it take to roll back a prompt change? A model version change? A knowledge base update? If any of these takes more than 30 minutes, what architectural change would reduce it?

WORKSHOP — Build the Runbook: Select one AI system in your organization. Build the complete operational runbook using the template from Section 24.5. Fill in every field with real information. Then give it to an engineer who did not build the system and ask them to simulate responding to each common issue. What information was missing that they needed?


Next: Module 25 — Commercial and Legal: Vendor Evaluation and AI Contracts