Skip to content

MODULE 16 — AI Integration Patterns & AI-Native Design

16.1 The Fundamental Design Question

When an organization wants to incorporate AI into an existing system or build a new one, there is a foundational design question that most teams never explicitly ask:

Is AI the core of this system, or an enhancement to it?

The answer determines the entire architecture. A system where AI is the core — where the primary value proposition depends on AI functioning — requires different design principles than a system where AI enhances an existing capability. The failure modes are different, the degradation strategy is different, the evaluation model is different, and the way users relate to the system is different.

Getting this distinction right is the architect's first job before any integration decision is made.


16.2 AI-Augmented vs. AI-Native: The Architectural Distinction

AI-Augmented Systems

AI-augmented systems are existing systems with AI features added. The system works without AI — AI makes it faster, better, or more convenient. When AI fails, the system degrades gracefully to its pre-AI functionality.

Examples: - A search engine that adds AI-generated summaries above organic results - A CRM with AI-suggested next actions - An email client with AI-assisted compose suggestions - A code editor with AI autocomplete

Architectural characteristics: - Core data model and business logic were designed before AI - AI operates on the edges of user workflows, not the center - The system has a natural fallback: "show the original results without the AI summary" - User trust in the system is not primarily contingent on AI quality - AI quality improvements are incremental enhancements

Design principle: AI-augmented systems should be designed so the AI layer can be switched off — for any individual user, for any feature, or for the entire system — without breaking the core value proposition.


AI-Native Systems

AI-native systems are designed from inception with AI as the primary interface or the primary value delivery mechanism. The system as designed does not work without AI — or the user experience without AI is so degraded that it no longer fulfills the system's purpose.

Examples: - Cursor (the IDE) — the entire design premise is AI-guided development - Perplexity — the product is AI-synthesized search, not a list of links - A clinical decision support system where AI analyzes patient data and generates recommendations — the system's value is the AI's analysis - An autonomous financial reporting system that extracts, analyzes, and narrates data — the narrative is the product

Architectural characteristics: - Data model, API design, and user flows are designed around AI capabilities and limitations - Latency, uncertainty, and quality variability are first-class design concerns, not edge cases - Evaluation infrastructure (how do you know the AI is working?) is as important as the AI itself - Human oversight is designed in, not added afterward - Failure modes of the AI are modeled as part of the system design

Design principle: AI-native systems must make the AI's limitations visible to users and must design the human oversight model before building the AI capability.


The Misclassification Problem

Teams frequently build AI-native systems while thinking they are building AI-augmented systems. This happens when the AI capability is so central to the value proposition that removing it would leave nothing useful — but the team hasn't explicitly acknowledged this and hasn't designed for it.

Signs of misclassification: - "We'll add a fallback later" — the fallback never gets built because there is no natural non-AI version - The entire user journey assumes AI results are correct — no path for the AI being wrong - No human escalation path exists — the system has no answer to "what do I do when the AI's output is wrong?" - The evaluation strategy is "we'll check if users complain" — which is not an evaluation strategy


16.3 The Strangler Fig Pattern for AI Integration

The Strangler Fig pattern (from Martin Fowler) is the standard approach for incrementally replacing legacy functionality. Applied to AI, it means: replace components of an existing system with AI-powered equivalents one at a time, allowing the system to continue operating throughout the migration.

The Classic Application

STRANGLER FIG FOR AI INTEGRATION

Phase 1: Route to AI alongside legacy (parallel operation)
  Incoming request
    ├── Legacy handler (processes request, returns result)
    └── AI handler (processes same request, returns result)

  Router sends both requests.
  Legacy result is shown to users.
  AI result is logged and compared against legacy result.
  Evaluate: does the AI result match or improve on legacy?

Phase 2: Shadow mode validation
  Continue for sufficient traffic volume to establish quality baseline.
  Compute: agreement rate, quality improvement metrics, failure rate.
  Identify edge cases where AI underperforms legacy.

Phase 3: Canary deployment
  10% of traffic goes to AI handler.
  90% remains on legacy.
  Monitor: user satisfaction, error rate, escalation rate.
  Compare against baseline metrics from shadow mode.

Phase 4: Gradual rollout
  Increase AI traffic percentage as confidence grows.
  Maintain ability to reverse to 0% AI at any point.

Phase 5: Legacy decommission
  Only when AI handler meets or exceeds legacy quality metrics
  across all major traffic segments.
  Legacy code is not deleted — it is archived for potential rollback.

The AI-Specific Complications

The classic Strangler Fig assumes deterministic replacement — the new system either works or doesn't, and you can compare outputs precisely. AI introduces complications:

Non-deterministic outputs. You cannot compare AI output to legacy output with an equality check. You need a quality evaluation function — which requires defining what "better" means before you start.

Distribution shift. AI models perform differently on different input distributions. The validation you did on Month 1 traffic may not predict performance on Month 6 traffic if user behavior changes. Build ongoing monitoring before decommissioning the legacy system.

Long tail failure modes. AI may outperform legacy on 95% of inputs and significantly underperform on 5%. The 5% may be the highest-stakes 5% — the edge cases the legacy system was specifically built to handle. Identify these before decommissioning legacy.

The evaluation problem. For some AI integrations, evaluating whether the AI result is "better" requires human judgment at scale. This requires a sampling-based evaluation approach, not automated comparison.


16.4 API Design for AI Services

When an AI capability is exposed as a service API — whether internal or external — the API design must account for properties that traditional APIs don't have: variable latency, uncertain outputs, streaming responses, and the need to convey confidence and provenance.

Streaming Response Design

Traditional APIs: request → wait → response. The wait time is bounded and predictable.

AI APIs: request → wait for first token → stream tokens as they're generated → end of stream. Latency to first token (TTFT) is much lower than total response time. Streaming lets the client start rendering while the model is still generating.

AI API STREAMING RESPONSE DESIGN

Response event stream (Server-Sent Events):

data: {"type": "start", "request_id": "req-abc123", "model": "claude-sonnet-5-5",
       "timestamp": "2024-03-15T14:23:11Z"}

data: {"type": "token", "content": "Your", "index": 0}
data: {"type": "token", "content": " annual", "index": 1}
data: {"type": "token", "content": " fee", "index": 2}
...

data: {"type": "citation", "source_doc_id": "fee-schedule-v3",
       "source_version": "2024-03-01", "section": "4.2"}

data: {"type": "tool_call", "tool_name": "get_account_balance",
       "status": "completed", "latency_ms": 145}

data: {"type": "end", "total_tokens": 47, "confidence": 0.91,
       "escalation_required": false, "finish_reason": "stop"}

Why each event type matters:
  start: client can render a loading indicator and bind request_id for audit
  token: incremental rendering — user sees response forming in real-time
  citation: client can render inline source references
  tool_call: client can show "checking your account..." UI feedback
  end: client has final metadata for audit logging; escalation flag triggers UI

Error handling in streaming:
data: {"type": "error", "code": "CONFIDENCE_BELOW_THRESHOLD",
       "message": "I couldn't find sufficient information to answer this question.",
       "escalation_path": "human_agent",
       "fallback_response": "Let me connect you with a specialist."}

Confidence and Uncertainty as First-Class API Fields

A traditional API either returns a result or throws an error. An AI API may return a result that is partially correct, uncertain, or based on outdated information — and the client needs to know.

AI API RESPONSE SCHEMA WITH UNCERTAINTY

{
  "response_id": "resp-abc123",
  "content": "Your annual fee is $0 for a basic checking account.",

  "quality": {
    "confidence": 0.91,           // 0-1 confidence in the answer
    "grounding": "full",          // full | partial | none
    "source_freshness": "current", // current | aging (>30d) | stale (>90d)
    "human_review_recommended": false,
    "escalation_required": false
  },

  "provenance": {
    "sources": [
      {
        "doc_id": "fee-schedule-v3",
        "title": "Account Fee Schedule",
        "version": "2024-03-01",
        "section": "4.2 Basic Checking Fees",
        "relevance_score": 0.94
      }
    ],
    "retrieval_strategy": "hybrid_bm25_dense",
    "model": "claude-sonnet-5-5",
    "prompt_template": "customer-support-v12",
    "prompt_version": "12.3"
  },

  "audit": {
    "request_id": "req-abc123",
    "session_id": "sess-xyz789",
    "timestamp": "2024-03-15T14:23:11Z",
    "user_id": "user-pseudo-456",  // pseudonymized
    "cost_usd": 0.0089
  }
}

Why these fields matter to downstream systems:

  • confidence and human_review_recommended: the UI layer can show a "this answer has been reviewed" badge when confidence is high, or a "please verify with a specialist" note when low
  • source_freshness: the client can show "based on policy from 3 months ago — [view current policy]" when the source is aging
  • provenance.sources: enables inline citations and allows compliance systems to trace every answer to its source
  • audit: required for model risk management, compliance, and debugging

The Citation Contract

For any AI service that makes factual claims, citations should be a required part of the API contract — not an optional enhancement. The citation contract defines:

CITATION CONTRACT

The AI service MUST:
  ├── Include at least one source for any factual claim
  ├── Not make factual claims without a retrievable source
  └── Flag uncertainty when sources are not available

The AI service MUST NOT:
  ├── Return a factual claim with confidence > 0.80 without a source
  ├── Mix information from sources with different effective dates
  │     without flagging the date discrepancy
  └── Return a source reference that cannot be retrieved by the client

The client MUST:
  ├── Render source references where the user can access them
  └── Display staleness warnings when source_freshness is not "current"

This is an API contract, not a prompt instruction.
The enforcement is in the API response schema validation,
not in the system prompt.

16.5 Progressive Enhancement with AI

Progressive enhancement is the design principle of building a working baseline that functions without AI, then layering AI capabilities on top in a way that enhances rather than replaces the baseline.

This is the architectural expression of the AI-augmented vs. AI-native distinction. Progressive enhancement explicitly designs the non-AI baseline and then makes AI a quality upgrade.

The Progressive Enhancement Pattern

PROGRESSIVE ENHANCEMENT ARCHITECTURE

BASELINE TIER (always works, no AI):
  ├── Core functionality available without AI
  ├── Deterministic, fast, no external model dependency
  └── What the user sees if AI is unavailable or confidence is low

ENHANCED TIER (AI-powered, best effort):
  ├── AI adds value on top of the baseline
  ├── When AI is unavailable: falls back to baseline seamlessly
  └── User sees a better experience, not a different experience

EXAMPLE: Document Search System

Baseline tier:
  User searches for "account fee waiver"
  System returns: keyword-matched document list, ranked by recency
  Always works, fast, no AI dependency

Enhanced tier:
  Same query, but:
  ├── AI generates a natural language summary answer at the top
  ├── AI highlights the most relevant passage in the top result
  └── AI suggests related queries the user might want to ask

  If AI is unavailable: user sees the baseline keyword results
  User experience degrades gracefully — they still find documents
  They don't see a blank page or an error

EXAMPLE: Customer Support Chat

Baseline tier:
  User submits a question
  System routes to a human agent queue
  Agent answers the question
  Always works, known quality

Enhanced tier:
  AI generates a suggested answer
  If confidence > 0.85: AI response shown to user with agent review option
  If confidence 0.65-0.85: AI draft shown to agent, agent edits and sends
  If confidence < 0.65: AI is not shown; route to agent queue (baseline)

  If AI is unavailable: all traffic routes to agent queue (baseline)
  Quality degrades (more agent time) but the service continues

Designing for Graceful Degradation

Graceful degradation requires explicit design, not just the absence of errors.

DEGRADATION DESIGN QUESTIONS (answer before building):

For each AI capability:
  1. What does the user see if this AI capability is unavailable?
     (server error is not an answer)

  2. What does the user see if AI confidence is below threshold?
     (blank response is not an answer)

  3. Is the degraded experience still useful?
     (useful means the user can accomplish their goal, not just
      that they see something other than an error)

  4. Is the degradation automatic or does it require operator action?
     (automatic is always better — if degradation requires someone
      to flip a switch, it won't happen fast enough in a real incident)

  5. Is the threshold for degradation calibrated?
     (too high: AI rarely activates; too low: AI activates when
      it shouldn't and quality suffers)

16.6 The AI Sidecar Pattern

The AI sidecar pattern places AI capability as a sidecar to an existing service — the AI runs alongside the service but is not embedded in it. The service can operate without the sidecar, and the sidecar can be updated independently.

This is the AI equivalent of the service mesh sidecar pattern (Envoy, Linkerd), applied to AI capability rather than networking.

AI SIDECAR PATTERN

WITHOUT SIDECAR:
  User Request → [Service A] → Response
  Service A: returns structured data

WITH AI SIDECAR:
  User Request → [Service A] → Structured Response
                    │
               [AI Sidecar]
               - Receives Service A's structured response
               - Adds AI-generated interpretation, summary, or recommendation
               - Returns enriched response

  The sidecar:
  ├── Has no shared state with Service A (separate process/container)
  ├── Can be deployed, updated, or removed independently
  ├── If it fails: Service A continues returning structured response
  └── Can be applied to multiple services without modifying them

REAL-WORLD APPLICATION:

A legacy insurance claims processing system returns:
  {claim_id: "CL-12345", status: "pending_review", age_days: 8}

With AI sidecar, the response becomes:
  {
    claim_id: "CL-12345",
    status: "pending_review",
    age_days: 8,
    ai_enrichment: {
      risk_flag: "high",          // AI assessment of claim risk
      recommended_action: "Escalate to senior adjuster — documentation incomplete",
      similar_claims: ["CL-11890", "CL-11234"],  // AI-found similar cases
      processing_sla_at_risk: true
    }
  }

The legacy system doesn't know about the AI.
The AI sidecar doesn't touch the legacy system.
The client uses whichever fields it can handle.

When the sidecar pattern is appropriate:

  • Legacy systems that cannot be modified but could benefit from AI enrichment
  • AI capabilities that should be applied broadly without embedding in each service
  • When AI governance requires a clear separation between deterministic business logic and AI reasoning
  • When the AI capability is being piloted and may not survive to production

The key constraint: The sidecar must be stateless and must have access only to the data in the response it is enriching — not to additional data sources that the originating service doesn't have access to. A sidecar that accesses more data than the service it enriches creates unexpected coupling.


16.7 The Event-Driven AI Pattern

AI as a synchronous service in an event-driven pipeline creates a bottleneck. The event-driven AI pattern treats the AI as an async processor that enriches events as they flow through the system.

Covered in depth in the Integration Patterns reference (Artifact 2, Pattern 12). Key architectural principles for this module:

Right-size the model to the event. Not all events require a frontier model. A fraud flag event that needs a one-sentence explanation goes to a cheap, fast model. A complex investigation summary goes to a capable model. Model routing within the event pipeline applies the same logic as the general routing architecture (Module 13).

AI in the event pipeline must be idempotent. Event pipelines often deliver events at-least-once. An AI enrichment step that produces a new action (send notification, update record) must be idempotent — processing the same event twice must not produce duplicate actions.

Design the DLQ with human resolution. When an AI enrichment step fails (model unavailable, confidence below threshold, timeout), the event goes to a Dead Letter Queue. Unlike traditional DLQs where the resolution is technical (retry, investigate the code), AI DLQ events often require human judgment — the AI couldn't determine the appropriate enrichment, so a human must. Design the human review interface for DLQ events as part of the AI pipeline design.


16.8 Data Mesh Integration with AI

The data mesh pattern — domain-owned, self-serve data products — intersects with AI integration at the question of how AI capabilities access and contribute to domain data.

AI as a Consumer of Data Products

In a data mesh, AI systems should be first-class consumers of data products. This means:

  • AI systems access domain data through the domain's published data product API — not through direct database access
  • The domain team's data product contract includes the quality and freshness guarantees that the AI system can rely on
  • When the data product's quality degrades, the AI system's quality degrades predictably rather than catastrophically
DATA MESH + AI INTEGRATION

Customer Domain publishes:
  customer_ai_context v1.0 (data product)
  ├── Customer tier, account status, risk flags
  ├── Excluded: PII fields not relevant to AI use cases
  ├── SLA: updated within 30 seconds of source change
  └── Quality SLA: completeness > 99%, freshness < 30 seconds

AI Customer Support System consumes:
  customer_ai_context v1.0
  ├── Injects into AI context window per conversation
  └── Fails gracefully when product SLA is not met
        (shows "account information temporarily unavailable")

This is correct. What is incorrect:

AI Customer Support System directly queries:
  customer_database.customers TABLE
  ├── No data product contract
  ├── AI system bypasses domain ownership
  └── Schema changes break AI without domain team knowing

AI as a Producer of Data Products

AI systems also produce data that other systems consume. The AI-generated enrichment should be treated as a data product with its own quality contract, not as ephemeral output.

AI AS DATA PRODUCT PRODUCER

AI Risk Scoring System produces:
  customer_risk_assessment v2.0 (data product)
  ├── customer_id, risk_score, risk_factors, model_version
  ├── Confidence interval per score
  ├── Assessment timestamp and model version used
  ├── Quality SLA: model drift detected within 72 hours
  └── Consumed by: loan origination, fraud detection, marketing

Other systems consume risk_assessment as a data product:
  ├── They receive the schema contract, not a raw model output
  ├── Model upgrades update the data product, not the downstream systems
  └── Data product version bumps are managed like API versions

16.9 Typed-Decision Models: Where They Fit in the Architecture

Module 2 (§2.2, Category 5) introduced typed-decision models: models that return a calibrated probability for each option in a schema you define, instead of generating text. This section covers the integration questions: where in a system they belong, how to use their probabilities, how to evaluate and adopt one, and where they should not be used.

Currency note. As of October 2026 the category is about three weeks old and every public detail comes from vendors and secondary coverage, so treat figures as (reported) and verify them. Public examples: Jev (TypeSafe AI, early access September 15, 2026; reported at $0.042 per million input tokens and 40–200× faster than frontier LLMs), Laya (Convai Innovations; open source; a ~421M-parameter bidirectional encoder, ~33 ms per local call), and Clef / Clef-flash (Cloudflare, October 1, 2026; open weights under Apache 2.0, post-trained from Qwen models of 27B and 9B parameters, with an API compatible with Jev). All are described as trained with RLCD (Reinforcement Learning for Calibrated Decisions), whose reward is reported to be a strictly proper scoring rule. TypeSafe has not published the RLCD method itself. See Appendix G for the current list. The placement and evaluation guidance below does not depend on which vendor wins.

What the Interface Looks Like

The caller sends the state to be judged and one or more typed questions. The model answers each question with probabilities and nothing else. The three question types reported across these models are a choice among enumerated options, a score on a bounded ordinal scale, and a yes/no with a probability. The shape below is illustrative and is not any vendor's real API:

REQUEST                                          RESPONSE
{                                                {
  "state": {                                       "answers": {
    "ticket": "I was charged twice ...",             "category": {
    "tier": "enterprise",                              "billing": 0.91, "outage": 0.04,
    "history": [...]                                   "access": 0.03, "other": 0.02 },
  },                                                 "urgency": { "1": 0.02, "2": 0.08,
  "questions": {                                       "3": 0.20, "4": 0.55, "5": 0.15 },
    "category": { "choice": [                        "needs_human": { "yes": 0.12 }
      "billing","outage","access","other"] },      },
    "urgency":  { "score": [1,5] },                "schema_version": "triage-v3",
    "needs_human": { "yes_no": true }              "model_version": "..."
  }                                                }
}

Three properties matter for architecture. The output cannot leave the schema, so there is no parsing or retry step. Latency is low and stable (tens of milliseconds, reported), because there is no decode loop. And the response is a probability distribution, which you can act on with thresholds, unlike a free-text answer.

Where They Belong

Typed-decision models fit decision points that were already bounded before you added AI.

Placement Decision it makes Why a typed-decision model fits Cross-reference
Request router Which model tier or pipeline handles this request? The answer set is small and fixed. The call must cost less than the work it routes Module 2 §2.3
Confidence gate Answer, escalate, or say "I don't know" A calibrated probability is the gate's input. An LLM's self-reported confidence is not Module 4 (confidence gate), §16.5
Agent step guard Is this tool call in scope? Does it need human approval? Runs on every step of an agent loop, so the per-call cost and latency add up fast Module 6
Triage and tagging at volume Category, priority, language, document type The classic high-volume bounded decision Module 13
Pre-filter for expensive pipelines Is this document worth GraphRAG extraction? Does this question need deep research? Spend the expensive path only where it pays off Modules 4, 36
Screening layer in a security design Does this input look like an injection or policy violation? Cheap enough to run on all traffic. Only one layer of several Module 9
Feature-flag decisions Should this user get the AI path? Bounded and frequent §16.10 Principle 6

Where not to use them: - Open-ended generation or explanation. Drafting, summarizing, multi-turn conversation, and reasoning you want to inspect all need a generative model. - Answer spaces you can't enumerate. If you can't list the possible answers in advance, this is not the tool. - Decisions with no labelled data. Calibration can only be checked against known-correct answers (see Evaluation, below). - Sole line of defence in adversarial settings. A classifier can be probed and evaded. Use it as one layer, never the only control (Module 9). - Consequential individual decisions with no human review. A probability is not an explanation. Where an adverse decision must be explained to the affected person, you still need an explanation capability (Module 11). - Rare or novel classes. Probabilities for classes the model saw few times are the least trustworthy.

A Quick Test

USE A TYPED-DECISION MODEL WHEN ALL FIVE ARE TRUE

  1. The answer set is fixed and known before the call
  2. You have (or can build) labelled examples from your real traffic
  3. Volume or latency makes an LLM call a real cost (e.g., on every request or agent step)
  4. You can measure calibration on YOUR data before trusting the probabilities
  5. A wrong answer has a defined fallback (cascade to an LLM, or a human)

If 2 or 4 is false: start with an LLM and collect labels.
If 5 is false: design the fallback first (Anti-Pattern 2 in §16.11).

Calibration Is the Product: Acting on Probabilities

The value of these models is that a probability can drive routing, so the architecture should be built around thresholds and fallbacks. The standard pattern is a three-zone cascade:

THREE-ZONE CASCADE

 p(top answer) ≥ τ_high     →  ACT automatically
 τ_low ≤ p < τ_high         →  CASCADE to a generative model (or deeper pipeline)
 p < τ_low                  →  ESCALATE to a human

 Choose τ_high and τ_low on a HELD-OUT set from your own traffic, using the cost of
 each kind of error — not a default like 0.5 or 0.9.

Illustrative economics (assumed numbers, not benchmarks). Take 1,000,000 support tickets a month. Suppose an LLM triage call costs $0.004, so LLM-only triage costs about $4,000. Suppose a typed-decision model handles every ticket for roughly $0.00002 (a 500-token ticket at the reported Jev input price of $0.042 per million tokens, about $21), 85% of tickets clear τ_high with acceptable precision, 10% cascade to the LLM ($0.004 × 100,000 = $400), and 5% go to humans. The cascade costs about $420 plus the human review, instead of $4,000. The saving depends entirely on the 85% figure. That figure must be measured on your data, at the precision your business needs, which is why the Evaluation subsection is not optional.

The calibration trap. A model is calibrated only on the kind of data it was evaluated on. Calibration can degrade for new customers, new languages, new product lines, or seasonal shifts, while accuracy on the old test set looks fine. Check calibration per segment, and re-check it on a schedule.

Integrating with the API Contract

Extend the response design from §16.4 (confidence as a first-class field) rather than inventing a parallel scheme:

  • Return and log the full probability vector, not only the top answer. The vector is your audit trail and your data for recalibration.
  • Version the schema. Adding, removing, or renaming an answer option changes what the probabilities mean. Treat it as a new version that needs re-evaluation (§16.10 Principle 1).
  • Record model_version, schema_version and the thresholds in force on every decision, so a later review can reconstruct why a request was auto-handled or escalated.
  • Keep thresholds in configuration, behind the feature-flag mechanism in Principle 6, so you can tighten them during an incident without a deployment.

Adopting One: Hosted, Open Weights, or Build

Option Examples (Oct 2026, reported) Strengths Watch for
Hosted API Jev Fastest start; no GPUs Data leaves your perimeter; vendor and pricing risk; the category is weeks old
Open weights, self-hosted Laya (small, local), Clef / Clef-flash (Apache 2.0) Data stays inside; fine-tune on your labels; no per-call vendor fee You operate it; verify each licence; benchmark on your data
Build / fine-tune your own Start from an open base model Fit to your schema and data Needs labelled data and a calibration evaluation. The RLCD recipe is unpublished

If you fine-tune, the practical path is: (1) start from an open typed-decision model or a small open base model; (2) train on your labelled examples; (3) run a cheap baseline first, which is ordinary supervised training followed by post-hoc recalibration (temperature scaling or isotonic regression) on a held-out set, because this often captures much of the calibration benefit; (4) move to a reinforcement-style objective with a proper scoring rule (such as the Brier score or log score) only if the baseline falls short. A rough guide for a single narrow decision is thousands to tens of thousands of labelled examples. That figure is an estimate, so measure your own learning curve. Module 34 §34.3 covers grader design and reward hacking for reinforcement fine-tuning, and the same cautions apply to any scoring-rule reward.

Portability. Several of these models share an API (Clef is reported to be compatible with Jev), which makes switching easy to wire. It does not make the models equivalent. Early independent comparisons report different strengths, for example one outlet reports Clef is better at routing and weaker at judgment than Jev (reported). Calibration profiles differ too. Run the behavioral contract and paired evaluation from Module 37 §37.5 before swapping, and include calibration metrics in the contract.

Evaluation

Evaluate a typed-decision model on your labels, before it makes any decision that matters.

TYPED-DECISION ACCEPTANCE CONTRACT (illustrative thresholds — set your own)

  Accuracy / F1 per class          ≥ current baseline − tolerance, on held-out real data
  Brier score or log loss          ≤ baseline (measures probability quality, not just the top answer)
  Calibration (ECE + reliability   within a stated tolerance overall AND for each key segment
    diagram)                         (language, customer tier, region, rare classes)
  Coverage at target precision     ≥ X% of traffic clears τ_high at ≥ Y% precision
                                     (this is the number your cost model depends on)
  Out-of-distribution behavior     probability falls (or "other" rises) on inputs unlike the training data
  Adversarial cases                injection-like and edge inputs do not silently receive high confidence
  Latency / cost                   P95 ≤ SLO; cost per decision including cascade calls
  • Use a held-out set the model, the vendor, and your tuning never touched (the "eval gap" problem in Module 34).
  • Treat vendor benchmark numbers as claims (Module 20). A vendor's calibration on its own benchmark says little about calibration on your traffic.
  • Monitor in production: sample decisions for human labelling, track calibration and coverage over time, and recalibrate on a schedule or when drift alerts fire (Module 12 §12.6).
  • Governance: a typed-decision model that scores or classifies people is still a model for risk-management purposes (Module 11), and it belongs in the model inventory with an owner.

Failure Modes and Anti-Patterns

  • Confidently wrong out of distribution. Inputs unlike the training data can still produce sharp probabilities. Test for it and keep a fallback.
  • Thresholds chosen on training data, or left at a default. The cascade then behaves differently from the cost model.
  • Schema drift. Someone adds an option "to make the dashboard nicer" and the old thresholds no longer mean anything.
  • Probability treated as explanation. "92% billing" says how sure the model is, not why.
  • Single-layer security. A screening classifier presented as the control against injection.
  • No fallback path. The cascade has no human or LLM tier, so low-confidence decisions go nowhere (Anti-Pattern 2, §16.11).
  • Believing the launch numbers. Speed and cost multiples are vendor-reported and the category is new.

Where the course states the concept: Module 2 §2.2 Category 5 defines the model class. Module 37 covers swapping one for another. This section is the integration guide between them.


16.10 AI-Native Application Design Principles

When building a new system where AI is the core — not a bolt-on — these principles guide the architecture.

Principle 1: Design for Model Versions, Not for Behaviors

AI-native systems depend on a model, and models change. The system must be designed so that model changes can be evaluated and managed.

  • Never hardcode model-specific behavior in application code
  • Define expected behaviors as tests (your eval suite), not as code that relies on specific model outputs
  • Design the system to be evaluable: given a model change, how do you know if the system is still working correctly?

Principle 2: Human Oversight Is Architecture, Not Policy

In AI-native systems, human oversight cannot be added after the fact. It must be designed into the data flow from the beginning.

  • Identify every consequential decision the system makes
  • For each: define whether it is auto-handled (with what confidence threshold?) or requires human approval
  • Design the human review interface as part of the core system design, not as an "admin panel" afterthought
  • Human approval events are first-class data — stored, audited, attributed to the approver

Principle 3: Uncertainty is a Product Feature

AI systems produce uncertain outputs. In an AI-augmented system, this uncertainty is hidden from users. In an AI-native system, uncertainty is visible and meaningful.

The user of Perplexity sees source links and knows the AI synthesized from those sources. The user of a clinical decision support system sees confidence scores and knows the AI's recommendation should be reviewed by a physician. The uncertainty is not a failure to hide — it is information the user needs.

Design principle: surface uncertainty in a way that is informative, not alarming. "Based on 3 sources, with high confidence" is informative. "Warning: AI output - 0.87 confidence score" is alarming and not actionable.

Principle 4: The Evaluation System Is Part of the Product

In traditional software, the test suite is engineering infrastructure. In AI-native systems, the evaluation system is product infrastructure — it tells you whether the product is working, and it must keep working as models change, prompts change, and knowledge bases update.

The eval suite must be: - Maintained with the same discipline as production code (version controlled, reviewed, updated from production failures) - Running continuously in production (online eval sampling) - Triggering alerts when quality degrades - Part of the deployment process (offline eval as a deployment gate)

This is not different from Module 12 on observability — it is the same principle applied to the product design layer: the eval system is designed in, not added on.

Principle 5: Design the Failure Narrative

In traditional software, failure means an error code. In AI-native systems, failure is more nuanced — the system may respond with incorrect information, with outdated information, with inappropriate content, or not at all. Each failure mode has a different user experience implication.

Design the failure narrative before building the happy path: - What does the user see when the AI is wrong? - What does the user see when the AI is uncertain? - What does the user see when the knowledge base is stale? - What does the user see when the AI system is unavailable? - What is the path from "AI got it wrong" to "human resolves it"?

Teams that don't design these failure narratives discover them in production, at the worst possible time, with the worst possible user impact.

Principle 6: AI Capability as a Feature Flag

AI capabilities should be deployable, toggleable, and scoped — not monolithically enabled or disabled.

AI FEATURE FLAG ARCHITECTURE

Feature flags for AI capabilities allow:
  ├── Per-user enablement (pilot users, beta users, premium tier)
  ├── Per-geography enablement (comply with regional AI regulations)
  ├── Per-data-classification (AI enabled for public data, not confidential)
  ├── Kill switch (disable AI instantly if quality degrades or incident occurs)
  └── Gradual rollout (10% → 25% → 50% → 100%)

Example feature flag configuration:
  ai.customer_support.answer_generation:
    enabled: true
    rollout_percentage: 75
    excluded_geographies: ["EU"]  // pending EU AI Act compliance
    excluded_data_classifications: ["HIGHLY_CONFIDENTIAL"]
    kill_switch: false
    fallback_behavior: "route_to_human_queue"

This is not the same as the Progressive Enhancement tier system.
Feature flags control whether a feature is available at all.
Progressive enhancement controls the experience when it is.

Principle 7: Version and Manage AI-Generated Stored Content

When AI generates content that gets stored — reports, summaries, recommendations, risk assessments, annotations — that stored content represents the AI's output at a specific model and prompt version. When the model or prompt changes, the stored content may be inconsistent with what the current system would generate.

AI-GENERATED CONTENT VERSIONING

Every stored AI-generated artifact must include:
  ├── ai_model_version: what model generated this
  ├── prompt_template_version: which prompt template was used
  ├── generation_timestamp: when it was generated
  ├── knowledge_base_version: (for RAG-based) which docs were current
  └── validity_period: when this content should be considered stale

When model or prompt changes:
  ├── Mark existing stored content as generated_by_prior_version: true
  ├── Display to users: "This summary was generated by an older model.
  │     [Regenerate]"
  └── For compliance-sensitive content: retain the old version in audit
        log; the new version is a separate record

Systems that don't do this serve users content that may be inconsistent
with the current system's quality level — sometimes better, sometimes
worse — without any way to distinguish.

16.11 Anti-Patterns in AI Integration

Anti-Pattern 1: The AI Black Box

The application calls the AI API, gets a response, and shows it to the user. No citations, no confidence scores, no provenance, no escalation path. When it's wrong, nobody knows why, nobody can trace it, and nobody has a path to resolution.

What makes this an anti-pattern: The first time a user acts on a wrong answer and the organization is asked to reconstruct what happened, there is no answer. In regulated industries, this is an examination finding. In any industry, it is a loss of user trust.


Anti-Pattern 2: The "We'll Add Fallbacks Later" Trap

The team builds an AI-native user experience with no non-AI fallback. When the model provider has an outage, the feature is completely unavailable. When quality degrades, there is no degraded mode — just a bad experience.

The fallback is never built because it is always "later" and the feature shipped without it.

What makes this an anti-pattern: Production incidents happen. Model providers have outages. Quality degrades without warning. A system that has no fallback has no resilience.


Anti-Pattern 3: AI in the Critical Transaction Path

The AI model call is in the synchronous critical path of a transaction — a payment, an order submission, a medical record write. When the model is slow (P99 latency spike), the transaction is slow. When the model is unavailable, the transaction fails.

What makes this an anti-pattern: AI inference is slow and variable. Transaction paths need to be fast and reliable. The solution is to separate the AI enrichment from the critical transaction path: commit the transaction, then enrich asynchronously.


Anti-Pattern 4: The Infinite Confidence Display

The AI response is shown to users with no indication of uncertainty — as if it were a database lookup, with complete confidence. Users trust the AI completely. When it is wrong, the trust violation is severe.

What makes this an anti-pattern: AI outputs are probabilistic. Users who understand this interpret them appropriately. Users who don't are set up for disappointment. Displaying confidence — even implicitly through source citations — calibrates user expectations.


Anti-Pattern 5: The Prompt-Only Quality Strategy

"We'll make the AI behave correctly by writing a better prompt." Prompts cannot enumerate all failure modes. Prompts cannot prevent injection. Prompts cannot guarantee structured output. Prompts are the starting point for AI behavior, not the guarantee.

What makes this an anti-pattern: Prompt instructions are soft constraints. The quality strategy must include: schema validation, eval pipelines, confidence gates, human review paths, injection defenses, and output monitoring. The prompt is one layer, not the whole stack.


16.12 Integration Patterns Governance Checklist

Typed-decision models (§16.9) - [ ] Used only where the answer set is fixed, labelled data exists, and a fallback is defined? - [ ] Thresholds (τ_high, τ_low) chosen on a held-out set from our own traffic, kept in configuration? - [ ] Calibration measured per segment (language, tier, region, rare classes), not only overall? - [ ] Full probability vector, schema version, model version and thresholds logged per decision? - [ ] Coverage at target precision measured, and the cost model built on that number? - [ ] Not the sole control in any security design; human review for consequential individual decisions? - [ ] Recalibration schedule and drift monitoring in place?

AI-augmented vs. AI-native - [ ] Explicitly classified: is AI the core of this system or an enhancement? - [ ] If AI-native: human oversight model designed before development started? - [ ] If AI-native: failure narratives designed (wrong, uncertain, unavailable)?

API design - [ ] Streaming implemented for long-running AI responses? - [ ] Confidence and quality metadata in response schema? - [ ] Provenance and citations in response schema (if factual claims)? - [ ] Audit fields in response schema (request_id, model, timestamp)? - [ ] Error responses include escalation path, not just error code?

Progressive enhancement - [ ] Non-AI baseline defined and implemented? - [ ] AI layer degrades gracefully to baseline? - [ ] Degradation is automatic (not operator-triggered)? - [ ] Threshold for degradation is calibrated and documented?

Integration architecture - [ ] AI sidecar isolated (stateless, fails independently of core service)? - [ ] Event-driven AI enrichment is idempotent? - [ ] DLQ for AI pipeline has human review interface? - [ ] Data mesh consumers/producers have schema contracts?

AI-native design - [ ] Eval system designed and built alongside the feature (not after)? - [ ] Model version abstracted (change doesn't require code change)? - [ ] Uncertainty surfaced to users in informative, non-alarming way? - [ ] Production feedback loop from failures to eval regression suite?


EXERCISE — Classify Your Systems: Take five AI-related features or systems in your organization. For each, honestly answer: is AI the core of this system, or an enhancement? For those where AI is the core — does the system have a defined non-AI fallback? Does it have a human escalation path? Does it surface uncertainty to users? Document the gaps.

PONDER — The API Contract: For the most consequential AI API in your organization: does its response include confidence, provenance, and citations? If a user acts on an incorrect response today, can the incident response team reconstruct: which document version the answer was based on, what the confidence was, which model produced it, and when? If not — what would it take to add that to the API response?

WORKSHOP — Design a Progressive Enhancement: Take an existing feature in your organization that currently has no AI. Design the progressive enhancement version: define the baseline tier (no AI, always works), the enhanced tier (AI-powered, best effort), and the degradation thresholds. For the enhanced tier: what is the API response schema that conveys confidence and uncertainty? What does the user see in each quality tier?

WORKSHOP — AI-Native Design Review: Review an AI-native feature that is planned or recently built. Apply the five AI-native design principles from Section 16.10. For each principle: is it satisfied? If not, what is the gap and what is the architectural change required? Apply the anti-pattern checklist from Section 16.10. Document findings as an architectural review.

WORKSHOP — Place a Typed-Decision Model: Pick one high-volume decision in your organization that currently uses an LLM call or a rules engine (ticket triage, routing, moderation, an agent step guard). Apply the five-point test in §16.9. If it passes, design the three-zone cascade: state the answer set, the labelled data you would use, the thresholds you would test, the fallback for each zone, and the acceptance contract (including per-segment calibration and coverage at target precision). Then estimate the monthly cost of the cascade against the current approach, and say which number in your estimate you have measured and which you have assumed.


Next: Module 17 — Platform Engineering for AI