Skip to content

MODULE 26 — AI Requirements Engineering

26.1 Why Standard Requirements Don't Work for AI

Traditional user stories follow the format: "As a [user], I want to [do something], so that [I get value]." This format assumes the system will either do the thing or not. It says nothing about quality, confidence, failure modes, or human oversight — the four dimensions that matter most for AI systems.

A team that writes AI requirements in the traditional format will ship AI systems with no acceptance criteria for quality, no designed response to failure, and no human oversight model. All of these will be discovered in production.


26.2 The AI User Story Format

The Extended Format

AI USER STORY FORMAT

AS A [user]
I WANT [the AI capability]
SO THAT [the value delivered]

QUALITY ACCEPTANCE CRITERIA:
  The AI [capability] must achieve [specific measurable quality bar]
  as measured by [specific evaluation method].
  Example: "The AI answer must be grounded in the retrieved
            context at least 85% of the time as measured by
            LLM-as-judge faithfulness scoring."

FAILURE MODE REQUIREMENTS:
  When [the AI cannot fulfill this story with sufficient confidence],
  the system must [specific behavior].
  Example: "When retrieval confidence is below 0.65, the system
            must respond with an escalation message rather than
            generating an answer."

HUMAN OVERSIGHT REQUIREMENTS:
  [Who reviews AI outputs, under what conditions, with what information]
  Example: "AI-generated loan summaries must be reviewed by a
            loan officer before the applicant receives them. The
            review interface must display: the AI summary, the
            source documents, and the confidence score."

OBSERVABILITY REQUIREMENTS:
  Every [AI interaction type] must be logged with:
  [list of required fields]
  Example: "Every AI response must be logged with: query_hash,
            retrieved_chunk_ids, model_version, prompt_template_id,
            confidence_score, response_hash, timestamp."

OUT-OF-SCOPE DECLARATION:
  This AI capability must NOT [specific prohibited behaviors].
  Example: "This AI must not provide investment advice,
            discuss competitor products, or make commitments
            about future rates or policy changes."

Worked Example

Traditional format (insufficient for AI):

As a customer, I want to ask questions about my account fees, so that I can understand what I'm being charged.

Extended AI format:

As a customer, I want to ask questions about my account fees, so that I can understand what I'm being charged.

Quality acceptance criteria: The AI must answer fee-related questions with verified accuracy at least 85% of the time as measured by RAGAS faithfulness scoring against the official fee schedule. The AI must cite the specific fee schedule version when providing fee information.

Failure mode requirements: When retrieval confidence is below 0.65 OR when the query is not fee-related, the AI must respond with: "I want to make sure I give you accurate information. Let me connect you with a fee specialist." and offer human escalation. When the fee schedule is more than 90 days old, the AI must add: "Please verify this information is current with our support team."

Human oversight requirements: No human review required for standard fee questions with confidence > 0.80. Human review required for: disputed charges (any query containing "dispute," "wrong," "charged incorrectly"), questions about fees > $500, and any response flagged with confidence < 0.80.

Observability requirements: Every interaction logged with: query_hash, retrieved_chunks, chunk_versions, confidence_score, fee_amounts_mentioned (for audit), escalation_triggered, model_version, prompt_template_id, timestamp.

Out-of-scope declaration: This AI must not discuss: account balances, transaction history, investment products, competitor fees, or provide financial advice. Out-of-scope queries must receive the escalation response.


26.3 Non-Functional Requirements for AI

Non-functional requirements (NFRs) for AI systems cover dimensions not present in traditional NFRs.

AI-Specific NFR Categories

Quality NFRs

Quality NFRs (examples):

Faithfulness: The system shall maintain a faithfulness score of
  ≥ 0.85 on the defined eval dataset, measured weekly.

Accuracy: The system shall achieve ≥ 90% accuracy on the
  classification task as measured on the held-out test set.

Consistency: For equivalent queries submitted within 1 hour,
  the system shall produce materially equivalent answers
  (same key facts, same recommended action) at least 95%
  of the time. [Tests for temperature misconfiguration]

Governance NFRs

Governance NFRs (examples):

Auditability: Every AI-influenced decision must be reproducible
  from the audit log for a period of [retention period].
  The audit log must contain sufficient information to reconstruct:
  the input, the retrieved context, the model version, the prompt
  version, and the output.

Explainability: For adverse decisions made using AI assistance,
  the system must generate a plain-language explanation of the
  factors that influenced the decision. This explanation must
  be available within 24 hours of a customer request.

Human override: Any AI-generated output or recommendation must
  be overridable by an authorized human without requiring a
  system deployment.

Safety NFRs

Safety NFRs (examples):

Scope containment: The system shall not produce content outside
  its defined authorization scope in more than 1 in 10,000
  interactions, as measured by adversarial testing.

Escalation reliability: When the AI triggers an escalation,
  the escalation must reach a human agent within [SLA].
  The escalation path must not depend on AI functioning correctly.

Irreversibility gate: No irreversible action (financial transaction,
  account modification, external communication) shall be taken
  by the AI system without explicit human approval. This requirement
  cannot be overridden by prompt instruction.

Cost NFRs

Cost NFRs (examples):

Unit cost: The cost per AI interaction must not exceed $[X].
  If this threshold is exceeded for > 5% of interactions,
  the system must alert the owning team automatically.

Monthly ceiling: Total AI cost for this system must not exceed
  $[Y] per month. The system must have a hard budget ceiling
  enforced at the gateway layer.

Cost attribution: 100% of AI interactions must be attributable
  to a specific team and feature in the cost reporting system.

26.4 The AI Definition of Done

The standard definition of done for software features needs extensions for AI features. Without these additions, "done" means the code works in a demo — not that the AI feature is production-ready.

AI FEATURE DEFINITION OF DONE

Standard criteria (from existing DoD):
  ✓ Code reviewed and merged
  ✓ Unit and integration tests passing
  ✓ Documentation updated
  ✓ Deployed to staging

AI-specific additions:
  ✓ Eval suite passing: minimum [N] cases, all mandatory assertions pass
  ✓ Prompt registered in prompt management system with version and owner
  ✓ Cost baseline documented: expected cost per interaction
  ✓ Observability configured: OTel spans emitting, cost tags active
  ✓ Alerts configured: quality, cost, and error rate alerts set
  ✓ Model version pinned (no floating alias in production config)
  ✓ Failure mode tested: does the escalation path work?
  ✓ Human oversight path tested: does the review workflow function?
  ✓ Security: Garak scan run, no P1/P2 findings unaddressed
  ✓ AI inventory: system registered with owner, tier, and risk classification
  ✓ Operational runbook: basic runbook exists for on-call engineers

For regulated AI features, additionally:
  ✓ Model risk owner assigned and notified
  ✓ Validation documentation prepared per tier requirements
  ✓ Compliance review completed (if required by tier)
  ✓ Adverse action explanation capability verified (if applicable)

26.5 The AI Acceptance Testing Checklist

When accepting an AI feature from a development team (for QA or architecture sign-off), use this checklist.

AI FEATURE ACCEPTANCE CHECKLIST

FUNCTIONAL
  □ Feature answers the core use cases correctly
    (run 10 representative scenarios — not developer-selected)
  □ Failure mode behavior is correct
    (what happens at low confidence, out-of-scope, model unavailable)
  □ Escalation path works end-to-end
    (simulate an escalation trigger; confirm it reaches a human)
  □ Human override works
    (simulate a human overriding an AI output; confirm it takes effect)

QUALITY
  □ Eval suite passes at defined thresholds
  □ Adversarial eval passes (at least 5 injection attempts)
  □ Quality is consistent across the range of user query types
    (not just the scenarios the team tested against)

GOVERNANCE
  □ Every interaction is logged with required fields
  □ Audit trail is complete and retrievable
  □ Model version is pinned and documented
  □ Prompt is in version control with owner

OPERATIONAL
  □ Cost per interaction is within expected range
  □ Latency meets SLA under normal load
  □ Alerts fire correctly (test by temporarily lowering thresholds)
  □ Runbook exists and is complete

SECURITY
  □ Garak scan completed with no unaddressed P1 findings
  □ PII handling is correct (prompt and response)
  □ Tool scope is minimum necessary (for agentic features)

26.6 Story Mapping for AI Workflows

AI features often span multiple user interactions. A story map helps visualize the complete user journey including the AI's role at each step, the failure modes, and the human touchpoints.

AI STORY MAP STRUCTURE

BACKBONE (top row): User journey stages
  [Searches for information] → [Receives AI answer] → [Acts on answer]
                                                            │
                                                    [Realizes answer was wrong]
                                                            │
                                                    [Seeks correction]

USER TASKS (middle rows): What the user does at each stage
  [Types query] → [Reads response] → [Clicks citation to verify]
                                  → [Escalates to human if uncertain]

AI BEHAVIORS (bottom rows): What the AI does at each stage
  [Retrieves context] → [Generates answer with citations] → [Logs interaction]
  [Assesses confidence] → [Triggers escalation if low confidence]

FAILURE FLOWS (separate track):
  [AI returns wrong answer] → [User acts on wrong answer] → [Dispute raised]
  [No AI escalation triggered] → [Human notified via complaint]
  [Audit pulls interaction log] → [Root cause identified]
  [Policy updated / document corrected]

The value of mapping the failure flow: most teams map the happy path and
assume failure handling will be built as needed. Mapping failures upfront
reveals design gaps before they become production incidents.

EXERCISE — Rewrite a User Story: Take any AI feature currently in development or recently shipped at your organization. Rewrite its user story using the extended AI format from Section 26.2. Add: quality acceptance criteria (with specific numbers), failure mode requirements, human oversight requirements, observability requirements, and out-of-scope declaration. Share the rewritten story with the team. What did they not know needed to be defined?

PONDER — The Missing NFRs: For any AI feature in production at your organization: does it have documented quality NFRs? Is there a faithfulness score or accuracy target that is tracked? Is there a cost ceiling enforced? If none of these exist, the feature has no measurable definition of quality — it is considered "done" if it doesn't crash.

WORKSHOP — AI Story Mapping: Choose an AI feature being planned in your organization. Build a story map that includes: the happy path, the failure flow, the escalation path, and the audit/compliance path. For each flow, identify: where the AI acts, where a human acts, what is logged, and what would happen if a step failed. The failure flow is the most important part of this exercise.


Next: Module 27 — Organizational Patterns: What Makes AI Initiatives Succeed or Fail