Skip to content

MODULE 11 — Compliance & Model Risk in Regulated Industries

⚠️ Currency note: This module is accurate as of October 2026. Regulatory dates and guidance change. The US banking agencies' planned request for information on AI could reshape §11.2–§11.3, and EU AI Act standards are still being cited. The current status of each lives in Appendix G §G.7 — Regulatory & Standards Status. Check there before you quote a date to a client or regulator.

11.1 The Regulatory Landscape Has Changed

The compliance landscape for AI in regulated industries shifted materially in early 2026. Architects working in financial services, healthcare, insurance, or any heavily regulated sector need to know the current state — not the pre-2026 state that most guidance documents still reference.

The most significant change: On April 17, 2026, the Federal Reserve (SR 26-2), the OCC (Bulletin 2026-13), and the FDIC (FIL-15-2026) issued revised interagency Supervisory Guidance on Model Risk Management. It supersedes SR 11-7 and its OCC and FDIC counterparts (OCC 2011-12, FIL-22-2017), and the 2021 interagency statement on model risk management for BSA/AML systems (SR 21-8 / FIL-27-2021). The new guidance is explicitly risk-based and principles-driven, tailored to each bank's model risk profile and size. Four facts matter most to architects:

  • Who it targets: it is "expected to be most relevant to banking organizations with over $30 billion in total assets." Smaller banks with significant model risk may still find it relevant.
  • How binding it is: it "does not set forth enforceable standards or prescriptive requirements," and non-compliance with the guidance alone will not draw supervisory criticism. Supervisory action can still follow from unsafe or unsound practices caused by poor model risk management.
  • What a model is: "a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates." It excludes simple spreadsheet arithmetic and deterministic rule-based processes.
  • What is out of scope: generative AI and agentic AI (see §11.2). The agencies plan a request for information (RFI) on model risk management and banks' use of AI, including generative and agentic AI. As of October 2026, check Appendix G for whether it has been issued.

For architects in financial services: if your governance framework is built around SR 11-7 terminology and structure, review it against the April 2026 interagency guidance. The core principles carry over (inventory, validation, ongoing monitoring, governance, vendor model oversight). The prescriptive SR 11-7 expectations have been replaced by risk-based tailoring, and you now have to justify your own choices.

For architects outside financial services: the principles of model risk management — inventory, validation, ongoing monitoring, governance — apply across regulated industries. This module covers those principles and their AI-specific application, using financial services as the primary reference with healthcare and legal considerations noted.


11.2 What Model Risk Management Means for AI

Model risk management (MRM) is the practice of identifying, assessing, and mitigating the risks that arise from using mathematical or algorithmic models to make or inform business decisions. It originated in financial services where models are used for credit scoring, market risk, fraud detection, and capital calculation — and where a model error can have consequences measured in billions of dollars.

The arrival of AI in regulated industries applies MRM to a new class of models that are harder to validate, more opaque, and more behaviorally complex than the regression and statistical models MRM was designed for.

What Makes AI Models Harder to Govern Under MRM

Traditional model: A logistic regression credit scoring model. Input features are explicit. Coefficients are documented. The model's behavior can be fully explained by examining its parameters. Conceptual soundness is evaluated by reviewing the theoretical basis for each feature and its coefficient.

AI model: An LLM-based customer interaction system. Input features are not explicit — the model processes the full context of a conversation, and its "features" are learned embeddings not interpretable by humans. The model's coefficients are 70 billion parameters. Conceptual soundness cannot be evaluated by reviewing parameters. The model is stochastic — the same input produces different outputs.

The April 2026 interagency guidance acknowledged this gap explicitly. In its words: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The same footnote adds two points. A bank's own risk management and governance practices should still determine the controls for tools the guidance does not cover. The guidance's principles do apply to "traditional statistical and quantitative models and non-generative, non-agentic AI models," so a gradient-boosted credit model is in scope and an LLM underwriting assistant is not. The agencies plan an RFI on AI, including generative and agentic AI. The implication: the regulatory framework for governing generative AI in banking decisions is still being developed. Organizations that wait for definitive guidance before building governance infrastructure will be behind. Out of scope does not mean ungoverned. Examiners can still act on unsafe or unsound practices.

The practical architect's position: apply the principles of the 2026 interagency guidance (inventory, validation appropriate to materiality, ongoing monitoring, governance) to AI systems, while acknowledging that the specific validation methodologies for generative AI are evolving and documenting your approach to that uncertainty.


11.3 The Model Inventory: Foundation of Compliance

You cannot manage risk in what you cannot see. The model inventory is the foundation of model risk management. It answers: what models are we running, who owns them, what decisions do they inform, and what is their risk tier?

What Constitutes a "Model" Under the 2026 Guidance

The 2026 guidance defines a model as "a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates." It explicitly excludes simple arithmetic calculations (such as spreadsheet formulas) and deterministic rule-based processes. Traditional statistical models and non-generative, non-agentic AI/ML models are in scope. Generative and agentic AI are out of scope, but they still need governance under the bank's own framework, and the prudent course is to apply MRM-equivalent rigor where outcomes are consequential.

In practice:

WHAT IS A MODEL (for MRM purposes)

Clearly in scope:
  ├── Credit scoring algorithms (traditional and AI-enhanced)
  ├── Fraud detection models (rule-based + ML)
  ├── Anti-money laundering transaction monitoring
  ├── Customer churn prediction models
  ├── Pricing models (insurance, loans, securities)
  ├── Risk rating models (counterparty, credit, market)
  └── AI recommendation systems that influence financial decisions

OUT of SR 26-2 scope, but govern with MRM-equivalent rigor
(your own framework; expect the planned RFI to address these):
  ├── LLM-based customer communication that materially
  │     affects customer outcomes (denial communications,
  │     product recommendations)
  ├── AI systems that automate or substantially influence
  │     underwriting, claims, or credit decisions
  └── Agentic systems that initiate financial transactions
        or take binding actions

Out of SR 26-2 scope, lighter internal governance:
  ├── Generative AI and agentic AI in general purpose use
  │     (productivity tools, drafting assistance)
  └── Determination: materiality and decision-consequence
        set the governance depth inside your own framework

NOT models (for MRM purposes):
  ├── Internal productivity tools (document summarization,
  │     meeting notes, email drafting)
  ├── AI that provides information but leaves all decisions
  │     to humans with full discretion
  └── Rule-based automation with no statistical estimation

The Model Inventory Schema

MODEL INVENTORY RECORD

Model Identity:
  ├── model_id: unique identifier
  ├── model_name: "Customer Churn Prediction v3.1"
  ├── model_type: classification | regression | LLM | agent | hybrid
  ├── model_version: "3.1.0" (semantic versioning)
  └── deployment_date: "2026-03-15"

Business Context:
  ├── purpose: what decision does this model inform or make?
  ├── business_owner: who is accountable for this model?
  ├── model_developer: who built and maintains it?
  ├── primary_users: who uses the model outputs?
  └── decision_authority: does the model decide, recommend, or inform?

Risk Classification:
  ├── tier: TIER_1 (high materiality) | TIER_2 | TIER_3 (low)
  ├── tier_rationale: documented justification for tier assignment
  ├── decision_consequence: financial | customer | regulatory | operational
  ├── estimated_annual_decision_impact: $X or N customers affected
  └── regulatory_applicability: [list of regulations that apply]

Technical Details:
  ├── model_architecture: logistic regression | gradient boosting |
  │     transformer | agent | ensemble
  ├── training_data_sources: [list with data classification]
  ├── input_features: [list]
  ├── output_description: "probability score 0-1, threshold 0.65 for action"
  └── infrastructure: [deployment environment]

Governance Status:
  ├── validation_status: current | overdue | pending | exempt
  ├── last_validation_date: "2026-03-01"
  ├── next_validation_due: "2027-03-01"
  ├── validation_findings: [summary of open findings]
  ├── ongoing_monitoring_status: active | not_configured | remediation_needed
  ├── model_risk_owner: named individual accountable for this model
  └── audit_history: [log of material changes and validations]

For AI/LLM-specific models, add:
  ├── model_provider: OpenAI | Anthropic | self-hosted | hybrid
  ├── model_version_pinned: dated snapshot ID, never a floating
  │     alias (yes/no + exact version string)
  ├── prompt_version: "customer-churn-v4.2"
  ├── eval_suite: [reference to evaluation dataset and metrics]
  ├── human_review_requirement: automated | human_review | human_decision
  └── explainability_approach: [how decisions are explained to stakeholders]

The Tiering Framework

Not all models require the same governance intensity. The 2026 guidance explicitly endorses a tiered, risk-proportionate approach:

MODEL RISK TIERING

TIER 1 — HIGH MATERIALITY (most intensive governance)
  Criteria:
    ├── Directly automates or substantially drives consequential
    │     decisions (credit approval, denial, pricing)
    ├── High financial impact (>$10M annual decision value) OR
    │     large customer population affected (>100K customers)
    ├── Regulatory scrutiny expected (ECOA, FCRA, HMDA)
    └── Model failure would have significant financial or reputational impact

  Governance requirements:
    ├── Independent validation before deployment
    ├── Annual validation (or after material change)
    ├── Quarterly ongoing monitoring review
    ├── Board-level reporting
    └── Model risk owner at VP level or above

TIER 2 — MEDIUM MATERIALITY
  Criteria:
    ├── Informs (but does not automate) consequential decisions
    ├── Moderate financial impact or customer population
    └── Moderate regulatory exposure

  Governance requirements:
    ├── Validation before deployment (may be self-validation for lower-risk)
    ├── Annual or biennial validation
    └── Quarterly monitoring metrics reviewed by model owner

TIER 3 — LOW MATERIALITY (lightest governance)
  Criteria:
    ├── Internal operational use only
    ├── No direct consumer-facing decisions
    ├── Low financial impact
    └── Easily reversed if model fails

  Governance requirements:
    ├── Basic documentation and owner assignment
    ├── Validation at deployment (proportionate depth)
    └── Annual self-assessment by model owner

11.4 Model Validation for AI: The Hard Problem

Traditional model validation evaluates: (1) conceptual soundness — is the theory behind the model sound? (2) data and process verification — is the model implemented correctly? (3) outcomes analysis — do the model outputs match expected behavior in practice?

For AI models, particularly LLMs and agentic systems, each of these has challenges that traditional validation methodology does not address.

Conceptual Soundness for LLMs

Traditional: "Explain why each input feature is expected to be predictive of the target variable. Show the coefficient signs are consistent with theory."

LLM: There are no explicit input features. The model has learned representations from training data in ways not interpretable by examining parameters. "Conceptual soundness" must be evaluated differently.

Approach for AI/LLM conceptual soundness:

Purpose-appropriateness evaluation: Is a language model the appropriate technology for this use case? Could a deterministic rule-based system, a simpler ML model, or human judgment achieve the same goal more reliably? The justification for using an LLM where a simpler model would suffice is itself a conceptual soundness question.

Grounding adequacy: For RAG-based systems, is the knowledge base adequate for the decision domain? Are the sources authoritative? Is the retrieval quality measured and documented?

Output calibration: For models used in consequential decisions, are the model's confidence signals calibrated — do high-confidence outputs actually correspond to more accurate outputs? An LLM that is "confident" but wrong is a poorly calibrated model.

Behavioral boundary testing: Does the model behave as expected at the edges of its intended domain? Testing with out-of-domain inputs, adversarial inputs, and edge cases is the AI equivalent of sensitivity analysis in traditional model validation.

Data and Process Verification for AI

Traditional: "Verify the training dataset is accurate and representative. Verify the model is implemented as documented."

AI/LLM: Training data is a black box for commercial models. Process verification focuses on the deployment configuration:

AI/LLM PROCESS VERIFICATION CHECKLIST

Prompt verification:
  ├── System prompt documented and version-controlled?
  ├── Prompt change control process exists and is followed?
  ├── Prompt tested against the validation dataset before deployment?
  └── Prompt scope limitations align with the model's documented purpose?

Configuration verification:
  ├── Model version pinned (not using a floating alias)?
  ├── Temperature and other hyperparameters documented and justified?
  ├── Context window limits enforced?
  └── Output validation (schema, length, content) implemented?

Integration verification:
  ├── Data inputs to the model reviewed for accuracy and completeness?
  ├── Tool integrations (for agents) tested for correct behavior?
  ├── Human review requirements implemented as designed?
  └── Audit logging captures all required fields?

Outcomes Analysis for AI

Traditional: "Compare model predictions to actual outcomes. Measure accuracy, precision, recall, Gini coefficient."

AI/LLM: Outcomes analysis for language models requires a different measurement approach because the output space is continuous text, not a discrete class.

AI OUTCOMES ANALYSIS FRAMEWORK

For classification/decision AI (fraud, churn, credit):
  ├── Standard ML metrics apply: accuracy, precision, recall, AUC
  ├── Fairness metrics by protected class (see Section 11.6)
  ├── False positive/negative rate monitoring by demographic group
  └── Population stability index (has the population changed?)

For language generation AI (recommendations, communications):
  ├── Faithfulness rate: what % of outputs are grounded in sources?
  ├── Human review escalation rate: trending up = quality degradation
  ├── User complaint/dispute rate on AI-informed decisions
  ├── LLM-as-judge eval scores: consistency, accuracy, appropriateness
  └── Subject matter expert review of sampled outputs (quarterly)

For agentic AI (autonomous workflows):
  ├── Task completion rate: % of tasks completed without human intervention
  ├── Escalation rate: % of tasks escalated (by escalation reason)
  ├── Error rate: % of tasks that resulted in a correctable error
  ├── Irreversible action error rate (highest severity)
  └── Cost per task: trending up may indicate quality/efficiency regression

11.5 Explainability: When, Why, and How

Explainability is a compliance requirement in specific contexts, not a general requirement for all AI. Architects must understand when it applies and what architecture supports it.

When Explainability Is Required

Regulatory mandates:

ECOA / Regulation B (US credit decisions): When an application for credit is denied or offered on terms less favorable than requested, the applicant must receive specific reasons for the adverse action. "The model said no" is not a compliant explanation. The reasons must be specific, principal, and actionable.

Fair Credit Reporting Act (FCRA): When a consumer credit score is used in an adverse action, the consumer must receive the key factors that adversely affected the score.

EU AI Act (high-risk systems): Providers must make high-risk systems transparent enough for deployers to interpret outputs, and affected persons have a right to an explanation of decisions taken with Annex III high-risk systems (Article 86). After the Digital Omnibus, Annex III high-risk obligations apply from December 2, 2027 (§11.10).

EU GDPR (Article 22): Individuals have the right not to be subject to solely automated decisions with legal or similarly significant effects, and the right to obtain an explanation of such decisions.

Internal governance mandates:

Audit and examination: Regulators and internal auditors expect institutions to be able to explain AI-driven decisions. "We cannot explain why the model made this recommendation" is an examination finding.

Customer escalations: When a customer disputes an AI-driven decision, the institution must be able to explain the decision to the customer and to supervisory bodies.

What Explainability Actually Means

There are two distinct things called "explainability," often confused:

Global explainability: Understanding the model's general behavior — what factors tend to increase or decrease the model's output. For traditional ML models (gradient boosting, etc.), SHAP values provide feature importance at the population level. For LLMs, global explainability is limited — the model's behavior is too complex to summarize in feature importances.

Local explainability: Understanding why the model produced a specific output for a specific input. This is what regulatory adverse action requirements demand.

EXPLAINABILITY ARCHITECTURE FOR REGULATED AI

For models using LLMs in consequential decisions:

APPROACH 1: Explanation-generation architecture
  LLM Decision System produces:
    ├── The decision or recommendation
    └── An explanation: a chain-of-thought reasoning trace
          that documents the factors the model considered
          and how they influenced the output

  The explanation is:
    ├── Stored with the decision record
    ├── Available to human reviewers
    ├── Validated for coherence (does the explanation
        actually describe the decision made?)
    └── Used in adverse action notices (with human review)

  Risk: The explanation may not accurately reflect the model's
        internal processing — it may be a post-hoc rationalization.
        For high-stakes decisions, human review of the explanation
        is required before it is used as a regulatory justification.

APPROACH 2: Hybrid architecture (most robust for regulated use)
  Separate the decision-informing step from the decision step:

    Step 1 (AI): Extract and score relevant factors from
                  the application/context (structured output)

    Step 2 (Deterministic): Apply a documented rule to the
                             scored factors to produce the decision

    Step 3 (AI, optional): Generate a plain-language explanation
                            of the factors and the rule that produced
                            the decision

  The decision logic (Step 2) is deterministic and fully explainable.
  The AI's role is factor extraction (Step 1), not decision-making.
  The explanation is grounded in the documented rule (Step 3).

  This architecture satisfies adverse action requirements because the
  decision is made by a documented rule, not a black box.

11.6 Fairness, Bias, and Fair Lending

AI models used in lending, insurance, employment, or housing decisions in the US are subject to fair lending laws (ECOA, HMDA, Fair Housing Act) and their disparate impact provisions. An AI model that produces disparate outcomes for protected classes — even without discriminatory intent — may violate these laws.

The Disparate Impact Test

Disparate impact occurs when a neutral policy or practice has a disproportionately adverse effect on members of a protected class. For AI:

DISPARATE IMPACT ASSESSMENT FOR AI MODELS

Protected classes (US lending context):
  Race, color, religion, national origin, sex, marital status,
  age, receipt of public assistance, familial status

Metrics to compute:
  ├── Approval rate by protected class
  │     Acceptable disparity: 80% Rule (4/5 rule)
  │     If approval rate for any group is < 80% of the highest
  │     group's approval rate, investigate for disparate impact
  │
  ├── Average score/output distribution by protected class
  │     Compare means and distributions
  │
  ├── False positive/negative rates by protected class
  │     Unequal error rates can indicate discrimination even
  │     if overall accuracy is similar
  │
  └── Adverse action reason distribution by protected class
        Are certain reasons concentrated in certain groups?
        Does that concentration have a legitimate business justification?

When disparate impact is found:
  1. Document the finding
  2. Assess: is there a legitimate business justification?
  3. Assess: is there a less discriminatory alternative that meets
             the business need equally well?
  4. If no business justification or a less discriminatory alternative
     exists: the model must be modified or replaced before deployment

Architecture for Bias Detection

Bias detection is not a one-time pre-deployment activity. It must be continuous in production because population composition changes, model behavior can drift, and the context in which the model is applied evolves.

BIAS MONITORING ARCHITECTURE

Data collection:
  ├── Log every model decision with:
  │     - Input features (or feature categories)
  │     - Output (decision, score, recommendation)
  │     - Protected class indicators (where legally permitted to collect)
  │     - Outcome (did the decision prove accurate over time?)
  └── Retain for the regulatory retention period

Monitoring computation:
  ├── Weekly: approval/denial rate by available demographic proxies
  ├── Monthly: full disparate impact analysis on decisions made
  ├── Quarterly: adverse action reason distribution analysis
  └── Triggered: any time a material change is made to the model

Alert thresholds:
  ├── 4/5 rule breach → immediate review required
  ├── Approaching 4/5 rule (>90% of limit) → monitor weekly
  └── Material shift in outcome distribution by group → investigate

Tools:
  ├── IBM AI Fairness 360 (open source, comprehensive bias metrics)
  ├── Google What-If Tool (visualization of model behavior by group)
  ├── Fairlearn (Microsoft, bias assessment and mitigation)
  └── Aequitas (fairness audit toolkit from UChicago)

11.7 Ongoing Monitoring: The MRM Requirement Most Teams Underinvest

Validation is a point-in-time assessment. Ongoing monitoring is continuous. The 2026 interagency guidance keeps ongoing monitoring as a core principle, with frequency and scope set by "the nature of the model, the availability of new data or modeling approaches, and model materiality." It is also where teams most often underinvest, because it has no launch date to force the work.

For AI models, ongoing monitoring must address risks that do not exist in traditional models:

AI MODEL ONGOING MONITORING FRAMEWORK

PERFORMANCE MONITORING (monthly minimum for Tier 1)
  ├── Output quality metrics (eval scores, human review rates)
  ├── Prediction accuracy vs. actual outcomes (where available)
  ├── Fairness metrics (Section 11.6)
  └── Compare to baseline at last validation

STABILITY MONITORING (monthly)
  ├── Population stability: has the input distribution changed?
  │     (PSI — Population Stability Index — for each key input)
  ├── Output distribution stability: has the distribution of
  │     outputs changed materially?
  └── Concept drift indicators: is model accuracy declining
        even though inputs appear stable?

AI-SPECIFIC MONITORING (continuous)
  ├── Prompt version in use: is the approved prompt deployed?
  ├── Model version in use: is the pinned model version active?
  │     (Alert if model version changes without validation)
  ├── Retrieval quality (for RAG-based models):
  │     Context precision and faithfulness metrics
  ├── Human escalation rate: trending up = quality issue
  └── Cost per decision: trending up may indicate model behavior change

THRESHOLD ALERTS
  ├── Any metric deviates > 10% from validated baseline → investigate
  ├── Model version changes without validation event → immediate alert
  ├── Prompt changes deployed without approved change process → alert
  └── Fairness metric approaches threshold → accelerate monitoring

REPORTING CADENCE
  ├── Model owner: monthly dashboard
  ├── Model risk committee: quarterly summary
  └── Board/audit committee: annual review (Tier 1 and 2)

11.8 The Model Risk Owner

One of the most common model risk management failures is unclear ownership. When a model causes harm, the question "who is responsible?" should have a clear, unambiguous answer. Often it does not.

The Model Risk Owner Role

The model risk owner is the individual who is: - Accountable for the model's fitness for purpose - Responsible for ensuring validation and monitoring are conducted - Empowered to suspend model use if risk thresholds are breached - The escalation point for material model issues

Who the model risk owner is NOT: - The model developer (they built it; they are not neutral on its quality) - IT (infrastructure is not the same as model governance) - A committee (committees diffuse accountability; a named individual concentrates it)

What the model risk owner does: - Reviews ongoing monitoring results monthly - Approves material changes to the model (prompt changes, data changes, threshold changes) - Escalates material findings to the model risk committee - Signs off that the model is operating within its validated parameters - Has the authority (and uses it) to suspend model use if monitoring indicates breach

For AI models specifically, the model risk owner must: - Understand what the model does and what it cannot reliably do - Know the model's validation status and monitoring results - Be able to explain the model's governance to a regulator - Not be the same person as the model developer or the primary business user

What Counts as a Material Change Requiring Re-validation

This is one of the most common model governance gaps: teams make changes to AI systems without triggering the validation process because the change is perceived as "minor." The regulatory expectation is explicit — material changes require re-validation.

MATERIAL CHANGE DEFINITION FOR AI MODELS

Changes that ALWAYS require re-validation (Tier 1):
  ├── Model version change (different model, even minor version bump)
  │     Rationale: Model updates change behavior in ways that are
  │     not predictable without testing
  ├── System prompt change that modifies scope, constraints,
  │     or decision logic
  ├── Knowledge base update that materially changes the
  │     information the model draws on for consequential decisions
  ├── Addition or removal of tools in an agentic model
  ├── Change to the decision threshold (e.g., score cutoff)
  └── Change to the human review requirement (reducing oversight)

Changes that require owner review (re-validate if scope impacts):
  ├── Minor prompt wording changes (behavior verified by eval)
  ├── New training data that does not change model scope
  └── Infrastructure changes that don't affect model behavior

Changes that do NOT require re-validation:
  ├── Infrastructure scaling (more compute, not different behavior)
  └── UI/UX changes with no impact on model inputs or outputs

Governance control: Every change to a Tier 1 or Tier 2 AI model
must be categorized as material or non-material by the model risk owner
before deployment. Material changes require validation sign-off.
Non-material changes require owner sign-off and documentation.
The categorization decision is itself auditable.

Third-Party Model Risk

When an organization uses a vendor's AI model — a commercial LLM API, an AI-powered SaaS feature, a third-party decisioning engine — the organization is still the model risk owner. The vendor's model risk is your model risk if you are using the model to make or inform decisions.

THIRD-PARTY AI MODEL RISK MANAGEMENT

Vendor assessment (before deployment):
  ├── Request the vendor's model documentation:
  │     - What training data was used?
  │     - What validation has been conducted?
  │     - What fairness testing has been done?
  │     - What are the known limitations?
  ├── Contractual requirements:
  │     - Data processing agreement (no training on your data)
  │     - Notification requirement for material model changes
  │     - Performance SLA and remediation process
  │     - Audit rights (can you inspect validation evidence?)
  └── Ongoing monitoring is YOUR responsibility:
        You cannot rely on the vendor's testing alone.
        You must monitor model performance in your context.

Vendor model changes:
  └── A vendor that silently upgrades the model version
        has triggered a material change to your model.
        Require contractual notification for model version changes.
        Your validation must run after any vendor model update.

11.9 Industry-Specific Compliance Considerations

Financial Services (Beyond Banking MRM)

FINRA and securities regulations: AI used in investment recommendations, suitability assessments, or order management is subject to FINRA's supervisory requirements. The "reasonable basis suitability" standard — knowing a product is suitable for at least some customers — and "customer-specific suitability" — knowing it is suitable for this specific customer — apply to AI-generated recommendations as they do to human recommendations.

Insurance: AI used in underwriting and claims decisions is subject to state insurance regulations, which vary widely. Some states have issued specific guidance on AI use (Colorado SB 21-169 on algorithmic discrimination in insurance). Insurance AI must document that it does not use protected class proxies (credit score used as a proxy for race is prohibited in some contexts).

AML/BSA: The 2021 interagency statement on model risk management for BSA/AML systems (SR 21-8) was superseded in April 2026 together with SR 11-7. The principles carry forward under the 2026 guidance for in-scope (non-generative, non-agentic) models. Key governance elements include model validation with performance benchmarking, documentation of detection logic, ongoing monitoring of false positive and false negative rates, and independent validation by staff not involved in model development.

Healthcare

HIPAA and AI: Any AI system that processes Protected Health Information (PHI) must comply with HIPAA's Privacy Rule and Security Rule. This affects: - The choice of model provider (no PHI to consumer AI tools; requires Business Associate Agreement) - Data retention in AI systems (prompts, responses, and audit logs containing PHI must be governed under HIPAA retention and disposal requirements) - Breach notification: unauthorized disclosure of PHI through an AI system constitutes a reportable breach

FDA AI/ML-based Software as a Medical Device (SaMD): AI systems that are used to make clinical decisions — diagnosis, treatment planning, therapeutic recommendations — may qualify as Software as a Medical Device under FDA regulations. SaMD requires premarket approval or clearance, including evidence of clinical validation. FDA finalized guidance on Predetermined Change Control Plans (PCCPs) for AI-enabled device software functions in December 2024. A PCCP lets a manufacturer make pre-authorized model updates within defined performance bounds without a new submission. FDA issued draft lifecycle guidance for AI-enabled device software functions in January 2025; check whether it has been finalized.

Clinical decision support (CDS) carve-out: Not all healthcare AI is a medical device. Clinical Decision Support software that provides information to clinicians who independently review it and make their own clinical decisions is carved out of device regulation. The architecture of human oversight (AI informs, human decides) is the governance pattern that preserves this carve-out.

Attorney-client privilege and AI: When lawyers use AI to draft documents, conduct research, or analyze cases using client-provided information, that information may be shared with the AI provider's infrastructure. Privilege may be at risk if confidential communications are disclosed to third parties. Law firms using AI must: - Use enterprise-grade AI with appropriate data processing agreements - Ensure client data does not go to training datasets - Review bar association guidance (which varies by jurisdiction) on AI use in legal practice

ABA Model Rules: The American Bar Association's Model Rules of Professional Conduct (particularly Rules 1.1, 1.6, and 5.3) apply to lawyer use of AI. Rule 1.1 (competence) requires understanding the benefits and risks of AI tools. Rule 1.6 (confidentiality) requires reasonable measures to prevent unauthorized disclosure of client information.


11.10 Management-System Standards and the EU AI Act Timeline

Sections 11.2–11.8 govern individual models. Regulators and enterprise customers increasingly also ask a different question: does the organization run a repeatable system for governing AI? The international answer is ISO/IEC 42001. The EU's answer is the AI Act, whose dates moved in 2026.

ISO/IEC 42001:2023 — The AI Management System Standard

What it is. ISO/IEC 42001, published in December 2023, is the first international standard for an AI management system (AIMS). It specifies requirements for establishing, implementing, maintaining, and continually improving how an organization governs the AI it develops, provides, or uses. It governs the organization's process. It does not certify that any single AI system is safe or compliant.

It is certifiable. Accredited certification bodies audit organizations against it, the same way ISO 27001 works for information security. ISO/IEC 42006:2025 sets the extra requirements those certification bodies must meet (building on ISO/IEC 17021-1), so that 42001 certificates are issued consistently. Enterprise procurement teams increasingly ask AI vendors for 42001 certification.

How it is built. Clauses 4–10 follow the same harmonized management-system structure as ISO 27001: context, leadership, planning (including AI risk assessment and AI system impact assessment), support, operation, performance evaluation, and improvement. Annex A provides 38 reference controls in nine areas. You select the applicable ones and justify the selection in a Statement of Applicability.

ISO/IEC 42001 ANNEX A CONTROL AREAS — WHAT THE ARCHITECT SUPPLIES

Control area                           │ Architect's evidence
───────────────────────────────────────┼──────────────────────────────────────
A.2  Policies related to AI            │ Policy-as-code rules (Module 10 §10.9)
A.3  Internal organization             │ RACI; model risk owner (§11.8)
A.4  Resources for AI systems          │ AI inventory + agent registry: models,
                                       │ data, tools, compute (§11.3; Module 10)
A.5  Assessing impacts of AI systems   │ Impact assessments (ISO/IEC 42005)
A.6  AI system life cycle              │ Design docs, eval suites, change
                                       │ control, release gates (§11.4, §11.8)
A.7  Data for AI systems               │ Data lineage, provenance, quality checks
A.8  Information for interested parties│ Model cards, disclosures, explanations
A.9  Use of AI systems                 │ Intended-use limits, human oversight,
                                       │ monitoring (§11.7)
A.10 Third-party and customer          │ Vendor model assessments, contracts
     relationships                     │ (§11.8 Third-Party Model Risk)

ISO/IEC 42005:2025 (published May 2025) gives guidance on performing an AI system impact assessment: how a system and its foreseeable uses may affect individuals, groups, and society, documented across the lifecycle. 42001 requires impact assessment. 42005 shows how to do it, and it is a useful template for EU AI Act fundamental-rights impact assessments.

How it relates to the other frameworks.

ISO/IEC 42001 vs. ISO 27001 vs. NIST AI RMF

               │ ISO/IEC 42001        │ ISO 27001           │ NIST AI RMF 1.0
───────────────┼──────────────────────┼─────────────────────┼──────────────────────
Scope          │ AI governance system │ Information security│ AI risk management
               │                      │ management system   │ framework
Certifiable?   │ Yes (third-party)    │ Yes (third-party)   │ No (voluntary)
Structure      │ Clauses 4–10 +       │ Clauses 4–10 +      │ Govern, Map, Measure,
               │ Annex A (38 controls)│ Annex A controls    │ Manage functions
Best use       │ Auditable proof of a │ Security baseline   │ Risk-identification
               │ repeatable AI program│ 42001 builds on     │ method inside the
               │                      │                     │ 42001 risk process

If you already hold ISO 27001, run 42001 as an integrated management system. Reuse document control, internal audit, management review, and corrective-action processes, and add the AI-specific risk assessment, impact assessment, and Annex A controls. Use NIST AI RMF as the risk method inside the 42001 risk process. The two are complementary, not alternatives.

What 42001 does and does not do for the EU AI Act. ISO/IEC 42001 is not a harmonised standard under the AI Act, so certification gives no presumption of conformity. The AI Act regulates each high-risk system as a product, while 42001 certifies an organization's process. The European standard written for the AI Act's quality-management requirement (Article 17) is EN 18286, Quality management system for EU AI Act regulatory purposes. It was reportedly approved by CEN/CENELEC on July 12, 2026 (reported). It gives a presumption of conformity only once it is cited in the Official Journal. Check its status in Appendix G. A 42001 program still produces much of the evidence the Act asks for:

HIGH-LEVEL MAPPING: EU AI ACT OBLIGATION → 42001 EVIDENCE
(approximate; a starting point for gap analysis, not a conformity claim)

AI Act obligation (high-risk)          │ Where 42001 helps
───────────────────────────────────────┼──────────────────────────────────
Art. 9   Risk management system        │ Clause 6 risk assessment; A.5
Art. 10  Data and data governance      │ A.7
Art. 11–12 Technical docs, logging     │ A.6 lifecycle records
Art. 13  Transparency to deployers     │ A.8
Art. 14  Human oversight               │ A.9 (partial; design is system-level)
Art. 15  Accuracy, robustness, cyber   │ A.6 verification + ISO 27001
Art. 17  Quality management system     │ Clauses 4–10 (partial; EN 18286 is
                                       │   the targeted standard)
Art. 26–27 Deployer duties, FRIA       │ A.9, A.10; A.5 + ISO/IEC 42005
Art. 50  Transparency (all systems)    │ A.8 (disclosure, content marking)

The EU AI Act Timeline After the Digital Omnibus

The Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on July 24, 2026 and entered into force on July 27, 2026. It amends the AI Act (Regulation (EU) 2024/1689). The main change is a delay to high-risk obligations. Transparency obligations were not delayed.

EU AI ACT — APPLICATION DATES (as of October 2026)

Feb 2, 2025  ● Prohibited practices apply (enforced)
Aug 2, 2025  ● General-purpose AI model obligations apply
Jul 27, 2026 ● Digital Omnibus (Reg. 2026/1744) in force
Aug 2, 2026  ● Article 50 transparency obligations apply — STILL ON
             │  ORIGINAL DATE: tell people they are interacting with AI;
             │  mark synthetic content; disclose deepfakes
Dec 2, 2026  ● Narrow grace ends: content-marking (watermarking) duties
             │  for systems already on the market before Aug 2, 2026
             ● New prohibitions apply: AI generating non-consensual
             │  intimate imagery and child sexual abuse material
Dec 2, 2027  ● Annex III stand-alone high-risk obligations apply
             │  (was Aug 2, 2026): employment, education, credit
             │  scoring, essential services, etc.
Aug 2, 2028  ● Annex I product-embedded high-risk obligations apply
             │  (AI in regulated products: machinery, medical devices…)

What this means for architects: - Do not read the delay as a pause. Article 50 already applies. Any customer-facing chatbot, voice agent, or content generator in scope needs AI-interaction disclosure and synthetic-content marking now. Systems deployed before August 2, 2026 have only until December 2, 2026 to add marking. - Use the extra time for Annex III systems to build evidence, not to defer design. Credit scoring and employment systems need risk management, data governance, logging, human oversight, and conformity assessment designed in. Retrofitting logging and oversight into a deployed system costs far more than building them in. - Tag every inventory entry with its applicable date (§11.3 schema, regulatory_applicability), so a dashboard can show which systems face which deadline. - Expect more change. Harmonised standards and Commission guidelines are due before December 2027. Recheck Appendix G quarterly.


11.11 The Compliance Architecture Decision Matrix

When building AI for regulated use cases, use this decision matrix to determine the governance requirements:

COMPLIANCE ARCHITECTURE DECISION MATRIX

Question 1: Does this AI inform or make decisions about individuals?
  NO → Lighter governance (internal operational use)
  YES → Continue to Question 2

Question 2: Could an adverse outcome affect a protected class?
  NO → Standard model risk management
  YES → Fairness monitoring required (Section 11.6)

Question 3: Is this a consequential decision (financial, health, employment)?
  NO → Tier 2 or Tier 3 governance
  YES → Tier 1 governance required

Question 4: Does the decision require an explanation to the individual?
  NO → Standard documentation
  YES → Explainability architecture required (Section 11.5)
         Adverse action notice capability required

Question 5: Is this a "model" under the applicable regulatory definition?
  NO → Operational governance only
  YES → Full MRM framework:
         ├── Model inventory registration
         ├── Validation before deployment
         ├── Ongoing monitoring
         ├── Model risk owner assignment
         └── Periodic review reporting

Question 6: Does it qualify as an AI Act high-risk system? (EU)
  NO → Limited or minimal risk requirements
  YES → EU AI Act compliance requirements (§11.10; Module 9)
         Annex III stand-alone: from Dec 2, 2027
         Annex I product-embedded: from Aug 2, 2028
  ALSO: Article 50 transparency (disclose AI interaction, mark
         synthetic content) applies from Aug 2, 2026 to ALL
         in-scope systems, high-risk or not

Question 7: Does it qualify as FDA SaMD? (healthcare)
  NO → HIPAA compliance only (if PHI)
  YES → FDA regulatory pathway required

11.12 The Compliance Governance Checklist

Model inventory - [ ] All AI models used in consequential decisions inventoried? - [ ] Each model tiered by materiality (Tier 1/2/3)? - [ ] Model risk owner assigned for each Tier 1 and 2 model? - [ ] AI-specific inventory fields completed (model version pinned, prompt version, eval suite)?

Model validation - [ ] Validation completed before production deployment? - [ ] Validation covers AI-specific elements: prompt testing, behavioral boundary testing, output calibration? - [ ] Independent validation (separate from model development team) for Tier 1? - [ ] Validation findings documented and tracked to resolution?

Ongoing monitoring - [ ] Monitoring dashboards active for Tier 1 and 2 models? - [ ] Alert thresholds configured and tested? - [ ] Fairness monitoring active for models affecting protected classes? - [ ] Model version and prompt version change alerts active? - [ ] Quarterly monitoring review scheduled?

Management system and EU AI Act - [ ] AI management system scope defined (ISO/IEC 42001 or equivalent), even if not seeking certification? - [ ] AI system impact assessments performed for systems affecting individuals (ISO/IEC 42005 as a template)? - [ ] Each system tagged with its EU AI Act deadline: Article 50 (Aug 2, 2026), Annex III (Dec 2, 2027), Annex I (Aug 2, 2028)? - [ ] Article 50 disclosure and content-marking in place for user-facing generative systems? - [ ] Regulatory dates rechecked against Appendix G each quarter?

Explainability and adverse action - [ ] For credit/lending AI: adverse action reason generation capability built and validated? - [ ] Explanation accuracy validated (does the explanation reflect the actual decision factors)? - [ ] Human review of AI-generated explanations before regulatory use?

Governance structure - [ ] Model risk committee with appropriate seniority? - [ ] Board-level reporting for Tier 1 models? - [ ] Model change control process documented and followed? - [ ] Audit trail for all model changes, validations, and monitoring events?


EXERCISE — Model Tier Assignment: Take three AI systems in your organization (or design three representative ones): (1) an AI assistant that helps employees draft internal emails, (2) an AI system that recommends product offers to retail banking customers, (3) an AI-powered fraud detection system that can freeze customer accounts. Assign a tier to each using the framework from Section 11.3. For the highest-tier system, complete the model inventory record schema. Identify the three most important governance gaps that would need to be closed before this model would pass a regulatory examination.

PONDER — The Model Risk Owner Question: For the most consequential AI system your organization operates: who is the model risk owner? Can they explain the model's governance, its validation status, and its monitoring results? Do they have the authority to suspend the model? If the answer to any of these is "no" or "I'm not sure" — that is the governance gap that regulators will find first.

WORKSHOP — Explainability Architecture for a Credit Decision: A bank wants to use an LLM to analyze customer narratives (support conversations, application notes) and incorporate that analysis into a loan decision recommendation. Design the hybrid explainability architecture from Section 11.5 for this use case: define what the AI extracts (Step 1), what the documented rule applies (Step 2), and what the plain-language explanation includes (Step 3). Identify the regulatory risks and the architectural controls that address each.

WORKSHOP — Disparate Impact Assessment: Design the ongoing monitoring architecture for a consumer lending AI model at a bank. Define: the protected class proxies available in the data, the metrics computed monthly, the alert thresholds, the investigation process when a threshold is breached, the reporting chain, and the documentation that would be provided to a regulator if asked to demonstrate fair lending compliance.


Next: Module 12 — AI Observability, Evals & Production Health