Skip to content

MODULE 28 — AI Technical Debt

28.1 A New Category of Debt

Technical debt in traditional software is well understood: shortcuts taken in implementation that reduce maintainability, shortcuts in testing that allow defects to accumulate, architectural decisions made for speed that create long-term coupling problems.

AI systems accumulate all of this — and six additional categories of debt that do not exist in traditional software. AI technical debt is more insidious than traditional technical debt because much of it is invisible until a production failure, a compliance audit, or a model deprecation makes it suddenly, expensively visible.


28.2 The Six Categories of AI Technical Debt

Category 1: Prompt Debt

Definition: System prompts that were written quickly, never reviewed, never tested properly, and have accumulated inconsistencies, redundancies, and outdated instructions over time.

How it accumulates: - Every production incident adds a new instruction to prevent it recurring - Every compliance feedback adds a new constraint - Nobody removes the old instructions that are no longer relevant - Nobody tests whether removing an instruction changes behavior - The prompt grows to 3,000 tokens. Half of it is dead code.

Signs of prompt debt: - System prompt exceeds 1,500 tokens for a simple feature - Nobody can explain why certain instructions are there - Contradictory instructions exist in the same prompt - Some instructions are written in response to long-resolved one-off incidents - No one knows which instructions still affect behavior

Cost: - Every extra token in a prompt costs money on every request (Module 13 — prompt bloat) - Contradictory instructions produce inconsistent outputs - New engineers cannot understand the system's intended behavior from reading the prompt - Changes become risky because nobody knows what else might change

Paying it down: - Prompt audit: review every instruction. Can you state why it's there? Is the reason still valid? - Behavioral testing: can you remove this instruction without failing any eval? If yes, remove it. - Prompt refactoring: rewrite the entire prompt from the current specification. Compare old vs. new on the eval suite. If the new version is cleaner and passes the same evals, ship it.


Category 2: Eval Debt

Definition: The gap between the tests that exist and the tests that are needed to confidently change the system.

How it accumulates: - The initial eval suite has 20 cases, sufficient for launch - The system grows to handle 50 different query types - New features are added without new eval cases - Production incidents occur but don't get added to the eval suite - After 12 months, the 20 original cases test about 5% of actual system behavior

Signs of eval debt: - Changes to the system "feel risky" because nobody is sure what might break - A production incident reveals a failure mode that isn't in any eval case - The eval suite passes but production quality is declining - Teams are reluctant to upgrade the model because "we don't know how it will behave" - No adversarial cases exist in the eval suite

Cost: - Changes cannot be made with confidence - Model upgrades are blocked or done with risk that materializes later - Production failures that were predictable (and testable) occur anyway

Paying it down: - Eval case backlog: every production incident, every near-miss, every "this feels wrong" interaction creates a new eval case - Coverage analysis: map the system's actual query distribution against eval coverage. Which query types have no eval cases? - Adversarial catch-up: run Garak against the current system. Every finding that survives adds an adversarial eval case.


Category 3: Model Debt

Definition: Accumulation of dependency on a specific model version that has been deprecated or is approaching deprecation, with no migration plan.

How it accumulates: - Model pinned at launch for stability - Model version starts appearing in deprecation notices - Migration requires re-running eval suite and validating behavior - This work competes with feature development and never gets prioritized - The deprecation deadline arrives. Emergency migration.

Signs of model debt: - Model version in use has a published deprecation date that is approaching - No one has run the eval suite against the replacement model - Teams are uncertain whether the new model will behave the same way - The system uses provider-specific features that aren't available on the replacement

Cost: - Emergency migrations are expensive (time pressure, inadequate testing, incident risk) - Provider-imposed timelines create externally-driven architectural work - Behavioral changes in new models create unexpected production issues

Paying it down: - Quarterly model version review: check deprecation timelines for all pinned versions - Migration eval policy: run the eval suite against the next major model version every quarter, whether or not migration is planned. Know what will break before you have to migrate. - Model portability design: ensure prompt behavior is testable against a model switch. Any prompt that only works on one model is fragile.


Category 4: Knowledge Base Debt

Definition: Stale, inconsistent, or ungoverned content in the RAG knowledge base that degrades AI quality over time.

How it accumulates: - Documents are added over time without a lifecycle management process - Old documents are never archived when superseded - New documents are added but existing related documents aren't updated - Access control on knowledge base documents drifts from access control on source systems - Nobody reviews the knowledge base for quality; it's assumed to be "the documents"

Signs of knowledge base debt: - The AI answers questions about policy that was changed months ago - Users report conflicting answers from the AI for similar questions - The knowledge base contains documents from teams that no longer exist - There is no process for a document owner to know their document is in the AI's knowledge base - Knowledge base size grows unboundedly; nobody has ever removed anything

Cost: - Quality degradation that appears as "the AI got worse" when actually the data got worse - Regulatory risk: AI citing outdated policy in regulated environments - Security risk: sensitive documents from departed teams may not be access-controlled correctly

Paying it down: - Quarterly knowledge base review: check the review_date on every document. Any past review_date is a debt item. - Source system audit: for each document in the knowledge base, verify the source document still exists and hasn't been superseded - Ownership audit: for each document, is there a named owner who is accountable for its accuracy?


Category 5: Governance Debt

Definition: AI systems operating in production without the governance controls that were planned (or should have been planned) — missing audit trails, missing model risk documentation, missing human oversight paths, missing compliance review.

How it accumulates: - "We'll add governance after we prove the concept" — the concept gets proven and governance gets deferred again - Compliance reviews were skipped during a fast-paced launch - The AI inventory was never updated after new features went live - Model risk owners were never assigned - Audit logs were designed to be "added later" and later never came

Signs of governance debt: - Cannot answer the question "which AI systems are we operating and who owns them?" - A compliance audit would find AI systems with no audit trail - Human oversight paths exist in design documents but are not implemented in production - Prompt changes are made without review or documentation

Cost: - Regulatory exposure: examination findings, remediation obligations - Incident response failures: no audit trail means no root cause analysis - Business risk: decisions made by AI systems with no human accountable for them

Paying it down: - AI inventory audit: build or update the complete inventory - Governance gap assessment: for each system, apply the pre-deployment checklist (Appendix C). Every unchecked item is a governance debt item. - Remediation sequencing: prioritize by risk (Tier 1 models, regulated use cases) before addressing lower-risk systems


Category 6: Data Pipeline Debt

Definition: AI data pipelines that are fragile, undocumented, not monitored, and dependent on tribal knowledge for operation.

How it accumulates: - Data pipelines were built quickly as part of the AI project - The engineer who built them has moved to another project - Pipeline failures are noticed when AI quality drops, not when the pipeline fails - There are no alerts on the pipeline - Data quality validation is absent; bad data reaches the AI silently

Signs of data pipeline debt: - Nobody knows what happens when the document ingestion pipeline fails - A pipeline failure was discovered because users reported wrong answers, not because the pipeline alerted - The ingestion pipeline depends on manual steps performed by one person - Data quality issues (corrupted documents, wrong versions) reach the AI without detection

Cost: - Silent quality degradation (the AI gets worse but nobody knows why) - Operational fragility (key-person dependency) - Audit failures (cannot trace which document version was active when)

Paying it down: - Pipeline documentation: every pipeline step documented with: purpose, inputs, outputs, failure mode, alert - Monitoring: data pipeline health metrics added to the AI health dashboard - Alert coverage: every pipeline failure alerting within 30 minutes


28.3 Measuring AI Technical Debt

Most technical debt is invisible until it causes a problem. The challenge with AI technical debt is that it is even more invisible than traditional debt — because the system keeps running, just with declining quality.

The AI Debt Assessment

Run this quarterly. Score each item 1 (no debt) to 5 (critical debt).

AI TECHNICAL DEBT ASSESSMENT

PROMPT DEBT
  Prompt age (time since last full review):
    1: < 3 months   3: 6-12 months   5: > 12 months

  Prompt length relative to necessity:
    1: Minimum necessary   3: Somewhat bloated   5: Clearly over-grown

  Prompt test coverage:
    1: Full eval suite   3: Partial   5: No systematic tests

EVAL DEBT
  Eval suite size relative to feature scope:
    1: Comprehensive coverage   3: Partial   5: Minimal or missing

  Adversarial coverage:
    1: Full injection testing   3: Basic   5: None

  Time since last new eval case added:
    1: This month   3: 3-6 months   5: > 6 months

MODEL DEBT
  Deprecation timeline for current model:
    1: > 12 months   3: 3-6 months   5: < 3 months or already deprecated

  Time since eval suite last run against next model version:
    1: Last quarter   3: > 6 months   5: Never

KNOWLEDGE BASE DEBT
  % of documents with review dates in the past:
    1: < 5%   3: 5-25%   5: > 25%

  % of documents with named owners:
    1: > 95%   3: 50-95%   5: < 50%

GOVERNANCE DEBT
  AI inventory completeness:
    1: Complete and current   3: Partially complete   5: Not maintained

  Audit trail completeness:
    1: All required fields   3: Partial   5: Missing or inadequate

DATA PIPELINE DEBT
  Pipeline monitoring coverage:
    1: Full alerting   3: Partial   5: No monitoring

TOTAL SCORE: [sum of all items] / [max possible]
  0.0-0.3: Low debt — maintainable state
  0.3-0.5: Moderate debt — address in next quarter
  0.5-0.7: High debt — prioritize remediation
  0.7+:    Critical debt — remediation before new features

28.4 The Debt Backlog and Prioritization

Not all debt can be paid simultaneously. Prioritization should be driven by risk exposure, not effort.

DEBT PRIORITIZATION FRAMEWORK

Priority 1 — Pay immediately (this quarter):
  ├── Model versions with deprecation in < 90 days
  ├── Governance debt on Tier 1 / regulated AI systems
  ├── Eval debt on customer-facing systems with no adversarial coverage
  └── Knowledge base debt causing active user-facing quality issues

Priority 2 — Address this half-year:
  ├── Prompt debt causing measurable cost overhead (> 20% above minimum)
  ├── Eval debt on important internal systems
  ├── Model versions with deprecation in 3-12 months
  └── Data pipeline debt with no monitoring

Priority 3 — Scheduled for future quarters:
  ├── Prompt refactoring for cleanliness (no functional impact)
  ├── Documentation improvements
  └── Eval expansion for low-risk features

28.5 Communicating AI Technical Debt to Leadership

Technical debt is difficult to fund because it is invisible and its cost is expressed in risk, not in current dollars. AI technical debt is even harder because leadership is less familiar with the failure modes.

The framing that works:

Frame debt as risk, not as work. Not: "We need to refactor our prompts." But: "Our prompt governance gap means we cannot safely change any of our AI features without risking production incidents. This creates a development slowdown of approximately X weeks per significant feature, and a production incident risk that we estimate at Y probability in the next 6 months."

Use the model deprecation as a forcing function:

Model deprecation is the one AI technical debt item with an external deadline that leadership understands. "Our AI systems use a model version that will stop working on [date]. Migration requires [effort] and carries [risk] if rushed. Budget approval by [date] allows an orderly migration; missing that date means emergency work at 3x cost."

The cost of debt vs. the cost of remediation:

For each significant debt item, estimate both the ongoing cost of the debt (quality degradation, operational fragility, risk of incident) and the one-time cost of remediation. When ongoing cost exceeds remediation cost within 6 months, remediation has a positive ROI. This calculation works for technical debt the same way it works for business cases.


EXERCISE — Debt Assessment: Run the AI technical debt assessment from Section 28.3 on the most important AI system in your organization. Score each item honestly. Calculate the total debt score. Identify the top 3 debt items by risk × current effort to remediate. Build a remediation plan: what would be done in the next quarter, the next two quarters, and the following six months?

PONDER — The Invisible Debt: Think about the AI system in your organization that has been running longest without architectural investment. What categories of debt have accumulated? If you had to bet, which category is most likely to cause the next production incident? What would it take to pay down that specific debt?

WORKSHOP — Communicating Debt to Leadership: For a significant AI technical debt item in your organization, build the leadership communication: the framing as risk (not work), the cost of the debt over the next 12 months, the cost of remediation, the ROI calculation, and the ask. Present it as if you had 5 minutes in a leadership meeting. The challenge: make it concrete enough that a non-technical executive can make a funding decision.


Next: Module 29 — AI Architecture Communication and Diagramming