MODULE 5 — AI Data Architecture¶
5.1 The Shift That Changes Everything¶
Traditional data architecture is built around structured data: rows, columns, schemas, relational integrity. ETL pipelines move structured data from operational systems to analytical systems. The schema is defined upfront. Data quality means rows conform to the schema.
AI systems consume primarily unstructured data: documents, emails, contracts, support tickets, transcripts, images, PDFs, code, web pages. Roughly 80% of enterprise data is unstructured and has historically been inaccessible to analytical systems because there was no schema to query against. AI changes this — LLMs can process natural language and extract meaning from unstructured content. But this capability is only as good as the pipeline that delivers that unstructured content to the AI in a clean, current, properly governed form.
Most enterprise AI failures trace back not to the model but to the data that reaches it: stale documents, PII that should have been stripped, inconsistently formatted content that confuses the chunker, no lineage tracking, no version management, no quality validation. The AI data architecture is the layer that prevents these failures.
This module covers the architectural decisions that make the difference between an AI system that works reliably in production and one that works in demos.
5.2 The AI Data Stack vs. the Traditional Data Stack¶
Traditional data stack:
Operational Systems → ETL → Data Warehouse → BI / Analytics
(structured) (transform) (structured) (SQL queries)
AI data stack:
Operational Systems ─────────────────────────────────────────┐
(structured) │
▼
Documents, PDFs, Emails ──► [Document AI Pipeline] [AI Data Layer]
(unstructured) (parse, extract, chunk) │
│
Databases, APIs ──────────► [Structured Extraction] [Vector Store]
(structured) (feature derivation) [Knowledge Graph]
[Feature Store]
External Feeds, Web ──────► [Web/API Ingestion] │
(semi-structured) (normalize, validate) ▼
[AI System]
(RAG / Agent / LLM)
Key differences:
The AI data layer is a first-class architectural component. In the traditional stack, the data warehouse is the central store. In the AI stack, the AI data layer — vector stores, knowledge bases, feature stores — is its own infrastructure tier with its own lifecycle management, quality requirements, and operational concerns. It cannot be an afterthought bolted onto the operational data infrastructure.
Three distinct AI data stores — often conflated: - Vector store: Stores document embeddings for semantic retrieval (RAG). Weaviate, Qdrant, pgvector. - Feature store: Stores pre-computed structured features about entities (users, products, accounts) that are injected into AI context at query time. Feast, Tecton, Hopsworks. Not for retrieval — for personalization and decision context. - Embedding store: Sometimes used interchangeably with vector store, but can also refer to a store of entity embeddings (user embeddings, product embeddings) used for recommendation or similarity tasks — distinct from document chunk embeddings used for RAG. Knowing which of these three you need is an architectural decision, not an implementation detail.
Unstructured data is a primary input. Traditional ETL discards or ignores unstructured fields. AI data pipelines must process them — which requires parsing (understanding document structure), extraction (pulling meaning from content), chunking (segmenting for retrieval), and embedding (converting to vector representations).
Multi-modal data requires pipeline extensions. Images, audio, and video require additional pipeline stages before they reach the AI data layer. Images embedded in documents: extract and caption them (a vision-capable model such as Claude, GPT, or Gemini) so they become searchable text. Audio recordings: transcribe (Whisper, AWS Transcribe) before chunking. Video: extract transcripts and key frames separately. Each modality extends the ingestion pipeline and has its own cost, latency, and quality implications. Design the pipeline for text first; extend to other modalities when the use case genuinely requires it.
Data freshness requirements are different. A BI dashboard can tolerate T+1 day data. An AI assistant that gives customers wrong information because its knowledge base is 24 hours behind may give wrong answers and create real liability. Data freshness in AI systems must be evaluated against the consequence of staleness, not just the convenience of real-time.
Quality is behavioral, not structural. Traditional data quality: does the row conform to the schema? AI data quality: does the content produce accurate, grounded AI responses? A document can be perfectly well-formed (correct format, valid fields) and still produce poor AI outputs because the content is ambiguous, the context is missing, or the information is outdated. AI data quality requires behavioral evaluation, not just schema validation.
5.3 The Document AI Pipeline¶
For most enterprise AI systems, the document ingestion pipeline is the most consequential data architecture decision. The quality of what comes out of this pipeline directly determines the quality of AI responses.
Document Types and Their Challenges¶
Different document types require different parsing strategies. A uniform approach applied to all documents degrades quality for documents it doesn't match.
DOCUMENT TYPE → PRIMARY CHALLENGE → APPROPRIATE TOOL/APPROACH
Native PDF (text-based)
Challenge: Complex multi-column layouts, headers, footers,
page numbers interrupting text flow
Approach: PDF parsing with layout analysis
Tools: Unstructured.io, LlamaParse, PyMuPDF
Scanned PDF (image-based)
Challenge: No extractable text — requires OCR
Variable scan quality, skewed pages, handwriting
Approach: OCR + layout reconstruction
Tools: AWS Textract, Azure Document Intelligence,
Google Document AI (handles tables and forms well)
Microsoft Word (.docx)
Challenge: Tracked changes, comments, embedded objects,
complex table structures
Approach: Direct DOCX parsing preserving structure
Tools: python-docx, Unstructured.io
Spreadsheets (.xlsx, .csv)
Challenge: Tabular data loses meaning when flattened to text
Column headers must accompany every row
Approach: Row-by-row serialization with column context
Tools: OpenPyXL, pandas with AI-aware serialization
HTML / Web content
Challenge: Navigation, ads, boilerplate dilute content
JavaScript-rendered content not in raw HTML
Approach: Content extraction (remove boilerplate)
Tools: Trafilatura, Jina Reader, Firecrawl
Emails
Challenge: Threading, quoted replies, signatures,
mixed HTML/plain text, attachments
Approach: Thread decomposition, deduplication of quoted text,
separate processing of attachments
Tools: Custom parsers + Unstructured.io for attachments
Code / Technical documentation
Challenge: Code blocks must not be treated as prose
Function signatures linked to descriptions
Approach: Structure-aware parsing, AST-level chunking for code
Tools: Tree-sitter for code structure, custom parsers
The Document Ingestion Pipeline¶
┌──────────────────────────────────────────────────────────────────┐
│ DOCUMENT INGESTION PIPELINE │
│ │
│ STAGE 1: INTAKE │
│ ├── Source detection: file type, encoding, language │
│ ├── Duplicate detection: hash-based deduplication │
│ ├── Source metadata extraction: │
│ │ doc_id, source_system, author, created_date, │
│ │ modified_date, classification, owner_team │
│ └── Route by document type to appropriate parser │
│ │
│ STAGE 2: PARSING │
│ ├── Document-type-appropriate parser (see table above) │
│ ├── Structure extraction: title, sections, headings, tables │
│ ├── Content extraction: body text per section │
│ ├── Metadata extraction: page numbers, footnotes, captions │
│ └── Quality gate: minimum content length, encoding validation │
│ │
│ STAGE 3: PII & CLASSIFICATION │
│ ├── PII detection: names, SSN, account numbers, emails, │
│ │ phone numbers, addresses, health identifiers │
│ │ Tools: Presidio (Microsoft), AWS Comprehend │
│ ├── Classification enforcement: │
│ │ If doc classification = CONFIDENTIAL: │
│ │ Route to restricted access vector namespace │
│ │ Tag all chunks with access_level = restricted │
│ │ If doc contains PII: │
│ │ Pseudonymize or redact before embedding │
│ │ Log: doc_id, PII types detected, action taken │
│ └── Compliance hold check: is this document under legal hold? │
│ If yes: ingest as read-only, flag for review │
│ │
│ STAGE 4: CHUNKING │
│ ├── Apply document-type-appropriate chunking strategy │
│ ├── Generate chunk metadata: │
│ │ chunk_id, parent_doc_id, chunk_index, section_title, │
│ │ start_char, end_char, chunk_type (text/table/code) │
│ └── Quality validation: chunk completeness, no orphan fragments │
│ │
│ STAGE 5: CONTEXTUAL ENRICHMENT │
│ ├── Generate context prefix for each chunk (Contextual │
│ │ Retrieval pattern — Module 4) │
│ ├── Extract named entities and store as metadata │
│ └── Add parent section summary to chunk metadata │
│ │
│ STAGE 6: EMBEDDING │
│ ├── Embed contextualized chunk text │
│ ├── Store: vector + full metadata + original text │
│ ├── Tag: embedding_model_name, embedding_model_version │
│ └── Status: active │
│ │
│ STAGE 7: LIFECYCLE REGISTRATION │
│ ├── Register in document registry: doc_id, version, status, │
│ │ owner, review_date, supersedes (prior version if update) │
│ ├── If update: mark prior version chunks status = archived │
│ └── Notify downstream systems: knowledge base updated │
└──────────────────────────────────────────────────────────────────┘
The Document Registry: The Most Underbuilt Component¶
Every enterprise RAG system that handles document updates needs a document registry — a dedicated store that tracks the lifecycle of every document in the knowledge base. Most teams build the vector store but not the registry, and this is where the stale knowledge problem (Failure Mode 2 from Module 4) lives.
DOCUMENT REGISTRY SCHEMA
{
doc_id: uuid,
canonical_name: "Fee Schedule 2024",
source_system: "Confluence | SharePoint | S3 | manual_upload",
source_uri: "s3://docs-bucket/policies/fee-schedule-2024-v3.pdf",
content_hash: sha256 of document content,
version: "3.0",
effective_date: "2024-03-01",
supersedes: "doc_id_of_prior_version",
superseded_by: null (null = currently active version),
status: "active | archived | retracted | pending_review",
owner_team: "Product Policy",
owner_contact: "policy-team@company.com",
review_date: "2024-09-01",
review_overdue: false,
classification: "internal | confidential | restricted | public",
jurisdictions: ["US", "EU"],
product_lines: ["retail_banking", "mortgage"],
chunk_count: 47,
embedding_model: "text-embedding-3-large",
embedding_model_version: "002",
last_ingested: "2024-03-02T09:15:00Z",
access_roles: ["customer_service", "compliance"],
audit_log: [
{"event": "created", "timestamp": "...", "actor": "..."},
{"event": "superseded_v2", "timestamp": "...", "actor": "..."},
{"event": "ingested_v3", "timestamp": "...", "actor": "..."}
]
}
Why the registry is non-negotiable: - Without it, there is no automated way to know when a document update requires re-ingestion - Without it, archival of old chunks cannot be automated — someone must remember to do it manually - Without it, the review_date monitoring that prevents stale knowledge is impossible - Without it, access control at the document level cannot be enforced - Without it, compliance audits ("show me every document that was in the knowledge base on March 1st") cannot be answered
Knowledge base snapshots for point-in-time compliance. Regulated environments require the ability to reconstruct what the AI knew at any historical point in time. The registry, combined with archived (not deleted) chunks and the response lineage record (Section 5.6), provides this capability. The key design principle: never delete historical data from the AI data layer — only change its status. Deletion makes point-in-time reconstruction impossible. Archival preserves the audit capability while excluding stale content from active retrieval.
5.4 Data Quality for AI: Different From What You Know¶
Traditional data quality frameworks (completeness, accuracy, consistency, timeliness) are necessary but insufficient for AI data. AI systems have additional quality requirements that traditional frameworks don't address.
The AI-Specific Data Quality Dimensions¶
Semantic coherence. Does the content make sense as a standalone unit? A chunk that contains a half-finished sentence or a table without column headers is structurally complete but semantically broken. Traditional quality checks do not catch this.
Retrieval fitness. Will this content retrieve correctly for the queries it should answer? A policy document written in dense legal language may contain the correct information but be retrieved poorly because the query language (natural, customer-facing) and the document language (formal, legal) are too dissimilar. Evaluating retrieval fitness requires running representative queries and checking whether the relevant content surfaces.
Contextual completeness. Does the chunk provide sufficient context to answer questions without requiring knowledge of adjacent chunks? A chunk that says "as described in the previous section" is contextually incomplete — the reference cannot be resolved in isolation. Contextual Retrieval (Module 4) is the architectural fix; detecting contextually incomplete chunks requires quality validation at ingestion time.
Temporal validity. Is the content currently accurate? A technically well-formed document with outdated content will produce wrong AI answers. Temporal validity requires tracking the effective date, the review date, and any source system changes that would make the content stale.
Density appropriateness. Is the information density in the document appropriate for AI consumption? Very sparse content (one sentence per page) or extremely dense content (tables with hundreds of data points compressed into minimal text) both degrade retrieval and generation quality.
Data Quality Validation Gates¶
QUALITY VALIDATION IN THE INGESTION PIPELINE
Gate 1: Structural validation (automated)
✓ File successfully parsed with no errors
✓ Minimum content length met (e.g., > 100 tokens)
✓ Character encoding valid (no corrupted characters)
✓ No empty sections after parsing
FAIL → Reject document, notify owner with specific error
Gate 2: Content quality checks (automated + LLM-assisted)
✓ No broken references ("see page X", "as above") in standalone chunks
✓ Tables have column headers in every chunk derived from them
✓ Code blocks are not split mid-function
✓ Language detection: content is in expected language(s)
FAIL → Flag for human review, do not auto-ingest
Gate 3: Compliance checks (automated)
✓ PII detection scan completed
✓ Classification tag applied
✓ Access control mapping validated
✓ Not under legal hold (check registry)
FAIL → Hold for compliance review, do not ingest
Gate 4: Freshness validation (automated + human trigger)
✓ effective_date is not in the future
✓ Document does not have a known superseding version
✓ Source system timestamp matches expected update frequency
WARN → Flag for human review, ingest with staleness warning tag
5.5 Synthetic Data: When and Why¶
Synthetic data — AI-generated data that approximates real data — is increasingly important in AI system development and has specific architectural applications.
Where Synthetic Data Helps¶
Evaluation dataset construction. Building a good RAG evaluation dataset (Module 4, RAGAS section) requires query/answer pairs. Getting domain experts to write 200 test cases is expensive and slow. Using an LLM to generate diverse test queries from existing documents accelerates evaluation dataset creation significantly. The process: feed document chunks to a capable LLM, instruct it to generate realistic user queries that the chunk should answer, have a domain expert review a sample (not all) for quality. The resulting dataset is imperfect but orders of magnitude faster than pure human generation.
Fine-tuning dataset augmentation. If you have 500 labeled examples for a classification task and need 5,000 to fine-tune a model, synthetic data generation — using an LLM to create varied paraphrases and examples of each class — can augment the dataset. Critical: synthetic data for fine-tuning must be reviewed for quality and must not introduce systematic biases that weren't in the original data.
Adversarial test case generation. Generating injection attempts, boundary condition inputs, and adversarial prompts for security testing. An LLM generating 1,000 variants of injection patterns is far more comprehensive than a human writing 20 examples.
Privacy-preserving development. When production data contains PII that cannot be used in development environments, synthetic data that matches the statistical distribution and structure of real data without containing real individuals' information enables development and testing without compliance exposure.
Where Synthetic Data Fails¶
As a substitute for real data in production RAG. AI-generated documents used as the knowledge base for a production RAG system will produce AI-generated answers — plausible, well-formatted, potentially wrong. Synthetic documents are useful for testing the pipeline but must never be the source of truth in a customer-facing system.
Without quality validation. LLMs generate synthetic data that sounds correct. They also generate synthetic data that contains subtle errors, anachronisms, or contradictions that are difficult to detect without expert review. Every synthetic dataset requires validation before use — the percentage of the dataset that requires human review is an architectural decision based on the risk of the downstream use.
For compliance evidence. Synthetic data used to demonstrate compliance with data handling policies is a red flag. Regulators expect real process evidence, not generated examples.
Synthetic Data Architecture¶
SYNTHETIC EVALUATION DATASET PIPELINE
Source documents ──► [LLM: generate N queries per chunk]
│
▼
[Query/Answer pairs]
│
[Expert sample review: 10% spot check]
│
┌──────┴──────────────┐
│ │
PASS FAIL
│ Adjust generation
▼ prompt, regenerate
[Eval dataset]
- Store with generation metadata
- Track: generated_by, reviewed_by, review_date
- Version controlled alongside prompts
- Not mixed with human-authored test cases
(keep separate for quality attribution)
5.6 Data Lineage for AI: The Audit Requirement¶
Traditional data lineage tracks the transformation path of data from source to analytical store. AI data lineage must additionally track: which documents influenced which AI outputs, which document version was active when a response was generated, and what the full context was for any given AI decision.
This is not a nice-to-have. In regulated industries — financial services, healthcare, legal — the ability to trace an AI output back to its source data is a compliance requirement.
The AI Lineage Data Model¶
AI RESPONSE LINEAGE RECORD
{
response_id: uuid,
timestamp: "2024-03-15T14:23:11Z",
query: {
query_id: "q-abc123",
query_hash: sha256, // not the raw query (privacy)
user_id: "u-pseudo-456", // pseudonymized
session_id: "s-789"
},
retrieval: {
retrieval_strategy: "hybrid_bm25_dense + rerank",
candidates_retrieved: 20,
chunks_after_rerank: 5,
chunks_used: [
{
chunk_id: "chunk-001",
doc_id: "doc-fee-schedule-v3",
doc_version: "3.0",
doc_effective_date: "2024-03-01",
relevance_score: 0.91,
rerank_score: 0.88
},
// ... other chunks
],
confidence_score: 0.88,
confidence_gate_triggered: false
},
generation: {
model: "claude-opus-4-20250514",
model_version: "claude-opus-4-20250514",
prompt_template_id: "customer-support-v12",
prompt_template_version: "12.3",
response_hash: sha256,
tokens_prompt: 2847,
tokens_completion: 184,
cost_usd: 0.0089,
latency_ms: 1240
},
quality: {
confidence_score: 0.88,
citations_included: true,
escalation_triggered: false,
human_reviewed: false,
human_reviewed_by: null
},
stored_at: "audit-log-s3://ai-audit/2024/03/15/"
retention_until: "2031-03-15" // 7-year compliance retention
}
Why every field matters:
- doc_version and doc_effective_date: Allows reconstruction of exactly what the AI knew when it answered
- prompt_template_version: Allows identification of whether a prompt change caused a behavioral shift
- model_version: Required to correlate behavioral changes with model updates
- cost_usd: Cost attribution per interaction for FinOps
- retention_until: Compliance retention enforcement
Data Lineage vs. Audit Log¶
These are related but different:
Audit log: Append-only record of every AI interaction. Designed for compliance audits. Cannot be modified. Retained for the compliance period. Queried rarely but must be available on demand.
Data lineage: The graph of data transformations — from source document through ingestion pipeline to vector store to retrieval to response. Designed for debugging and root cause analysis. Allows you to trace "why did the AI say X?" back to specific document versions and prompt templates.
Both are required in production AI systems. They should be implemented as separate stores with different access patterns and retention policies.
5.7 Real-Time Data Integration Patterns for AI¶
The choice between batch, near-real-time, and real-time data for AI systems has significant implications for cost, architecture complexity, and the quality of AI responses.
The Freshness-Complexity Trade-off¶
FRESHNESS OPTIONS FOR AI DATA
BATCH (T+hours to T+days)
──────────────────────────────────────────────────────────────
Trigger: Nightly job, weekly job
Mechanism: ETL to flat files → bulk ingestion pipeline
Cost: Low
Complexity: Low
Appropriate for: Reference data that changes infrequently
(product catalogs, general FAQs, static policies)
Not appropriate for: Customer account data, prices, availability,
compliance flags, time-sensitive information
NEAR-REAL-TIME (T+seconds to T+minutes)
──────────────────────────────────────────────────────────────
Trigger: Change Data Capture (CDC) events
Mechanism: Debezium → Kafka → AI data consumer
Cost: Medium
Complexity: Medium-High
Appropriate for: Customer context (profile, account status),
compliance flags, product availability,
operational knowledge base updates
Architecture: See Module 12 (AI in Event-Driven Systems)
in the Integration Patterns reference (Artifact 2)
REAL-TIME (T+milliseconds)
──────────────────────────────────────────────────────────────
Trigger: Every transaction, every state change
Mechanism: Direct API call at query time (tool call, not RAG)
Cost: High (latency + infrastructure)
Complexity: High
Appropriate for: Live prices, real-time availability,
current account balance, live status
Not appropriate for: Historical documents, policy content,
general knowledge
Architecture: This is not a RAG pattern — it is a tool call.
The AI agent calls a live API to get the data
rather than retrieving from a knowledge base.
The critical distinction: RAG vs. Tool Calls for Data Access
RAG is appropriate for knowledge that changes infrequently (hours to months). Tool calls are appropriate for data that changes frequently (seconds to minutes). Confusing these two patterns creates architectures where: - RAG is used for live prices → stale data, wrong answers - Tool calls are used for policy documents → unnecessary API calls, latency overhead, no caching benefit
The decision rule: if the data's correctness depends on it being current at the moment of the AI interaction, use a tool call. If the data is knowledge that has a lifecycle (effective dates, version history, supersession), use RAG.
The Context Window Data Injection Pattern¶
For data that is user-specific and should influence the AI's responses but is not knowledge-base content — customer tier, account status, user preferences, recent interaction history — inject directly into the context window rather than retrieving via RAG.
BAD: Adding customer-specific data to the vector store
Problem: Retrieval brings up OTHER customers' data
Privacy violation. Data isolation failure.
Scale: vector store size grows with every customer
CORRECT: Direct context injection
At query time, the application layer fetches:
- Customer tier and status (from CRM API or cache)
- Relevant account summary (from account service)
- Open support cases (from ticketing system)
Injected into the prompt context as typed fields:
{
customer_tier: "premium",
account_status: "active",
outstanding_balance: "$0",
open_cases: 0
}
The AI uses this context without it ever entering the
vector store. The data stays in the transactional system.
RAG is for knowledge. Context injection is for personalization.
5.8 The Data Mesh Intersection with AI¶
Data mesh — the organizational pattern of domain-owned, self-serve data products — intersects with AI architecture at the question of knowledge ownership. Who owns the AI knowledge base? Who is responsible for keeping it current?
The naive answer: a central AI platform team owns all knowledge. They ingest all documents, maintain the vector store, and are responsible for accuracy. This does not scale. The platform team becomes a bottleneck for every document update across every domain.
The data mesh answer: each domain owns its AI knowledge products the same way it owns its data products. The customer domain owns the customer policy knowledge base. The product domain owns the product knowledge base. The compliance domain owns the regulatory knowledge base.
DATA MESH AI KNOWLEDGE ARCHITECTURE
Retail Banking Domain
├── Owns: [Retail Policy Documents]
├── Publishes: [retail_knowledge_product]
│ - Schema: standardized (chunk, metadata, embeddings)
│ - SLA: updated within 24h of source document change
│ - Quality: min context precision = 0.70 (validated weekly)
│ - Access: retail_agents, customer_support_ai
└── Maintains: ingestion pipeline for their own docs
Compliance Domain
├── Owns: [Regulatory Documents, Policy Mandates]
├── Publishes: [compliance_knowledge_product]
│ - Schema: standardized (same as retail)
│ - SLA: updated within 4h of regulatory change
│ - Quality: min context precision = 0.85 (higher bar)
│ - Access: all_ai_systems (read-only)
└── Maintains: ingestion pipeline with legal review gate
AI Platform Team (enables, does not own)
├── Provides: standard ingestion pipeline library
├── Provides: vector store infrastructure
├── Provides: quality validation framework
├── Defines: knowledge product schema standard
├── Defines: access control model
└── Does NOT own: individual domain knowledge bases
The data contract for AI knowledge products:
Each domain's knowledge product must publish a contract: - What content is in scope (and what is explicitly excluded) - The update SLA: how quickly after a source document changes does the knowledge base reflect it? - The quality SLA: what RAGAS context precision floor is maintained? - The access model: which AI systems can read this knowledge - The review process: how are stale documents identified and removed
Without this contract, the AI platform team has no recourse when a domain's stale knowledge causes incorrect AI outputs. With it, the domain team has accountability for the quality of their knowledge product.
5.9 PII in the AI Data Pipeline: The Architecture of Prevention¶
PII handling in AI data pipelines is more complex than traditional systems because PII can appear in unexpected places (embedded in documents, in conversation history, in retrieved context) and because the LLM itself can surface PII from context in unexpected ways.
The PII Architecture Principle¶
PII should be stripped, pseudonymized, or access-controlled before it enters any AI processing layer. Not after. Not by asking the LLM to ignore it. Before.
Why prompting the LLM to "not use" PII is insufficient: - The PII is in the context window, which means it was sent to the LLM API (already a potential data residency issue) - LLMs do not reliably "ignore" content in their context — they may reference it in unexpected ways - The audit log must not contain PII — but if the prompt contains PII, the audit log entry does too
PII Handling Architecture by Data Type¶
Document content PII:
At ingestion time:
1. Run Presidio (or equivalent) on raw document text
2. For customer-facing knowledge bases: redact or pseudonymize detected PII
"John Smith's account number 12345678" →
"[PERSON]'s account number [ACCOUNT_NUMBER]"
3. For internal-only bases: tag chunks containing PII with access restriction
Chunks tagged requires_pii_clearance = true
Not retrievable by AI systems without appropriate access level
4. Log: document_id, PII types detected, action taken, timestamp
Conversation history PII:
User messages typically contain PII (names, account numbers, addresses).
Do not store raw conversation history in any AI context store.
Pattern:
Raw user message → PII detection → Pseudonymized summary
"My account number is 12345678 and I can't log in" →
{intent: "account_access_issue", account_identifier: "[REDACTED]"}
The AI receives the structured summary, not the raw message.
The raw message is stored in the transactional system with
appropriate access controls, not in the AI layer.
Tool call data:
When an agent fetches customer data via a tool call, the response
may contain PII fields. The tool call response should:
- Return only fields the AI task requires (minimum necessary)
- Pseudonymize identifiers the AI doesn't need in raw form
- Not be stored in the audit log verbatim
Tool: get_account_summary(account_id)
Returns to AI: {tier: "premium", status: "active", has_open_cases: false}
NOT: {account_holder_name: "John Smith", ssn: "...", address: "..."}
Tools: Microsoft Presidio (open source, extensible entity detection), AWS Comprehend (managed PII detection), Azure AI Language (PII extraction with confidence scores), Google Cloud DLP (enterprise data loss prevention)
5.10 Data Readiness Assessment Framework¶
Before building an AI system that depends on a specific data source, run a data readiness assessment. Teams skip this and build elaborate AI architectures on data foundations that are not ready — and discover it six months in when the AI is producing poor results that trace back to data quality, not model quality.
The Data Readiness Checklist¶
Availability - [ ] Is the data accessible via a stable API or data pipeline? - [ ] Is the data available in a format that can be parsed (not locked in proprietary binary formats)? - [ ] Is access controlled — can the AI system be granted appropriate permissions? - [ ] Is the data available in the environments needed (dev, staging, production)?
Quality - [ ] Has the data been profiled? What is the completeness rate for key fields? - [ ] Is there a known error rate? What causes errors? - [ ] Is the content semantically coherent (not generated by another broken process)? - [ ] For documents: are they human-readable, consistently formatted, and complete?
Freshness - [ ] What is the current update frequency of the source data? - [ ] Is that update frequency sufficient for the AI use case? - [ ] Is there a mechanism to detect when the source data has changed? - [ ] What is the acceptable staleness window, and does the current update cycle fit within it?
Governance - [ ] Is there a clear data owner for this source? - [ ] Is there a process for the owner to notify AI systems of material changes? - [ ] What is the retention policy? Will this data still be available in 3 years? - [ ] Are there regulatory restrictions on using this data for AI purposes?
PII and Classification - [ ] Has the data been scanned for PII? What types were found? - [ ] Is there an approved handling procedure for the PII types present? - [ ] What is the data classification? Is AI use permitted under that classification? - [ ] Can the required PII handling be implemented without breaking the AI use case?
Volume and Scale - [ ] What is the current volume? What is the projected volume in 12 months? - [ ] Does the volume require batch ingestion, streaming, or both? - [ ] What is the infrastructure cost of ingesting and storing this volume in the AI data layer? - [ ] Are there volume spikes that would overwhelm the ingestion pipeline?
A failing data readiness assessment is architectural information. It tells you what must be fixed in the data infrastructure before the AI use case can succeed. A team that proceeds with a failing assessment and blames the AI model when it produces poor results has misidentified the root cause.
EXERCISE — Document Registry Design: Your organization has 3,000 policy documents spread across SharePoint, Confluence, and a legacy document management system. They are updated by 12 different teams on no consistent schedule. Design the document registry schema and the ingestion pipeline trigger mechanism. How does the AI system know when a document has been updated? Who is responsible for triggering re-ingestion? What happens if a team updates a document without triggering the pipeline?
PONDER — The Stale Data Question: For an AI system you are currently closest to: what is the oldest piece of information in its knowledge base? How do you know? What would it take to find out? Is there any monitoring in place that would alert you if content in the knowledge base became stale? If not, what is the risk of a user receiving an answer based on outdated information, and who is accountable for that?
WORKSHOP — Data Readiness Assessment: Choose a data source that a team in your organization is planning to use for an AI system. Run the data readiness checklist from Section 5.10. For each failing item: quantify the risk (what goes wrong in the AI system if this is not addressed), estimate the effort to fix it, and recommend whether to fix it before building or accept the risk and mitigate in the AI layer. Present the assessment to the team.
WORKSHOP — PII Architecture Design: Map the data flows for an AI customer support system. Identify every point where PII could enter the AI processing layer: user messages, retrieved documents, tool call responses, conversation history, audit logs. For each entry point: design the PII handling — is it redacted, pseudonymized, access-restricted, or acceptable to pass through? Document the decisions as an architectural decision record.
Next: Module 6 — Agentic Systems Architecture