MODULE 29 — AI Architecture Communication & Diagramming¶
29.1 Why Communication Is Architecture Work¶
An architectural decision that cannot be communicated is an architectural decision that will not be implemented consistently. AI architecture has a specific communication challenge: the systems are harder to visualize than traditional software because the interesting complexity is in the behavior of the AI, not in the structure of the code.
A database schema diagram shows you everything you need to know about a database. An AI system diagram that shows only the service topology misses the most important things: what the AI can and cannot do, how it behaves at its limits, and where humans stay in the loop.
This module covers how to diagram AI systems effectively and how to write about AI architecture for different audiences.
29.2 The C4 Model Extended for AI Systems¶
The C4 model (Context → Containers → Components → Code) is the standard architectural visualization framework. For AI systems, it needs extensions at each level to capture AI-specific concerns.
Level 1: System Context Diagram (AI Extension)¶
At the context level, the AI extension adds one requirement: show the human touchpoints explicitly. Traditional context diagrams show users and external systems. AI context diagrams must also show: where humans review AI outputs, where humans can override AI decisions, and where humans receive escalations from the AI.
STANDARD CONTEXT DIAGRAM:
[User] → [System] → [External Systems]
AI-EXTENDED CONTEXT DIAGRAM:
[User] ────────────────────────────────────► [AI System]
│
▼
[Compliance/Audit Team] ◄──── audit log ──── [Audit Store]
[Human Agent] ◄──── escalation ──────────── [AI System]
reviews, overrides
[Model Risk Owner] ◄──── monthly report ─── [Quality Monitoring]
The human touchpoints are as architecturally significant as the
technical integrations. They must be visible at the context level.
Level 2: Container Diagram (AI Extension)¶
At the container level, the AI extension requires showing: - The LLM as an external dependency (not just "AI service") - The knowledge base / vector store as a distinct container - The prompt management system if it exists - The eval infrastructure (even though it's not in the runtime path)
AI-EXTENDED CONTAINER DIAGRAM
┌────────────────────────────────────────────────────────────────┐
│ AI SYSTEM (your organizational boundary) │
│ │
│ ┌─────────────┐ ┌──────────────┐ ┌──────────────────┐ │
│ │ Gateway │ │ Orchestrator│ │ Knowledge Base │ │
│ │ (LiteLLM) │◄───│ (LangGraph) │◄───│ (Weaviate) │ │
│ └──────┬──────┘ └──────────────┘ └──────────────────┘ │
│ │ │ │ │
│ │ ┌────▼────┐ ┌────────▼──────────┐ │
│ │ │ Prompt │ │ Document Registry │ │
│ │ │ Mgmt │ │ (PostgreSQL) │ │
│ │ └─────────┘ └───────────────────┘ │
│ │
│ [Eval Infrastructure] [Observability] [Audit Log] │
│ (offline — not in runtime path but architecturally significant) │
└──────────────────────────────────────┬─────────────────────────┘
│
┌───────▼─────────┐
│ External LLM │
│ (Claude API) │
│ [data leaves │
│ your perimeter]│
└─────────────────┘
The critical annotation: Mark clearly which container boundaries data crosses. This is the PII and data residency concern made visible. Any arrow from inside your boundary to the external LLM is a data flow that requires policy approval.
Level 3: Component Diagram (AI Extension)¶
At the component level for RAG systems, the diagram should show the retrieval pipeline stages explicitly — not just "retrieval component."
RAG PIPELINE COMPONENT DIAGRAM
[Query] → [Query Classifier] → [Query Expander (HyDE)]
│
┌───────▼──────────┐
│ Hybrid Search │
│ ┌─────────────┐ │
│ │Dense Search │ │
│ │(Weaviate) │ │
│ └─────────────┘ │
│ ┌─────────────┐ │
│ │BM25 Search │ │
│ │(Weaviate) │ │
│ └─────────────┘ │
└───────┬──────────┘
│
[Score Fusion (RRF)]
│
[Cross-Encoder Reranker]
│
[Confidence Gate]
│ │
PASS │ │ FAIL
▼ ▼
[LLM Gen] [Escalation]
This diagram communicates something that a simple "retrieval component" box does not: the pipeline has stages, each stage can be independently configured or replaced, and there is a quality gate that routes to humans when confidence is insufficient.
29.3 The Agent Architecture Diagram Pattern¶
Agent architectures are harder to diagram because the interesting behavior is dynamic (the agent's loop) not static (service topology). The pattern that communicates most effectively:
AGENT ARCHITECTURE DIAGRAM PATTERN
Show THREE things:
1. The state machine (what states exist, what triggers transitions)
2. The tool manifest per phase (what tools are available when)
3. The human touchpoints (when/how humans interact)
DIAGRAM STRUCTURE:
[PHASE 1: DATA COLLECTION] [PHASE 2: ANALYSIS]
Tools available: Tools available:
✓ get_customer_data ✓ calculate_risk
✓ get_document ✓ generate_summary
✗ send_email (not available) ✗ send_email (not available)
✗ execute_transaction ✗ execute_transaction
Data flows to →
[HUMAN REVIEW]
Human sees: analysis results
Human can: approve/modify/reject
Approved ↓ Rejected ↓
[PHASE 3: ACTION] [BACK TO PHASE 1]
Tools available:
✓ send_email (NOW available)
✓ create_record
✗ execute_transaction (still not available)
The phase-tool relationship is the security boundary made visible.
29.4 Writing Architecture Documents for Different Audiences¶
For Engineering Teams: The Architecture Decision Record (ADR)¶
Already covered in Appendix D. The key principle: ADRs record the decision-making context, not just the decision. Engineers six months later need to know WHY, not just WHAT.
For Technical Leaders and Principal Engineers: The Architecture Brief¶
A 2-3 page document that covers:
ARCHITECTURE BRIEF STRUCTURE (2-3 pages)
1. Problem Statement (half page)
What problem are we solving? Why does it need an AI solution
specifically? What constraints shape the solution?
2. Proposed Architecture (1 page)
Container-level diagram with key design decisions called out.
Each major design decision with brief rationale.
NOT a comprehensive component-level design — save that for ADRs.
3. Trade-offs Accepted (half page)
What are we giving up with this approach?
What risks are we accepting and why?
What would cause us to revisit this architecture?
4. Open Questions (quarter page)
What is not yet decided?
What do we need to learn before the architecture can be finalized?
LENGTH: Never more than 3 pages. If it's longer, the design isn't
clear yet. Clarity in architecture comes from clarity of thought,
not from comprehensive documentation.
For Executives and Business Leaders: The Architecture One-Pager¶
Executives do not need architecture diagrams. They need answers to: what does it do, what can go wrong, what does it cost, and who is responsible.
AI SYSTEM EXECUTIVE ONE-PAGER
WHAT IT DOES (2-3 sentences):
This system [does X for Y users] using [AI approach].
It handles approximately [volume] interactions per month.
It is used for [business purpose].
WHAT COULD GO WRONG AND HOW WE HANDLE IT:
Risk: The AI provides incorrect information.
Mitigation: Quality monitoring with automated alerts; all
high-stakes responses reviewed by human agent before sending.
Risk: Customer data is exposed to external services.
Mitigation: No personally identifiable information is sent
to external AI providers.
WHAT IT COSTS:
Technology cost: $X/month at current volume
Human oversight cost: [N] agent-hours/week for review
Total cost per interaction: $Y
WHO IS RESPONSIBLE:
Business owner: [Name, Title]
Technical owner: [Name, Title]
Escalation for serious issues: [Name, Title]
GOVERNANCE STATUS:
✓ Audit trail in place — all interactions logged for [retention period]
✓ Human review implemented for [specific categories]
✓ [Regulatory requirement] compliant
⚠ [Outstanding item] — being addressed by [date]
For Compliance and Legal: The Compliance Narrative¶
Compliance reviewers are not looking for architecture diagrams. They are looking for evidence that specific controls exist. The compliance narrative maps controls to requirements.
COMPLIANCE NARRATIVE STRUCTURE
For each applicable requirement:
Requirement: [Specific regulatory or policy requirement]
How we satisfy it: [Specific technical or process control]
Evidence: [Where to find evidence that the control is operating]
Example:
Requirement: GDPR Article 13 — individuals must be informed when
automated decision-making is used.
How we satisfy it: All customer-facing AI interactions include
a disclosure: "This response was generated with AI assistance."
This disclosure is hardcoded in the UI layer and cannot be
disabled by configuration.
Evidence: Screenshot of disclosure in UI; code review confirming
disclosure cannot be disabled; test cases verifying disclosure
appears on all AI responses.
The compliance narrative is a bridge document — it translates the
architecture into the language of the control framework.
29.5 Diagram Anti-Patterns¶
Anti-Pattern 1: The God Diagram¶
Everything in one diagram. 15 services, 30 arrows, unreadable at any resolution. Shows the system's complexity without communicating anything useful.
Fix: One diagram per concern. The context diagram shows the big picture. The container diagram shows the components. The component diagram shows the RAG pipeline. Never try to put all three in one diagram.
Anti-Pattern 2: The "AI Box"¶
A diagram that shows [Service A] → [AI] → [Service B]. The AI is a single opaque box. The viewer has no idea what the AI does, what data it uses, what its failure modes are, or where humans are involved.
Fix: Open the AI box. Show the gateway, the prompt, the retrieval pipeline, the model, the confidence gate, the escalation path. The AI is not a black box architecturally — diagram it with the same specificity as any other component.
Anti-Pattern 3: Leaving Out the Failure Paths¶
The diagram shows the happy path. Every arrow represents success. There are no branches for failure, no escalation paths, no confidence gates.
Fix: Add failure paths explicitly. The failure paths are architecturally as important as the success paths. For AI systems, the failure path (low confidence → escalation → human) is often the most safety-critical path.
Anti-Pattern 4: No Data Classification Annotations¶
A diagram that shows data flows without indicating the classification of the data flowing in each direction. A reviewer cannot tell whether PII is crossing the perimeter.
Fix: Annotate data flows with: the data type, the classification (internal/confidential/PII), and whether the destination is inside or outside the organizational perimeter. Any PII flowing to an external system should be highlighted.
EXERCISE — Diagram an Existing System: Take an AI system in your organization that has no architecture diagram. Create a C4 context diagram and container diagram using the AI extensions from Section 29.2. Add: explicit human touchpoints, data classification annotations on all cross-boundary flows, failure paths, and the phase-tool security boundary for any agentic components. Share with the team. What did the team learn about their own system from seeing it diagrammed?
WORKSHOP — Write for Three Audiences: Take the same AI system. Write three documents about it: (1) the architecture brief for technical leaders (2 pages max), (2) the executive one-pager, (3) the compliance narrative for one specific regulatory requirement. The challenge: the same system, three audiences, three completely different documents. Note which document was hardest to write and why.
Next: Module 30 — Career and Professional Development for AI Architects