APPENDICES¶
AI Architect's Comprehensive Reference¶
APPENDIX A — Tool Reference Matrix 2026¶
How to use: This matrix covers 70+ tools across 12 categories. For each tool: what problem it solves, when to choose it, and what to consider. Tools marked ⭐ are the category default in 2026 for most enterprise use cases.
A.1 LLM Gateways & Proxy¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| LiteLLM ⭐ | Routes to 100+ LLM providers via unified API, cost tracking, fallback chains | Open source, multi-provider, self-hosted; default for most | Requires operational investment; less polished UI |
| Portkey | Gateway + semantic caching + guardrails + observability in one | Teams wanting managed gateway with built-in caching | Cost at high volume |
| Helicone | Lightweight proxy, per-user cost, prompt versioning | Quick observability setup, minimal ops overhead | Less feature-rich than Portkey |
| Kong AI Gateway | Enterprise API gateway with AI-specific plugins | Organizations already on Kong | Heavier setup |
| Azure APIM | Microsoft API management with AI gateway features | Azure-first organizations | Vendor lock-in |
A.2 LLM Orchestration & Agent Frameworks¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| LangChain | Broad ecosystem, many integrations, chains and agents | Teams needing wide tool/service integrations; large community | Can be over-abstracted; complexity grows fast |
| LangGraph ⭐ | Graph-based agent orchestration, explicit state, conditional routing | Complex agent workflows needing state machines | Steeper learning curve; requires graph thinking |
| LlamaIndex ⭐ | RAG-first framework, strong retrieval primitives, data connectors | RAG-heavy workloads, production document pipelines | Less strong for non-RAG agent patterns |
| AutoGen | Multi-agent conversations, Microsoft-maintained | Collaborative agent systems, Azure environments | Research-oriented; production maturity catching up |
| CrewAI | Role-based agent teams, crew coordination | Document processing pipelines, multi-role workflows | Less granular control than LangGraph |
| Temporal | Durable workflow execution, long-running workflows | Production workflows spanning hours/days; requires persistence | Infrastructure overhead; separate from LLM tools |
| Prefect / Airflow | Workflow orchestration (not AI-specific) | Orchestrating batch AI jobs; teams with existing workflow infra | Not designed for agentic LLM patterns |
A.3 Vector Databases¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| Weaviate ⭐ | Hybrid search built-in (dense + sparse), schema-based, open source | Production RAG with hybrid search requirement; open source preference | More complex setup than pure vector stores |
| Qdrant | High-performance, Rust implementation, open source | Performance-sensitive or air-gapped deployments | Smaller ecosystem than Weaviate |
| pgvector | Postgres extension for vector similarity | Already on Postgres; smaller scale; want single DB | Performance limitations at very large scale |
| Pinecone | Managed, simple ops, good at scale | Low ops overhead priority, medium-to-large scale | Cost at high volume; vendor lock-in |
| Chroma | Lightweight, developer-friendly, in-memory or persistent | Development, prototyping, small-scale | Not production-grade for large deployments |
| Milvus | High-scale open source, distributed | Very large scale (billions of vectors), self-hosted | Complex operational setup |
| Redis Vector | Vector search in existing Redis | Already using Redis, low-latency vector search | Less feature-rich than dedicated vector DBs |
A.4 Embedding Models¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| text-embedding-3-large (OpenAI) | 3072-dim, strong general-purpose | Default for OpenAI-native stacks | Vendor lock-in; cannot self-host |
| nomic-embed-text ⭐ | Open source, 8192 context, strong performance | Self-hostable, no vendor lock-in, good quality | Slightly below top-tier at longer contexts |
| bge-m3 (BAAI) | Multilingual, 100+ languages, strong hybrid | Multilingual content, dense+sparse in one model | Larger model; requires GPU for production throughput |
| e5-mistral-7b | Instruction-following embedding, strong retrieval | Complex retrieval requiring instruction tuning | Large model; expensive to run |
| Cohere embed-english-v3 | Strong for English business text | Cohere-first stacks; business document retrieval | Vendor lock-in |
| Jina embeddings | Supports code, multilingual, long context | Code+text hybrid corpus | Less widely benchmarked |
A.5 RAG Infrastructure¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| Unstructured.io ⭐ | Document parsing: PDF, DOCX, HTML, tables, images, scanned | Production ingestion pipeline; complex document types | Managed API costs at scale; open source version available |
| LlamaParse | Advanced PDF parsing, table extraction, LLM-powered | Complex financial/legal PDFs with tables and charts | LlamaIndex ecosystem; per-page cost |
| AWS Textract | OCR + structure extraction from scanned documents | AWS-native teams; scanned doc heavy workloads | AWS lock-in; cost at scale |
| Azure Document Intelligence | Form/document extraction, pre-built models | Azure-first; form processing, invoice extraction | Azure lock-in |
| Cohere Rerank ⭐ | Cross-encoder re-ranking API | Production RAG quality improvement; managed API | Cost at high query volume; API dependency |
| BGE Reranker | Open-source cross-encoder, self-hostable | Self-hosted re-ranking; cost-sensitive | Requires GPU for production throughput |
| RAGAS ⭐ | RAG evaluation metrics (faithfulness, relevance, precision, recall) | Any production RAG system needing quality measurement | Requires LLM calls for judge-based metrics |
| TruLens | LLM app evaluation, feedback functions | Teams wanting RAG + agent evaluation in one framework | LangChain-centric history |
| GPTCache | Semantic caching for LLM responses | High-volume RAG with repetitive queries | Adds infrastructure complexity |
A.6 AI Observability & Evaluation¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| LangSmith ⭐ | LangChain tracing, eval pipeline, dataset management, human annotation | LangChain shops; teams wanting eval + observability together | LangChain-centric; cost at high volume |
| Arize Phoenix | Open source, vendor-neutral, RAGAS integration, drift detection | Teams wanting open source; RAG-heavy; vendor-neutral | Less polished UI than commercial options |
| Weave (W&B) | ML experiment to production LLM tracking, lineage | Teams using W&B for ML training; experiment continuity | W&B ecosystem dependency |
| Helicone | Lightweight proxy observability, cost per user, prompt versioning | Fast setup; cost-sensitive; simple observability needs | Limited eval capabilities |
| Datadog AI Observability | OTel-native, correlate AI + infra health, enterprise | Organizations on Datadog; unified observability | Expensive at scale; vendor lock-in |
| Grafana + OTel Collector | DIY observability stack, vendor-neutral | Organizations with existing Grafana; infrastructure teams | More setup; no AI-specific out-of-box dashboards |
| PromptFoo ⭐ | Prompt regression testing, CI/CD gate, security assertions | Prompt change validation, adversarial testing, CI/CD | Requires test case design; not a full observability platform |
| DeepEval | AI testing framework, 14+ metrics, CI/CD integration | Teams wanting pytest-style AI testing | Newer tool; ecosystem still maturing |
A.7 AI Security Tools¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| Garak ⭐ | LLM vulnerability scanner, 100+ attack probes, automated | Initial security baseline; CI/CD security gates | Output interpretation requires security expertise |
| PyRIT (Microsoft) | Adversarial attack orchestration, Azure integration | Azure-native; deep agentic system testing; customizable | More setup than Garak; Python-heavy |
| PromptBench | Prompt robustness evaluation framework | Research-oriented prompt robustness testing | Less production-oriented |
| DeepTeam | AI red teaming platform, agentic attack patterns | Agentic system testing; high-volume attack variants | Newer tool |
| LlamaGuard 3 ⭐ | Input/output safety classification, open source | Self-hosted content safety; regulated environments | Requires GPU for production throughput |
| NeMo Guardrails | Programmable safety rails, NVIDIA-maintained | NVIDIA stack; complex safety logic; dialog management | NVIDIA ecosystem preference |
| Presidio ⭐ | PII detection and anonymization, open source | Any system handling PII in prompts; Microsoft-backed | Custom entity types require configuration |
| OPA ⭐ | Policy-as-code authorization engine | Tool/action authorization for agents; enterprise policy | Learning curve; Rego policy language |
| Lakera Guard | Real-time LLM input/output filtering API | Managed safety API; quick integration | External API dependency |
A.8 Inference Infrastructure¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| vLLM ⭐ | Production LLM serving, PagedAttention, high throughput, OpenAI-compatible API | Self-hosted open-weight model production serving | NVIDIA CUDA only; requires GPU |
| Ollama | Local model running, Apple Silicon support, developer-friendly | Developer workstations; edge; single-user applications | Not for concurrent production serving |
| TensorRT-LLM | NVIDIA-optimized inference, maximum throughput | Maximum throughput on NVIDIA hardware | Model-specific compilation; NVIDIA lock-in |
| SGLang | Structured generation focus, PagedAttention-based | Structured output heavy workloads; JSON-mode optimization | Smaller community than vLLM |
| LMDeploy | Efficient inference, TurboMind engine | Production inference with quantization support | Less community support than vLLM |
| Unsloth | Memory-efficient fine-tuning, LoRA/QLoRA | Fine-tuning open-weight models efficiently | Training-focused; not serving |
| Axolotl | Flexible fine-tuning framework | Experimentation and training diverse model types | Less optimization than Unsloth |
A.9 AI Coding Assistants¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| GitHub Copilot Enterprise ⭐ | AI autocomplete + enterprise features, M365 integration | Enterprise standardization; GitHub-centric orgs | 9% developer satisfaction vs. alternatives; data agreements needed |
| Cursor ⭐ | AI-native IDE, codebase context, agentic modes | Senior engineers; complex refactoring; best satisfaction (18%) | Privacy mode required for enterprise use; not an Eclipse/IntelliJ replacement |
| Claude Code | CLI agentic coding, SOTA quality (46% satisfaction) | Complex multi-file refactoring; architecture exploration | CVE-2025-59536: review .claude/ configs; system-level access |
| Windsurf (Codeium) | AI-native IDE, enterprise self-hosting option | Data residency requirements; enterprise deployment | Less ecosystem maturity than Cursor/Copilot |
| Codeium | Free tier coding assistant, multi-IDE | Budget-conscious teams; IntelliJ/Eclipse users | Less capable than Cursor/Claude Code at top tier |
A.10 Cloud AI Platforms (Managed)¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| AWS Bedrock ⭐ | Hosted Claude, OpenAI models (2026), Llama, Mistral + others inside AWS VPC | AWS-first orgs; data residency needs; frontier quality | Models slightly behind public API release |
| Microsoft Foundry (incl. Azure OpenAI) ⭐ | OpenAI GPT models plus Claude and other catalog models inside Azure tenant | Microsoft-first; strict data sovereignty; enterprise compliance | Rebranded from Azure AI Foundry (Nov 2025); catalog and regional availability vary |
| Google Gemini Enterprise Agent Platform (formerly Vertex AI) | Gemini + partner models (including Claude) inside GCP | GCP-first organizations | Smaller model ecosystem vs. Bedrock |
| NVIDIA AI Enterprise | Enterprise AI software stack, GPU orchestration | NVIDIA hardware-first deployments | Narrow use case |
A.11 Prompt Management¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| LangSmith Prompt Hub | Prompt versioning + eval integration | LangChain teams | LangChain ecosystem dependency |
| Humanloop | Prompt management + human feedback loop | Teams iterating rapidly with user feedback | Cost at scale |
| Portkey Prompt Library | Prompt versioning in gateway | Teams using Portkey gateway already | Less feature-rich than dedicated tools |
| Custom (Git + CI/CD) ⭐ | Prompts as code, version-controlled, review process | Organizations with strong engineering culture | Requires building eval integration |
A.12 MCP Ecosystem¶
| Tool | What It Does | When to Choose | Consider |
|---|---|---|---|
| Anthropic MCP SDK ⭐ | TypeScript + Python SDKs for building MCP servers/clients | Building MCP servers; the reference implementation | Maturing spec; some attributes experimental |
| GitHub MCP Server | GitHub operations via MCP (PRs, issues, repos) | Code review agents; development automation | Limited to GitHub-hosted repos |
| Postgres MCP Server | Database operations via MCP | Database-connected agents; data analysis | Scope carefully to prevent SQL injection risk |
| Slack MCP Server | Slack messaging via MCP | Notification agents; conversation search | OAuth scope management required |
| Filesystem MCP Server | File operations via MCP | Local development agents; document processing | Security: restrict to project directory |
APPENDIX B — OWASP Top 10 for LLM Applications (2025): Architect's Annotation¶
For each risk: the architectural response, not just the description.
| # | Risk | Architect's Response |
|---|---|---|
| LLM01 | Prompt Injection — Malicious input overrides model behavior | Structural content/instruction separation; XML delimiters; phase-based tool access limits blast radius; output monitoring for anomalous patterns. Never rely on string matching alone. |
| LLM02 | Sensitive Information Disclosure — Model reveals sensitive content from context or training | PII pseudonymization before prompts reach external models; no secrets in system prompts; context isolation between users; output PII scanning via Presidio |
| LLM03 | Supply Chain Vulnerabilities — Compromised model weights, datasets, dependencies | Model SHA256 verification; pin all dependencies; private PyPI mirror; MCP server governance registry; Skills/plugin security review before deployment |
| LLM04 | Data and Model Poisoning — Contaminated training/knowledge base produces wrong outputs | Never let user-submitted content flow directly to vector store; content provenance tagging; retrieval anomaly detection; separate user-generated content from authoritative knowledge |
| LLM05 | Improper Output Handling — LLM output passed unsanitized to downstream systems | Output schema validation; HTML sanitization (DOMPurify); never auto-execute AI-generated code; URL allowlist before fetching; CSV formula injection prevention |
| LLM06 | Excessive Agency — Agent has too much access, autonomy, or functionality | Phase-based tool manifests; minimum necessary tools per task; human approval gates for irreversible actions; code-enforced (not prompt-enforced) authorization |
| LLM07 | System Prompt Leakage — Attacker extracts system prompt contents | Assume system prompts will eventually be extracted — design accordingly; no secrets in prompts; prompt protection instructions; monitor for extraction attempt patterns |
| LLM08 | Vector and Embedding Weaknesses — RAG access control failures, embedding attacks | Metadata filtering enforces tenant isolation (not optional); access control at retrieval layer not just application layer; embedding model version pinned and stored with vectors |
| LLM09 | Misinformation — Confident incorrect outputs cause real harm | Citations required for factual claims; confidence gates; human review for high-stakes outputs; treat AI misinformation as an incident category requiring post-mortems |
| LLM10 | Unbounded Consumption — API abuse, cost attacks, DoS via token exhaustion | Rate limiting per user; token budget per request at gateway; per-task cost ceiling code-enforced for agents; cost anomaly detection and alerts |
Additional: OWASP Agentic AI Top 10 (December 2025)
| # | Risk | Architect's Response |
|---|---|---|
| AA01 | Agent Behavior Hijacking — Adversarial inputs redirect agent planning | Constrained tool scope; structured output from planning step; human review of action plans before execution |
| AA02 | Tool Misuse — Agent calls tools with unintended parameters or sequences | Typed tool input schemas; parameter range validation; idempotency keys prevent duplicate actions |
| AA03 | Identity and Privilege Abuse — Agent operates with more privilege than warranted | Agent identity separate from user identity; agent scope ⊂ user scope always; OPA authorization per tool call |
| AA04 | Goal Manipulation — Long-running agent's goal redirected through context accumulation | Periodic context summarization with constraint re-injection; session monitoring for goal drift |
APPENDIX C — AI Architecture Review Checklist¶
Use at design review for any AI system. Organized by concern area.
C.1 Foundation¶
- [ ] Is AI the core of this system or an enhancement? (documented)
- [ ] Has the consequence of wrong AI output been defined? (who is affected, what is the impact)
- [ ] Is there a non-AI fallback or baseline? (for AI-augmented) or a failure narrative? (for AI-native)
- [ ] Has the human oversight model been designed? (who approves what, at what confidence threshold)
- [ ] Has the model risk tier been assessed? (Module 11 tiering for regulated uses)
C.2 LLM Access and Gateway¶
- [ ] Are all LLM calls routed through the organizational gateway? (no direct API calls from services)
- [ ] Does the gateway have per-service token budgets and cost attribution tags?
- [ ] Is PII scanning active before prompts leave the perimeter?
- [ ] Are model versions pinned (not using floating aliases)?
- [ ] Is there a fallback chain when the primary model is unavailable?
C.3 Prompts and Knowledge¶
- [ ] Is the system prompt registered in the prompt management system?
- [ ] Is the prompt version-controlled and has a designated owner?
- [ ] Has the prompt passed the offline eval suite before production deployment?
- [ ] Is the knowledge base version-managed? (old chunks archived when documents update)
- [ ] Is the embedding model version pinned and stored with each vector?
- [ ] Is there a confidence gate for RAG retrieval? (defined threshold, business-approved)
C.4 Agents and Tools¶
- [ ] Is the tool manifest scoped to the task? (not all tools for all tasks)
- [ ] Are irreversible tool calls behind human approval gates? (code-enforced)
- [ ] Is the orchestrator deterministic? (routing decisions are code, not LLM judgment)
- [ ] Does every agent task have a hard iteration limit and cost ceiling? (code-enforced)
- [ ] Is every tool call logged with: input, output, caller, timestamp?
- [ ] Are agent tasks idempotent? (retries don't create duplicate side effects)
C.5 Security¶
- [ ] Is PII pseudonymized before entering any AI processing layer?
- [ ] Is user-submitted content structurally separated from instructions? (XML delimiters or typed fields)
- [ ] Has red team testing been completed? (Garak + manual adversarial for injection)
- [ ] Are plugins/skills/MCP servers in the governance registry and reviewed?
- [ ] Is the system prompt free of secrets, API keys, and internal URLs?
- [ ] Is output PII scanning active before responses are returned?
C.6 Observability and Evaluation¶
- [ ] Are OTel GenAI spans instrumented on all LLM calls?
- [ ] Is cost attribution active? (team, feature, workflow tags on every call)
- [ ] Is there an offline eval suite with minimum 20 cases (including adversarial)?
- [ ] Does the eval suite gate production deployments? (failing evals block promotion)
- [ ] Is online eval sampling active? (2-5% of production traffic)
- [ ] Are drift detection alerts configured? (model version, prompt version, retrieval quality)
C.7 Compliance and Governance¶
- [ ] Is the system registered in the AI inventory?
- [ ] Is the EU AI Act risk tier assessed? (for systems with EU users/data)
- [ ] For regulated use cases: is a model risk owner assigned?
- [ ] Is the audit trail append-only and retention-compliant?
- [ ] Is the data processing agreement with model providers in place?
- [ ] For agentic systems: is the agent identity separate from user identity?
APPENDIX D — ADR Templates for Common AI Decisions¶
D.1 Model Selection ADR¶
Decision: Which LLM for [use case]
Context: - Use case: [description] - Latency SLA: [P99 target] - Data sensitivity: [classification — can it leave the perimeter?] - Volume: [requests/day, tokens/day] - Quality requirement: [what does acceptable output look like]
Options Considered:
| Option | Model | Estimated Cost/Month | Data Residency | Quality Assessment |
|---|---|---|---|---|
| Option A | [model] | $X | [in-VPC / external] | [eval score or human judgment] |
| Option B | [model] | $Y | [in-VPC / external] | [eval score or human judgment] |
Decision: [Chosen option]
Rationale: [Why this model, what constraints drove the decision]
Consequences: - Cost: [monthly at current volume, at 3x volume] - Data handling: [what data goes where, under what agreement] - Deprecation risk: [model version, deprecation date if known, fallback] - Reversibility: [what would be required to switch models]
Review trigger: Model deprecated, cost exceeds $X/month, quality falls below baseline
D.2 RAG Confidence Threshold ADR¶
Decision: Minimum retrieval confidence score to generate an answer vs. escalate
Context: - Use case: [what is the AI answering] - Consequence of wrong answer: [regulatory exposure, customer harm, financial impact] - Escalation cost: [human agent time, user friction]
Business sign-off required: [yes — list stakeholders who must approve]
Proposed threshold: [0.XX]
Rationale: [Evidence for this value — pilot data, comparable implementations]
Consequences of threshold too high: [excessive escalation, underutilized AI, higher support cost]
Consequences of threshold too low: [wrong answers, regulatory risk, customer harm]
Monitoring: [How threshold is evaluated in production — weekly review, adjustment triggers]
Owner: [Who is accountable for this threshold decision — not engineering]
D.3 Human-in-the-Loop Placement ADR¶
Decision: Which agent actions require human approval before execution
Context: - Agent workflow: [description] - Tools available to agent: [list]
Irreversibility Analysis:
| Tool / Action | Reversible? | Consequence of Error | Human Approval Required? |
|---|---|---|---|
| [tool name] | Yes/No | [description] | Yes/No |
Approval mechanism: - Approval interface: [what the human sees] - Timeout behavior: [what happens if no response in T hours] - Override process: [can approval be bypassed, under what conditions] - Audit record: [what is logged with the approval decision]
Decision: [List of tools that require human approval]
Rationale: [Why these specific tools — irreversibility, consequence, regulatory requirement]
Exception process: [Under what conditions can this be changed, who approves]
D.4 Prompt Change Control ADR¶
Decision: Process for making changes to production system prompts
Context: - System: [AI feature name] - Current prompt template: [ID and version] - Change frequency: [estimated] - Risk level: [customer-facing, internal, regulated]
Change process: 1. Change proposed in prompt management system by [role] 2. Automated eval suite runs against the proposed change 3. If eval score < [threshold]: change rejected, returned to proposer with failure report 4. If eval score ≥ [threshold]: change available for staging 5. Staging deployment: [how validated in staging] 6. Production promotion: requires approval from [role] 7. Rollback procedure: [how quickly, who can trigger]
Emergency override process: [for urgent production fixes — higher accountability, documented]
Audit requirements: [every change logged with: proposer, reviewer, eval results, rationale]
APPENDIX E — Startup Landscape Map¶
Current as of mid-2026. Verify revenue and valuation data before citing — these change frequently.
E.1 Vertical AI¶
| Category | Company | Problem | Status | Lesson |
|---|---|---|---|---|
| Legal | Harvey | Contract review, legal research, discovery | $11B valuation, ~$200M ARR, ~50% of Am Law 100 | Domain data moat; embedded legal engineering |
| Legal | Legora | European legal AI, GDPR-compliant | $5.55B valuation, €550M raised | Geographic regulatory moat |
| Healthcare | Abridge | Ambient clinical scribe, EHR integration | $5.3B valuation, mainstream US health system adoption | EHR integration = distribution moat |
| Healthcare | Hippocratic AI | Clinical support agents, patient communication | Hospital partnerships in production | Human oversight architecture; clinical non-decision-maker |
| Finance | Sierra | Customer operations AI, enterprise chatbots | Bret Taylor-founded; enterprise contracts | Escalation-first architecture |
| Finance | Klarna AI | Customer service automation | ~$40M/year savings (company-reported) | High-volume structured workflows |
E.2 AI Search & Knowledge¶
| Company | Problem | Status | Lesson |
|---|---|---|---|
| Glean | Enterprise search across all tools | >$100M ARR | Integration breadth is the moat; RAG at org scale |
| Perplexity | AI-native web search with citations | $9B+ valuation | Citation-first design is the trust architecture |
| OpenEvidence | Medical literature search | Growing adoption in clinical settings | Authoritative corpus restriction = trust in regulated domains |
E.3 AI Coding¶
| Company | Problem | Status | Lesson |
|---|---|---|---|
| Cursor | AI-native IDE | $2B ARR, Q1 2026 | Redesign for AI vs. bolt-on |
| GitHub Copilot | AI coding in existing IDE | 4.7M paid, 90% Fortune 100 | Distribution moat via GitHub |
| Claude Code | CLI agentic coding | 46% developer satisfaction | Agentic capability; system-level access risk |
| Replit | AI coding for non-developers | $150M ARR annualized | Expanding the developer population |
E.4 AI Infrastructure¶
| Company | Problem | Status | Lesson |
|---|---|---|---|
| Groq | Custom LPU inference, low latency | Production at scale | Hardware moat; custom silicon for AI |
| Weaviate | Vector database with hybrid search | Open source + managed; widely deployed | Hybrid search built-in differentiates |
| Qdrant | High-performance vector search | Growing enterprise adoption | Rust performance; open source |
| LiteLLM | LLM gateway / proxy | De facto open-source standard | Protocol abstraction value |
| Unstructured.io | Document parsing for AI | Production standard for RAG ingestion | Infrastructure layer for unstructured data |
| Scale AI | Data labeling and AI evaluation | Used by major model labs | Human eval as infrastructure |
| ElevenLabs | Voice AI infrastructure | Enterprise adoption growing | Platform layer beats application layer |
| Deepgram | Domain-specific speech AI | Proprietary speech models | Data moat in domain-specific audio |
E.5 AI Security¶
| Company | Problem | Status | Lesson |
|---|---|---|---|
| Garak (NVIDIA) | LLM vulnerability scanning | Open source standard | New attack surfaces need new tools |
| Lakera | Real-time LLM guardrails | Enterprise adoption growing | Managed safety as infrastructure |
| Knostic | AI data visibility / oversharing | Growing with Copilot deployment | Oversharing is the #1 enterprise AI risk |
APPENDIX F — Glossary¶
Agent (AI agent): A system that receives a goal, plans steps, takes actions via tool calls, observes results, and adapts until the goal is achieved or a limit is reached. Distinguished from a chatbot by its ability to take actions with real-world effects.
AGI (Artificial General Intelligence): Hypothetical AI with general reasoning capability equal to or exceeding human intelligence across all domains. Not currently demonstrated; timeline highly uncertain.
A2A (Agent-to-Agent Protocol): Google-initiated open protocol for inter-agent communication and task delegation. Standardizes how one AI agent delegates work to another.
Agentic loop: The cycle of reason → act → observe → reason that characterizes agent behavior. Each loop iteration makes a tool call and observes its result.
Chunking: The process of splitting source documents into segments for vector embedding and retrieval. Chunk strategy (size, overlap, semantic vs. fixed) significantly affects RAG quality.
Context window: The maximum amount of text (measured in tokens) that a model can process in a single inference call. Content outside the context window is not available to the model.
Cross-encoder: A re-ranking model that evaluates query and document together, producing a direct relevance score. More accurate than bi-encoder (vector) retrieval but computationally expensive at query time.
Embedding: A numerical vector representation of text, capturing semantic meaning. Similar texts produce similar (close) vectors. Used for semantic search in RAG systems.
Eval (evaluation): A test for AI behavior — asserting that a given input produces an appropriate output according to defined criteria. Analogous to unit tests but for probabilistic AI behavior.
Fine-tuning: Training a pre-trained model on domain-specific data to improve performance on domain-specific tasks. Creates a new model variant that requires ongoing governance.
Guardrails: Architectural controls that constrain AI system inputs and outputs: input filtering (blocking harmful prompts), output filtering (blocking harmful responses), and action scope enforcement (limiting what an agent can do).
Hallucination: When an LLM generates factually incorrect information with apparent confidence. A structural property of LLMs, not a bug to be fixed — managed through grounding, citations, and confidence gating.
Human-in-the-loop: Architectural pattern where a human provides judgment, approval, or oversight at defined decision points in an AI workflow. Not the same as "a human can be involved if needed."
Idempotency: The property of an operation that produces the same result regardless of how many times it is executed with the same input. Required for safe retry logic in agent workflows.
KV cache (Key-Value cache): Memory that stores intermediate computations (attention keys and values) from prior tokens in the context. Reusing the KV cache avoids recomputing these for repeated prompt prefixes (prefix caching).
LoRA (Low-Rank Adaptation): A parameter-efficient fine-tuning technique that trains a small adapter layer rather than full model weights, enabling domain adaptation at a fraction of the cost and compute.
LLM (Large Language Model): A neural network trained on large text corpora to predict the next token given prior context. The foundation of modern AI systems including GPT, Claude, Gemini, and Llama.
MCP (Model Context Protocol): Anthropic-originated open protocol for agent-to-tool connectivity. Standardizes how AI agents access tools, data, and prompts via MCP servers. Now governed by the Agentic AI Foundation under the Linux Foundation.
Model drift: Changes in model behavior over time — either because the model provider updated the model, or because the input distribution shifted. Requires ongoing monitoring to detect.
Model risk: The risk that a model produces incorrect outputs that cause financial, regulatory, or operational harm. Subject to formal governance requirements (SR 11-7 successor in financial services).
Orchestrator: The component that coordinates agents in a multi-agent system. Should be deterministic code (not an LLM) for routing and sequencing decisions.
PagedAttention: vLLM's technique for managing GPU memory for KV caches in non-contiguous pages, like virtual memory in an operating system. Eliminates KV cache fragmentation, dramatically increasing concurrent request capacity.
Perplexity (in AI context): A measure of how well a language model predicts a sample — lower perplexity indicates better model performance. Distinct from the company Perplexity.
Prompt injection: An attack where malicious instructions in user input, retrieved documents, or external data override the model's intended behavior.
Quantization: Reducing model weight precision (from FP32 to FP16 to INT8 to INT4) to reduce memory requirements and increase inference speed at some quality cost. Common formats: GGUF (Ollama, CPU), AWQ and GPTQ (GPU production), FP8 (H100 native).
RAG (Retrieval-Augmented Generation): The pattern of retrieving relevant content from a knowledge base before generation, grounding the LLM's output in specific information.
Reasoning model: An LLM that uses extended test-time computation (a "thinking" phase) before producing a final response. Examples: OpenAI GPT-6 with reasoning effort, Claude adaptive thinking, Gemini 3.x thinking levels (see Appendix G, G.2).
Semantic caching: Caching LLM responses indexed by semantic similarity, serving cached responses for queries that are similar (not identical) to previously answered queries. Reduces API calls for repetitive queries.
Shadow AI: AI tools used by employees outside official IT oversight — consumer ChatGPT, personal AI accounts, unauthorized browser extensions — often with corporate data.
SLM (Small Language Model): Models under approximately 10B parameters, optimized for efficiency. Examples: Phi-4-mini (3.8B), Gemma 4 small variants (E2B/E4B), Llama 3.2 1B/3B. Suitable for edge deployment and on-device inference.
Strangler fig: Incremental replacement pattern — new functionality gradually replaces old, with both running in parallel during the transition. Applied to adding AI to legacy systems.
System prompt: Instructions given to an LLM at the beginning of a context that define its role, constraints, and behavior. Processed by the model but not part of the user-visible conversation.
Temperature: A hyperparameter controlling LLM output randomness. Temperature 0 = deterministic (same input → same output). Temperature 1 = full randomness. Production systems typically use low temperature for consistency.
Token: The unit of text that LLMs process. Approximately 0.75 English words. Pricing for LLM APIs is denominated in tokens (per million input tokens, per million output tokens).
Token budget: A hard limit on the number of tokens consumed by a request, an agent task, or a user session. Enforced in code (not in prompt instructions) to prevent cost overruns.
vLLM: Open-source LLM inference serving engine, the production standard for self-hosted open-weight models. Features PagedAttention, continuous batching, and OpenAI-compatible API. NVIDIA CUDA only.
Vector database: A database optimized for storing and querying high-dimensional vectors (embeddings). Supports similarity search operations that traditional databases cannot perform efficiently.
This concludes the AI Architect's Comprehensive Reference. Artifact 1 — The Strategic & Conceptual Layer Companion: Artifact 2 — AI Integration Patterns: A Deep Reference