Skip to content

APPENDICES

AI Architect's Comprehensive Reference


APPENDIX A — Tool Reference Matrix 2026

How to use: This matrix covers 70+ tools across 12 categories. For each tool: what problem it solves, when to choose it, and what to consider. Tools marked ⭐ are the category default in 2026 for most enterprise use cases.


A.1 LLM Gateways & Proxy

Tool What It Does When to Choose Consider
LiteLLM ⭐ Routes to 100+ LLM providers via unified API, cost tracking, fallback chains Open source, multi-provider, self-hosted; default for most Requires operational investment; less polished UI
Portkey Gateway + semantic caching + guardrails + observability in one Teams wanting managed gateway with built-in caching Cost at high volume
Helicone Lightweight proxy, per-user cost, prompt versioning Quick observability setup, minimal ops overhead Less feature-rich than Portkey
Kong AI Gateway Enterprise API gateway with AI-specific plugins Organizations already on Kong Heavier setup
Azure APIM Microsoft API management with AI gateway features Azure-first organizations Vendor lock-in

A.2 LLM Orchestration & Agent Frameworks

Tool What It Does When to Choose Consider
LangChain Broad ecosystem, many integrations, chains and agents Teams needing wide tool/service integrations; large community Can be over-abstracted; complexity grows fast
LangGraph ⭐ Graph-based agent orchestration, explicit state, conditional routing Complex agent workflows needing state machines Steeper learning curve; requires graph thinking
LlamaIndex ⭐ RAG-first framework, strong retrieval primitives, data connectors RAG-heavy workloads, production document pipelines Less strong for non-RAG agent patterns
AutoGen Multi-agent conversations, Microsoft-maintained Collaborative agent systems, Azure environments Research-oriented; production maturity catching up
CrewAI Role-based agent teams, crew coordination Document processing pipelines, multi-role workflows Less granular control than LangGraph
Temporal Durable workflow execution, long-running workflows Production workflows spanning hours/days; requires persistence Infrastructure overhead; separate from LLM tools
Prefect / Airflow Workflow orchestration (not AI-specific) Orchestrating batch AI jobs; teams with existing workflow infra Not designed for agentic LLM patterns

A.3 Vector Databases

Tool What It Does When to Choose Consider
Weaviate ⭐ Hybrid search built-in (dense + sparse), schema-based, open source Production RAG with hybrid search requirement; open source preference More complex setup than pure vector stores
Qdrant High-performance, Rust implementation, open source Performance-sensitive or air-gapped deployments Smaller ecosystem than Weaviate
pgvector Postgres extension for vector similarity Already on Postgres; smaller scale; want single DB Performance limitations at very large scale
Pinecone Managed, simple ops, good at scale Low ops overhead priority, medium-to-large scale Cost at high volume; vendor lock-in
Chroma Lightweight, developer-friendly, in-memory or persistent Development, prototyping, small-scale Not production-grade for large deployments
Milvus High-scale open source, distributed Very large scale (billions of vectors), self-hosted Complex operational setup
Redis Vector Vector search in existing Redis Already using Redis, low-latency vector search Less feature-rich than dedicated vector DBs

A.4 Embedding Models

Tool What It Does When to Choose Consider
text-embedding-3-large (OpenAI) 3072-dim, strong general-purpose Default for OpenAI-native stacks Vendor lock-in; cannot self-host
nomic-embed-text ⭐ Open source, 8192 context, strong performance Self-hostable, no vendor lock-in, good quality Slightly below top-tier at longer contexts
bge-m3 (BAAI) Multilingual, 100+ languages, strong hybrid Multilingual content, dense+sparse in one model Larger model; requires GPU for production throughput
e5-mistral-7b Instruction-following embedding, strong retrieval Complex retrieval requiring instruction tuning Large model; expensive to run
Cohere embed-english-v3 Strong for English business text Cohere-first stacks; business document retrieval Vendor lock-in
Jina embeddings Supports code, multilingual, long context Code+text hybrid corpus Less widely benchmarked

A.5 RAG Infrastructure

Tool What It Does When to Choose Consider
Unstructured.io ⭐ Document parsing: PDF, DOCX, HTML, tables, images, scanned Production ingestion pipeline; complex document types Managed API costs at scale; open source version available
LlamaParse Advanced PDF parsing, table extraction, LLM-powered Complex financial/legal PDFs with tables and charts LlamaIndex ecosystem; per-page cost
AWS Textract OCR + structure extraction from scanned documents AWS-native teams; scanned doc heavy workloads AWS lock-in; cost at scale
Azure Document Intelligence Form/document extraction, pre-built models Azure-first; form processing, invoice extraction Azure lock-in
Cohere Rerank ⭐ Cross-encoder re-ranking API Production RAG quality improvement; managed API Cost at high query volume; API dependency
BGE Reranker Open-source cross-encoder, self-hostable Self-hosted re-ranking; cost-sensitive Requires GPU for production throughput
RAGAS ⭐ RAG evaluation metrics (faithfulness, relevance, precision, recall) Any production RAG system needing quality measurement Requires LLM calls for judge-based metrics
TruLens LLM app evaluation, feedback functions Teams wanting RAG + agent evaluation in one framework LangChain-centric history
GPTCache Semantic caching for LLM responses High-volume RAG with repetitive queries Adds infrastructure complexity

A.6 AI Observability & Evaluation

Tool What It Does When to Choose Consider
LangSmith ⭐ LangChain tracing, eval pipeline, dataset management, human annotation LangChain shops; teams wanting eval + observability together LangChain-centric; cost at high volume
Arize Phoenix Open source, vendor-neutral, RAGAS integration, drift detection Teams wanting open source; RAG-heavy; vendor-neutral Less polished UI than commercial options
Weave (W&B) ML experiment to production LLM tracking, lineage Teams using W&B for ML training; experiment continuity W&B ecosystem dependency
Helicone Lightweight proxy observability, cost per user, prompt versioning Fast setup; cost-sensitive; simple observability needs Limited eval capabilities
Datadog AI Observability OTel-native, correlate AI + infra health, enterprise Organizations on Datadog; unified observability Expensive at scale; vendor lock-in
Grafana + OTel Collector DIY observability stack, vendor-neutral Organizations with existing Grafana; infrastructure teams More setup; no AI-specific out-of-box dashboards
PromptFoo ⭐ Prompt regression testing, CI/CD gate, security assertions Prompt change validation, adversarial testing, CI/CD Requires test case design; not a full observability platform
DeepEval AI testing framework, 14+ metrics, CI/CD integration Teams wanting pytest-style AI testing Newer tool; ecosystem still maturing

A.7 AI Security Tools

Tool What It Does When to Choose Consider
Garak ⭐ LLM vulnerability scanner, 100+ attack probes, automated Initial security baseline; CI/CD security gates Output interpretation requires security expertise
PyRIT (Microsoft) Adversarial attack orchestration, Azure integration Azure-native; deep agentic system testing; customizable More setup than Garak; Python-heavy
PromptBench Prompt robustness evaluation framework Research-oriented prompt robustness testing Less production-oriented
DeepTeam AI red teaming platform, agentic attack patterns Agentic system testing; high-volume attack variants Newer tool
LlamaGuard 3 ⭐ Input/output safety classification, open source Self-hosted content safety; regulated environments Requires GPU for production throughput
NeMo Guardrails Programmable safety rails, NVIDIA-maintained NVIDIA stack; complex safety logic; dialog management NVIDIA ecosystem preference
Presidio ⭐ PII detection and anonymization, open source Any system handling PII in prompts; Microsoft-backed Custom entity types require configuration
OPA ⭐ Policy-as-code authorization engine Tool/action authorization for agents; enterprise policy Learning curve; Rego policy language
Lakera Guard Real-time LLM input/output filtering API Managed safety API; quick integration External API dependency

A.8 Inference Infrastructure

Tool What It Does When to Choose Consider
vLLM ⭐ Production LLM serving, PagedAttention, high throughput, OpenAI-compatible API Self-hosted open-weight model production serving NVIDIA CUDA only; requires GPU
Ollama Local model running, Apple Silicon support, developer-friendly Developer workstations; edge; single-user applications Not for concurrent production serving
TensorRT-LLM NVIDIA-optimized inference, maximum throughput Maximum throughput on NVIDIA hardware Model-specific compilation; NVIDIA lock-in
SGLang Structured generation focus, PagedAttention-based Structured output heavy workloads; JSON-mode optimization Smaller community than vLLM
LMDeploy Efficient inference, TurboMind engine Production inference with quantization support Less community support than vLLM
Unsloth Memory-efficient fine-tuning, LoRA/QLoRA Fine-tuning open-weight models efficiently Training-focused; not serving
Axolotl Flexible fine-tuning framework Experimentation and training diverse model types Less optimization than Unsloth

A.9 AI Coding Assistants

Tool What It Does When to Choose Consider
GitHub Copilot Enterprise ⭐ AI autocomplete + enterprise features, M365 integration Enterprise standardization; GitHub-centric orgs 9% developer satisfaction vs. alternatives; data agreements needed
Cursor ⭐ AI-native IDE, codebase context, agentic modes Senior engineers; complex refactoring; best satisfaction (18%) Privacy mode required for enterprise use; not an Eclipse/IntelliJ replacement
Claude Code CLI agentic coding, SOTA quality (46% satisfaction) Complex multi-file refactoring; architecture exploration CVE-2025-59536: review .claude/ configs; system-level access
Windsurf (Codeium) AI-native IDE, enterprise self-hosting option Data residency requirements; enterprise deployment Less ecosystem maturity than Cursor/Copilot
Codeium Free tier coding assistant, multi-IDE Budget-conscious teams; IntelliJ/Eclipse users Less capable than Cursor/Claude Code at top tier

A.10 Cloud AI Platforms (Managed)

Tool What It Does When to Choose Consider
AWS Bedrock ⭐ Hosted Claude, OpenAI models (2026), Llama, Mistral + others inside AWS VPC AWS-first orgs; data residency needs; frontier quality Models slightly behind public API release
Microsoft Foundry (incl. Azure OpenAI) ⭐ OpenAI GPT models plus Claude and other catalog models inside Azure tenant Microsoft-first; strict data sovereignty; enterprise compliance Rebranded from Azure AI Foundry (Nov 2025); catalog and regional availability vary
Google Gemini Enterprise Agent Platform (formerly Vertex AI) Gemini + partner models (including Claude) inside GCP GCP-first organizations Smaller model ecosystem vs. Bedrock
NVIDIA AI Enterprise Enterprise AI software stack, GPU orchestration NVIDIA hardware-first deployments Narrow use case

A.11 Prompt Management

Tool What It Does When to Choose Consider
LangSmith Prompt Hub Prompt versioning + eval integration LangChain teams LangChain ecosystem dependency
Humanloop Prompt management + human feedback loop Teams iterating rapidly with user feedback Cost at scale
Portkey Prompt Library Prompt versioning in gateway Teams using Portkey gateway already Less feature-rich than dedicated tools
Custom (Git + CI/CD) ⭐ Prompts as code, version-controlled, review process Organizations with strong engineering culture Requires building eval integration

A.12 MCP Ecosystem

Tool What It Does When to Choose Consider
Anthropic MCP SDK ⭐ TypeScript + Python SDKs for building MCP servers/clients Building MCP servers; the reference implementation Maturing spec; some attributes experimental
GitHub MCP Server GitHub operations via MCP (PRs, issues, repos) Code review agents; development automation Limited to GitHub-hosted repos
Postgres MCP Server Database operations via MCP Database-connected agents; data analysis Scope carefully to prevent SQL injection risk
Slack MCP Server Slack messaging via MCP Notification agents; conversation search OAuth scope management required
Filesystem MCP Server File operations via MCP Local development agents; document processing Security: restrict to project directory

APPENDIX B — OWASP Top 10 for LLM Applications (2025): Architect's Annotation

For each risk: the architectural response, not just the description.

# Risk Architect's Response
LLM01 Prompt Injection — Malicious input overrides model behavior Structural content/instruction separation; XML delimiters; phase-based tool access limits blast radius; output monitoring for anomalous patterns. Never rely on string matching alone.
LLM02 Sensitive Information Disclosure — Model reveals sensitive content from context or training PII pseudonymization before prompts reach external models; no secrets in system prompts; context isolation between users; output PII scanning via Presidio
LLM03 Supply Chain Vulnerabilities — Compromised model weights, datasets, dependencies Model SHA256 verification; pin all dependencies; private PyPI mirror; MCP server governance registry; Skills/plugin security review before deployment
LLM04 Data and Model Poisoning — Contaminated training/knowledge base produces wrong outputs Never let user-submitted content flow directly to vector store; content provenance tagging; retrieval anomaly detection; separate user-generated content from authoritative knowledge
LLM05 Improper Output Handling — LLM output passed unsanitized to downstream systems Output schema validation; HTML sanitization (DOMPurify); never auto-execute AI-generated code; URL allowlist before fetching; CSV formula injection prevention
LLM06 Excessive Agency — Agent has too much access, autonomy, or functionality Phase-based tool manifests; minimum necessary tools per task; human approval gates for irreversible actions; code-enforced (not prompt-enforced) authorization
LLM07 System Prompt Leakage — Attacker extracts system prompt contents Assume system prompts will eventually be extracted — design accordingly; no secrets in prompts; prompt protection instructions; monitor for extraction attempt patterns
LLM08 Vector and Embedding Weaknesses — RAG access control failures, embedding attacks Metadata filtering enforces tenant isolation (not optional); access control at retrieval layer not just application layer; embedding model version pinned and stored with vectors
LLM09 Misinformation — Confident incorrect outputs cause real harm Citations required for factual claims; confidence gates; human review for high-stakes outputs; treat AI misinformation as an incident category requiring post-mortems
LLM10 Unbounded Consumption — API abuse, cost attacks, DoS via token exhaustion Rate limiting per user; token budget per request at gateway; per-task cost ceiling code-enforced for agents; cost anomaly detection and alerts

Additional: OWASP Agentic AI Top 10 (December 2025)

# Risk Architect's Response
AA01 Agent Behavior Hijacking — Adversarial inputs redirect agent planning Constrained tool scope; structured output from planning step; human review of action plans before execution
AA02 Tool Misuse — Agent calls tools with unintended parameters or sequences Typed tool input schemas; parameter range validation; idempotency keys prevent duplicate actions
AA03 Identity and Privilege Abuse — Agent operates with more privilege than warranted Agent identity separate from user identity; agent scope ⊂ user scope always; OPA authorization per tool call
AA04 Goal Manipulation — Long-running agent's goal redirected through context accumulation Periodic context summarization with constraint re-injection; session monitoring for goal drift

APPENDIX C — AI Architecture Review Checklist

Use at design review for any AI system. Organized by concern area.

C.1 Foundation

  • [ ] Is AI the core of this system or an enhancement? (documented)
  • [ ] Has the consequence of wrong AI output been defined? (who is affected, what is the impact)
  • [ ] Is there a non-AI fallback or baseline? (for AI-augmented) or a failure narrative? (for AI-native)
  • [ ] Has the human oversight model been designed? (who approves what, at what confidence threshold)
  • [ ] Has the model risk tier been assessed? (Module 11 tiering for regulated uses)

C.2 LLM Access and Gateway

  • [ ] Are all LLM calls routed through the organizational gateway? (no direct API calls from services)
  • [ ] Does the gateway have per-service token budgets and cost attribution tags?
  • [ ] Is PII scanning active before prompts leave the perimeter?
  • [ ] Are model versions pinned (not using floating aliases)?
  • [ ] Is there a fallback chain when the primary model is unavailable?

C.3 Prompts and Knowledge

  • [ ] Is the system prompt registered in the prompt management system?
  • [ ] Is the prompt version-controlled and has a designated owner?
  • [ ] Has the prompt passed the offline eval suite before production deployment?
  • [ ] Is the knowledge base version-managed? (old chunks archived when documents update)
  • [ ] Is the embedding model version pinned and stored with each vector?
  • [ ] Is there a confidence gate for RAG retrieval? (defined threshold, business-approved)

C.4 Agents and Tools

  • [ ] Is the tool manifest scoped to the task? (not all tools for all tasks)
  • [ ] Are irreversible tool calls behind human approval gates? (code-enforced)
  • [ ] Is the orchestrator deterministic? (routing decisions are code, not LLM judgment)
  • [ ] Does every agent task have a hard iteration limit and cost ceiling? (code-enforced)
  • [ ] Is every tool call logged with: input, output, caller, timestamp?
  • [ ] Are agent tasks idempotent? (retries don't create duplicate side effects)

C.5 Security

  • [ ] Is PII pseudonymized before entering any AI processing layer?
  • [ ] Is user-submitted content structurally separated from instructions? (XML delimiters or typed fields)
  • [ ] Has red team testing been completed? (Garak + manual adversarial for injection)
  • [ ] Are plugins/skills/MCP servers in the governance registry and reviewed?
  • [ ] Is the system prompt free of secrets, API keys, and internal URLs?
  • [ ] Is output PII scanning active before responses are returned?

C.6 Observability and Evaluation

  • [ ] Are OTel GenAI spans instrumented on all LLM calls?
  • [ ] Is cost attribution active? (team, feature, workflow tags on every call)
  • [ ] Is there an offline eval suite with minimum 20 cases (including adversarial)?
  • [ ] Does the eval suite gate production deployments? (failing evals block promotion)
  • [ ] Is online eval sampling active? (2-5% of production traffic)
  • [ ] Are drift detection alerts configured? (model version, prompt version, retrieval quality)

C.7 Compliance and Governance

  • [ ] Is the system registered in the AI inventory?
  • [ ] Is the EU AI Act risk tier assessed? (for systems with EU users/data)
  • [ ] For regulated use cases: is a model risk owner assigned?
  • [ ] Is the audit trail append-only and retention-compliant?
  • [ ] Is the data processing agreement with model providers in place?
  • [ ] For agentic systems: is the agent identity separate from user identity?

APPENDIX D — ADR Templates for Common AI Decisions


D.1 Model Selection ADR

Decision: Which LLM for [use case]

Context: - Use case: [description] - Latency SLA: [P99 target] - Data sensitivity: [classification — can it leave the perimeter?] - Volume: [requests/day, tokens/day] - Quality requirement: [what does acceptable output look like]

Options Considered:

Option Model Estimated Cost/Month Data Residency Quality Assessment
Option A [model] $X [in-VPC / external] [eval score or human judgment]
Option B [model] $Y [in-VPC / external] [eval score or human judgment]

Decision: [Chosen option]

Rationale: [Why this model, what constraints drove the decision]

Consequences: - Cost: [monthly at current volume, at 3x volume] - Data handling: [what data goes where, under what agreement] - Deprecation risk: [model version, deprecation date if known, fallback] - Reversibility: [what would be required to switch models]

Review trigger: Model deprecated, cost exceeds $X/month, quality falls below baseline


D.2 RAG Confidence Threshold ADR

Decision: Minimum retrieval confidence score to generate an answer vs. escalate

Context: - Use case: [what is the AI answering] - Consequence of wrong answer: [regulatory exposure, customer harm, financial impact] - Escalation cost: [human agent time, user friction]

Business sign-off required: [yes — list stakeholders who must approve]

Proposed threshold: [0.XX]

Rationale: [Evidence for this value — pilot data, comparable implementations]

Consequences of threshold too high: [excessive escalation, underutilized AI, higher support cost]

Consequences of threshold too low: [wrong answers, regulatory risk, customer harm]

Monitoring: [How threshold is evaluated in production — weekly review, adjustment triggers]

Owner: [Who is accountable for this threshold decision — not engineering]


D.3 Human-in-the-Loop Placement ADR

Decision: Which agent actions require human approval before execution

Context: - Agent workflow: [description] - Tools available to agent: [list]

Irreversibility Analysis:

Tool / Action Reversible? Consequence of Error Human Approval Required?
[tool name] Yes/No [description] Yes/No

Approval mechanism: - Approval interface: [what the human sees] - Timeout behavior: [what happens if no response in T hours] - Override process: [can approval be bypassed, under what conditions] - Audit record: [what is logged with the approval decision]

Decision: [List of tools that require human approval]

Rationale: [Why these specific tools — irreversibility, consequence, regulatory requirement]

Exception process: [Under what conditions can this be changed, who approves]


D.4 Prompt Change Control ADR

Decision: Process for making changes to production system prompts

Context: - System: [AI feature name] - Current prompt template: [ID and version] - Change frequency: [estimated] - Risk level: [customer-facing, internal, regulated]

Change process: 1. Change proposed in prompt management system by [role] 2. Automated eval suite runs against the proposed change 3. If eval score < [threshold]: change rejected, returned to proposer with failure report 4. If eval score ≥ [threshold]: change available for staging 5. Staging deployment: [how validated in staging] 6. Production promotion: requires approval from [role] 7. Rollback procedure: [how quickly, who can trigger]

Emergency override process: [for urgent production fixes — higher accountability, documented]

Audit requirements: [every change logged with: proposer, reviewer, eval results, rationale]


APPENDIX E — Startup Landscape Map

Current as of mid-2026. Verify revenue and valuation data before citing — these change frequently.

E.1 Vertical AI

Category Company Problem Status Lesson
Legal Harvey Contract review, legal research, discovery $11B valuation, ~$200M ARR, ~50% of Am Law 100 Domain data moat; embedded legal engineering
Legal Legora European legal AI, GDPR-compliant $5.55B valuation, €550M raised Geographic regulatory moat
Healthcare Abridge Ambient clinical scribe, EHR integration $5.3B valuation, mainstream US health system adoption EHR integration = distribution moat
Healthcare Hippocratic AI Clinical support agents, patient communication Hospital partnerships in production Human oversight architecture; clinical non-decision-maker
Finance Sierra Customer operations AI, enterprise chatbots Bret Taylor-founded; enterprise contracts Escalation-first architecture
Finance Klarna AI Customer service automation ~$40M/year savings (company-reported) High-volume structured workflows

E.2 AI Search & Knowledge

Company Problem Status Lesson
Glean Enterprise search across all tools >$100M ARR Integration breadth is the moat; RAG at org scale
Perplexity AI-native web search with citations $9B+ valuation Citation-first design is the trust architecture
OpenEvidence Medical literature search Growing adoption in clinical settings Authoritative corpus restriction = trust in regulated domains

E.3 AI Coding

Company Problem Status Lesson
Cursor AI-native IDE $2B ARR, Q1 2026 Redesign for AI vs. bolt-on
GitHub Copilot AI coding in existing IDE 4.7M paid, 90% Fortune 100 Distribution moat via GitHub
Claude Code CLI agentic coding 46% developer satisfaction Agentic capability; system-level access risk
Replit AI coding for non-developers $150M ARR annualized Expanding the developer population

E.4 AI Infrastructure

Company Problem Status Lesson
Groq Custom LPU inference, low latency Production at scale Hardware moat; custom silicon for AI
Weaviate Vector database with hybrid search Open source + managed; widely deployed Hybrid search built-in differentiates
Qdrant High-performance vector search Growing enterprise adoption Rust performance; open source
LiteLLM LLM gateway / proxy De facto open-source standard Protocol abstraction value
Unstructured.io Document parsing for AI Production standard for RAG ingestion Infrastructure layer for unstructured data
Scale AI Data labeling and AI evaluation Used by major model labs Human eval as infrastructure
ElevenLabs Voice AI infrastructure Enterprise adoption growing Platform layer beats application layer
Deepgram Domain-specific speech AI Proprietary speech models Data moat in domain-specific audio

E.5 AI Security

Company Problem Status Lesson
Garak (NVIDIA) LLM vulnerability scanning Open source standard New attack surfaces need new tools
Lakera Real-time LLM guardrails Enterprise adoption growing Managed safety as infrastructure
Knostic AI data visibility / oversharing Growing with Copilot deployment Oversharing is the #1 enterprise AI risk

APPENDIX F — Glossary

Agent (AI agent): A system that receives a goal, plans steps, takes actions via tool calls, observes results, and adapts until the goal is achieved or a limit is reached. Distinguished from a chatbot by its ability to take actions with real-world effects.

AGI (Artificial General Intelligence): Hypothetical AI with general reasoning capability equal to or exceeding human intelligence across all domains. Not currently demonstrated; timeline highly uncertain.

A2A (Agent-to-Agent Protocol): Google-initiated open protocol for inter-agent communication and task delegation. Standardizes how one AI agent delegates work to another.

Agentic loop: The cycle of reason → act → observe → reason that characterizes agent behavior. Each loop iteration makes a tool call and observes its result.

Chunking: The process of splitting source documents into segments for vector embedding and retrieval. Chunk strategy (size, overlap, semantic vs. fixed) significantly affects RAG quality.

Context window: The maximum amount of text (measured in tokens) that a model can process in a single inference call. Content outside the context window is not available to the model.

Cross-encoder: A re-ranking model that evaluates query and document together, producing a direct relevance score. More accurate than bi-encoder (vector) retrieval but computationally expensive at query time.

Embedding: A numerical vector representation of text, capturing semantic meaning. Similar texts produce similar (close) vectors. Used for semantic search in RAG systems.

Eval (evaluation): A test for AI behavior — asserting that a given input produces an appropriate output according to defined criteria. Analogous to unit tests but for probabilistic AI behavior.

Fine-tuning: Training a pre-trained model on domain-specific data to improve performance on domain-specific tasks. Creates a new model variant that requires ongoing governance.

Guardrails: Architectural controls that constrain AI system inputs and outputs: input filtering (blocking harmful prompts), output filtering (blocking harmful responses), and action scope enforcement (limiting what an agent can do).

Hallucination: When an LLM generates factually incorrect information with apparent confidence. A structural property of LLMs, not a bug to be fixed — managed through grounding, citations, and confidence gating.

Human-in-the-loop: Architectural pattern where a human provides judgment, approval, or oversight at defined decision points in an AI workflow. Not the same as "a human can be involved if needed."

Idempotency: The property of an operation that produces the same result regardless of how many times it is executed with the same input. Required for safe retry logic in agent workflows.

KV cache (Key-Value cache): Memory that stores intermediate computations (attention keys and values) from prior tokens in the context. Reusing the KV cache avoids recomputing these for repeated prompt prefixes (prefix caching).

LoRA (Low-Rank Adaptation): A parameter-efficient fine-tuning technique that trains a small adapter layer rather than full model weights, enabling domain adaptation at a fraction of the cost and compute.

LLM (Large Language Model): A neural network trained on large text corpora to predict the next token given prior context. The foundation of modern AI systems including GPT, Claude, Gemini, and Llama.

MCP (Model Context Protocol): Anthropic-originated open protocol for agent-to-tool connectivity. Standardizes how AI agents access tools, data, and prompts via MCP servers. Now governed by the Agentic AI Foundation under the Linux Foundation.

Model drift: Changes in model behavior over time — either because the model provider updated the model, or because the input distribution shifted. Requires ongoing monitoring to detect.

Model risk: The risk that a model produces incorrect outputs that cause financial, regulatory, or operational harm. Subject to formal governance requirements (SR 11-7 successor in financial services).

Orchestrator: The component that coordinates agents in a multi-agent system. Should be deterministic code (not an LLM) for routing and sequencing decisions.

PagedAttention: vLLM's technique for managing GPU memory for KV caches in non-contiguous pages, like virtual memory in an operating system. Eliminates KV cache fragmentation, dramatically increasing concurrent request capacity.

Perplexity (in AI context): A measure of how well a language model predicts a sample — lower perplexity indicates better model performance. Distinct from the company Perplexity.

Prompt injection: An attack where malicious instructions in user input, retrieved documents, or external data override the model's intended behavior.

Quantization: Reducing model weight precision (from FP32 to FP16 to INT8 to INT4) to reduce memory requirements and increase inference speed at some quality cost. Common formats: GGUF (Ollama, CPU), AWQ and GPTQ (GPU production), FP8 (H100 native).

RAG (Retrieval-Augmented Generation): The pattern of retrieving relevant content from a knowledge base before generation, grounding the LLM's output in specific information.

Reasoning model: An LLM that uses extended test-time computation (a "thinking" phase) before producing a final response. Examples: OpenAI GPT-6 with reasoning effort, Claude adaptive thinking, Gemini 3.x thinking levels (see Appendix G, G.2).

Semantic caching: Caching LLM responses indexed by semantic similarity, serving cached responses for queries that are similar (not identical) to previously answered queries. Reduces API calls for repetitive queries.

Shadow AI: AI tools used by employees outside official IT oversight — consumer ChatGPT, personal AI accounts, unauthorized browser extensions — often with corporate data.

SLM (Small Language Model): Models under approximately 10B parameters, optimized for efficiency. Examples: Phi-4-mini (3.8B), Gemma 4 small variants (E2B/E4B), Llama 3.2 1B/3B. Suitable for edge deployment and on-device inference.

Strangler fig: Incremental replacement pattern — new functionality gradually replaces old, with both running in parallel during the transition. Applied to adding AI to legacy systems.

System prompt: Instructions given to an LLM at the beginning of a context that define its role, constraints, and behavior. Processed by the model but not part of the user-visible conversation.

Temperature: A hyperparameter controlling LLM output randomness. Temperature 0 = deterministic (same input → same output). Temperature 1 = full randomness. Production systems typically use low temperature for consistency.

Token: The unit of text that LLMs process. Approximately 0.75 English words. Pricing for LLM APIs is denominated in tokens (per million input tokens, per million output tokens).

Token budget: A hard limit on the number of tokens consumed by a request, an agent task, or a user session. Enforced in code (not in prompt instructions) to prevent cost overruns.

vLLM: Open-source LLM inference serving engine, the production standard for self-hosted open-weight models. Features PagedAttention, continuous batching, and OpenAI-compatible API. NVIDIA CUDA only.

Vector database: A database optimized for storing and querying high-dimensional vectors (embeddings). Supports similarity search operations that traditional databases cannot perform efficiently.


This concludes the AI Architect's Comprehensive Reference. Artifact 1 — The Strategic & Conceptual Layer Companion: Artifact 2 — AI Integration Patterns: A Deep Reference