THE AI ARCHITECT'S FIELD CARDS — PART 4¶
Cards 34-37 (Advanced Topics)¶
Part 4 of 4. Covers the Advanced Topics: Fine-Tuning (Card 34), Multimodal (Card 35), GraphRAG (Card 36), and Model Portability (Card 37). Part 1 covers Acts I-IV (Cards 1-11). Part 2 covers Acts V-VI (Cards 12-22). Part 3 covers the Practical Layer, Cloud, and Responsible AI (Cards 23-33). Each card follows the same six-section structure: Essence → Core Insight → Key Framework → Decision Rule → Red Flags → The One Question.
---¶
ADVANCED TOPICS¶
CARD 34 — Fine-Tuning & Model Customization¶
THE ESSENCE Fine-tuning teaches behavior, not knowledge. If the gap is facts, the answer is RAG. If the gap is a prompt nobody has seriously worked on, the answer is the prompt.
THE CORE INSIGHT Fine-tuning is the most misused technique in enterprise AI. It exists to make a behavior reflexive: a format, a style, a narrow structured-prediction task, or a large model's behavior distilled into a small, cheap one. Reinforcement fine-tuning (RFT) is the newer option for reasoning quality on tasks with checkable answers: you write a grader, and the model is rewarded for passing it. The grader is the critical artifact, and reward hacking is the main risk. Teams reach for it to make a model "know" company documents. That knowledge goes stale, cannot be cited, and cannot be updated without retraining. Done right, fine-tuning comes after a locked baseline, a serious prompt-engineering effort, and 500+ high-quality examples, and it is evaluated for catastrophic forgetting as well as task gains.
THE KEY FRAMEWORK — The Customization Spectrum + Decision Gate
Prompting → Few-shot → RAG → LoRA/QLoRA → RFT (grader-based) → Full fine-tune → Pre-training
(cheaper, faster to update ◄──────────────────────► deeper internalization)
GATE: knowledge gap? → RAG, stop. Prompt engineered ≥3 days? no → keep prompting.
Gap remains + ≥500 examples? → pick by driver:
FORMAT/STYLE → LoRA · COST/LATENCY → distillation (first price the efficient tier) ·
CHECKABLE REASONING → RFT with a programmatic grader ·
DOMAIN VOCAB → fine-tune + RAG together · TASK UNDEFINED → stop
THE DECISION RULE - Lock a base-model baseline on the full eval suite before the first training step - Use RAG for knowledge and fine-tuning for behavior. For domain-heavy work, use both together - Evaluate on held-out real data and run a regression suite for catastrophic forgetting - Treat every fine-tuned model as a governed asset: registry, data provenance, license check, rollback - Before distilling, price a cheaper off-the-shelf model: the efficient tier may beat the student with no training - Check platform risk: a fine-tune on a closed platform can't be exported, and OpenAI announced in May 2026 that it is winding down its fine-tuning platform
RED FLAGS - "Fine-tune it on our documents so it knows our business" - No baseline, so nobody can say whether training helped - Eval set drawn from the same synthetic distribution as the training data (the eval gap) - A fine-tuned model with no owner, no training-data lineage, and no rollback path - An RFT grader that was never tested for ways to score high without doing the task
THE ONE QUESTION "Is this gap about what the model knows or how it behaves, and do I have the baseline that would prove fine-tuning closed it?"
---¶
CARD 35 — Multimodal Architecture¶
THE ESSENCE Multimodal is the baseline in 2026, not a premium. The risk isn't the cost of a vision model. It's a system that silently discards the 30% of information living outside the text.
THE CORE INSIGHT Multimodal breaks three text-era assumptions. Token math: images and scanned pages cost an order of magnitude more than text intuition suggests, and doubling image resolution roughly quadruples tokens. Latency: each modality has its own latency profile, and voice has a hard real-time budget. Evals: RAGAS and ROUGE cannot tell you whether a chart value or a spoken proper noun was read correctly. Every non-text modality must be transformed into text or embeddings before it can join RAG. The architecture decision is when that happens: at ingestion, at retrieval, or at generation.
THE KEY FRAMEWORK — Transform Timing + Modality-Specific Evals
WHEN to transform: ingestion (batch, cached — cheapest per query)
retrieval (on demand) · generation (inline — costliest)
EVAL per modality: docs → cell precision/recall (>90% / >85% digital PDFs)
charts → relative error (<5% clean, <15% scanned)
voice → WER (<5% clear, <10% telephony), P95 latency <800ms
vision → visual grounding (judge must SEE the image)
THE DECISION RULE - Budget tokens and latency per modality, not with one system-wide number - Transform at ingestion whenever content is queried more than once - Build a human-annotated eval set per modality (50–200 cases) before production - Any LLM judge must receive the same image or audio the system received
RED FLAGS - Text-only OCR pipeline on documents where tables, charts, and stamps carry the meaning - Cost model built from text-RAG intuition (wrong by ~10× for document intelligence) - A text-only judge "evaluating" visual answers - Voice pipeline with P95 above ~1.2s presented as "real-time"
THE ONE QUESTION "What information in our inputs lives outside the text, and how would I know if the system were silently getting it wrong?"
---¶
CARD 36 — GraphRAG & Knowledge Graph Architecture¶
THE ESSENCE Vector search finds text that is similar. Some answers live between documents, in the relationships connecting entities. When the answer is in the traversal, you need a graph.
THE CORE INSIGHT Flat RAG structurally fails on three query types: multi-hop relationships (customer → contract → entity → jurisdiction), entity-network conjunctions ("expiring in 90 days AND an open Sev-1"), and corpus-wide themes. GraphRAG adds entity identity, directed relationships, and global summaries. It also costs 5-20× more to index and brings a production data asset that needs an owner: extraction, entity resolution, freshness, access control, and PII erasure. Its value is fewer failures on a specific query class, not lower cost.
THE KEY FRAMEWORK — The GraphRAG Decision + Adoption Path
Entities with cross-document relationships? no → vector RAG
Key queries are about relationships? no → vector RAG
>1,000 entity-dense documents? no → wait
Multi-hop queries failing today (>10%)? no → monitor
Team can OWN a graph in production? no → fix capacity first
→ YES: Graph-Augmented RAG → Graph-First → Hybrid Router (in that order)
THE DECISION RULE - Invest only on a documented, recurring, business-consequential failure of vector RAG - Start with Graph-Augmented RAG; add graph-first retrieval once extraction quality is proven - Entity resolution is the hard part: budget for it, measure it, and govern the ontology - Treat the graph as a production data asset: freshness SLA, node-level access control, PII erasure, read-only query path
RED FLAGS - GraphRAG adopted because it's interesting, with no failing query class behind it - No owner for the graph, so it goes stale and gives confident wrong answers - Text2Cypher queries executed without validation (or with write access) - PII nodes with no tested deletion path
THE ONE QUESTION "Which real, recurring user questions fail today because the answer lives in relationships across documents, and who will own the graph that answers them?"
---¶
CARD 37 — Model Portability¶
THE ESSENCE The model is commoditizing; your integration with it is not. The model is a replaceable part only behind a contract you own.
THE CORE INSIGHT A gateway normalizes the API call, not the model's behavior. Swaps break at ten layers: transport, parameters, tool calling, structured output, reasoning controls, limits, tokenizer, behavior, caching/state, and commercial access. Many of those breaks are silent, such as a changed default or a different token count. Same-vendor version upgrades break too, and providers have withdrawn access for business reasons (Windsurf 2025, Cursor 2026). Portability comes from four things: a capability interface, a per-model profile, a behavioral contract that defines "equivalent," and a fallback you have actually exercised.
THE KEY FRAMEWORK — The Portability Stack + Readiness Levels
App → CAPABILITY interface (triage_ticket(), not chat()) →
MODEL ADAPTER (profile: params, features, limits, eligibility) →
GATEWAY (transport, policy) → providers
CONTRACT per capability: hard invariants · quality vs. baseline · SLOs · style
FALLBACK readiness: HOT (keep-alive traffic) · WARM (eval-passed) · COLD (named)
Tier A → hot · Tier B → warm · Tier C → cold acceptable
THE DECISION RULE - Application code calls capabilities, never models or provider SDKs - Always set reasoning depth and limits explicitly; never rely on a model's default - Cap input size at what your weakest acceptable fallback can handle - Decide swaps on paired comparison over a sufficient dataset, measuring cost per completed task - Run a quarterly candidate eval that includes the next version of your current primary - Record deliberate lock-in in a ledger with an exit cost and a revisit trigger
RED FLAGS - "We use LiteLLM, so we're portable" (transport ≠ semantics) - A fallback that has never served real traffic, or with quota sized only for keep-alive - Prompts never run on the fallback model; policy text living in per-model overlays - Conversation history stored only in the provider's raw format - No idea what a cache-cold failover would do to the bill
THE ONE QUESTION "If my largest model provider gave me ten days' notice tomorrow, which capabilities would move by configuration, and which would become a project?"
---¶
End of Field Cards. 37 cards, one per module. Each is a standalone artifact: usable as a study aid, a teaching summary, a content unit, or an exam-prep sheet. The "One Question" on each is the diagnostic that separates understanding from familiarity.