Skip to content

MODULE 34 — Fine-Tuning & Model Customization

⚠️ Currency note: This module is accurate as of October 2026. Model names, GPU and token prices, and provider fine-tuning offerings change quarterly, and in 2026 one major provider began shutting its fine-tuning platform down (see §34.1). Volatile facts live in Appendix G — Current Landscape. Every dollar figure below is illustrative. Re-run the examples with current prices before you quote them.

34.1 Why Fine-Tuning Exists (and When It's the Wrong Answer)

Fine-tuning is the most misused technique in enterprise AI. Teams reach for it when they should be engineering better prompts. They spend weeks and thousands of dollars training a model to do something a well-written system prompt would accomplish in an afternoon. Understanding why fine-tuning exists requires understanding the full customization spectrum — and where each technique belongs on it.

The fundamental problem fine-tuning solves is behavioral internalization: making the model respond in a particular way as a reflex, not as the result of parsing instructions at inference time. This is distinct from the knowledge problem (which is RAG's job) and distinct from the capability problem (which requires pre-training).

The customization spectrum:

CUSTOMIZATION SPECTRUM — cost vs. update-frequency vs. data requirement

TECHNIQUE          │ Cost     │ Data Needed    │ Update Freq │ Best For
───────────────────┼──────────┼────────────────┼─────────────┼─────────────────────
Prompting          │ $0       │ 0 examples     │ Instant     │ Behavior guidance,
                   │          │                │             │   persona, format
───────────────────┼──────────┼────────────────┼─────────────┼─────────────────────
Few-shot           │ $0       │ 3–20 examples  │ Instant     │ Output format,
                   │          │   (in-context) │             │   classification schema
───────────────────┼──────────┼────────────────┼─────────────┼─────────────────────
RAG                │ $–$$     │ Any corpus     │ Real-time   │ Factual knowledge,
                   │          │   (retrieval)  │             │   current information
───────────────────┼──────────┼────────────────┼─────────────┼─────────────────────
Fine-tuning        │ $$$      │ 500–100K+      │ Hours–days  │ Behavior, style,
(LoRA/QLoRA)       │          │   examples     │             │   domain tone,
                   │          │                │             │   format consistency
───────────────────┼──────────┼────────────────┼─────────────┼─────────────────────
Full fine-tuning   │ $$$$     │ 10K–1M+        │ Days–weeks  │ Deep capability
                   │          │   examples     │             │   acquisition, new
                   │          │                │             │   reasoning patterns
───────────────────┼──────────┼────────────────┼─────────────┼─────────────────────
Pre-training       │ $$$$$    │ Billions of    │ Months      │ New domain at
                   │          │   tokens       │             │   foundation level

The table encodes the two most important rules: techniques to the left cost less and update faster. Techniques to the right internalize deeper patterns. The correct choice is determined by what is actually causing the performance gap, not by what the team finds technically interesting.

When fine-tuning is the right answer: - The model produces the correct information in the wrong format, consistently, even with explicit format instructions - The model needs to adopt a highly specific stylistic register that cannot be described in a system prompt - Inference latency or token cost is the constraint and the task is well-defined enough to be learned — so you want to distill a large model's behavior into a smaller fine-tuned model - The model must perform a structured prediction task (classification, extraction, slot-filling) at very high accuracy that few-shot examples do not achieve - A proprietary reasoning pattern, evaluation rubric, or domain-specific decision logic needs to be internalized

When fine-tuning is the wrong answer: - The gap is about knowledge the model doesn't have (the answer: RAG) - The team has fewer than 200–500 high-quality labeled examples (too little data) - The task definition is still being refined (you will need to retrain every time the task changes) - The problem is a bad prompt (try harder before training) - There is no evaluation baseline (you cannot measure whether fine-tuning helped)

The most common and expensive mistake: teams fine-tune a model on proprietary documents to make it "know" the company's internal information. This is the knowledge problem. The model memorizes information at training time that will become stale, cannot be updated without retraining, cannot be cited, and cannot be audited. RAG solves the knowledge problem. Fine-tuning does not.

The 2026 market signal. On May 7, 2026, OpenAI announced it is winding down its self-serve fine-tuning platform. Organizations that had never run fine-tuning lost access that day. From July 2, 2026, organizations with no inference on a fine-tuned model in the previous 60 days also lost the ability to create jobs. On January 6, 2027, no customer can create new fine-tuning jobs. Existing fine-tuned models stay available for inference until their base models are deprecated. OpenAI's stated reason was that newer base models such as GPT-5.5 follow instructions and formats much better, and that prompt-based approaches are now cheaper and faster (reported). Two lessons for architects follow. First, the bar for "prompting cannot do this" keeps rising, which reinforces everything above. Second, a model fine-tuned on a closed provider's platform is an artifact you cannot export. When the provider changes course, you cannot retrain it and you cannot move it. Fine-tune open-weight models, or treat closed-platform fine-tunes as deliberate lock-in (Module 37 §37.10).


34.2 Fine-Tuning Techniques: Which Parameters You Train

Every fine-tuning job makes two independent choices. The first is which parameters you update: all of them (full fine-tuning) or a small adapter (LoRA, QLoRA, other PEFT methods). This section covers that choice. The second is what signal you train toward: imitating examples (supervised fine-tuning, SFT), preferring one answer over another (preference optimization, e.g., DPO), or maximizing a score from a grader (reinforcement fine-tuning, RFT). §34.3 covers that second choice. The two combine freely. For example, GRPO-based RFT with a LoRA adapter on a 9B open-weight model is a common 2026 pattern.

Full Fine-Tuning

Full fine-tuning updates all parameters in the model during training. Every weight that was set during pre-training is eligible for gradient updates. For a 7B parameter model, this means updating 7 billion floating-point values per training step.

The computational requirements scale directly with parameter count. Full fine-tuning of a 7B model requires at least 2 × 80GB A100s (for model + optimizer states + gradients). Full fine-tuning of a 70B model requires a multi-GPU cluster. At the extreme end, fully fine-tuning a 400B-class dense model (e.g., Llama 3.1 405B) requires 16+ H100s and is outside the budget of most organizations. Today's largest open-weight models are 1T+ parameter mixture-of-experts models (Appendix G), which puts full fine-tuning of them even further out of reach.

When it is justified: when the task requires deep restructuring of the model's behavior — for example, converting a general-purpose chat model into a specialized code-generation model with a fundamentally different output distribution. For most enterprise tasks, full fine-tuning's cost is not justified by the marginal capability gain over LoRA.


LoRA (Low-Rank Adaptation)

LoRA is the production standard for fine-tuning. Rather than modifying the existing weight matrices directly, LoRA adds small trainable adapter matrices alongside the frozen base model weights. The key insight is that the weight updates needed for a specific task have a much lower intrinsic dimensionality than the full weight matrix — they can be represented as the product of two small matrices.

FULL FINE-TUNING vs. LoRA ARCHITECTURE

Full Fine-Tuning:
  Input → [W + ΔW] → Output
            ↑
            All 7B parameters updated
            Optimizer state: 2× parameter count in memory
            Total memory: ~112GB for 7B model (fp32 AdamW)

LoRA:
  Input → [W (frozen)] + [A × B] → Output
                          ↑   ↑
                          Low-rank adapter matrices
                          A: (d × r),  B: (r × d)
                          where r = rank (typically 8–64)
                          r << d (d may be 4096)

  For a 4096×4096 weight matrix with r=16:
    Full matrix: 4096 × 4096 = 16,777,216 parameters
    LoRA:        (4096×16) + (16×4096) = 131,072 parameters
    Reduction:   ~99.2% fewer trainable parameters

LoRA adapters are trained while the base model weights are frozen. Only the adapter parameters receive gradient updates. This reduces memory requirements by 60–80% compared to full fine-tuning and makes fine-tuning feasible on single-GPU hardware for models up to 13B parameters.

Rank selection: rank r controls the expressiveness of the adapter. Higher rank = more parameters = more expressiveness = more data required. Typical production values: r=8 for simple style/format tasks, r=16 for moderate complexity, r=64 for complex domain adaptation. There is no free lunch — higher rank with insufficient data overfits.

Which layers to apply LoRA to: by convention, LoRA is applied to the attention weight matrices (q_proj, v_proj, k_proj, o_proj) and sometimes to the feed-forward layers (up_proj, down_proj, gate_proj). Applying to all linear layers maximizes adaptation but increases parameter count. Start with attention layers only.


QLoRA

QLoRA combines LoRA with quantization of the base model. The frozen base model weights are loaded in 4-bit precision (using NF4 or FP4 quantization) rather than 16-bit. LoRA adapters remain in 16-bit (bfloat16). This achieves the parameter efficiency of LoRA on top of a dramatically smaller memory footprint for the base model.

MEMORY COMPARISON: Fine-tuning a 70B model

Full fine-tuning (bf16):
  Model weights:    70B × 2 bytes = 140GB
  Optimizer states: 70B × 8 bytes = 560GB (Adam momentum + variance)
  Gradients:        70B × 2 bytes = 140GB
  Total:            ~840GB  →  10+ H100 80GB GPUs

LoRA (bf16 base + bf16 adapters):
  Model weights:    70B × 2 bytes = 140GB  (frozen, no optimizer states)
  Adapter params:   ~200M × 2 bytes = 0.4GB
  Adapter optimizer: ~200M × 8 bytes = 1.6GB
  Total:            ~142GB  →  2 × H100 80GB GPUs

QLoRA (4-bit base + bf16 adapters):
  Model weights:    70B × 0.5 bytes = 35GB  (4-bit NF4)
  Adapter params:   ~200M × 2 bytes = 0.4GB
  Adapter optimizer: ~200M × 8 bytes = 1.6GB
  Total:            ~37GB  →  1 × A100 80GB GPU

QLoRA made 70B fine-tuning accessible on a single A100. This is why it became a standard in the open-source community starting in 2023. The quality trade-off from quantizing the base model is modest — typically 1–3% degradation on task metrics compared to full-precision LoRA — and is acceptable for most enterprise use cases.


PEFT (Parameter-Efficient Fine-Tuning)

PEFT is the umbrella term for all fine-tuning techniques that update only a small fraction of the model's parameters. LoRA and QLoRA are PEFT methods. Other PEFT methods include:

  • Prefix tuning: prepend trainable "soft tokens" to the input sequence. The soft tokens are optimized to condition the model's behavior without modifying weights. More limited than LoRA for most tasks.
  • Prompt tuning: a simplified prefix tuning applied only at the embedding layer. Works for classification tasks with large models (>10B), less effective for smaller models.
  • IA³ (Infused Adapter by Inhibiting and Amplifying Inner Activations): scales internal activations with learned vectors. Even fewer parameters than LoRA. Best for very low-resource settings.

Decision rule for technique selection: - Default production choice: QLoRA (data ≤ 50K examples, GPU budget ≤ 2 × A100) - Sufficient GPU memory and large dataset: LoRA in bf16 (marginally better quality than QLoRA) - Maximum quality, large dataset, cluster available: full fine-tuning - Extremely low data regime (<200 examples): prompt tuning or few-shot only (fine-tuning will overfit) - Multiple tasks on a single base model: LoRA with per-task adapters (served simultaneously, see §34.8) - The gap is reasoning correctness on a task with checkable answers: reinforcement fine-tuning (§34.3), usually with a LoRA adapter


34.3 Reinforcement Fine-Tuning (RFT) for Reasoning Models

Supervised fine-tuning teaches a model to imitate examples: here is the input, here is the ideal output, make your output look like this. Reinforcement fine-tuning teaches a model to score well. You supply prompts and a grader. The model generates several candidate answers per prompt, the grader scores each one, and training shifts the model toward the behaviors that scored higher. No one writes the ideal answer. You only need a reliable way to tell a better answer from a worse one.

This is the training approach behind the 2025–2026 generation of reasoning models, and it has moved from research labs to managed cloud services. It matters to architects because it fits a gap SFT handles badly: tasks where the final answer is checkable but the reasoning path to it is hard to write down as examples.

The Three Training Signals

TRAINING SIGNAL COMPARISON

             │ SFT                  │ Preference (DPO)       │ RFT (RL with a grader)
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Data         │ Input → ideal output │ Input → chosen output  │ Input (+ reference answer
             │ pairs                │ + rejected output      │ the grader can check)
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Signal       │ "Copy this"          │ "This beats that"      │ "This scored 0.8 out of 1"
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Who writes   │ Humans or a teacher  │ Humans or a judge      │ Nobody. The model explores;
the answers  │ model                │ ranking two outputs    │ the grader scores
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Typical size │ 500–50K examples     │ 1K–50K pairs           │ Dozens to ~1,000 prompts
             │ (§34.4)              │                        │ + a validated grader
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Best for     │ Format, style, tone, │ Tone, helpfulness,     │ Reasoning correctness on
             │ extraction schema    │ "which answer is       │ checkable tasks: codes,
             │                      │ better" judgments      │ math, SQL, policy rules
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Main failure │ Overfits; copies     │ Learns judge biases    │ Reward hacking: learns
mode         │ errors in examples   │ (e.g., length)         │ to game the grader
─────────────┼──────────────────────┼────────────────────────┼──────────────────────────
Compute      │ Lowest               │ Low (offline, no       │ Highest: K samples per
             │                      │ sampling loop)         │ prompt per step + grading

DPO (Direct Preference Optimization, Rafailov et al., 2023) trains directly on chosen/rejected pairs without a separate reward model or an online sampling loop. That makes it cheap and stable, and HuggingFace TRL supports it directly (§34.5). Use it when the target is "more like this answer, less like that one" and you can collect pairs. Use RFT when the target is "correct" and a program can check correctness.

Verifiable Rewards (RLVR)

Reinforcement learning with verifiable rewards (RLVR), a term introduced in AI2's Tülu 3 work (November 2024), replaces the learned reward model of classic RLHF with a deterministic check: correct answer gets reward, incorrect gets none. A verifiable reward is cheap, consistent, and much harder to game than a judge model. Use one whenever the task allows it.

TASKS WITH VERIFIABLE REWARDS (programmatic graders)

Task                          │ Grader checks
──────────────────────────────┼─────────────────────────────────────────────
Medical / billing coding      │ Predicted code set vs. reference set (F1)
Structured extraction         │ JSON schema valid AND field values match
Text-to-SQL                   │ Query executes; result rows match reference
Code generation               │ Hidden unit tests pass; no test-file edits
Math / financial calculation  │ Final number matches within tolerance
Policy / eligibility rules    │ Decision + cited rule ID match reference
Classification / routing      │ Label matches; confusion-weighted penalty
Tool-call planning            │ Tool sequence and arguments match reference

NOT VERIFIABLE (needs a model grader; higher hacking risk):
  Open-ended writing quality, empathy, persuasiveness, "helpfulness"

GRPO: Why RFT Became Practical

Group Relative Policy Optimization (GRPO) was introduced in the DeepSeekMath paper (February 2024) and popularized by DeepSeek-R1 (January 2025). DeepSeek trained R1-Zero with large-scale RL using only rule-based rewards: an accuracy reward (is the final answer correct?) and a format reward (is the reasoning in the required structure?). Classic PPO-based RLHF needs a separate value model about the size of the policy. GRPO drops it. For each prompt it samples a group of answers and scores each answer relative to the group's average. Removing a whole model from memory is a large part of why RL fine-tuning now fits on hardware that teams already use for LoRA.

THE RFT LOOP (GRPO-style)

  ┌──────────────┐   K samples per prompt (e.g., K = 8)
  │ Prompt batch │──────────────────────────────┐
  └──────────────┘                              ▼
                                      ┌───────────────────┐
                                      │  Policy model     │
                                      │  (base + adapter) │
                                      └───────────────────┘
                                                │ answers a1..aK
                                                ▼
                                      ┌───────────────────┐
                                      │  GRADER           │  ← the critical artifact
                                      │  score each 0–1   │
                                      └───────────────────┘
                                                │ r1..rK
                                                ▼
                         advantage_i = (r_i − mean(r)) / std(r)
                                                │
                                                ▼
                         Update: raise probability of above-average
                         answers, lower below-average ones
                         (with a KL penalty to stay near the base model)
                                                │
                         Every N steps ─────────┴──► held-out eval
                                                     (different grader
                                                      where possible)

The consequence architects must know: if every sample in a group gets the same score, the advantage is zero and that prompt teaches nothing. Prompts the base model always fails, or always passes, are wasted compute. The useful training set is prompts the base model solves sometimes. Measure base pass rates before training and filter on them.

Managed RFT Offerings (as of October 2026)

PROVIDER-MANAGED RFT (as of October 2026 — verify before use)

Provider / service          │ Models                     │ Graders / rewards        │ Status notes
────────────────────────────┼────────────────────────────┼──────────────────────────┼──────────────────────────
OpenAI fine-tuning API      │ o4-mini only (o-series     │ string_check,            │ Platform winding down:
                            │ reasoning models)          │ text_similarity,         │ closed to new orgs
                            │                            │ score_model (LLM judge), │ May 7, 2026; no new
                            │                            │ python, multi (weighted) │ jobs after Jan 6, 2027
                            │                            │                          │ (§34.1)
────────────────────────────┼────────────────────────────┼──────────────────────────┼──────────────────────────
Microsoft Foundry           │ o4-mini                    │ Grader-based; GPT-4.1 /  │ Active Apr 2026 (Global
                            │                            │ 4.1-mini / 4.1-nano as   │ Training added). Agentic
                            │                            │ model graders (Apr 2026) │ RFT in private preview
                            │                            │                          │ (reported). Whether
                            │                            │                          │ OpenAI's wind-down
                            │                            │                          │ affects Foundry: verify
────────────────────────────┼────────────────────────────┼──────────────────────────┼──────────────────────────
Amazon Bedrock              │ Amazon Nova 2 Lite         │ Rule-based graders, AI   │ Launched Dec 3, 2025;
                            │ (Dec 2025); Qwen3-32B and  │ judges, built-in         │ open-weight support
                            │ gpt-oss-20B (Feb 2026)     │ templates, custom Lambda │ Feb 17, 2026 with
                            │                            │ graders                  │ OpenAI-compatible APIs
────────────────────────────┼────────────────────────────┼──────────────────────────┼──────────────────────────
Google Gemini Enterprise    │ Gemini Flash-class models  │ User-defined reward      │ Reported as pre-GA
Agent Platform              │ (check current list)       │ functions                │ (preview); verify
────────────────────────────┼────────────────────────────┼──────────────────────────┼──────────────────────────
Self-managed (open weights) │ Any open-weight model      │ Anything you can code    │ HuggingFace TRL
                            │ (e.g., Qwen 3.5 9B)        │                          │ (GRPOTrainer) and other
                            │                            │                          │ open-source RL stacks

AWS claims RFT on Bedrock delivers "66% accuracy gains on average over base models." That is a vendor average across its own selected tasks. Treat it as a reason to run your own pilot, not as a forecast for your task.

The portability rule from §34.1 applies with extra force here. A managed RFT job on a closed model produces weights you never hold. If the platform closes, as OpenAI's did, your investment in the model is lost. Your investment in the grader and prompt set is not: those are portable and can retrain an open-weight model elsewhere. Version and store them as first-class artifacts.

Data Needs

RFT needs far fewer examples than SFT, but each one must be gradable.

  • Size: OpenAI's guidance is to start with "several dozen to a few hundred" examples to test whether RFT helps (its limits are 50,000 training and 1,000 validation examples). In practice: 100–1,000 prompts for training and a separate 100–300 for evaluation.
  • Each prompt carries a reference the grader can check: the correct code set, the expected SQL result, the hidden unit tests.
  • Difficulty band: keep prompts the base model solves sometimes. A practical filter is to sample 8 answers per prompt from the base model and drop prompts scoring 0/8 or 8/8.
  • Distribution: prompts must look like production inputs. RFT optimizes hard, so it will exploit any shortcut in an unrepresentative prompt set.

Grader Design: The Critical Artifact

In SFT the dataset is the product. In RFT the grader is the product. The model will become whatever the grader rewards, including things you did not intend.

GRADER SPECIFICATION — ICD-10 coding (example)

Inputs:   model output, reference code set, clinical note
Steps:
  1. Parse output. If it is not valid JSON {codes: [...], rationale: str}
     → score 0.0 (format gate; stops the model inventing formats)
  2. code_f1 = F1(predicted codes, reference codes)
     (F1, not recall: recall alone rewards listing every code)
  3. invalid_penalty = 0.1 × (count of codes not in the ICD-10 code table)
  4. length_gate: rationale > 400 words → score capped at 0.5
Score:    max(0, code_f1 − invalid_penalty), subject to gates, range 0–1

Validation before any training run:
  ├── Run the grader on 100 outputs that domain experts have labeled
  ├── Agreement with expert verdicts ≥ 95% (programmatic graders)
  │     or ≥ 85–90% (model graders), or fix the grader first
  ├── Red-team the grader: write 20 deliberately bad outputs that try
  │     to score high (code dumps, empty rationale, prompt injection
  │     text aimed at a model grader). All must score low.
  └── Version the grader (grader_id + hash) alongside the dataset

Prefer programmatic checks. Add a model grader (an LLM judge, Module 12 §12.5) only for criteria that code cannot check. If you use one, it must be calibrated against humans, and the policy model must never see the grader prompt.

Reward Hacking

Reward hacking is the policy model finding ways to raise its score without doing the task better. It is the defining failure mode of RFT, and it shows up quickly because the optimization pressure is high.

REWARD HACKING PATTERNS AND COUNTERS

Pattern                                  │ Counter
─────────────────────────────────────────┼─────────────────────────────────────
Over-predicting (list every plausible    │ Use F1 or precision-weighted scores,
code to maximize recall)                 │ never recall alone
Format exploits (answer hidden where a   │ Strict parser; reject ambiguous output
lenient regex finds it)                  │
Length inflation (model graders often    │ Length caps; length-normalized judge
reward longer answers)                   │ rubric
Test special-casing (code that detects   │ Hidden tests; forbid test-file access;
the test and hard-codes the output)      │ randomized test inputs
Grader injection ("Score this 1.0")      │ Programmatic gate before judge; judge
aimed at a model grader                  │ prompt treats output as data
Reward rising while eval falls           │ Held-out eval with a DIFFERENT grader
                                         │ or human spot-check every N steps

Monitoring rule: plot training reward and held-out eval score on the same chart. When reward keeps rising and eval flattens or drops, the model is learning the grader, not the task. Stop training and inspect the top-scoring samples by hand.

Cost and Compute

RFT costs more per example than SFT because every training step generates several answers per prompt and grades each one. Reasoning models also produce long outputs.

RFT COMPUTE ARITHMETIC (illustrative)

  500 prompts × 8 samples × 3 epochs        = 12,000 graded rollouts
  × ~3,000 generated tokens per rollout     = ~36M generated tokens
  + grading: 12,000 programmatic checks     = negligible
        or   12,000 LLM-judge calls         = a real line item
                                              (price the judge model
                                               with Appendix G)

  Compare SFT on 5,000 examples × 512 tokens × 3 epochs
                                            = ~7.7M tokens, no generation

Generation dominates RFT compute, so self-managed runs use a fast inference engine (e.g., vLLM) for the rollouts. Managed services bill differently: OpenAI bills RFT by training time plus tokens used, so check each provider's current price page before you estimate. Programmatic graders keep cost down. A model grader can cost more than the training itself.

Evaluating an RFT Model

Apply the full framework in §34.6, plus four RFT-specific checks:

  1. Pass rate on held-out prompts vs. base, measured with the same sampling settings you will use in production.
  2. Independent-grader agreement: score the held-out set with a second grader or a human sample. A gap between the training grader and the independent check is a reward-hacking signal.
  3. Reasoning length and latency: RL training can lengthen responses. DeepSeek reported that R1-Zero's average response length grew steadily during RL. Longer reasoning means more tokens and higher latency per call (Module 13).
  4. Out-of-task regression: RL optimizes one objective hard. Run the regression suite in §34.6 and the safety evaluation every time.

When to choose RFT: the task has checkable answers, prompting and few-shot have plateaued, the base model succeeds on some of the hard cases but not reliably, and you can build and validate a grader. If you cannot write the grader, you are not ready for RFT. You may not yet have a well-defined task (§34.9, final branch).


34.4 Data Requirements for Fine-Tuning

Data quality is the primary determinant of fine-tuning success. This is not a platitude — it is the specific finding from every large-scale fine-tuning study published between 2022 and 2025. A model trained on 500 meticulously curated examples will consistently outperform the same model trained on 10,000 scraped, unfiltered examples on the same task.

Minimum Viable Dataset Sizes

MINIMUM EXAMPLES BY TECHNIQUE AND TASK TYPE

Technique        │ Task Type                  │ Minimum  │ Recommended
─────────────────┼────────────────────────────┼──────────┼─────────────
LoRA/QLoRA       │ Format/style adaptation    │ 200–500  │ 1,000–5,000
LoRA/QLoRA       │ Domain tone/vocabulary     │ 500–1K   │ 3,000–10,000
LoRA/QLoRA       │ Structured extraction      │ 500–2K   │ 5,000–20,000
LoRA/QLoRA       │ Domain Q&A                 │ 1K–5K    │ 10,000–50,000
Full fine-tuning │ Capability acquisition     │ 10K      │ 50,000–500K
Full fine-tuning │ RLHF-style alignment       │ 50K      │ 500K+
RFT (LoRA/full)  │ Checkable reasoning task   │ 50–100   │ 200–1,000 prompts
                 │  (see §34.3)               │ prompts  │  + validated grader

Below the minimum, you are not fine-tuning — you are overfitting. The model memorizes the training set and fails on any example outside that distribution.

JSONL Instruction/Response Format

The standard fine-tuning data format is JSONL (JSON Lines), one training example per line. The structure follows the instruction-following format that matches how modern models are trained:

{"messages": [
  {"role": "system", "content": "You are a medical coding specialist. Extract ICD-10 codes from clinical notes with high precision."},
  {"role": "user", "content": "Patient presents with type 2 diabetes mellitus without complications, managed with oral medication."},
  {"role": "assistant", "content": "E11.9 — Type 2 diabetes mellitus without complications\n\nRationale: No mention of complications (neuropathy, nephropathy, retinopathy). Oral medication indicates not insulin-dependent. E11.9 is the correct billable code."}
]}
{"messages": [
  {"role": "system", "content": "You are a medical coding specialist. Extract ICD-10 codes from clinical notes with high precision."},
  {"role": "user", "content": "Patient with hypertensive chronic kidney disease, stage 3, with hypertension."},
  {"role": "assistant", "content": "I13.10 — Hypertensive chronic kidney disease with stage 3 through stage 5 chronic kidney disease, or end stage renal disease, without heart failure\n\nRationale: I13.10 is the combination code that captures both the hypertension and CKD stage 3. Do not code I10 and N18.3 separately when the relationship is documented."}
]}

Quality criteria for each training example: - The response is correct (verified by a domain expert, not inferred) - The response is complete (not truncated, not missing steps) - The response demonstrates exactly the behavior you want internalized - The input distribution matches real production inputs (same length, same vocabulary, same ambiguity level) - Negative examples (wrong outputs) are not included in the training set unless you are doing DPO/preference training

Data Contamination

Data contamination is the presence of evaluation data in the training set. A model that has seen its test cases during training produces inflated metric scores that evaporate in production. This is not a theoretical concern — it is a documented failure mode in published model evaluations.

Preventing contamination: - Split training and eval data before any data collection begins, not after - Hash each example; run deduplication across train and eval splits - If using synthetic data (see below), generate training and eval examples with different seed queries so they don't share surface-level overlap

Synthetic Data Generation

Synthetic data generation uses a larger, more capable model to generate training examples for the smaller model you intend to fine-tune. This is the most practical path to dataset construction when you don't have labeled production data.

SYNTHETIC DATA PIPELINE

Step 1: Collect seed inputs
  → 50–200 real production-style inputs (no labels needed)
  → These define the input distribution

Step 2: Generate outputs with teacher model
  → Current frontier model (e.g., Claude Opus 5.5, GPT-6 Astra,
    Gemini 3.1 Pro — current list in Appendix G)
  → Confirm the provider's terms permit training on its outputs (§34.10)
  → Prompt: "You are a [domain expert]. Given this input,
    produce the ideal output in the following format: [format spec]"
  → Generate 10–20 outputs per seed input with varied temperature

Step 3: Filter for quality
  → Manual review of random 10% sample by domain expert
  → Automated filtering: length, format compliance, perplexity
  → Target rejection rate: 20–40% (if rejecting less, filters are weak)

Step 4: Augment seed inputs
  → Use teacher model to generate additional input variations
  → Paraphrase, rephrase, add noise, change complexity level
  → Goal: 5–10× expansion of seed set

Step 5: Final deduplication
  → Exact match dedup
  → Near-dedup via MinHash on n-grams
  → Semantic dedup: cluster embeddings, keep one per cluster

Distillation-style synthetic generation is appropriate when: the task is well-specified, the teacher model produces high-quality outputs, and manual labeling at scale is not feasible. It is not appropriate when the teacher model itself makes errors on the task — you will train the student model to replicate the teacher's errors.

Data curation is the hardest part. Most practitioners underestimate this. Collecting seed inputs, running generation, filtering for quality, reviewing samples, and iterating takes 2–4 weeks for a well-run fine-tuning project. The model training itself takes hours. The ratio of data work to training work is roughly 90:10.


34.5 Fine-Tuning Infrastructure and Cost

Hardware Requirements

GPU SELECTION BY MODEL SIZE AND TECHNIQUE

Model Size │ Technique   │ Minimum GPU      │ Recommended         │ VRAM Used
───────────┼─────────────┼──────────────────┼─────────────────────┼───────────
7B         │ QLoRA       │ 1× RTX 4090 24GB │ 1× A10G 24GB        │ ~18GB
7B         │ LoRA bf16   │ 1× A10G 24GB     │ 1× A100 40GB        │ ~22GB
13B        │ QLoRA       │ 1× A10G 24GB     │ 1× A100 40GB        │ ~20GB
13B        │ LoRA bf16   │ 1× A100 40GB     │ 1× A100 80GB        │ ~35GB
70B        │ QLoRA       │ 1× A100 80GB     │ 1× H100 80GB        │ ~40GB
70B        │ LoRA bf16   │ 2× A100 80GB     │ 2× H100 80GB        │ ~75GB
70B        │ Full FT     │ 8× H100 80GB     │ 16× H100 80GB       │ ~640GB+
405B       │ QLoRA       │ 8× H100 80GB     │ 16× H100 80GB       │ ~210GB

Training Frameworks

Unsloth: The fastest option for single-GPU fine-tuning. Achieves 2–5× training speed compared to naive HuggingFace implementations through manual Triton kernel implementations. Memory-efficient: fits larger batch sizes into the same VRAM. The default choice for 7B–13B LoRA/QLoRA on a single GPU.

Axolotl: The most flexible framework. Supports virtually every PEFT method, every model family, multi-GPU distributed training, and custom data formats. Configured via YAML. The production choice when you need fine-grained control over training hyperparameters or are running distributed training.

HuggingFace TRL (Transformer Reinforcement Learning): The reference implementation. Supports SFT (supervised fine-tuning), DPO (direct preference optimization), PPO (for RLHF), and GRPO (for reinforcement fine-tuning, §34.3). More verbose than Axolotl but better documented and more actively maintained. The choice for preference-based and reinforcement fine-tuning on open-weight models.

A different training target — calibration instead of preference. DPO and PPO/RLHF above both optimize toward what a human (or a preference model) prefers. A separate paradigm, RLCD (Reinforcement Learning for Calibrated Decisions — not the same as an older, same-acronym 2023 paper on contrastive preference distillation), optimizes instead for whether a model's stated confidence matches its actual hit rate on a fixed set of outcomes. It's the training method behind the non-generative "typed-decision" models covered in Module 2 §2.2 (Category 5) — a different model category entirely, not a fine-tuning technique for the generative models this module covers, but worth knowing the two targets (preference vs. calibration) are distinct so you don't reach for DPO when what you actually want is a calibrated confidence score.

Cloud Platform Options

CLOUD COST COMPARISON (August 2025 approximate rates — carried forward,
                       NOT re-verified for October 2026; check current
                       provider pricing before estimating)

Provider      │ GPU            │ $/hr  │ Best For
──────────────┼────────────────┼───────┼──────────────────────────────────
Lambda Labs   │ A100 80GB      │ $1.89 │ Single-run jobs, spot-style pricing
Lambda Labs   │ H100 80GB      │ $2.49 │ Large model fine-tuning
Modal         │ A100 80GB      │ $2.20 │ Serverless, pay-per-second billing
RunPod        │ A100 80GB      │ $1.74 │ Cheapest single GPU option
AWS SageMaker │ ml.p4d.24xl    │ $32.77│ Enterprise, managed, audit trail
              │ (8× A100 40GB) │       │
Google Cloud   │ a2-highgpu-8g  │ $26.30│ GCP ecosystem, managed training (Gemini Enterprise Agent Platform, formerly Vertex AI)
              │ (8× A100 40GB) │       │
Azure ML      │ NC96ads v4     │ $27.20│ Microsoft ecosystem, compliance
              │ (4× A100 80GB) │       │

Managed platforms (AWS, GCP, Azure) cost 10–15× more per GPU-hour than bare-metal cloud (Lambda, RunPod, Modal). The premium buys: managed job orchestration, automatic checkpointing, enterprise compliance certifications (SOC2, HIPAA BAA), IAM integration. For regulated industries, the premium is often worth it. For a research or prototype fine-tune, it is not.

Real Cost Examples

FINE-TUNING COST ESTIMATES (illustrative; 2025 GPU rates carried forward)

Task: Style/format adaptation on an 8–9B dense model (e.g., Qwen 3.5 9B)
  Dataset: 1,000 examples
  Hardware: 1× A100 80GB on Lambda ($1.89/hr)
  Training time: ~1.5–2 hours (3 epochs)
  Total cost: $3–4

Task: Domain Q&A on an 8–9B dense model
  Dataset: 10,000 examples
  Hardware: 1× A100 80GB on Lambda ($1.89/hr)
  Training time: ~8–12 hours (3 epochs)
  Total cost: $15–23

Task: Domain extraction on a 70B-class dense model (QLoRA)
  Dataset: 10,000 examples
  Hardware: 1× H100 80GB on Lambda ($2.49/hr)
  Training time: ~24–36 hours (3 epochs)
  Total cost: $60–90

Task: Full fine-tuning of 7B model
  Dataset: 100,000 examples
  Hardware: 4× H100 80GB on Lambda (4 × $2.49 = $9.96/hr)
  Training time: ~40–60 hours
  Total cost: $400–600

Note: these are training costs only. Add data preparation costs
(engineer time) and evaluation costs (LLM-as-judge inference).

Training time rule of thumb: 1,000 tokens/second throughput on a single A100 80GB with Unsloth/QLoRA on a 7B model. A dataset of 10,000 examples × 512 tokens/example × 3 epochs = 15.4M tokens → ~4.3 hours.


34.6 Evaluation of Fine-Tuned Models

The Baseline Requirement

The single most common evaluation mistake: starting to fine-tune before measuring the base model's performance. Without a pre-fine-tuning baseline, you cannot determine whether the fine-tuned model is better, the same, or worse. You cannot quantify the improvement. You cannot justify the investment.

The baseline must be measured before a single training step. Run the base model on your complete eval suite, record every metric, and lock those numbers as the baseline. This takes one afternoon. Skipping it costs weeks of ambiguity after training.

Catastrophic Forgetting

Catastrophic forgetting is the tendency of fine-tuned models to degrade on tasks outside the fine-tuning distribution as they update weights to optimize for the training task. A model fine-tuned on medical coding will perform worse on general reasoning than the base model. A model fine-tuned on a specific writing style will lose range. This is expected and not always a problem — unless the deployment context requires the general capability.

Measuring catastrophic forgetting requires a regression eval suite: a set of benchmark tasks (coding, reasoning, instruction-following) that you evaluate on both the base model and the fine-tuned model. If the fine-tuned model regresses more than 5–10% on tasks that matter for deployment, the fine-tuning approach needs adjustment (less aggressive training, lower learning rate, fewer epochs).

The Full Evaluation Framework

FINE-TUNED MODEL EVALUATION DIMENSIONS

Dimension 1: Task Performance
  → Does the fine-tuned model perform better on the target task?
  → Metrics: task-specific (F1, exact match, BLEU, human rating)
  → Compare: base model vs. fine-tuned model on held-out eval set
  → Minimum acceptable improvement: >10% over base on primary metric

Dimension 2: General Capability (Regression Check)
  → Does the fine-tuned model retain general capability?
  → Benchmarks: internal regression suite first; public sets (MMLU,
    HumanEval, MT-Bench) are saturated or contaminated for current
    models and are only a coarse smoke test
  → Acceptable regression: <5% on benchmarks relevant to deployment
  → If regression > 10%: reduce epochs, add regularization

Dimension 3: Safety & Alignment
  → Does fine-tuning degrade safety behaviors?
  → Test: refusal of harmful requests, privacy preservation,
    factual accuracy on sensitive topics
  → Safety regression is a hard blocker — do not ship

Dimension 4: Distribution Shift Robustness
  → Does the model generalize to inputs outside the training distribution?
  → Test: slightly different phrasing, longer inputs, edge cases,
    adversarial inputs, out-of-domain queries
  → The eval gap: models that overfit to training distribution
    fail on adjacent cases with different surface form

The Eval Gap Problem

The eval gap occurs when the evaluation set shares the same distribution as the training set, producing high scores that do not generalize to production. A model fine-tuned and evaluated on the same type of synthetic data will appear to perform well in eval but fail when exposed to real user inputs that differ in phrasing, domain, or complexity.

Mitigations: - Hold out real production data (if any exists) for a "gold eval" that is never used for hyperparameter selection - Include adversarial and paraphrased versions of eval examples - After deployment, collect production failures and add them to the regression eval suite


34.7 Model Distillation

Knowledge distillation is the process of training a smaller student model to replicate the behavior of a larger teacher model. The student learns not just the correct output labels but the teacher's full output distribution (the "soft targets" — the probability mass across all tokens). This transfers the teacher's uncertainty representation and reasoning patterns, not just its final answers.

Distillation is architecturally important because it enables a specific trade: pay the frontier model cost once (at distillation time), then serve the student model at a fraction of the cost for the specific task domain.

DISTILLATION ECONOMICS
(illustrative, October 2026 list prices — see Appendix G §G.3)

Workload: 10M calls/month × (~1,500 input + ~500 output tokens)
        = 15B input tokens + 5B output tokens per month

Frontier teacher: Claude Opus 5.5 ($4.00 in / $20.00 out per 1M)
  Input:   15,000 × $4.00   = $60,000
  Output:   5,000 × $20.00  = $100,000
  Total:                     ~$160,000/month (before prompt caching)

Distilled student: Qwen 3.5 9B (Apache 2.0, dense), LoRA fine-tuned,
self-hosted
  Serving: 4 A100/H100-class GPUs (load + headroom + redundancy;
           a capacity ASSUMPTION — benchmark your own traffic)
  Cost:    4 × ~$2,000–3,000/month = ~$8,000–12,000/month
  Quality: target 85–95% of teacher on the specific task domain
           (a typical target, not a guarantee)
  Latency: usually well below the teacher's; measure p50/p95

Saving vs. teacher: ~$148,000–152,000/month
Annual: ~$1.8M vs. ~$10–30K fine-tuning investment + ongoing ops

THE 2026 BASELINE YOU MUST ALSO PRICE — an efficient-tier API,
no training at all:
  GPT-6 Luna ($0.10 in / $0.50 out):
           15,000 × $0.10 + 5,000 × $0.50 = ~$4,000/month
  → If an efficient-tier model passes your eval, it beats the
    distilled student on cost AND removes the training/ops burden.
  → Distill when it does NOT pass, or when you need self-hosting
    (data residency, air-gap), latency control, or independence
    from provider pricing and access changes (Module 37).

The Distillation Data Pipeline

Standard fine-tuning trains on human-labeled examples. Distillation trains on teacher model outputs on your actual production distribution. This distinction matters: the data captures the teacher's behavior on exactly the inputs your system will encounter.

DISTILLATION DATA PIPELINE

Step 1: Capture production inputs
  → Sample real queries from production logs (or simulate them)
  → Minimum 5,000 inputs, target 50,000+
  → The distribution must match your real production traffic,
    not a benchmark or curated dataset

Step 2: Generate teacher outputs
  → Run each input through the frontier teacher model
  → Capture full token logits if possible (for soft-target training);
    closed frontier APIs generally do not expose full logits (some
    expose limited top-k log-probabilities), so full soft-target
    distillation usually needs an open-weight teacher
  → If logits unavailable, capture top-k log-probabilities where offered,
    or sample multiple completions per input (temperature > 0)
  → This is the expensive step: 50K calls × frontier pricing

Step 3: Quality filtering
  → Remove examples where teacher output is clearly wrong
    (domain expert spot-check, automated format validation)
  → Remove examples where teacher output is ambiguous
  → Target: 85–90% pass rate

Step 4: Student training
  → Standard SFT if only text outputs available
  → Soft-target distillation (KL divergence on logits) if
    token probabilities available — higher quality transfer
  → Use QLoRA if student is 7B–13B, LoRA if 7B and GPU-rich

Step 5: Distillation-specific eval
  → Compare student vs. teacher on held-out production inputs
  → Metric: response similarity (semantic similarity + format match)
  → Quality gate: student achieves ≥90% of teacher's quality
    score on domain eval before any deployment consideration

Student Model Selection

The student model must be capable enough to learn the task. A 1B parameter model cannot replicate a 70B model's reasoning on complex multi-step problems, regardless of how much distillation data you provide. The minimum capable student size scales with task complexity:

  • Simple classification/extraction: 1B–3B parameters is sufficient
  • Structured generation with moderate reasoning: 7B–8B
  • Multi-step reasoning, analysis tasks: 13B–14B minimum, 70B preferred
  • Complex judgment tasks requiring broad knowledge: full fine-tune of 70B or don't distill

Distillation vs. Fine-Tuning

Use distillation when: the goal is cost and latency reduction for a well-defined, stable task that a frontier model already handles well.

Use fine-tuning directly when: there is no clear teacher model (proprietary task, internal style), you have human-labeled data and don't want to pay frontier model inference costs for data generation, or the task requires behavior that no existing model exhibits.

The most powerful combination: use a frontier model to generate synthetic training data (distillation-style data collection), then fine-tune the student model on that data. This is distillation without requiring logit access and is now the standard production pipeline.


34.8 Deployment of Fine-Tuned Models

Adapter Merging vs. Adapter Serving

After LoRA training, you have two deployment options for the adapter weights:

Merge the adapter into the base model: The adapter matrices A and B are mathematically combined with the original weight matrix W: W_merged = W + A×B. The result is a single model file with no architectural difference from the base model. No adapter loading overhead at inference time. Simpler serving setup. Use when: you have one task and one adapter.

Serve the adapter separately: The base model is loaded once. Adapter weights are loaded on top at request time (or per-request if multi-LoRA serving). The base model memory footprint is shared across all adapters. Use when: you have multiple fine-tuned tasks on the same base model — multi-tenant applications, multiple departments on a single serving stack.

vLLM supports multi-LoRA serving natively since v0.3. You can load up to 8–16 LoRA adapters alongside a single base model instance and route requests to the correct adapter. This amortizes the base model memory cost across all variants.

Quantization for Deployment

Fine-tuned adapters are typically merged and then quantized for production serving. The quantization format depends on the serving infrastructure:

QUANTIZATION FORMAT SELECTION

Format    │ Runtime           │ Use Case
──────────┼───────────────────┼────────────────────────────────────────
GGUF      │ llama.cpp         │ CPU serving, edge devices, local dev
          │                   │ Quantization levels: Q4_K_M (4-bit),
          │                   │ Q5_K_M (5-bit), Q8_0 (8-bit)
          │                   │ Rule: Q4_K_M for memory-constrained,
          │                   │ Q5_K_M for balanced, Q8_0 for quality
──────────┼───────────────────┼────────────────────────────────────────
GPTQ      │ vLLM, AutoGPTQ    │ GPU serving, calibrated quantization
          │                   │ Good at 4-bit with calibration dataset
          │                   │ Slower to quantize, good quality
──────────┼───────────────────┼────────────────────────────────────────
AWQ       │ vLLM, AutoAWQ     │ GPU serving, activation-aware
          │                   │ Typically better quality than GPTQ at
          │                   │ same bit-width; faster inference
          │                   │ Production default for GPU deployment
──────────┼───────────────────┼────────────────────────────────────────
FP8       │ vLLM (H100)       │ H100-native precision, near-fp16 quality
          │                   │ Best quality/speed tradeoff on H100
          │                   │ Use when H100 is available
──────────┼───────────────────┼────────────────────────────────────────
bfloat16  │ vLLM, TGI         │ No quantization, highest quality
          │                   │ Use when VRAM is not the constraint

Model Registry and Versioning

Fine-tuned models are software artifacts and must be versioned like software. The model registry is the source of truth for what model versions exist, what they are trained on, and what their eval metrics are.

Every fine-tuned model artifact in the registry must record: - Base model name, version, and exact repository revision hash (e.g., Qwen 3.5 9B at a pinned revision) - Adapter type and hyperparameters (LoRA rank, alpha, target modules) - Training dataset: name, version, record count, hash - Training run ID: links to the compute logs, hyperparameter config - Eval results: all metrics on all eval dimensions (§34.6), with comparison to base - Training timestamp and engineer responsible - Deployment status: candidate → staging → production → deprecated

HuggingFace Hub, MLflow Model Registry, and Weights & Biases Artifact Registry are all viable options. The specific tool matters less than the discipline of using it.

Rollback

Every fine-tuned model deployment must have a documented rollback path to the base model or the previous fine-tuned version. Rollback should be executable in under 15 minutes. At the container/orchestration layer, this means maintaining the previous model version image alongside the current version so rollback is a traffic routing change, not a new build.

Rollback triggers: monitored automatically; if any of these thresholds breach in production, rollback initiates: - Error rate on production requests exceeds 5% above baseline - Mean response latency degrades more than 30% above baseline - Model quality metric (sampled production eval) drops more than 10% from staging eval - Any safety evaluation regression vs. the rollback target


34.9 The Fine-Tuning Decision Framework

Before committing to fine-tuning, exhaust the cheaper alternatives. Each exit condition is specific — it is not "try prompting and see if it works," it is a measured performance gate.

FINE-TUNING DECISION TREE

Start: Production system has a performance gap
        │
        ▼
┌───────────────────────────────────────────────────────┐
│ STEP 1: Is the gap about knowledge (facts, documents, │
│         current information)?                         │
└───────────────────────────────────────────────────────┘
  │ YES → Implement or improve RAG. Stop here.
  │ NO  ↓
  ▼
┌───────────────────────────────────────────────────────┐
│ STEP 2: Have you spent ≥3 days engineering the        │
│         system prompt (examples, format spec,         │
│         persona, chain-of-thought)?                   │
└───────────────────────────────────────────────────────┘
  │ NO  → Improve the system prompt. Return to Step 2.
  │ YES ↓
  ▼
┌───────────────────────────────────────────────────────┐
│ STEP 3: After prompt engineering, what is the gap?    │
│         Measure with your eval suite.                 │
└───────────────────────────────────────────────────────┘
  │ Gap closed → Done. No fine-tuning needed.
  │ Gap remains ↓
  ▼
┌───────────────────────────────────────────────────────┐
│ STEP 4: Do you have ≥500 high-quality labeled         │
│         examples (or can you generate them)?          │
│         For RFT: ≥100–200 gradable prompts AND a      │
│         validated grader instead (§34.3)              │
└───────────────────────────────────────────────────────┘
  │ NO  → Fine-tuning is premature. Build the dataset first.
  │ YES ↓
  ▼
┌───────────────────────────────────────────────────────┐
│ STEP 5: What is the PRIMARY driver of the gap?        │
└───────────────────────────────────────────────────────┘
  │
  ├─ FORMAT/STYLE inconsistency
  │    → LoRA fine-tuning of the serving model
  │    → 500–5K examples, r=8–16
  │
  ├─ LATENCY / COST at acceptable quality
  │    → First price an efficient-tier API against your eval (§34.7)
  │    → If it fails: distillation to smaller model (§34.7)
  │    → Capture frontier model outputs on production distribution
  │
  ├─ DOMAIN VOCABULARY / entity types not in base model
  │    → Fine-tuning + RAG together
  │    → Fine-tune for domain fluency, RAG for factual grounding
  │    → Do not choose one; use both
  │
  ├─ REASONING QUALITY on checkable tasks (base model right
  │  sometimes, not reliably: codes, SQL, calculations, rules)
  │    → RFT with a programmatic grader (§34.3)
  │    → 100–1,000 prompts with reference answers
  │    → Open-weight + LoRA + GRPO, or a managed RFT service
  │    → Gate: grader validated against experts BEFORE training
  │
  ├─ DEEP REASONING PATTERN not exhibited by any base model
  │    → Full fine-tuning or large-scale LoRA
  │    → Requires 10K+ examples demonstrating the reasoning trace
  │    → Most expensive and highest risk path
  │
  └─ TASK NOT YET WELL-DEFINED
       → Stop. Define the task first.
       → Fine-tuning an ill-defined task produces an ill-defined model.
       → No amount of compute substitutes for a clear task specification.

The combined fine-tuning + RAG architecture: for domain-heavy applications (legal, medical, financial), the most effective production pattern is LoRA fine-tuning for domain fluency combined with RAG for factual grounding. The fine-tuned model understands domain vocabulary, reasoning patterns, and output format. RAG provides the specific facts, documents, and current information. Neither alone achieves what both together accomplish.


34.10 Fine-Tuning Governance

Fine-tuned models are production AI systems. The governance requirements are higher than for API-accessed frontier models, not lower, because you own the model artifact and are accountable for its behavior.

Inventory and Ownership

Every fine-tuned model in production must be in a model inventory with a named owner. The owner is accountable for the model's behavior, its eval results, and its decommissioning. "The AI team owns all models" is not an ownership assignment — it is an ownership vacuum.

Model inventory record (minimum fields):

model_id:          ft-medrec-qwen35-9b-v3
base_model:        Qwen3.5-9B (Apache 2.0), revision pinned by hash
task:              Medical record extraction (ICD-10 coding assistance)
owner:             Clinical Informatics Team, Jane Smith
deployed_at:       2026-06-15
training_data_v:   medrec-training-v3 (8,432 examples, hash: sha256:a3f...)
eval_scores:       F1=0.94 (target task), MMLU=67.2 (regression: -1.8%)
production_volume: ~45,000 inferences/day
status:            active
next_review:       2026-12-15

Training Data Provenance

You must be able to answer the following questions for any fine-tuned model in production: - Where did the training data come from? - Is the training data licensed for this use? - Does the training data contain PII? If so, how was it handled? - If synthetic: which model generated it? Is use of that model's outputs for training data permitted under the model provider's terms of service? - Can the training data be reproduced if needed for a legal or regulatory inquiry?

This last question is not hypothetical. OpenAI's terms prohibit using its outputs to develop models that compete with OpenAI, and Anthropic, Google, and xAI have comparable restrictions (reported). Whether your fine-tune "competes" is a legal question, not an engineering one. Get the answer in writing before you generate training data with a closed model. Open-weight licenses differ too: Qwen 3.5 is Apache 2.0, while Llama models use Meta's own license. These terms change. Track them, and record the terms version in the audit trail.

IP and License Exposure

Training on customer data creates questions of IP ownership. If Customer A provides 10,000 labeled examples to fine-tune your model, and you then use that fine-tuned model to serve Customer B, you have transferred Customer A's IP (embedded in the model weights) to serve Customer B's use case. This is not a theoretical concern — it is an active area of litigation and contract negotiation.

Governance rules: - Multi-tenant fine-tuning (training on data from multiple customers in one model) requires explicit contractual permission from all customers whose data is used - Single-tenant fine-tuning for an enterprise customer typically requires the trained model artifact to be owned by or exclusively licensed to that customer - In doubt: separate models per customer, and get a legal opinion before training

Regulatory Validation

Fine-tuned models used for regulated decisions — credit scoring, medical diagnosis assistance, insurance underwriting, hiring — are subject to the same validation requirements as any other algorithmic decision system. In many jurisdictions this means: - Bias testing across protected classes before deployment - Documented performance validation on representative test sets - Explainability requirements (which may conflict with the black-box nature of neural models — a tension you need to resolve in your architecture, not ignore) - Change management: retraining is a change to a regulated system and requires the same approval process as any other system change

The Fine-Tuning Audit Trail

The complete audit trail for a fine-tuned model consists of:

FINE-TUNING AUDIT TRAIL

1. Training data record
   - Dataset name and version
   - Record count and hash of the training file
   - Source of each example (human-labeled, synthetic, production log)
   - PII handling documentation
   - License clearance record

2. Training run record
   - Training framework and version (e.g., Axolotl, TRL — pin exact versions)
   - For RFT: grader ID, grader version/hash, grader validation results
   - Hyperparameters: learning_rate, num_epochs, batch_size,
     lora_r, lora_alpha, lora_target_modules, optimizer
   - Hardware: GPU type, count, cloud provider
   - Training duration and cost
   - Final training loss, validation loss (overfitting check)
   - Checkpoints saved at intervals (for rollback to any epoch)

3. Evaluation record
   - Eval suite version
   - All metric scores: task performance + regression + safety
   - Comparison to base model and previous fine-tuned version
   - Sign-off: who reviewed and approved the eval results

4. Deployment record
   - Deployment date and deploying engineer
   - Serving infrastructure: model format, quantization, framework
   - Rollback target: which version to roll back to and how
   - Traffic routing: percentage of production traffic on this version

5. Monitoring record (ongoing)
   - Sampled production quality scores
   - Error rates and latency metrics
   - Scheduled review date

This audit trail is not optional if the model is used in regulated decisions. For unregulated use cases, maintain it anyway — the cost of maintaining it is hours, and the cost of not having it when you need it is weeks.


34.11 The Fine-Tuning Checklist

A checklist is only useful if every item is a hard gate. Each item below blocks progress to the next stage if not satisfied.

Pre-Training Gates - [ ] Task is clearly defined: input format, output format, success metric - [ ] Knowledge vs. behavior gap analyzed: RAG ruled out as sufficient solution - [ ] System prompt engineering exhausted (documented effort, before/after eval scores) - [ ] Eval suite built: minimum 100 examples, covers task performance + regression + safety - [ ] Baseline eval run on base model, all metrics recorded and locked - [ ] Training dataset assembled: ≥500 examples (LoRA style/format), ≥1K (domain Q&A) - [ ] Training/eval split enforced: no contamination between sets - [ ] Data quality review: 10% random sample reviewed by domain expert - [ ] Data provenance documented: sources, licenses, PII handling - [ ] Training infrastructure identified: GPU type, framework, cloud provider - [ ] Cost estimate completed: training + eval + serving cost over 12 months - [ ] Governance: model owner assigned, inventory entry prepared - [ ] Platform risk checked: can you export the tuned weights? If not (closed-platform fine-tune), lock-in accepted in writing and an open-weight fallback plan exists (§34.1)

RFT Gates (reinforcement fine-tuning only) - [ ] Task answers are checkable; grader specification written and versioned (grader ID + hash) - [ ] Grader validated: ≥95% agreement with expert labels on 100 outputs (≥85–90% for model graders) - [ ] Grader red-teamed: 20+ deliberately gamed outputs all score low - [ ] Prompt set filtered by base pass rate: prompts scoring 0/8 and 8/8 removed - [ ] Held-out eval uses a different grader or human spot-check, not only the training grader - [ ] Reward and held-out eval plotted together; stop rule defined for reward/eval divergence - [ ] Reasoning length and latency change measured against the base model

Training Gates - [ ] Hyperparameters selected with justification (LoRA rank, learning rate, epochs) - [ ] Training run monitored: loss curves reviewed, no divergence - [ ] Validation loss tracked: no significant overfitting (val_loss ≤ 1.2× train_loss) - [ ] Checkpoints saved at each epoch for rollback capability - [ ] Training run record complete: all hyperparameters and costs logged - [ ] For RFT: top-scoring samples inspected by hand at each checkpoint for reward hacking

Evaluation Gates - [ ] Task performance: fine-tuned model exceeds baseline by ≥10% on primary metric - [ ] Regression eval: general capability degradation < 5% on relevant benchmarks - [ ] Safety eval: no regression on refusal/alignment behaviors - [ ] Distribution shift test: model tested on paraphrased/adversarial eval inputs - [ ] Eval results reviewed and signed off by task owner

Deployment Gates - [ ] Adapter merged and quantized in target format (AWQ/GGUF/FP8) - [ ] Model artifact committed to model registry with full metadata - [ ] Serving infrastructure validated: latency p50/p95 within SLA - [ ] Rollback procedure documented and tested - [ ] Monitoring configured: error rate, latency, quality sampling - [ ] Rollback triggers defined with specific thresholds - [ ] Traffic ramp plan defined: 5% → 25% → 100% with hold periods

Governance Gates - [ ] Full audit trail complete and stored in the model registry - [ ] Training data provenance documented and accessible - [ ] IP/license review completed for training data - [ ] If regulated use case: validation documentation prepared for compliance review - [ ] Review schedule set: next evaluation date on calendar


EXERCISE — The Customization Audit: Take an AI system you know (internal or production). Work through the decision framework in §34.9 systematically. Have you actually exhausted prompting — measured before and after with an eval suite, not just "we tried a few prompts"? Have you considered RAG — and if you rejected it, what specifically made it insufficient? Is the remaining performance gap about knowledge (RAG's job) or about behavior (fine-tuning's job)? What is the minimum viable dataset to attempt fine-tuning responsibly — where do those examples come from, and how long does it take to get them? What would a 100-case eval suite look like that covers task performance, regression, and safety for this system?

PONDER — The Distillation Question: In the §34.7 example (illustrative October 2026 prices), a frontier teacher costs ~$160,000/month for 10M calls, and a distilled 9B student self-hosted on four GPUs costs ~$8,000–12,000/month. That is roughly $150,000/month in savings. An efficient-tier API would cost ~$4,000/month with no training at all, so first ask what the student does that the cheap API cannot. But the math captures only the inference cost. What does the organization need to be true operationally to capture that saving? Think about: who monitors the fine-tuned model's quality drift over time (and what does "drift" even mean when there is no ground truth label for each production call)? What is the retraining cadence when the frontier model improves and the student falls behind? What happens to the user experience when the frontier model handles a novel input elegantly and the fine-tuned student returns a malformed response? At what call volume does the operational burden of maintaining a fine-tuned model outweigh the inference cost savings? Write a break-even analysis that captures both sides of this trade.

WORKSHOP — Build a Fine-Tuning Plan: For a real or hypothetical domain (legal contract review, medical coding, internal IT helpdesk, or your own choice), produce a complete fine-tuning plan with five components. (1) Dataset specification: JSONL format definition with two representative examples, target size with justification by task type from §34.4, quality criteria (what makes an example pass or fail your filter), and a synthetic generation plan if human labeling is not feasible — which teacher model, what prompting approach, what rejection criteria. (2) Technique selection: choose between LoRA, QLoRA, and full fine-tuning with explicit justification — cite the model size, GPU budget, dataset size, and the decision criteria from §34.2. Then choose the training signal (SFT, DPO, or RFT, §34.3); if RFT, write the grader specification and three ways a model could game it. (3) Baseline eval suite: write 20 specific test cases — not category descriptions but actual input/expected-output pairs — covering normal cases, edge cases, adversarial cases, and out-of-scope cases. (4) Hardware and cost plan: select a GPU tier from §34.5, estimate training time using the tokens/hour rule of thumb, and compute a total training cost. Then compute the monthly serving cost for your expected inference volume and determine at what volume the fine-tuned model becomes cheaper than the frontier model API — and compare it with an efficient-tier API as well (§34.7). (5) Deployment and rollback plan: specify the serving format and quantization from §34.8, the rollback trigger thresholds, and the monitoring metrics that would indicate the model is degrading in production.


Next: Module 35 — Multimodal Architecture