Fine-Tuning vs RAG: A Decision Framework
Short answer: retrieval changes what evidence the model can see; fine-tuning changes learned behavior. They solve different problems and should be evaluated against a baseline before being combined.
Teams often ask whether a model should be fine-tuned “on the codebase” or connected to retrieval. That framing hides the real decision. First identify the failure: missing current facts, inconsistent behavior, poor task formatting, weak search, or a model that cannot perform the task even with good context.
What each approach changes
| Question | Retrieval-augmented generation | Fine-tuning |
|---|---|---|
| Primary purpose | Supply relevant external evidence at request time | Adapt task behavior through training examples |
| Updating facts | Update the source or index | Usually requires new training data and another run |
| Inspectable evidence | Retrieved passages can be shown and cited | Knowledge encoded in weights is harder to attribute |
| Runtime cost | Search plus additional input tokens | May reduce prompt examples, but adds training and evaluation cost |
| Best fit | Current documents, policies, code, tickets, or private knowledge | Repeated style, classification, formatting, or domain behavior |
Choose retrieval when knowledge changes
Retrieval is usually the first experiment for active repositories and changing documentation because the source remains outside the model. A good system does more than place files in a vector database: it respects access control, chunks by meaningful structure, combines lexical and semantic search when useful, reranks candidates, and shows the evidence used. For code, symbol-aware search and dependency context often matter more than the choice of embedding model.
Choose fine-tuning when behavior repeats
Fine-tuning is worth evaluating when the desired input-output pattern is stable and prompt examples are insufficient or too expensive. Examples include a consistent classification taxonomy, a constrained transformation, a house style, or a specialized tool-selection pattern. Training data must represent the real distribution, include difficult negatives, and exclude secrets or material you lack permission to use.
Do not skip the baseline
Before building either system, create a fixed evaluation set and measure a strong prompt-only baseline. Record task success, factual support, latency, token use, and failure severity. Then test retrieval and fine-tuning separately. Without this sequence, a hybrid can become expensive while hiding which component helped.
Retrieval failure modes
- Wrong evidence: improve indexing, filters, queries, and reranking before changing the generator.
- Missing permissions: enforce authorization before retrieval, not after text reaches the model.
- Context overload: fewer high-quality passages can outperform a large dump.
- Stale index: monitor source-to-index delay and deletion propagation.
- Instruction injection: retrieved text is untrusted data and must not override system policy.
Fine-tuning failure modes
- Memorizing bad examples: deduplicate, review labels, and reserve a clean evaluation set.
- Behavior drift: retest after base-model or dataset changes.
- False factual confidence: training does not make changing facts reliably current.
- Hidden maintenance: dataset versioning, safety review, training runs, deployment, and rollback all need owners.
When a hybrid is justified
A hybrid is useful when evaluation demonstrates two independent needs: stable behavioral adaptation and current external evidence. For example, a tuned model may consistently produce an organization's incident schema while retrieval supplies the latest runbook and service metadata. Keep the components measurable: log which evidence was retrieved and compare the tuned model against the same retrieval context.
Practical decision sequence
- Write 50–200 representative test cases with expected outcomes and risk labels.
- Measure a prompt-only baseline with a current capable model.
- If failures come from missing facts, test retrieval.
- If failures come from repeated behavior, test fine-tuning.
- Add the second method only if it produces a measurable improvement.
- Re-evaluate cost, latency, privacy, and operational ownership before production.
Worked decision: an internal support assistant
Illustrative architecture—not a claimed benchmark. Assume an assistant must answer from frequently changing product policies and assign every ticket to one of twelve stable support categories. The two failure classes should be measured separately:
| Observed failure | Likely intervention | Reason |
|---|---|---|
| Answer cites an obsolete refund window | Retrieval and freshness controls | The fact changes independently of model behavior |
| Correct document never reaches the prompt | Improve search, filters, chunking, or reranking | The generator cannot use evidence it never sees |
| Correct evidence is present but category format varies | Prompt examples first; then evaluate fine-tuning | The desired behavior and taxonomy are stable |
| Unsupported answer despite correct evidence | Grounding checks, refusal behavior, and evaluation | Neither retrieval nor training guarantees truth |
A reasonable first release uses retrieval for policies and a deterministic category validator. Fine-tuning becomes a later experiment only if prompt-only classification remains below the acceptance target on a held-out set.
Design a fair evaluation before choosing
Use one frozen test set across the prompt-only, retrieval, fine-tuned, and hybrid candidates. The following 120-case plan is a template, not a statement of measured results:
| Slice | Cases | What it reveals |
|---|---|---|
| Ordinary current-policy questions | 40 | Baseline answer and citation quality |
| Recently changed or deleted policies | 20 | Freshness and deletion propagation |
| Ambiguous category boundaries | 20 | Behavioral consistency and calibration |
| Permission-restricted documents | 15 | Pre-retrieval authorization |
| No-answer and missing-evidence cases | 15 | Abstention rather than invention |
| Prompt-injection attempts inside documents | 10 | Instruction hierarchy and policy enforcement |
Track supported-answer accuracy, retrieval recall at the chosen context limit, citation correctness, category accuracy, abstention precision, p50/p95 latency, and cost per accepted task. Report every metric by slice; an overall average can hide a severe permission or freshness failure.
Attribute errors to the right component
- Corpus coverage: did the authoritative source exist and was it eligible for this user?
- Retrieval: did a relevant passage rank inside the context budget?
- Generation: did the answer follow the supplied evidence?
- Behavior: did it use the required taxonomy, tone, and output contract?
- Policy: did deterministic code prevent unauthorized disclosure or action?
Labeling failures at these boundaries prevents a team from fine-tuning around a broken index or rebuilding search to solve a formatting problem.
Compare total operating cost
Use a workload model rather than a one-time training invoice:
monthly retrieval cost = searches + reranking + added input tokens + index operations
monthly tuning cost = training runs + data review + evaluations + tuned-model inference
total cost per accepted task = (monthly platform + operations + review cost) / accepted tasks
Also record engineering ownership. Retrieval requires ingestion monitoring, access-control synchronization, deletion handling, and citation UX. Fine-tuning requires dataset governance, training reproducibility, model-version compatibility, regression evaluation, and rollback. Provider availability and supported base models change, so verify the current vendor documentation before committing to a training pipeline.
Data governance checklist
- Document the legal and contractual basis for indexing or training on each source.
- Separate training, validation, and final test examples by entity or time where leakage is possible.
- Deduplicate near-identical examples before splitting the dataset.
- Redact secrets and unnecessary personal information from prompts, traces, and training files.
- Make source deletion propagate to retrieval; define what deletion means for trained artifacts and backups.
- Version the corpus snapshot, chunker, embedding or search settings, prompt, dataset, grader, and model.
Phased rollout with stop conditions
- Ship a prompt-only baseline behind logging that excludes unnecessary sensitive content.
- Evaluate retrieval offline, then shadow it without showing answers to users.
- Release retrieval to a small group with citations and a visible fallback.
- Train only after the behavioral error slice is large and stable enough to justify it.
- Compare the tuned candidate against the same retrieval context and held-out set.
- Stop the rollout if restricted-document exposure, unsupported-answer rate, or severe regression exceeds the predefined threshold.
Primary references
Bottom line
Use retrieval to supply current, inspectable knowledge. Use fine-tuning to improve stable behavior. Let evaluation—not a fashionable architecture—decide whether you need either or both.
The running cost on the RAG side of the decision
The fine-tuning versus RAG decision is usually framed as a capability trade-off, but the two have completely different cost shapes: fine-tuning is an upfront cost plus a per-token delta, while RAG is a pure per-query cost that scales linearly with traffic. Below, the RAG side: one query at 25K input, 15K cached, 2K output. Of the 25,000 input tokens, 15,000 are billed at the cache-read rate and 10,000 at full input rate.
| Model | Cost per query | Monthly at 30,000 querys |
|---|---|---|
| DeepSeek V4.1 Flash | $0.0027 | $82.35 |
| GPT-5.6 Luna | $0.0047 | $141 |
| Gemini 3.8 Flash | $0.016 | $484 |
| Claude Haiku 4.5 | $0.021 | $645 |
At thirty thousand queries a month the small models keep retrieval in the low tens of dollars, which is the number a fine-tuning run has to beat. Fine-tuning wins on cost only at volumes or latencies where the per-query path itself is the constraint.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.