RAG
RAG vs Fine-Tuning
These are not two ways of doing the same thing. One changes what a model can look up; the other changes how it behaves. Choosing between them is easy once you know which of those you actually need — and most business cases need the first.
Drafted with AI assistance and edited by the Bralak engineering team. No client examples, performance figures or benchmark comparisons appear in these pieces — every technical claim is one a reader can check independently.
The question is usually asked as though these were competing implementations of one capability, and answered with a comparison of accuracy. That framing is the reason so many teams pick wrong.
Fine-tuning changes how a model behaves. Retrieval changes what it can see. Confusing the two is how a knowledge problem gets solved with a training run.
Knowledge versus behaviour
Start with what each mechanism physically does.
Fine-tuning adjusts the weights
You take a base model and continue training it on examples of the input-output behaviour you want. The weights shift. What comes back is a model with different defaults — a different register, a different structure to its replies, a stronger grip on a specialised vocabulary, a reliable output format. In practice this rarely means updating all of the weights: LoRA freezes the pre-trained ones and trains a small pair of low-rank matrices injected into each layer instead, and QLoRA quantises the frozen base to 4 bits so that a 65-billion-parameter model can be tuned on a single 48GB GPU. That is what made fine-tuning cheap enough to reach for. It is not what makes it the right tool.
What it is genuinely good at is teaching a manner. Classify these support tickets using our internal taxonomy. Write in the clipped house style our clinicians expect. Always return this exact JSON shape. These are patterns, learnable from examples, and hard to specify exhaustively in a prompt.
Retrieval changes the input
The weights are untouched. At query time you find the relevant material and put it in the context. The model reads it and answers from it. Knowledge lives in a store you control, and updating it is an edit to a document rather than a training run.
What each costs to build and to keep current
The build cost is where the comparison usually stops, and it is the less important half. The cadence cost is what decides whether the thing is still working in a year.
| RAG | Fine-tuning | |
|---|---|---|
| Main build effort | Ingestion, chunking, retrieval quality, permissions | Assembling and cleaning a labelled example set |
| Updating knowledge | Edit the document, re-index that document | Assemble new examples and retrain; there is no partial edit |
| Time for a correction to take effect | Minutes | A training and evaluation cycle |
| Per-query cost | Higher — retrieval, plus roughly 3,750 added input tokens a query | Lower — no retrieved context to pay for |
| Traceability | Every answer carries its sources | None; the output is a property of the weights |
| Failure when out of scope | Can be made to refuse, on a checkable signal | Answers anyway, in the trained style |
| Access control | Enforceable in the retrieval query, per user | Not expressible; the model knows what it knows |
The per-query line is worth making concrete, because it is usually assumed to be worse than it is. Eight chunks of 400 tokens is 3,200 tokens of retrieved context; add a query of about 50 and a system prompt of about 500 and a retrieval-augmented call carries roughly 3,750 input tokens more than a bare one. At Claude Opus 5’s published rate of $5.00 per million input tokens that is about $0.019 a query; at Claude Sonnet 5’s $2.00 per million, about $0.0075. Those are list prices as published by Anthropic and read on 18 September 2026 — rates change, so check that date before leaning on the figure, and substitute your own chunk count and prompt length, because every term in the sum is one you control.
That last pair is decisive more often than anything about quality. If different people are entitled to see different things, retrieval can enforce it and a fine-tuned model structurally cannot — what is in the weights is available to everyone who can reach the endpoint.
A decision table
| If the requirement is… | Then… |
|---|---|
| Answers grounded in documents that change | RAG. This is the case it exists for. |
| Answers that must cite a source | RAG. Nothing else can produce a real citation. |
| Different answers depending on who is asking | RAG, with permissions in the retrieval query. |
| A consistent output format or house style | Fine-tuning, once prompting has genuinely failed. |
| A specialised classification the base model handles poorly | Fine-tuning, on real labelled examples. |
| Lower cost or latency on a narrow, stable, high-volume task | Fine-tuning a smaller model, if the volume justifies it. |
| Grounded answers in a specialist register | Both. See below. |
When both apply
They compose cleanly, because they act on different parts of the system. A clinical assistant might be fine-tuned so its summaries follow the structure and register clinicians expect, while every clinical fact in those summaries comes from retrieval over current guidance, cited.
The sequencing is what matters. Build retrieval first and run it. It is faster to stand up, it produces the evaluation set you will need anyway, and it tells you whether the residual complaints are about knowledge — which more retrieval work fixes — or about manner, which is the only thing fine-tuning was going to address.
Doing it the other way round means training against a specification you have not tested, then discovering the problem was retrieval, and having a model to maintain that solved nothing.
Why most business cases are retrieval problems
Look at what organisations actually ask for. Answer questions from our documentation. Draft replies using our current policies. Summarise this account before the call. Find the clause that applies here.
Every one of those is a request about information the model does not have. None is a request to change how the model writes. The base models are already fluent, already competent at summarising and drafting; what they lack is your specifics. That is the definition of a retrieval problem.
There is also a quieter reason the retrieval answer keeps winning. Your policies change. Your prices change. Your process changed last week and somebody updated the wiki. A system where correcting a wrong answer means editing a document — and where you can see which document produced it — is one your organisation can actually own. A model that has to be retrained to be corrected is one that is permanently somebody else’s dependency.
This piece supports our RAG Development page, which covers how we design, build and run these systems in production.
Your Next Intelligent System Starts Here.
Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.