RAG
What Is RAG?
Retrieval-augmented generation closes the gap between what a model was trained on and what your organisation actually knows. The interesting part is not the generation. It is that a well-built RAG system changes what happens when it does not know.
Drafted with AI assistance and edited by the Bralak engineering team. No client examples, performance figures or benchmark comparisons appear in these pieces — every technical claim is one a reader can check independently.
A language model knows what was in its training data, and that data has a date on it — Claude Opus 5’s runs to May 2026, Claude Sonnet 5’s to January 2026. It does not know your pricing, your policies, the decision your leadership took last quarter, or which of your two conflicting process documents is the current one. Ask it anyway and it will answer, because producing a plausible continuation is what it does. That is the gap.
Retrieval-augmented generation closes it by changing the question. Instead of asking the model what it knows, you find the relevant material first, put it in front of the model, and ask it to answer from that. The model stops being the source of knowledge and becomes the thing that reads and explains.
The approach and the name both come from Lewis et al., published at NeurIPS 2020, which paired a pre-trained generator with a dense vector index of Wikipedia and a neural retriever. The parts have been replaced many times since; the division of labour it proposed is what survived.
The part that matters is the refusal
Most explanations stop at it looks things up first, which undersells it. The real change is to the failure mode.
A model answering from its weights has no way to distinguish knowing something from producing something that has the shape of knowing. There is no internal signal to check, so a wrong answer arrives with exactly the same confidence as a right one.
A retrieval system has such a signal. If nothing sufficiently relevant comes back, that is a fact about the corpus, available before the model is called at all. The system can say this is not in the documentation I have — and can say it for a checkable reason.
The value is not that it answers more questions. It is that it can tell the difference between the ones it should answer and the ones it should not.
The pipeline, end to end
- InputDocumentsPDFs · wikis · tickets · databases
- SystemIngestionIncremental re-indexing
- SystemChunkingStructure-aware splitting
- SystemEmbeddings
- RetrievalVector storepgvector · Pinecone
- RetrievalHybrid retrievalVector + keyword, re-ranked
- ActionGrounded responseWith citations
- Human escalationHonest refusalBelow threshold, it says so
Pull the source material in — PDFs, wikis, ticket histories, contracts, database records — and normalise it to text while keeping the structure that carries meaning. A table flattened to a wall of numbers has lost the thing that made it a table.
Split documents into retrievable units. This decision has more effect on final quality than the choice of model, and it is routinely made by accepting a default character count. Chunks of 256 to 512 tokens are the range that holds: below roughly 200 there is too little context left to interpret the passage, and above roughly 800 the embedding averages too many topics together and precision falls. Overlap of 10 to 20 per cent — 50 to 100 tokens on a 512-token chunk — stops a sentence that straddles a boundary from belonging to neither side. Split recursively on the document’s own structure first, headings, list items, code fences and table rows, and fall back to fixed character windows only where there is no structure left to cut on. Splitting every thousand characters regardless cuts arguments in half and produces fragments that are individually meaningless. For PDFs none of this is better than the parse it rests on, which is why layout-aware extraction comes before any of it.
Convert each chunk to a vector that positions it by meaning, and store it in an index built for nearest-neighbour search. HNSW is the honest default — a navigable small-world graph,
mof 16 to 32 links per node,ef_constructionof 128 to 200 while building andef_searchof 64 to 128 at query time, each of them trading recall against latency (Malkov & Yashunin). DiskANN suits corpora too large to hold in memory; IVF-PQ suits a tight memory budget. Keep the original text and its source reference alongside the vector — that reference is what a citation is made of later.At query time, find the candidate chunks. Vector similarity handles paraphrase well and exact tokens badly, which is why serious systems run BM25 alongside it and merge the two lists. A part number is a keyword problem wearing a semantic disguise. Merge with Reciprocal Rank Fusion, which scores a document by summing
1 / (k + rank)over the lists it appears in, withk = 60— the constant its authors fixed during a pilot investigation and never altered (Cormack, Clarke & Büttcher, SIGIR 2009). RRF reads only ranks, so two retrievers on entirely unrelated score scales need nothing normalised between them. Weighted score fusion is the alternative, usually around 0.6 to 0.7 dense against 0.3 to 0.4 sparse, and it has to be retuned per corpus.Take 50 to 150 candidates from the union of the two lists and score them properly with a cross-encoder, which reads the query and the passage together rather than comparing two vectors computed apart. Keep the 5 to 20 that survive. Retrieval is tuned to be broad and cheap; re-ranking is where precision comes from, and skipping it means padding the prompt with near-misses that give the model room to pick the wrong one.
Give the model the question and the surviving passages, with instructions to answer only from them and to cite what it used. This is the last step and the least of the engineering.
Two of those stages have defaults worth knowing you have accepted. BM25 is almost always left at k1 = 1.2 and b = 0.75, which are Lucene’s implementation defaults rather than a finding — the survey that defines the model is more careful, offering 1.2 < k1 < 2 and 0.5 < b < 0.8 as ranges that work in many circumstances while noting that the best values depend on the documents and the queries you actually have (Robertson & Zaragoza, 2009).
The other is that a chunk is embedded alone, stripped of the document it came from. Contextual retrieval is the fix: generate a sentence or two situating each chunk in its parent document and prepend it before embedding, so a clause that says the limit is £5,000 carries the fact that it is the excess clause of the 2026 motor policy. Anthropic published the technique in September 2024 along with its own measurement of it — combining contextual embeddings with contextual BM25 moved their top-20 retrieval failure rate from 5.7 per cent to 2.9 per cent, and adding re-ranking took it to 1.9 per cent. Those are their figures on their corpora, which is exactly how to read them: evidence the approach is worth testing, not a number to expect.
Why retrieval quality dominates generation quality
There is a hard ceiling in this architecture, and it is worth stating plainly: the model cannot answer from a passage that retrieval did not return. If the right paragraph is not in the context window, no amount of model capability recovers it, and a larger window does not repeal the rule — Claude Opus 5 and Claude Sonnet 5 both carry a million tokens of context, enough to hold a small corpus outright, and a passage retrieval left behind is still absent from all of it. The generation step can only lose information relative to what it was given.
This has a practical consequence that surprises teams. When a RAG system answers badly, the instinct is to change the model or rewrite the prompt. Neither addresses the common cause. Pull the retrieved chunks for the failing question and look at them: most of the time the answer plainly is not there, and you have been tuning the wrong stage.
It also reframes what to measure. Answer quality is the thing you care about and the thing that is slow and subjective to score. Retrieval quality is neither: recall@k asks whether the passage containing the answer came back at all within the top k, nDCG@10 asks how highly it ranked among the ten the model will actually read, and both are arithmetic over a labelled set. Measure them separately or you will not know which half is broken. They are also the measures the public benchmarks report — BEIR for retrieval across 18 datasets, MTEB for embedding models across 58 — which is how to shortlist a model before running your own comparison rather than after.
Citations and refusals are designed, not hoped for
Both are commonly treated as prompt instructions. Cite your sources. Say you do not know if you are unsure. Both are more reliable as system properties.
Citations
A citation is only worth something if it survives a click. That means carrying the source reference through the pipeline as structured data — document, section, revision — and rendering it from that record, rather than asking the model to write the reference into its prose where it can be approximated. A plausible-looking citation to a document that does not say what was claimed is worse than none, because it converts a reader’s scepticism into misplaced trust.
Refusals
The refusal should be a decision made before generation. If the best retrieved passage scores below a threshold you set, the system does not call the model to write an answer — it returns the honest response and, ideally, what it did find and a route to a human. Leaving the judgement to the model asks the component with no calibrated uncertainty to report its uncertainty.
Do not set that threshold on raw cosine similarity. An absolute cosine value is uncalibrated: it moves with the embedding model and with the corpus, so 0.7 means one thing on one deployment and something else entirely on the next, and a cutoff copied from a tutorial means nothing on yours. Threshold on the re-ranker’s score instead — a relevance judgement on a 0 to 1 scale after a sigmoid — and calibrate it against a labelled set of your own queries. Operating points usually land between 0.3 and 0.5, but the number is an output of that calibration rather than an input to it.
Use two gates rather than one. The retrieval gate runs before generation and is the one just described: nothing clears the threshold, nothing is written. The groundedness gate runs after, and checks that the claims in the drafted answer are actually entailed by the passages it was handed — because a model given good passages can still assert something none of them supports, and that failure is invisible to the first gate. RAGAS formalises the second as a measurable property rather than an impression, which is what makes it something you can hold a threshold against.
Tuning that threshold is a real product decision rather than a technical one: it trades coverage against trust. A system that refuses too readily gets abandoned; one that never refuses gets believed when it should not be. Pick the point deliberately, and revisit it with evidence.
Common failure modes
Three versions of the expenses policy are indexed and two are obsolete. The system retrieves faithfully and answers wrongly, and it is right to. RAG makes your documentation load-bearing, which is why these projects so often surface a content problem that predates them.
The clause says the limit is defined in Schedule 2. Schedule 2 is a different chunk. Retrieved separately, neither is an answer. Overlap and structure-aware splitting mitigate this; nothing eliminates it entirely for cross-referencing documents.
Filtering results the user may not see, after using them, still leaks — through summaries, through the shape of the answer, through what the system declines to discuss. Permissions belong in the retrieval query, so restricted material is never a candidate.
The index was built once. Documents changed. Nothing failed, no error appeared, and the answers quietly drifted out of date. Re-indexing needs to be an automatic consequence of the source changing, not a task somebody remembers. State it as a service level rather than a schedule — the index reflects the source within N minutes — and pick the mechanism from how fast the corpus actually moves: change-data-capture on write for volatile sources, landing in seconds to minutes, and scheduled batch for slow-moving ones, where nightly is the common floor. The version that bites is not a document but the embedding model: change it and every vector in the index is invalidated at once, because vectors from two model versions cannot be compared and cannot share an index. That is a full re-index, so version the index and dual-write through the migration rather than upgrading in place.
A document containing text addressed to the model is a live injection vector once it is retrieved into the prompt. OWASP lists prompt injection as LLM01, the first entry in its Top 10 for LLM Applications, and a retrieval corpus is the delivery route that entry describes: the attacker needs no access to your prompt, only to a document you will index. Treat retrieved content as untrusted input, because that is exactly what it is.
None of these are reasons not to build one. They are the list to design against on day one, and each is far cheaper to handle in the architecture than to retrofit after a system has been trusted for six months.
This piece supports our RAG Development page, which covers how we design, build and run these systems in production.
Your Next Intelligent System Starts Here.
Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.