
Image: Olkeri
By Olkeri.space
What Is RAG? Retrieval-Augmented Generation Explained for Business
How retrieval-augmented generation lets AI systems answer from your own documents, why it reduces hallucination, and how to build one that works.
Read this story in: Deutsch · Español · Français
Retrieval-augmented generation, universally shortened to RAG, is the most widely deployed pattern in enterprise artificial intelligence. Almost every system that answers questions about a company's own documents, policies, products or records uses some version of it.
The idea is straightforward, and understanding it well is the difference between a system people trust and one they quietly abandon.
The problem RAG solves:
A language model knows only what was in its training data, which stopped at a fixed date and never included your internal documents. Ask it about your refund policy and it will produce something that sounds like a refund policy, invented from patterns rather than retrieved from your actual policy.
You could retrain the model on your documents, but that is expensive, slow, must be repeated whenever anything changes, and still provides no guarantee the model will reproduce the text accurately.
RAG takes a different route. Instead of teaching the model your information, you find the relevant passages at the moment a question is asked and place them directly into the model's context. The model then answers using text it can actually see rather than patterns it half-remembers. It converts a recall problem into a reading-comprehension problem, which is what these models are genuinely good at.
How retrieval works:
The system has two phases: preparation, done once and updated as content changes, and retrieval, done on every question.
In preparation, documents are split into chunks of a few hundred words. Each chunk is passed through an embedding model, which converts text into a list of numbers, a vector, positioned so that passages with similar meaning sit near each other. These vectors are stored in a database built for finding nearest neighbours quickly.
At question time, the question is embedded the same way. The system finds the chunks closest to it, takes the best few, and builds a prompt: here is the question, here are the relevant passages, answer using only this material and say so if the answer is not present.
Because meaning is compared rather than keywords, a question about "time off" can retrieve a policy that only ever says "annual leave".
Why hybrid search usually wins:
Pure semantic search has a predictable weakness: exact identifiers. Product codes, error numbers, invoice references and surnames are precisely where meaning-based matching struggles, because the vector for "SKU-88421" carries little semantic signal.
Production systems therefore run both semantic search and traditional keyword search, then merge the results. Keyword search nails exact terms; semantic search catches paraphrases. Combining them consistently outperforms either alone, and the failure cases they cover are largely complementary.
A second stage, reranking, improves results further. A cheap search returns perhaps fifty candidate chunks, then a slower, more accurate model scores those fifty for genuine relevance and keeps the best five. This two-stage approach is far more accurate than retrieving five directly, and costs little because the expensive model only ever sees a shortlist.
Chunking decides more than people expect:
How documents are split has an outsized effect on quality, and it is where most disappointing systems go wrong.
Chunks that are too small lose context: a passage saying "this does not apply to contractors" is dangerous when separated from what "this" refers to. Chunks that are too large dilute meaning, so the embedding represents a vague average of several topics and matches nothing precisely.
Splitting on structure, by section and heading rather than by fixed character count, works far better than mechanical division. Overlapping chunks slightly prevents answers from falling between boundaries. Attaching metadata, the document title, section heading, date and source, lets the system filter by recency or department and, critically, cite where an answer came from.
Citations are not optional:
The single most important design decision in a RAG system is requiring the model to cite which retrieved passage supports each claim, with a link back to the source document.
Citations do three things at once. They let users verify answers instead of trusting them, which is what makes the system usable for anything consequential. They make hallucination visible, because an unsupported claim has no citation. And they turn the tool from an oracle into a research assistant, which is a far more defensible position.
Systems that answer confidently with no traceable source are the ones that lose user trust after the first serious error.
Where RAG systems fail:
The most common failure is not the model at all. It is retrieval. If the right passage was never fetched, no model can produce a correct answer. When a RAG system disappoints, inspect the retrieved chunks before blaming the model; the fault is usually there.
The second failure is stale content. If the underlying documents contradict each other because obsolete versions were never removed, the system will confidently surface outdated policy. RAG inherits the quality of the source material exactly.
The third is questions requiring synthesis across many documents. Retrieval finds passages similar to a question, which works well for "what is our policy on X" and poorly for "summarise every complaint this quarter and identify patterns". That is an analytics problem wearing a chat interface.
RAG, fine-tuning, or a longer context:
These are often presented as competitors. They solve different problems.
RAG supplies knowledge: facts, documents, current information. Fine-tuning shapes behaviour: tone, format, domain vocabulary, consistent structure. If the complaint is that the model does not know something, that is RAG. If it knows but responds in the wrong style, that is fine-tuning.
Very long context windows have led some to suggest simply pasting everything in. For a moderate set of documents, that is now genuinely viable and much simpler. But it costs more per query, gets slower as content grows, and models still attend unevenly to material buried in the middle of very long inputs. Retrieval remains the practical choice at scale, and the two combine well: retrieve broadly, then let a large window hold more of what you found.
Building one that works:
Start narrow. A system covering one well-maintained document set for one clearly defined group of users will succeed where an attempt to index everything at once will not.
Measure retrieval separately from generation, using a fixed set of real questions with known correct sources. Most improvement comes from better chunking, hybrid search and reranking rather than from switching models.
Above all, keep the source of truth clean. RAG does not fix a disorganised document repository. It exposes it.