Retrieval-augmented generation, usually shortened to RAG, is a way of getting a language model to answer questions using your organization's own information rather than only what it learned during training. Before the model writes an answer, the system searches your documents for the most relevant passages and gives them to the model along with the question. The model then answers from those passages and can point to where each statement came from.
If you remember one thing: RAG changes what the model knows at the moment of answering, without changing the model itself.
How it works, in plain language
A RAG system has three stages.
- Preparing the knowledge. Documents such as policies, manuals, contracts, tickets or product pages are collected, split into sensible sections and indexed so they can be searched by meaning, not only by exact words.
- Retrieving at question time. When someone asks a question, the system finds the sections most likely to contain the answer, filtered by what that person is allowed to see.
- Generating the answer. The model receives the question and the retrieved sections, writes an answer based on them and cites the sources. If the sources do not contain the answer, a well-built system says so.
The quality of the answer depends heavily on the first two stages. A capable model given the wrong passages will produce a confident, wrong answer.
When RAG is the right choice
RAG tends to fit when:
- Knowledge changes often. Policies, prices, product details and procedures are updated regularly. With RAG, updating the documents updates the answers.
- Answers must be traceable. People need to check the source, whether for compliance, customer trust or their own confidence.
- Access varies by user. Different teams, customers or roles are allowed to see different information.
- The knowledge is specific to you. Internal procedures, contracts and product documentation are not in any public model's training data.
RAG or fine-tuning?
Fine-tuning means further training a model on your examples so its behavior changes. It is often discussed as an alternative to RAG, but the two solve different problems.
| Question | RAG | Fine-tuning |
|---|---|---|
| What does it change? | What the model can see when answering | How the model behaves by default |
| Good for | Facts, documents, frequently updated knowledge | Consistent format, tone, classification, specialized tasks |
| Keeping it current | Update the documents and re-index | Retrain when the desired behavior changes |
| Citing sources | Built in | Not inherent |
| Access control per user | Applied at retrieval | Difficult; the model cannot forget selectively |
In our experience, most knowledge-access projects should start with RAG. Fine-tuning becomes worth considering when the model needs to follow a specialized format or style reliably, handle a narrow classification task cheaply at high volume, or use domain language that general models handle poorly. The two can also be combined: a fine-tuned model that is better at the task, answering from retrieved documents.
Permission-aware retrieval
This is the part decision makers should ask about first. If a RAG system indexes documents without respecting who can see them, it can surface confidential information to the wrong person, simply by answering a question well.
A permission-aware system:
- Carries source permissions into the index, so every section knows which users or groups may see it.
- Filters before the model sees anything. Documents a user cannot access are never retrieved for that user, rather than being hidden after the answer is written.
- Keeps permissions in sync. When access changes in the source system, the index reflects it promptly.
- Logs what was retrieved for each answer, so access can be audited.
Questions to ask a vendor or internal team: How are source permissions captured? How quickly does a permission change take effect? Can you show which documents were used for a given answer?
Keeping answers fresh
A RAG system is only as current as its index. Freshness is an operational concern, not a one-time setup.
- Update frequency. Decide how quickly a change in a source document must appear in answers, and design the indexing schedule or change triggers to match.
- Retiring old content. Superseded documents should be removed or marked, otherwise the system may answer from an outdated version.
- Version awareness. For some questions, such as what the policy was on a given date, the system needs to know which version applied when.
- Ownership. Someone owns each source collection and is responsible for its accuracy. RAG makes gaps and contradictions in documentation more visible, which is useful but needs a person to act on it.
What can go wrong
- Poor source material. Duplicated, contradictory or outdated documents lead to inconsistent answers.
- Weak retrieval. The right passage exists but is not found, often because documents were split badly or the search does not handle your terminology.
- Answering beyond the sources. The model fills gaps from general knowledge. Clear instructions and evaluation reduce this, and the interface should show sources so people can check.
- No evaluation. Without a set of real questions with known good answers, you cannot tell whether a change made the system better or worse.
Questions to ask before you start
- Which questions do people ask most, and where do they look for answers today?
- Which document collections hold those answers, and who owns them?
- How are permissions managed in those sources?
- How current must answers be?
- How will you measure whether answers are correct?
Where to go next
For how we design and build these systems, see RAG & Knowledge Systems. For a ready-shaped starting point, see the Knowledge Assistant solution.