Agentic RAG is a RAG setup where the AI model does more than write the answer. Before it writes anything, it decides where to look, and whether it should answer at all. For a small business, that second part is what matters: it's the difference between a support bot that invents a refund policy and one that says "I'll pass this to a person."
Below is a plain-language explanation of both ideas, what the agentic version costs, and a small version of it I already run in Make.
What is RAG?
RAG stands for retrieval-augmented generation. A language model on its own only knows what it learned in training. It has never seen your price list, your cancellation rules or your service areas. RAG fixes that by handing the model the right piece of your own material right before it answers.
A standard RAG pipeline works like this:
- A customer (or an app) sends a question.
- The system searches your documents for the parts that match it. Larger setups use a vector database for this.
- The matching text is added to the prompt as context.
- The model writes its answer based on that context.
The model is called once, at the end. Everything before that is a fixed path: same search, same source, every time.
Where standard RAG falls short
A fixed path is fine when you have one source and every question belongs to it. Real inboxes don't look like that.
Some questions are about your policies. Some are how-to questions. Some have nothing to do with your business. And some, like a double charge or a complaint, shouldn't get an automatic reply at all. Standard RAG can't tell these apart. It searches the one place it knows, takes whatever comes back, and writes an answer anyway. When the search returns something only loosely related, you get a confident reply built on the wrong material.
What makes RAG "agentic"
In agentic RAG, the model also acts as a decision-maker inside the pipeline. It reads the question first and chooses what happens next: which source to search, whether to search at all, or whether to stop and hand off.
Take a fictional cleaning company with two sources:
- Policies: prices, cancellation rules, service areas.
- Cleaning guides: stain removal, safe products, care instructions.
Here's how an agent would route four incoming messages:
| Customer question | Agent's decision |
|---|---|
| "Can I cancel tomorrow's booking without a fee?" | Search the policies |
| "How do I get red wine out of a wool rug?" | Search the cleaning guides |
| "Can you recommend a plumber?" | Neither source covers it: reply that it's outside what you offer |
| "You charged me twice this month." | Billing problem: no AI reply, send to a person |
The agent isn't picking at random. It uses the model's reading of the question, the same skill that makes it good at writing, to decide where the question belongs.
Why the "I don't know" route matters most
The routing to two sources gets the attention. The fallback does the real work.
A standard pipeline always answers. An agentic one can recognize that nothing it has fits the question and say so, or pass the message to a human. That one decision is what makes an AI reply safe to send to a customer. A wrong answer about a refund costs you more than a slow answer from a person.
Standard RAG vs agentic RAG
| Standard RAG | Agentic RAG | |
|---|---|---|
| Who decides where to search | Fixed in advance | The model, per question |
| Number of sources | Usually one | Several, plus outside tools if needed |
| Model calls per question | One | Two or more |
| Off-topic question | Answers anyway | Can decline or hand off |
| Cost and speed | Lower, faster | Higher, slower |
What it costs
Every decision is another model call. More calls mean more tokens, more waiting, and more moving parts that can break. IBM notes that the extra compute makes agentic RAG a better fit when you actually need to query several data sources. Neo4j's guide makes the same point more bluntly: the agentic approach comes with real costs and should be used for questions that need routing or multi-step reasoning, not as a default.
Skip it if you have one short FAQ. If all your answers fit in a single document, put that document in the prompt and add a "needs a human" path. You get most of the benefit with one model call.
A small version you can build in Make
My customer support tutorial already uses the core idea. Gmail picks up the email, a Google Doc holds the approved FAQ, DeepSeek reads the message and returns a category, and a Router sends it to one of three paths: draft a reply, mark as spam, or flag for a human.
To be honest about it: that scenario isn't full RAG. There's no vector database and no search step. The whole FAQ goes into the prompt, which works because the FAQ is short. The agentic part is the decision. The model chooses the path, and the scenario follows it.
To grow it into two sources, the structure looks like this:
- First AI call: classify the message as
policies,guides,out_of_scopeorhuman, and return only that as JSON. - Router: one path per category, with
humanas the fallback. - Fetch the source: on the policies and guides paths, a Google Docs module gets the matching document.
- Second AI call: write the reply using only that document, and say so if the answer isn't in it.
- Log it: add a row in Google Sheets with the category and the action taken, so you can check the routing later.
Lesson from building the support scenario: every Router path needs at least one module before you test. A path with a filter but nothing connected doesn't throw an error. Make skips it, and the message falls through to the fallback. I spent a while thinking the AI was misclassifying emails before I found that.
The cost here is one extra AI module run per message, plus the provider's token price for the second call. For a few dozen emails a day, that's small. Check it against your Make plan before you turn it on for a busy inbox.
Where agentic RAG is used
- Customer support: sending each question to the right knowledge base, and passing harder cases to staff.
- Legal work: internal briefs for one question, public case records for another.
- Internal operations: HR policy, IT how-tos and finance rules kept in separate places, with one assistant in front of all of them.
The agent can also call outside tools, like a live order lookup or a calendar, instead of searching documents. That's where it starts to overlap with the AI agents covered in BPA vs RPA vs AI agents.
When to use it, and when not to
Use agentic RAG when your answers live in more than one place, when some questions should never get an automatic reply, or when wrong answers are expensive. Skip it when one document covers everything and the questions are predictable. A simpler pipeline is cheaper and easier to fix.
Either way, keep a person in the loop for money, complaints and anything legal. The agent's job is to sort and draft. Deciding what your business promises a customer is still yours.
FAQ
What is the difference between RAG and agentic RAG?
Standard RAG always searches the same source and always answers. Agentic RAG adds a decision step: the model chooses which source to search, or decides not to answer and hands the question off.
Do I need a vector database for agentic RAG?
Not for a small setup. If each source is a short document, you can pass the whole document to the model. Vector databases help when you have too much material to fit in a prompt.
Can I build agentic RAG without code?
A simple version, yes. In Make, one AI module classifies the question, a Router sends it down the matching path, and a second AI module writes the reply from the right document.
Is agentic RAG more expensive than standard RAG?
Yes. Each decision is an extra model call, so you pay more tokens and wait longer per question. It's worth it when you have several sources or when wrong answers are costly.