Compare · Agentic RAG
Agentic RAG vs Agent Memory: What Is the Difference?
Agentic RAG makes retrieval smarter by letting the agent decide what to search for and when; agent memory changes what there is to search, by writing down what the interaction produced. One improves how a corpus is read, the other creates a store that did not exist before. Systems that need both are common, and confusing them is why some teams build a retrieval pipeline and still have an agent that forgets.
Three stages
Definition
What is agentic RAG?
Agentic RAG is retrieval-augmented generation where the agent controls the retrieval: it decides whether to search, what query to issue, whether the results were good enough, and whether to search again. Classic RAG runs one retrieval before the model answers, with the query fixed by the application.
The difference is control rather than storage. Both read from a corpus somebody else maintains, both use the same vector store, and both inject passages into a prompt. What changes is that the agent can reformulate a bad query, search a second source when the first returns nothing useful, and decide that a question needs no retrieval at all.
That control unlocks multi-step questions. A question whose answer requires combining two documents defeats a single retrieval, because no one query is close to the whole answer. An agent that can search, read, and search again on the basis of what it read can assemble it.
What agentic RAG does not do is write anything down. When the conversation ends, the corpus is exactly as it was, and the next session begins knowing nothing about this one. Classic RAG is covered on RAG explained.
That last point is the whole distinction from memory: what separates agentic RAG from agent memory.
The distinction
How is agentic RAG different from agent memory?
Agentic RAG is a read strategy over knowledge that already exists; memory is a write path that creates knowledge from the interaction. An agent can have the most sophisticated retrieval in the world and still greet a returning user as a stranger.
Three tests separate them quickly. Who wrote the content? A corpus is authored by the organisation; memories are produced by this user’s interactions. Does anything persist after the session? Agentic RAG leaves no trace; memory is the trace. Would two users get different answers? A corpus should give both the same answer; memory should give each of them a different one.
The overlap that causes confusion is real. Both search a vector store, both rank results, both inject passages, and both are often built by the same team on the same infrastructure in the same sprint. It is entirely reasonable to run one Weaviate or pgvector instance holding a documentation collection and a memory collection side by side.
But the write paths could not be more different. A RAG corpus is ingested in batch: chunk, embed, index, done. A memory store is written on every turn, has to decide what is worth keeping, has to deduplicate against what it already holds, and has to resolve contradictions when a new fact supersedes an old one. That is a much harder loop, described on how agents write and store memories.
Framed as a progression, the two are consecutive rather than competing: how the field moved from RAG to memory.
The progression
How did the field move from RAG to agent memory?
In three steps, each one addressing a limit of the last: fixed retrieval, then agent-controlled retrieval, then agent-written storage. Reading them in order makes the current landscape much easier to navigate.
Classic RAG solved a real problem: models could not answer questions about documents they were never trained on. One retrieval, one answer. Its limit is that a single query cannot serve a question requiring several lookups, and it cannot recover from retrieving the wrong thing.
Agentic RAG addressed that by giving the agent the search. It can iterate, reformulate and check, which handles multi-hop questions and reduces the failure where a badly phrased query returns nothing useful. Its limit is that nothing it learns survives the conversation.
Agent memory addressed that in turn by adding a write path. The agent now produces knowledge as well as consuming it, and the store grows through use. Its limits are the ones this site covers at length: extraction quality, conflicting facts, and stores that degrade without consolidation.
Each step added capability and a maintenance burden. That is worth stating plainly, because it explains why teams keep discovering that a memory system needs more ongoing care than the retrieval pipeline it grew out of. The mechanism in full is on how AI memory works.
Knowing the progression makes the practical question easy to state: which one your system needs.
Selection
Which does your system need?
Ask what has to be true after the session ends. If the answer is nothing, you need retrieval. If the answer is a list of facts about this user, you need memory.
You need agentic RAG when questions require several lookups across a corpus, when a single query frequently retrieves the wrong passage, or when the agent must decide between multiple sources. These are properties of the knowledge base, not of the relationship with the user.
You need memory when the same person returns and expects to be known, when a task spans sessions, or when the agent should stop repeating a mistake. These are properties of the relationship, and no retrieval strategy supplies them.
You need both in most production assistants, and they divide cleanly: the corpus supplies organisational knowledge, memory supplies personal and session state. A support agent needs the refund policy from one and the customer’s history from the other, and neither store can do the other’s job.
The mistake worth avoiding is building an increasingly clever retrieval pipeline in response to complaints that are actually about memory. If users are frustrated that the assistant does not remember them, better search over documentation will not help, however sophisticated it becomes. The distinction is set out further on AI memory versus RAG.
Where both are needed, the assembly matters: running agentic RAG and memory together.
Together
How do agentic RAG and memory work together?
Retrieve memory first, then let it shape what the agent searches for in the corpus. Running them in the other order means retrieving documents for a generic version of the user and then discovering who they are.
The sequence that works is short. Retrieve the relevant memories for this user, inject them, and let the agent formulate its corpus queries with that context available. Knowing a customer is on the enterprise plan means the agent searches for the enterprise policy rather than the general one, which is a better query than any reformulation loop would have produced without it.
Memory can also terminate the retrieval loop early. If the store already holds the answer, because this question was resolved last month, the agent does not need to search the corpus at all, which is both faster and more consistent than re-deriving it.
The budget question is the one to watch. Both compete for the same context window, and injecting ten document chunks alongside ten memories leaves a diluted prompt, which is the effect documented in “Lost in the Middle” by Liu et al. (arXiv:2307.03172). Explicit caps on each, rather than letting whichever pipeline runs first consume the space, is what keeps behaviour predictable.
The implementation is on how to combine RAG and memory in an agent, and the tools that provide the memory half on the best AI memory tools.
One claim deserves direct treatment, because it circulates widely: whether RAG and memory are the same thing.
The disagreement
Is memory just RAG over your conversations?
It is a defensible simplification and it breaks on the write path. The claim has genuine support among practitioners, and it is worth taking seriously rather than dismissing, because what it gets right explains the confusion.
What the claim gets right: the read paths really are the same. Embed a query, search a vector store, rank, inject. If you only ever observed retrieval, memory would be indistinguishable from RAG over a corpus of transcripts, and a system built that way does work, at first.
Where it breaks is everything before storage. RAG over conversations means storing conversations, and a store of raw transcripts retrieves badly, because every chunk looks equally relevant to a semantic search and nothing distinguishes a durable preference from small talk. Memory systems exist because extraction, deduplication and conflict resolution turn out to be necessary rather than optional refinements.
The second break is maintenance. A document corpus is versioned by whoever owns it. A memory store has to supersede its own facts when they change, merge its own duplicates, and evict its own stale entries, or it degrades. Nothing in a RAG pipeline does any of that, which is covered on memory consolidation and conflicting memories.
So the honest summary is that memory is not RAG over conversations, but a system built that way is a reasonable first version that will need the write path adding later. Teams who describe memory this way are usually describing the version they have rather than the one they will need.
FAQ
Frequently asked questions
The questions that follow: whether one store can serve both, and which to build first.
Can one vector store serve both agentic RAG and memory?
Yes, commonly as two collections in the same instance. They differ in the write path rather than the storage: a corpus is ingested in batch, while memory is written per turn with extraction, deduplication and conflict resolution.
Which should you build first?
Whichever answers your users' actual complaints. If people cannot get answers about your documentation, build retrieval. If they are frustrated that the assistant does not remember them, no amount of retrieval sophistication will help.
Is agentic RAG the same as multi-hop RAG?
Multi-hop retrieval is one thing agentic RAG enables, but the broader property is that the agent controls whether, what and how often to search, including deciding that a question needs no retrieval at all.
Does agentic RAG remove the need for memory?
No. It improves how an existing corpus is read and writes nothing down, so the next session starts knowing nothing about this one. Persistence requires a write path, which is what memory adds.
Where does GraphRAG fit?
It is a retrieval technique using a knowledge graph rather than flat chunks, so it sits on the RAG side of this comparison. Graph structure also appears in memory systems for a different reason, to record when a fact was valid, covered on knowledge graphs for AI memory.