Guides · RAG and memory

How to Combine RAG With Agent Memory

Running both means two retrievals in one turn, and the interesting decisions are all in how they meet: which runs first, how the prompt budget is divided, how each source is labelled, and which one wins when they disagree. The one thing that does not work is pointing your existing RAG retriever at a memory store and expecting it to behave the same way.

One turn

1
Memory
Who is asking
2
Query
Shaped by it
3
Docs
Retrieved
4
Compose
Labelled

The pairing

Why do RAG and memory belong in the same system?

Because a useful answer usually needs both what is documented and what is true of the person asking. RAG supplies the first and cannot supply the second, since a corpus shared by every user contains nothing about any of them.

The failure of each on its own is easy to picture. An agent with documents and no memory gives every user the same correct, impersonal answer, and asks for the same context every time. An agent with memory and no documents knows exactly who it is talking to and cannot tell them how the product works.

They are also retrieved differently for a reason. A document corpus is authored, shared and updated on a publishing schedule. A memory store is produced by interactions, partitioned per user, and updated continuously, which makes a mistake in the first a wrong answer and a mistake in the second a disclosure. The full comparison is on memory versus RAG, and the agentic variant on agentic RAG versus agent memory.

Since most teams arrive here already running RAG, the first practical question is whether the machinery transfers: whether you can reuse your RAG pipeline for memory.

The wrong shortcut

Can you reuse your RAG pipeline for memory?

The infrastructure transfers, the retrieval strategy does not. Memory is not a corpus of independent passages, it is a correlated stream in which the same fact recurs in many near-identical forms, and similarity top-k handles that badly.

A document corpus of independent passages compared with a memory store of correlated near-duplicate statements.
Figure 1. The retriever has not changed. The data has, and top-k assumes a diversity that a memory store does not have.

The mechanism is worth stating precisely. If a user has expressed the same preference across five sessions, a top-k search for it returns five near-duplicates, all genuinely relevant, and the entire retrieval budget goes on one fact while four other facts that mattered are left outside. Research on retrieval for agent memory from King’s College London (Hu et al., 2026) describes exactly this collapse: RAG assumes a large heterogeneous corpus of diverse passages, while a memory store is a bounded, coherent stream of highly correlated spans. Their answer is to organise memories into a hierarchy of intact units so retrieval returns a shorter, more sufficient context.

Two practical consequences follow before any of that research is adopted. Deduplicate at write time so the store does not contain five phrasings of one preference, as covered on how agents write and store memories. And diversify at read time, preferring one memory per distinct claim over the top five by score.

The second mismatch is scoring. Document retrieval is dominated by relevance; memory retrieval also needs recency and importance, and applied wrongly recency buries a preference stated once and still true. The scoring differences are on memory scoring.

With the retrieval strategy corrected, the two searches have to be arranged in a turn: how to run both retrievals together.

Arrangement

How do you run both retrievals in one turn?

Read memory first, use it to shape the document query, then search the corpus. Running both in parallel is simpler and gives up the main advantage of having memory at all, which is that it changes what is worth searching for.

One turn with two retrievals: read memory, shape the query with it, search the document corpus, then compose a labelled prompt.
Figure 2. Sequential costs one extra round trip. It buys a document search that already knows who is asking.

The memory read is cheap enough to make this affordable: a small, fixed set of durable facts fetched by user ID rather than by similarity, which is a key lookup rather than a search. The expensive retrieval is the second one, and it benefits from running with more information.

Parallel retrieval remains the right choice in two cases: where latency is genuinely critical and the extra round trip is unaffordable, and where memory holds nothing that could change the document query, which is true of preference-only stores. Otherwise sequential wins on quality.

What the memory read does to the second search is the part worth designing deliberately: whether memory should shape the RAG query.

Query shaping

Should memory shape the RAG query?

Yes, and it is the highest-value thing memory does in a combined system. The same question from two different users should retrieve different documents, and only memory knows why.

Consider “why did my export fail”. Retrieved against the corpus as written, it returns the general troubleshooting article. Rewritten with what memory holds, that the user is on the free plan and exporting a format the free plan does not support, it retrieves the plan limits page, which is the actual answer.

Three uses are worth implementing, in order of payoff. Disambiguation, where memory resolves what a pronoun or a bare noun refers to, since “the integration” means a different thing for each customer. Filtering, where memory supplies metadata constraints such as version, plan or region, so the search is narrowed rather than reworded. Expansion, where the product names a user actually owns are added to the query text.

Filtering is the safest of the three and disambiguation the most valuable. Expansion needs care: adding too much from memory drifts the query away from what the user asked, and an agent that retrieves documents about the wrong topic confidently is worse than one that retrieves a general answer.

Both retrievals now return material, and the prompt cannot hold all of it: dividing the prompt budget between the two.

Budget

How do you split the prompt budget between documents and memories?

Give memory a small fixed reservation and let documents have the remainder. Never rank the two together in one list. Similarity scores from two different indexes are not comparable, so a combined ranking is decided by whichever index happens to score higher, which is an accident rather than a decision.

Splitting the prompt budget: a fixed small memory block, the remainder for documents, and why one ranking across both indexes fails.
Figure 3. A reservation rather than a competition, because the two sources are not substitutes for one another.

The reservation is small on purpose. Four to six memories cover who the user is and what they have already tried, and beyond that the marginal memory displaces a document passage that would have carried more of the answer. Memory is context about the question; documents are usually the substance of it.

Total budget matters as much as the split. Filling the window because it is available degrades the answer, the effect described on context rot, and the degradation applies to both sources equally.

One asymmetry is worth building in: when memory returns nothing, give the whole budget to documents rather than padding with low-scoring memories to fill a quota. A quota met with weak memories is worse than an unmet one.

Whatever the split, the model has to be able to tell the two apart: labelling documents and memories in the prompt.

Composition

How should documents and memories be labelled in the prompt?

In separate named blocks, each saying what the content is and where it came from. Unlabelled context is the most common defect in combined systems, and it produces failures that look like hallucination and are formatting.

An agent handed a paragraph with no marker cannot tell a documented policy from something the user said in March from something the user said thirty seconds ago. It will then attribute all three the same way, telling a user “as you mentioned” about a line from the manual, or citing a stale memory as though it were documentation.

A workable structure is three blocks: what is known about this user, from memory, each with the date it was recorded; relevant documentation, from the corpus, each with its source; and the current conversation. Dates on memories matter, because they let the model hedge appropriately about something recorded a year ago.

The instruction that goes with the blocks is short and does real work: state which source takes precedence for which kind of claim. That is not a formatting question but a truth question, and it is the last one: what happens when a memory contradicts a document.

Precedence

What happens when a memory contradicts a document?

Decide by what the claim is about: documents win on the product, memory wins on the person, and where both describe the person, the newer source wins. Without a stated rule the model picks, and it picks by fluency rather than by authority.

Precedence rules when a memory contradicts a document, by whether the claim is about the product or about the person.
Figure 4. Four cases, and the fourth is the one worth implementing deliberately: some conflicts should be surfaced rather than resolved.

Documents win on the product. Pricing, policy and how a feature works live in the corpus, and a memory that says otherwise is a stale copy of a document, usually captured months ago. This case is common and is the reason not to write documentation into memory in the first place.

Memory wins on the person. Preferences, entitlements and what has already been tried are not in any document, and no document should be consulted about them.

The hard case is a person-fact appearing in both, such as a CRM record that disagrees with what the user said last week. Compare timestamps and prefer the more recent, with a stated exception: a system of record beats a memory on anything commercial, since an inferred entitlement is not an entitlement. The resolution rules are on handling conflicting memory updates.

The fourth case deserves building rather than resolving. Where a conflict is material and cannot be settled from what is retrieved, surface both and say they disagree, because a system that silently picks the wrong one is indistinguishable from a system that is simply wrong.

All of which needs verifying rather than assuming: evaluating a system that retrieves from both.

Verification

How do you evaluate a system that retrieves from both?

Evaluate the two retrievals separately, then the composition, then the whole. A combined system fails in four places and an end-to-end score cannot tell you which one moved.

A build-side walkthrough of the retrieval half, useful for the pipeline this guide arranges memory around.

Test the document retrieval as you would any RAG system, on questions with a known answering passage. Test the memory retrieval separately on questions whose answer is a stored fact, including the near-duplicate case above: if a user has stated one preference five times, a good retrieval returns it once and spends the rest of the budget elsewhere.

Then test the two together on the cases only a combined system can fail: a question needing one fact from each source, a question where memory should have changed the document query, and a question where the two sources disagree. Those three are where combined systems actually break, and none of them appears in a single-source test set.

The ablation that matters here is narrower than usual: run with query shaping disabled and see whether the retrieved documents change. If they do not, memory is not doing the job this architecture exists for. The general method is on how to evaluate agent memory, and the metrics on memory evaluation metrics.

FAQ

Frequently asked questions

The questions that follow when a RAG system gains a memory layer: indexes, ordering and what to do about cost.

Should documents and memories live in the same index?

No. Keep them in separate indexes with separate budgets. Their scores are not comparable, their partitioning requirements differ, and a memory index needs a hard per-user boundary that a shared document index does not.

Which should be retrieved first, documents or memories?

Memories, so that what you know about the user can shape the document query. Parallel retrieval is only preferable when latency is critical or when memory holds nothing that could change what documents to search for.

How many memories should go into the prompt alongside documents?

Four to six is the working range. Beyond that each extra memory displaces a document passage that usually carries more of the answer, and both sources degrade once the window is padded.

Does combining RAG with memory increase cost?

Slightly per turn, through one extra lookup and a few hundred tokens. It usually lowers total cost, because a query shaped by memory retrieves fewer and better passages than a broad one. See reducing token cost with memory.

What if a memory says something the documentation contradicts?

For claims about the product, the documentation wins and the memory is a stale copy. For claims about the person, memory wins. Where both describe the person, prefer the more recent source, with a system of record beating memory on anything commercial.

Can one vector database hold both?

Physically yes, logically no. Use separate collections with separate retrieval settings, since sharing one collection makes cross-user leakage a filter away rather than a boundary away.