Evaluation · Cluster hub
How to Evaluate AI Agent Memory
Evaluating memory means testing four separate competencies, not one: whether the agent finds a stored fact, learns during the interaction, combines evidence across sessions, and answers with the current version of something that changed. Public benchmarks cover the first well and the fourth badly, which is why a strong published score and a disappointing production system are entirely compatible.
Four competencies
Definition
What does evaluating agent memory actually mean?
Measuring whether the memory system changed the agent’s answers for the better, which is a different question from whether the agent performed well. An agent can score highly with no memory at all, and a memory system can be excellent while sitting behind an agent that never uses what it returns.
That separation is why general agent evaluation does not transfer. Task-completion benchmarks measure reasoning, planning and tool use, and a memory system contributes to those indirectly at best. Memory evaluation asks a narrower question: given that something was established earlier, does the system use it correctly now.
There are two ways to ask it. Direct evaluation tests the store on its own: given this query, are the right memories returned, in what order, and how fast. Indirect evaluation tests the agent’s answers with and without memory and attributes the difference. Direct measurement localises faults, and indirect measurement is the one that tells you whether any of it mattered.
Both should be applied at each stage of the memory loop, because a failure at any of the three, write, maintain and read, presents identically to a user. A fact that was never extracted, one that was stored but superseded and never retired, and one that exists but ranks tenth all produce the same complaint.
Before choosing an instrument, it is worth knowing what a complete evaluation has to cover: the competencies a memory evaluation must include.
Coverage
Which competencies does a memory evaluation have to cover?
Four: accurate retrieval, test-time learning, long-range understanding and conflict resolution. The set comes from the research introducing MemoryAgentBench (Hu et al., 2025), which identified them precisely because existing datasets tested the first and largely ignored the rest.
Accurate retrieval is the ability to return a fact that was stated earlier when a later question needs it. It is the competency every benchmark measures and the one in-house evaluations almost always test exclusively.
Test-time learning is acquiring something during the interaction and applying it, rather than only recalling what was supplied up front. Long-range understanding is combining evidence spread across many sessions, which a single similarity search returns fragments of rather than answers to.
Conflict resolution is answering with the current version of a fact that changed. It is the least tested and the most commercially consequential, because a confidently stated superseded fact is worse than an admission of not knowing. The mechanics of the failure are on handling conflicting memories.
Each public benchmark covers a different subset of the four, which is the practical reason to know what each one measures: the benchmarks and what each tests.
The instruments
Which agent memory benchmarks exist, and what does each test?
Four matter in practice, and they sit on a difficulty gradient from recall inside one conversation to memory inside a working agent. Vendor claims cite the first two almost exclusively.
LOCOMO tests recall within very long conversations, using dialogues averaging around 300 turns. It is the most reported number in the field and the narrowest, covering accurate retrieval and little else. The dataset, the tasks and the published scores are on the LOCOMO benchmark.
LongMemEval tests recall across separate sessions, including questions whose answer changed between them, which brings conflict resolution partly into scope. It is the harder benchmark and the closer analogue to how products are used, covered on the LongMemEval benchmark.
MemoryAgentBench was built around the four competencies and evaluates them through incremental multi-turn interaction rather than static question answering. STATE-Bench, released open source by Microsoft in 2026, goes further in the same direction by measuring what memory does for an agent completing real tasks. Both are newer and less widely reported, which is precisely why a score on either carries more information than another LOCOMO figure.
Knowing what a benchmark tests is only half of reading a claim, because the same benchmark yields different numbers in different hands: why vendors report different scores.
Reading claims
Why do vendors report different scores on the same benchmark?
Because a benchmark score is a measurement of a configuration, not of a product, and three parts of that configuration are rarely stated. Two vendors can both be honest and still publish numbers that cannot be compared.
The retrieval budget is the largest lever. A system allowed to place twenty memories in the prompt will beat one restricted to four, and the score does not show the token cost of the difference. Any comparison without a stated budget is a comparison of budgets.
The judge is next. These benchmarks are graded by a model deciding whether an answer matches a reference, and changing the judge or its prompt moves the number without anything in the memory system changing. The subset is third: category difficulty varies sharply, so a headline average over a chosen subset is not comparable to an average over another.
Three questions make a published claim readable: how many memories were retrieved per question, which model judged the answers, and what the per-category breakdown was. A vendor who reports all three is telling you something; one who reports a single percentage is telling you very little, which is also true of the figures cited on this site, each of which names its source beside it.
Even a perfectly reported score answers a narrower question than it appears to: what benchmark scores fail to predict.
Limits
What do benchmark scores fail to predict?
Almost everything that breaks in production: stale facts, scope leaks between users, latency at load, cost per conversation, and how the system behaves when a memory is simply wrong. Benchmarks run on curated dialogues with clean ground truth, and production has none of those properties.
The most instructive result in this area came from Letta, who benchmarked a plain filesystem against specialised memory tools on LOCOMO and found the filesystem competitive. The honest reading is not that memory tools are unnecessary. It is that the benchmark was measuring something a filesystem also does well, which is retrieval over a bounded conversation, and not the things a memory system is bought for.
Four gaps recur. Staleness is untested because benchmark conversations end before facts have time to go out of date. Isolation is untested because benchmarks run one user at a time, so a scope leak scores perfectly. Operational cost is invisible because no benchmark reports the extraction calls the system made to reach its score. And failure behaviour is unmeasured, since a benchmark records a wrong answer identically whether the system hedged or asserted.
None of this makes benchmarks useless. It makes them a filter rather than a decision: a system scoring poorly can be excluded, and a system scoring well still has to be evaluated on your own data. The production-side numbers that fill the gap are on memory evaluation metrics.
Which raises the practical question of what to watch once the system is running: the metrics to track in production.
In production
Which metrics should you track in production?
Three families: whether retrieval returns the right memories, what it costs in latency and tokens, and whether the store itself is healthy. Most teams instrument the second family only, because it is the one their existing monitoring already collects.
Retrieval quality is precision first and recall second. The asymmetry is deliberate: a wrong memory in the prompt actively misleads, while a missing one merely leaves the agent where it would have been without memory at all. Precision is also the metric that degrades quietly as a store grows.
Cost and latency are the retrieval call at the p99 rather than the average, plus the tokens memory adds per turn and the extraction calls made per closed conversation. The last of those is usually the largest line and the least monitored.
Store health is the family nobody watches: how many candidates were written versus discarded, what share of writes were merges rather than appends, how many memories have never been retrieved, and how many contradict another. A store where nothing merges is a store that is accumulating duplicates. Definitions, targets and the observability setup are on memory evaluation metrics.
Metrics tell you the state of the system. Deciding whether a change improved it needs a fixed set of cases: building your own memory evaluation set.
Your own set
How do you build your own memory evaluation set?
From real traces, not from imagined questions, and tagged by which of the four competencies each case exercises. A set written from scratch tests the failures you thought of, which are by definition not the ones that surprised you.
Start by logging the parts of a turn that are usually thrown away: which memories were retrieved, what scores they had, which ones entered the prompt, and what the agent answered. Without that record a failure cannot be diagnosed, because there is no way to tell a retrieval failure from an injection failure from a reasoning failure.
Then convert failures into fixed cases. Each one is a conversation prefix, a question, an expected answer, and a tag naming the competency it tests. Fifty such cases drawn from real use are worth more than a thousand generated ones, and the tags are what stop the set drifting into fifty variations of accurate retrieval.
Include the awkward cases deliberately: a fact stated then corrected, a question needing two memories from different sessions, a question whose answer is genuinely not stored, and a memory belonging to a different user that must not appear. The last two matter most, since a system that invents rather than admitting absence, or that leaks across a scope, fails in ways no accuracy score reports.
With a fixed set in place, one question remains, and it is the one most evaluations never answer: how to know the memory caused the improvement.
Attribution
How do you know the memory caused the improvement?
Run the same cases with memory retrieval disabled and compare. The gap between the two runs is the contribution of memory, and without it every other number is compatible with the memory system doing nothing at all.
This is not a formality. The Letta filesystem result exists because someone ran the comparison, and it showed that a benchmark presented as measuring memory tooling was substantially measuring something else. The same surprise is available in any system where nobody has checked.
Three ablations are worth running rather than one. Memory off establishes the floor. Retrieval widened, returning more memories, separates a ranking problem from a coverage problem: if the answer improves with twenty memories, the facts were stored and the ranker was not finding them. Memory replaced with the full transcript gives the ceiling, since a system that cannot beat pasting in the whole history is not yet earning its complexity.
In production the same logic becomes a held-out slice: memory disabled for a small share of traffic, with the support or product metrics compared across the two groups. Anything weaker is a before-and-after comparison that a seasonal shift can explain away.
Ablation is one step of a cycle rather than a one-off exercise: what a repeatable evaluation loop looks like.
The loop
What does a repeatable memory evaluation loop look like?
Collect traces, convert failures into cases, run the set on every change, and ablate. The value is in the repetition: a memory system has enough interacting parts that a fix in one reliably breaks something in another.
Run the set against every change to extraction, scoring, retrieval or injection, because those four are coupled more tightly than they look. Tightening a salience threshold improves precision and quietly removes the fact a later case depends on, and only a fixed set catches it.
Keep a human in the loop on a sample. Automated grading misses the failures that matter most in memory systems, particularly an answer that is correct but reveals something the agent should not have surfaced, and an answer that is confidently stale. Both read as passes to a judge model comparing against a reference.
Two habits make the loop pay for itself. Every production incident becomes a case, so the set grows in exactly the places the system is weak. And every vendor claim gets re-run locally before it informs a decision, since the configuration behind a published number is rarely yours. The tool comparison that follows from doing this is on the best AI memory tools, and the build side on the memory build guides.
FAQ
Frequently asked questions
The questions that come up when reading a vendor claim or starting an evaluation of your own.
How do you evaluate AI agent memory?
With two measurements taken together: a direct test of whether the store returns the right memories for a query, and an indirect test comparing the agent's answers with memory enabled and disabled. The second is what attributes an improvement to memory rather than to everything else that changed.
What is the difference between LOCOMO and LongMemEval?
LOCOMO tests recall inside one very long conversation. LongMemEval tests recall across separate sessions, including questions whose answer changed in between, which makes it harder and closer to product use. See LOCOMO and LongMemEval.
Is a high benchmark score enough to choose a memory system?
No. Benchmarks do not test staleness, scope isolation between users, operational cost or how the system behaves when a memory is wrong. Use a score to exclude candidates, then evaluate the survivors on your own traces.
Which benchmark tests contradiction handling in agent memory?
LongMemEval covers it partly through questions whose answer changed between sessions, and MemoryAgentBench treats conflict resolution as one of its four core competencies. Most other benchmarks do not test it at all.
How many evaluation cases do you need?
Around fifty drawn from real failures beats a thousand generated ones, provided they are tagged by competency. The tags are what stop a set becoming fifty variations of the same retrieval test.
How often should a memory evaluation run?
On every change to extraction, scoring, retrieval or injection, because those four are coupled: tightening a write threshold routinely removes a fact a retrieval case depends on.