Evaluation · Metrics
How to Measure Agent Memory Quality
A memory system that scores well on LoCoMo or LongMemEval has only proven it can retrieve a fact; it has not proven it can learn a new one mid-session, hold context across a long conversation, or resolve a contradiction. Testing memory quality means checking all four competencies, measuring context completeness before answer correctness, and treating every vendor’s self-reported benchmark score as a claim to verify, not a leaderboard entry to trust.
Four competencies
The framework
What actually needs to be true for a memory system to work?
Four separate competencies, and most teams only ever test the first one: accurate retrieval, test-time learning, long-range understanding, and conflict resolution. A system can pass the first with a strong benchmark score and still fail in production on any of the other three.
This framework comes from MemoryAgentBench (arXiv:2507.05257), a peer-reviewed benchmark paper that names it directly and states that no prior benchmark covers all four at once. Accurate retrieval, extracting the correct stored snippet in response to a query, is what LoCoMo and LongMemEval mostly measure. Test-time learning is a different skill: does the agent incorporate a fact given mid-session and use it correctly before that fact has even been written to permanent storage? A failure here looks specific: a user updates a shipping address, the agent acknowledges it, then the next tool call in the same session still returns the old one. Long-range understanding is about tracking how earlier facts constrain later ones across a long conversation, not just recalling isolated facts on demand. Conflict resolution, detecting and resolving a contradiction between a stored fact and new information, is the competency every tested system struggles with most, covered in detail below.
A practitioner guide from Label Studio independently arrives at the same four-competency structure, citing the same paper, which is a stronger signal than either source alone: an academic benchmark and a production-evaluation practice converging on the same list. Deciding what to measure is only half the job; the other half is deciding in what order to measure it. What should you actually measure, and in what order?
Measurement order
What should you actually measure, and in what order?
Context completeness first, because it isolates the memory system from the language model’s generation step. Then answer correctness, retrieval latency, and token cost, measured together rather than traded against each other.
Context completeness answers one narrow question: did retrieval surface every fact needed to answer, regardless of what the model then does with them? If the right facts were never retrieved, no amount of prompt engineering downstream will fix the answer, which is why this metric comes before answer correctness rather than after it. Answer correctness then measures the end-to-end result, whether the final response matched the expected answer, but it conflates two failure modes: the memory system retrieved the wrong context, or it retrieved the right context and the model still answered incorrectly. Separating the two, by grading retrieval and generation as distinct steps, is what tells you which part of the system to fix.
Retrieval latency and token cost are not secondary in the sense of mattering less; they are gates rather than scores. A memory system with excellent accuracy that adds a second of latency to every turn, or that burns tens of thousands of tokens of context to retrieve a handful of facts, has traded one kind of correctness for another kind of failure. Reporting accuracy without reporting the latency and token cost it took to get there is an incomplete result, not a favorable one.
Knowing what to measure still leaves the practical problem of how to generate the test cases in the first place, since most teams do not start with a labeled evaluation set. How do you build a test set that catches real failures?
Test construction
How do you build a test set that catches real failures?
Start from a small number of real target interactions, expand them into cases with a defined golden answer, spread the answers across multiple sessions, and deliberately include facts that change over time. A test set built any other way tends to test the easy case, not the one that actually breaks in production.
Begin with three to five target interactions: what a user might ask, and what a correct agent should answer for your domain specifically. Expand each into ten or more variations and related questions, with a golden answer describing exactly what a correct response must contain. Author multi-session conversations that contain the answers spread naturally across several sessions rather than in one tidy prompt, since fragmented, realistic context is what a memory system is actually built to handle. Then add the case most test sets skip entirely: seed a fact, change it later in the timeline, and verify the agent returns the current value, and ideally the historical one when asked what was true as of an earlier date. Burying the relevant facts inside a larger volume of background data and unrelated noise completes the picture, since retrieval quality under a toy dataset with no distractors says little about retrieval quality under a realistic one.
A test set built this way surfaces failures automated scoring is well suited to catch. It is deliberately not built to catch every failure mode; some of what matters most has no labeled ground truth to grade against at all. What can’t an automated score catch?
Beyond automated scoring
What can’t an automated score catch?
Conflict resolution with no single correct answer, implicit memory learned from observation rather than stated explicitly, and domain-specific correctness that depends on rules a general scorer doesn’t know. These three failure modes need a human reviewing the agent’s actual trace, not a grader comparing a response to a fixed answer.
Conflict resolution often has no labeled ground truth for “the user’s current preference given their entire history,” because the correct resolution depends on judgment a static dataset cannot encode. A domain expert reading the trace can see the contradiction and judge whether the agent resolved it sensibly; an automated scorer, comparing against one fixed expected answer, cannot. Implicit memory, an agent picking up a user’s preferred communication style through observation rather than an explicit instruction, has the same problem: there is no ground-truth label for tone and pacing to grade against, only a human judgment of whether the agent’s behavior reflects what the user actually demonstrated. Domain-specific accuracy compounds this further: a fact can be retrieved correctly and still be applied wrongly given a domain’s own rules, in legal, clinical or financial contexts specifically, in ways a generic accuracy check has no way to catch.
The practical answer is not to replace automated scoring with human review, but to run both on a cadence: automated benchmark regression checks after every model or architecture change, since these run cheaply and catch retrieval degradation immediately, and a sample of production traces routed to human review on a fixed schedule, weekly or per release, specifically for the failure modes above that automated grading structurally cannot see. Each stage catches a different failure: automated checks catch the read stage, trace review catches the manage stage, and neither substitutes for the other.
Running both consistently produces a real number for your own system. It’s worth understanding why the industry’s most commonly cited numbers, the ones vendors publish about their own products, are so often impossible to compare against each other in the first place. Why do vendors report such different benchmark numbers?
Reading vendor claims
Why do vendors report such different benchmark numbers?
Because they are not run the same way. Different question counts, different scoring methodology, and different years, published by the vendor about its own product, on its own harness. The right response to seeing a high score is to ask how it was produced, not to accept it as a ranking against competitors.
Zep reports 94.7% accuracy on LoCoMo and 90.2% on LongMemEval from its own evaluation harness, without publishing the question count on the same page as the headline figures. Mem0 reports 92.5 on LoCoMo across 1,540 questions and 94.4 on LongMemEval across 500 questions, a different scale and a different scoring shape entirely. Letta, MemGPT’s original research team, ran a plain filesystem-search agent, no specialized memory tool at all, against LoCoMo and scored 74.0%, which it published as explicitly higher than Mem0’s own reported 68.5% for the same benchmark, adding that Mem0 did not respond to Letta’s request to clarify how its number was produced. None of these three figures share a comparable methodology, and stacking them into a ranked list would manufacture a winner none of the underlying evaluations actually support.
The deeper problem Letta’s own writeup raises is structural, not just a methodology mismatch: an agent’s measured memory performance depends heavily on the surrounding agent framework’s tool-calling ability, not on the memory mechanism in isolation. A more capable retrieval tool can still score poorly if the agent framework around it prompts and calls tools poorly, which means comparing two memory products through a benchmark score is comparing two entire systems, not two memory layers. The scoring pages at the LoCoMo benchmark and the LongMemEval benchmark go deeper into each benchmark’s own construction and limitations; this page’s point is narrower, that self-reported numbers from different vendors are not safely comparable to each other at all.
Self-reported numbers are one kind of evidence. A different, independently controlled kind of evidence exists too, and it says something more specific about where every system, regardless of vendor, actually struggles. What does a controlled academic evaluation actually show?
Independent evidence
What does a controlled academic evaluation actually show?
That every memory system tested, commercial and open source alike, fails multi-hop conflict resolution at no better than 6% accuracy, even when those same systems score well on retrieval-only tasks. This is the one figure on this page that comes from a controlled, peer-reviewed evaluation rather than a vendor’s own report.
MemoryAgentBench evaluated a range of systems, from simple context-based and retrieval-augmented agents to commercial memory agents including Mem0 and MemGPT, across all four competencies under one consistent protocol rather than letting each vendor run its own harness. The paper’s finding on conflict resolution is specific and stated plainly: every method it tested failed the multi-hop case, achieving at most 6% accuracy, while only long-context agents that keep the entire history in the prompt achieved reasonable results on the simpler single-hop version of the same task. The same evaluation also found that commercial memory agents underperform relative to long-context baselines on long-range understanding and test-time learning specifically, which the paper attributes in part to systems like Mem0 discarding substantial original content during fact extraction, making the underlying information harder to reconstruct for tasks that need more than an isolated fact.
This is exactly the gap a strong LoCoMo score hides. LoCoMo weights toward accurate retrieval on relatively low-information-density conversational data, a case several commercial systems handle well, while providing little signal about conflict resolution at all. A team that adopts a memory system on the strength of a published LoCoMo number, without separately testing whether it resolves an updated fact correctly, is extrapolating from the one competency benchmarks cover well to three others the same number says nothing about.
Individual metrics, vendor claims to read carefully, and one controlled academic finding are the pieces. Putting them together into something you can actually run against your own system is the last step. How should you actually run an eval, end to end?
Putting it together
How should you actually run an eval, end to end?
Build a test set covering all four competencies, score context completeness before answer correctness, report latency and token cost alongside accuracy, and route conflict resolution and implicit-memory cases to human review rather than an automated grader. Then repeat it on a cadence rather than treating a single run as a permanent verdict.
Public benchmarks are still useful, as a directional first filter and as a way to sanity-check that a system handles the basic retrieval case at all, following the standard-benchmark guidance covered on the individual benchmark pages linked above. What they cannot substitute for is testing on your own data, since your actual queries, your actual conflict patterns, and your actual session lengths are the shape a benchmark score was never built to predict. Run the same evaluation loop, test-set construction, completeness-then-correctness scoring, human review for the cases automated scoring cannot judge, on a fixed cadence: after every model or memory-architecture change for the automated portion, and on a weekly or per-release schedule for the human-reviewed portion.
Teams that want this evaluation loop, and the underlying extraction and conflict-resolution logic it is testing, handled for them rather than assembled from separate benchmark harnesses and review tooling can look at a managed option such as Engram as an alternative to building the full pipeline in-house. Whichever path is chosen, the standard to hold it to is the one this page opened with: a memory system has to be tested on all four competencies, not just the one industry benchmarks happen to make easiest to score.
FAQ
Frequently asked questions
The evaluation decisions that follow once the framework above is understood.
Is a high LoCoMo score enough to trust a memory system in production?
No. LoCoMo mostly tests accurate retrieval on relatively low-information-density conversational data. A controlled academic evaluation (MemoryAgentBench, arXiv:2507.05257) found systems that score well on retrieval-only tasks still fail multi-hop conflict resolution at no better than 6% accuracy. A strong LoCoMo score says nothing about that.
Should I compare Zep, Mem0 and Letta by their published LoCoMo scores?
Not directly. Each reports its number from a different harness, question count and scoring methodology, and Letta's own comparison notes that a memory score depends heavily on the surrounding agent framework's tool-calling ability, not the memory mechanism alone. Treat each figure as a claim about that vendor's own setup, not a ranked leaderboard.
How do I test whether an agent handles a fact that changes over time?
Seed a fact, change it later in the conversation timeline, then ask the question. A correct system returns the current value, and ideally the historical one when asked what was true as of an earlier date. Standard benchmarks like LoCoMo rarely test this case.
What is context completeness, and why measure it before answer correctness?
Context completeness measures whether retrieval surfaced every fact needed to answer, independent of what the language model then does with them. Measuring it first isolates a retrieval failure from a generation failure; if the facts were never retrieved, no downstream prompting will fix the answer.
Can an automated LLM-as-judge score catch a conflict-resolution failure?
Only partially. Conflict resolution and implicit memory often have no single labeled ground truth to grade against, since the correct resolution depends on judgment. A human reviewing the agent's actual trace can catch cases an automated grader, built to compare against one fixed expected answer, structurally cannot.
How often should a memory evaluation actually be re-run?
Run automated benchmark regression checks after every model or memory-architecture change, since these are cheap and catch retrieval degradation immediately. Route a sample of production traces to human review on a fixed cadence, weekly or per release, for the conflict-resolution and domain-specific cases automated scoring misses.