Evaluation · Benchmarks
The LongMemEval Benchmark, and Why Its Scores Disagree
LongMemEval tests five long-term memory abilities across 500 curated questions, and it is genuinely one of the better benchmarks in the field. It is also the clearest case on this site of vendors publishing self-reported top scores that conflict with each other, at different dates, with different judge models, and in at least one documented case, a single vendor publishing two different numbers for itself. This page describes the benchmark from its own paper and tells you how to read any score claimed against it.
Five abilities
Definition
What does LongMemEval actually measure?
Five distinct long-term memory abilities: information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention. The benchmark was introduced by Wu, Wang, Yu, Zhang, Chang and Yu at ICLR 2025 (arXiv:2410.10813), and it is designed specifically to go beyond simple needle-in-a-haystack recall.
Information extraction asks a system to recall a specific detail stated once, by either party, somewhere in an extensive history. Multi-session reasoning requires synthesising or comparing information that only makes sense combined across separate conversations, which a system that merely searches for the single most relevant passage will fail. Knowledge updates test whether a system recognises that a later statement supersedes an earlier one about the same fact, exactly the supersession problem covered on conflicting memories. Temporal reasoning tests whether a system uses timestamps and explicit time references correctly rather than treating every stored fact as equally current. Abstention, the fifth and most often overlooked ability, tests whether a system correctly declines to answer when the history genuinely does not contain the answer, rather than confidently inventing one.
That fifth category is worth pausing on. A benchmark that only rewards answering confidently would make a system’s failure mode worse, not better, since a system optimised purely for recall has every incentive to guess. Testing abstention directly is one of the reasons this benchmark is considered a serious instrument rather than a simple recall test.
Testing five abilities well requires test data built for the purpose rather than scraped from real conversations, which raises a fair question about scale and construction: how is the test data built, and how large is it?
Construction
How is the test data built, and how large is it?
500 questions, each embedded in a purpose-built, timestamped chat history, released in two sizes. The authors describe an attribute-controlled pipeline, inspired by the needle-in-a-haystack test format, that compiles a coherent and extensible history around each question rather than relying on found conversational data.
LongMemEval S embeds each question in a history of roughly 30 to 40 sessions, around 115,000 tokens. LongMemEval M scales that to roughly 500 sessions, around 1.5 million tokens, a scale that exceeds what most models can hold in a single context window at all, which forces a genuine memory mechanism rather than allowing brute-force inclusion of the full history. The code and data are published on the authors’ own GitHub repository, which is the authoritative source for reproducing either variant.
The construction pipeline matters as much as the question count, because a benchmark built from scraped conversations would struggle to guarantee that every question has a verifiable, locatable answer somewhere in its history. Purpose-building the histories around each question is what makes scoring an answer as correct or incorrect unambiguous, which is a precondition for the whole exercise being meaningful. Given a benchmark built this carefully, its own reported results are worth taking seriously: why is this benchmark considered hard, and what did the original paper find?
Original findings
Why is this benchmark considered hard, and what did the original paper find?
Long-context language models show a reported 30 to 60 percent accuracy drop on the easier LongMemEval S variant, and the paper’s own experiments point to specific, actionable design choices rather than only a difficulty score. This is worth citing precisely because it comes from peer-reviewed research published at ICLR, the benchmark’s own study of the problem, not from a party being measured by it.
The authors propose a three-stage framework for any long-term memory system: indexing, retrieval, and reading. Working within that framework, their own experiments produced four specific findings worth naming. Storing history at the level of a conversational round, rather than a whole session or an individual compressed fact, gave the best overall performance, though compressing further into individual facts specifically helped multi-session reasoning despite losing some information. Expanding the search index’s keys with facts extracted from the raw text, rather than indexing on the raw text alone, improved recall by roughly four points and downstream answer accuracy by roughly five. A time-aware indexing and query expansion strategy, narrowing the search range using temporal cues, improved recall on temporal reasoning questions by roughly seven to eleven points. And applying structured reasoning techniques at the point of reading retrieved results, rather than only at retrieval, improved reading accuracy by as much as ten points across three different models.
Those findings describe how to build a better memory system in the abstract. What they do not settle is which specific product, at any given moment, has actually implemented these ideas best, which is where the published landscape around this benchmark gets genuinely confusing: why do vendor scores on this benchmark disagree with each other?
The central problem
Why do vendor scores on this benchmark disagree with each other?
Because every published score in this landscape is self-reported, run at a different date, often scored by a different judge model, and none of them has been run head to head under one shared protocol. This is the single most important thing to understand before treating any number here as settled.
Within roughly a year, four different published posts each claimed the highest score recorded on this benchmark at the time of writing, each using its own methodology description and, in several cases, a different judge model scoring the answers as correct or incorrect. Each claim was true in a narrow sense at the moment it was published. None of them constitutes a stable ranking, because the next post in the sequence made the same claim again shortly afterward.
The clearest evidence for treating every number here cautiously does not come from a competitor trying to discredit anyone. A benchmark roundup published by a vendor with its own score in the same comparison table explicitly warns its readers that several of the listed numbers are disputed by the vendors named next to them, and documents a specific case of one company publishing two different, mutually inconsistent scores for itself across different posts. A source naming that problem while still appearing in its own table is a stronger signal than an outside party making the same accusation would be.
None of this means the benchmark is broken or that the numbers are meaningless. It means a published score is a claim about one run, on one date, under one methodology, and reading it correctly requires knowing what those three things were. Two details in particular change what a score is actually claiming: what do “oracle” and “full context” mean as comparison points?
Baselines
What do “oracle” and “full context” mean as comparison points?
An oracle baseline is given only the sessions that actually contain the answer; a full-context baseline is given the entire history and has to find the answer itself. These are opposite ends of a spectrum, and where a system’s score sits relative to each tells you something a raw percentage alone does not.
Beating a full-context baseline is a meaningful but modest claim: it says a system’s retrieval and reasoning outperform simply dumping everything into the model and hoping it copes. Beating, or even matching, an oracle baseline is a stronger claim, because the oracle is not really being tested on retrieval at all; it has already been handed the answer’s location and only has to read it correctly. A system that matches an oracle’s accuracy while actually having to find the relevant information first, among many sessions of irrelevant history, is demonstrating something oracle comparisons are specifically designed to reveal.
Reading which baseline a claimed score is compared against, and by how much, is more informative than the headline percentage on its own. A score with no baseline comparison stated at all is the weakest kind of claim, since there is nothing to judge it against. A related question matters just as much as the number itself: does a system that scores well also behave well in ways a single accuracy percentage cannot capture? does a higher score always mean a better memory system?
Beyond the number
Does a higher score always mean a better memory system?
No. Two systems with similar accuracy can differ substantially in properties a single percentage does not capture, and at least one of those properties matters as much as raw accuracy for a production system.
Whether the context a memory system produces is stable and reusable across turns is one such property. A system that dynamically retrieves and re-injects different context on every single turn produces a prompt that changes shape constantly, which defeats provider-side prompt caching entirely, covered on reducing token cost with memory. A system that produces a more stable, predictable context, even at comparable accuracy, can be meaningfully cheaper to run in production for reasons the accuracy score never measures.
Latency per question is another dimension published alongside some scores and not others, and it can vary by several times between approaches that score similarly. A system a full point behind another on accuracy but several times faster is not automatically the worse choice for a latency-sensitive product. None of these properties invalidate the accuracy score; they are additional axes a single number cannot represent, and a fair comparison between two systems has to look at more than one column.
Given how much a claimed score can vary in meaning depending on all of this, it is worth having a concrete, repeatable way to evaluate one whenever you encounter it: how do you evaluate a claimed LongMemEval score critically?
The checklist
How do you evaluate a claimed LongMemEval score critically?
Check the judge model, the dataset variant, the baseline compared against, and the publication date, in that order, before comparing any two numbers to each other. Two scores are only directly comparable when all four of these match.
The judge model matters because scoring an open-ended answer as correct or incorrect is itself a judgement call delegated to another model, and a stricter or more lenient judge shifts every score it touches, independent of the system being evaluated. The dataset variant matters because S and M differ by roughly thirteen times in context length, and a strong score on the smaller variant says nothing directly about performance on the larger one. The baseline matters for the reasons above: beating a weak full-context baseline is a different and lesser claim than approaching an oracle. And the publication date matters because these scores climb over time as techniques and underlying models improve; an older number is not wrong, it is simply measuring an earlier point in a field that is still moving.
Applying all four checks to any two scores usually resolves an apparent disagreement into “these are not measuring the same thing,” which is a more useful conclusion than trying to decide which vendor is telling the truth. If the four checks do line up and the scores still disagree, that is the rarer and more interesting case, and it is worth reading both sources’ full methodology sections rather than trusting either headline number alone.
All of this evaluation assumes trust in numbers other people ran. The most reliable way to know how a system performs on your own use case remains running the benchmark yourself: can you run the benchmark yourself?
Reproducing it
Can you run the benchmark yourself?
Yes. The dataset, the evaluation code, and instructions for constructing custom chat histories are published on the authors’ own GitHub repository. This is the authoritative source for the benchmark itself, separate from any vendor’s own fork or wrapper around it.
Running it directly against your own memory system, rather than relying on a vendor’s published number, sidesteps every problem raised above at once: you control the judge model, you choose the dataset variant appropriate to your use case, you decide the baseline to compare against, and the result is current by definition because you just ran it. It costs real compute and engineering time, which is exactly why so many published scores exist in the first place, but for a decision that actually matters, it is worth more than any comparison table.
Whichever way you arrive at a score, the same discipline applies to any other benchmark you encounter for agent memory, including LoCoMo, and to evaluation generally, covered on evaluating agent memory: read the methodology before the headline number, and treat a percentage with no stated judge model, dataset variant, baseline or date as an incomplete claim rather than a finished one.
FAQ
Frequently asked questions
The practical questions that come up once you are looking at a specific claimed score.
Which is harder, LongMemEval S or M?
M, substantially. It scales the same 500 questions to roughly 1.5 million tokens of history per question, versus roughly 115,000 for S, which exceeds most models' context windows entirely and forces a genuine memory mechanism rather than allowing the full history to simply be included.
Is a LongMemEval score comparable to a LoCoMo score?
No. They are different benchmarks testing different things with different question sets and different scoring methodologies. A strong score on one says nothing directly about performance on the other; see the LoCoMo benchmark for what it measures instead.
Why would a vendor publish two different scores for the same system?
Typically because the two posts used different judge models, different dataset variants, or were run at different times as the underlying system improved. One documented case in this landscape shows exactly this pattern, which is why checking methodology details matters more than comparing headline numbers directly.
Does a system that beats the oracle baseline mean it is better than perfect retrieval?
Not quite. It means the system, working from the full history, answered as well as or better than a system handed only the relevant sessions in advance. That can happen when the oracle's narrow context loses some useful surrounding information the full system still has access to, not because the system exceeded what perfect information access would allow.
How often do LongMemEval scores get updated by vendors?
Frequently enough that several different posts claimed the highest recorded score within about a year of each other in the set gathered for this page. Treat any specific number as a snapshot from its publication date, not a permanent ranking.
Should I trust a benchmark score more than my own testing?
No. A published score reflects one run under one methodology, possibly on data unlike yours. Running the benchmark yourself, or testing your own realistic queries against your own system, gives you a result you can actually verify and that reflects your specific use case rather than someone else's.