Evaluation · Benchmarks
What Does the LOCOMO Benchmark Measure?
LOCOMO evaluates whether an AI system can still answer questions about a conversation that ran for hundreds of turns across dozens of sessions, using dialogues averaging 300 turns and 9,000 tokens over as many as 35 sessions. It is the benchmark most memory vendors quote, and reading their numbers correctly requires knowing what it does and does not test.
The dataset
Definition
What is the LOCOMO benchmark?
LOCOMO is an evaluation for very long-term conversational memory, introduced in “Evaluating Very Long-Term Conversational Memory of LLM Agents” by Maharana, Lee, Tulyakov, Bansal, Barbieri and Fang (arXiv:2402.17753). It exists because the conversations used to test memory before it were far too short to distinguish a system that remembers from one that simply has a large context window.
The dialogues are constructed rather than scraped, grounded in personas and temporal event graphs so that the conversation has an internal history that can be asked about, and then verified by humans for consistency. That construction matters: the questions have checkable answers because the events behind them were designed.
The benchmark defines three tasks rather than one. Question answering over the conversation is the one everybody quotes. Event summarisation tests whether a system can describe what happened over time rather than retrieve a single fact. Multi-modal dialogue generation extends the evaluation beyond text.
Knowing what it contains makes the vendor numbers readable: what systems actually score on LOCOMO.
Results
What do systems score on LOCOMO?
The most-cited result is Mem0’s, which reports an LLM-as-a-judge score of 66.9 for its extract-and-store pipeline and 68.4 with graph memory, against 72.9 for a full-context baseline that costs roughly fifteen times as many tokens (Chhikara et al., 2025). Read plainly, the memory system scores slightly lower and costs dramatically less.
The efficiency figures from the same paper are the ones worth carrying into a design decision: a median search latency of 0.148 seconds, a p95 total latency of 1.44 seconds against 17.1 seconds for the full-context baseline, and roughly 1,800 tokens per query instead of 26,000.
That shape of result recurs across the category and is the honest summary of where memory systems stand. Retrieval rarely beats simply sending everything on accuracy; it wins on cost, latency, and the fact that sending everything stops being possible. A vendor quoting only the accuracy number is telling you the least interesting half.
One caution applies to every figure here. Both the tool and the score usually come from the same team, and no independent leaderboard exists for this category. Treat published numbers as claims with citations, which is how they are presented on the best AI memory tools.
LOCOMO is not the only benchmark, and the differences matter: how LOCOMO compares with LongMemEval and DMR.
The set
How does LOCOMO compare with LongMemEval and DMR?
Three benchmarks dominate the category, and they test different things: LOCOMO tests very long conversations, LongMemEval tests long-horizon recall across specific ability types, and DMR tests whether an agent can answer a question that depends on one earlier session.
| Benchmark | What it tests | Introduced by |
|---|---|---|
| LOCOMO | Question answering, event summarisation and multi-modal generation over conversations of up to 35 sessions | Maharana et al., arXiv:2402.17753 |
| DMR | Consistency: one question answerable only from an earlier session | Packer et al., arXiv:2310.08560 |
| LongMemEval | Long-horizon recall separated into distinct ability types | See LongMemEval |
DMR is the narrowest of the three, which is why scores on it cluster so tightly at the top: MemGPT reports 93.4% and Zep reports 94.8%, a gap of 1.4 points. A benchmark where everyone scores in the nineties has stopped discriminating, which is part of why the category moved toward LOCOMO and LongMemEval.
For a team choosing a tool, the useful move is to check which benchmark a vendor quotes and whether it resembles your workload. A system optimised for single-fact consistency is not automatically good at summarising six months of interaction. The metric set is covered on which metrics matter for agent memory.
All three share limitations that are worth naming before you rely on any of them: what LOCOMO does not measure.
Limits
What does LOCOMO not measure?
It does not measure whether memory stays correct as facts change, how a system behaves after months of accumulated writes, or anything about your domain. Those are the three failure modes that actually break production systems.
- Fact churn. The dialogues are internally consistent by construction. Real users change jobs, preferences and circumstances, and the resulting need to supersede a stored fact is not what LOCOMO scores. That problem is covered on conflicting memories.
- Store decay over time. A benchmark run starts from an empty store. It cannot detect a system that works well at 500 memories and retrieves badly at 50,000, which is the failure described on memory consolidation.
- Your vocabulary. Constructed persona conversations do not resemble support tickets, code review or clinical notes. Retrieval quality is sensitive to phrasing, so a domain gap moves results more than most people expect.
The practical conclusion is that benchmarks are a filter, not a decision. Use them to eliminate systems that cannot do the basic job, then measure the survivors on a set of your own questions whose answers depend on stored facts. That evaluation is described in how to add memory to an AI agent.
FAQ
Frequently asked questions
The questions that follow: whether LOCOMO is public, and how to run it yourself.
What is the LOCOMO benchmark?
LOCOMO tests long-conversation memory — whether agents recall facts from thousands of tokens earlier in the same chat. Standard eval for comparing memory frameworks.
Who published LOCOMO?
LOCOMO is widely cited in agent-memory research. The Mem0 paper (Chhikara et al., 2025) publishes comparative LOCOMO J scores for Mem0, Zep, LangMem and full-context baselines.
What is Mem0's LOCOMO score?
LOCOMO J 66.9 (Chhikara et al., 2025). Median search 0.148 s; ~1,800 vs ~26,000 tokens per query.
What is Letta's LOCOMO score?
N/A on LOCOMO publicly. The MemGPT paper (Packer et al., 2023) reports 93.4% on Deep Memory Retrieval (DMR) — a different benchmark.
LOCOMO vs LongMemEval?
LOCOMO = long single conversation recall. LongMemEval = cross-session recall with temporal gaps. Use both. See LongMemEval.
Can you pass LOCOMO with long context only?
Full-context baseline scored 72.9 J but uses ~26,000 tokens and 17.1 s p95 latency vs Mem0's ~1,800 tokens and 1.44 s p95 (Chhikara et al., 2025).
What is Engram's LOCOMO score?
N/A publicly as of July 2026 — Engram reached GA in June 2026. See Engram explained.
Where is the LOCOMO leaderboard?
No single official leaderboard — scores are published in framework papers. We track published results on best AI memory tools.