Evaluation · Benchmarks

What Does the LOCOMO Benchmark Measure?

LOCOMO evaluates whether an AI system can still answer questions about a conversation that ran for hundreds of turns across dozens of sessions, using dialogues averaging 300 turns and 9,000 tokens over as many as 35 sessions. It is the benchmark most memory vendors quote, and reading their numbers correctly requires knowing what it does and does not test.

The dataset

300
Turns
On average
35
Sessions
Up to
9K
Tokens
On average

Definition

What is the LOCOMO benchmark?

LOCOMO is an evaluation for very long-term conversational memory, introduced in “Evaluating Very Long-Term Conversational Memory of LLM Agents” by Maharana, Lee, Tulyakov, Bansal, Barbieri and Fang (arXiv:2402.17753). It exists because the conversations used to test memory before it were far too short to distinguish a system that remembers from one that simply has a large context window.

What the LOCOMO benchmark contains: very long conversations averaging 300 turns over up to 35 sessions, and the three tasks it scores.
Figure 1. The scale is the design. A conversation short enough to fit in a prompt cannot separate memory from context.

The dialogues are constructed rather than scraped, grounded in personas and temporal event graphs so that the conversation has an internal history that can be asked about, and then verified by humans for consistency. That construction matters: the questions have checkable answers because the events behind them were designed.

The benchmark defines three tasks rather than one. Question answering over the conversation is the one everybody quotes. Event summarisation tests whether a system can describe what happened over time rather than retrieve a single fact. Multi-modal dialogue generation extends the evaluation beyond text.

Knowing what it contains makes the vendor numbers readable: what systems actually score on LOCOMO.

Results

What do systems score on LOCOMO?

The most-cited result is Mem0’s, which reports an LLM-as-a-judge score of 66.9 for its extract-and-store pipeline and 68.4 with graph memory, against 72.9 for a full-context baseline that costs roughly fifteen times as many tokens (Chhikara et al., 2025). Read plainly, the memory system scores slightly lower and costs dramatically less.

Published benchmark results for agent memory tools, each with the paper that reported it.
Figure 2. LOCOMO sits alongside DMR in most vendor claims. Both are usually reported by the authors of the system being measured.

The efficiency figures from the same paper are the ones worth carrying into a design decision: a median search latency of 0.148 seconds, a p95 total latency of 1.44 seconds against 17.1 seconds for the full-context baseline, and roughly 1,800 tokens per query instead of 26,000.

That shape of result recurs across the category and is the honest summary of where memory systems stand. Retrieval rarely beats simply sending everything on accuracy; it wins on cost, latency, and the fact that sending everything stops being possible. A vendor quoting only the accuracy number is telling you the least interesting half.

One caution applies to every figure here. Both the tool and the score usually come from the same team, and no independent leaderboard exists for this category. Treat published numbers as claims with citations, which is how they are presented on the best AI memory tools.

LOCOMO is not the only benchmark, and the differences matter: how LOCOMO compares with LongMemEval and DMR.

The set

How does LOCOMO compare with LongMemEval and DMR?

Three benchmarks dominate the category, and they test different things: LOCOMO tests very long conversations, LongMemEval tests long-horizon recall across specific ability types, and DMR tests whether an agent can answer a question that depends on one earlier session.

The three agent memory benchmarks most often quoted, and what each one is built to detect.
BenchmarkWhat it testsIntroduced by
LOCOMOQuestion answering, event summarisation and multi-modal generation over conversations of up to 35 sessionsMaharana et al., arXiv:2402.17753
DMRConsistency: one question answerable only from an earlier sessionPacker et al., arXiv:2310.08560
LongMemEvalLong-horizon recall separated into distinct ability typesSee LongMemEval

DMR is the narrowest of the three, which is why scores on it cluster so tightly at the top: MemGPT reports 93.4% and Zep reports 94.8%, a gap of 1.4 points. A benchmark where everyone scores in the nineties has stopped discriminating, which is part of why the category moved toward LOCOMO and LongMemEval.

For a team choosing a tool, the useful move is to check which benchmark a vendor quotes and whether it resembles your workload. A system optimised for single-fact consistency is not automatically good at summarising six months of interaction. The metric set is covered on which metrics matter for agent memory.

All three share limitations that are worth naming before you rely on any of them: what LOCOMO does not measure.

Limits

What does LOCOMO not measure?

It does not measure whether memory stays correct as facts change, how a system behaves after months of accumulated writes, or anything about your domain. Those are the three failure modes that actually break production systems.

  • Fact churn. The dialogues are internally consistent by construction. Real users change jobs, preferences and circumstances, and the resulting need to supersede a stored fact is not what LOCOMO scores. That problem is covered on conflicting memories.
  • Store decay over time. A benchmark run starts from an empty store. It cannot detect a system that works well at 500 memories and retrieves badly at 50,000, which is the failure described on memory consolidation.
  • Your vocabulary. Constructed persona conversations do not resemble support tickets, code review or clinical notes. Retrieval quality is sensitive to phrasing, so a domain gap moves results more than most people expect.
Short-term context window memory compared with the long-term external store, by speed, size and whether it survives the session.
Figure 3. LOCOMO’s conversations are built to exceed what the left-hand store can hold, which is what forces a system to use the right-hand one.

The practical conclusion is that benchmarks are a filter, not a decision. Use them to eliminate systems that cannot do the basic job, then measure the survivors on a set of your own questions whose answers depend on stored facts. That evaluation is described in how to add memory to an AI agent.

FAQ

Frequently asked questions

The questions that follow: whether LOCOMO is public, and how to run it yourself.

What is the LOCOMO benchmark?

LOCOMO tests long-conversation memory — whether agents recall facts from thousands of tokens earlier in the same chat. Standard eval for comparing memory frameworks.

Who published LOCOMO?

LOCOMO is widely cited in agent-memory research. The Mem0 paper (Chhikara et al., 2025) publishes comparative LOCOMO J scores for Mem0, Zep, LangMem and full-context baselines.

What is Mem0's LOCOMO score?

LOCOMO J 66.9 (Chhikara et al., 2025). Median search 0.148 s; ~1,800 vs ~26,000 tokens per query.

What is Letta's LOCOMO score?

N/A on LOCOMO publicly. The MemGPT paper (Packer et al., 2023) reports 93.4% on Deep Memory Retrieval (DMR) — a different benchmark.

LOCOMO vs LongMemEval?

LOCOMO = long single conversation recall. LongMemEval = cross-session recall with temporal gaps. Use both. See LongMemEval.

Can you pass LOCOMO with long context only?

Full-context baseline scored 72.9 J but uses ~26,000 tokens and 17.1 s p95 latency vs Mem0's ~1,800 tokens and 1.44 s p95 (Chhikara et al., 2025).

What is Engram's LOCOMO score?

N/A publicly as of July 2026 — Engram reached GA in June 2026. See Engram explained.

Where is the LOCOMO leaderboard?

No single official leaderboard — scores are published in framework papers. We track published results on best AI memory tools.