Architecture · Compression

Summarizing and Compressing Agent Memory

Memory summarization compresses long conversation histories into shorter representations, fitting more effective context into the window and cutting token cost, with a trade-off in fine-grained detail.

Raw history

26,000 tokens · 40 messages

Summary

~1,800 tokens · key facts preserved

Motivation

Why summarize agent memory?

Context limits, token cost and context rot: raw history eventually exceeds what models can use effectively.

Mem0’s selective retrieval uses ~1,800 tokens per LOCOMO query vs ~26,000 full-context, over 90% fewer (Chhikara et al., 2025). Summarization is the main compression lever before selective retrieval kicks in.

→ Context window problem · Reduce token cost with memory

Techniques

What techniques actually compress agent memory?

Seven techniques in production use, from simple truncation to an agent that decides for itself when to compress.

  • Rolling window summary: summarize every N turns, keep only the latest summary plus recent messages
  • Recursive hierarchical summary: summarize summaries for multi-session histories
  • Anchored incremental summarization: maintain one persistent structured summary and extend it per eviction span rather than regenerating from scratch; named by Mem0’s own engineering research as the current state of the art over plain rolling summary, and the underlying idea behind Factory’s, Anthropic’s and OpenAI’s own production compression endpoints
  • Structured handoffs: a fixed schema for the compressed summary itself, with dedicated fields for files modified, tools called, decisions made, in-progress state, constraints and preferences, so a later compression pass updates the existing document instead of regenerating it from scratch and silently dropping a field nobody remembered to re-derive
  • Extractive: pull key sentences verbatim (lower hallucination risk)
  • Abstractive: LLM rewrites in fewer tokens (higher compression, higher risk)
  • Embedding-based compression: store historical turns as dense embeddings and reconstruct only semantically relevant segments; Mem0’s own engineering research reports 80 to 90% token reduction this way, at which point the line between aggressive compression and external memory with retrieval becomes definitional
  • Agent-driven autonomous compression: the agent itself decides when to compress rather than an external scheduler; see below

Trade-off: more compression leads to higher risk of losing retrievable facts. Validate with LOCOMO recall@k after summarization.

Autonomous compression

Can the agent decide for itself when to compress?

A 2026 research paper tested it directly: an agent that compresses its own history on its own schedule, not a fixed threshold.

A 2026 study (Verma, arXiv:2601.07190) built Focus, an agent equipped with 2 tools, start_focus and complete_focus, that let it checkpoint its own exploration, summarize what it learned, and delete the raw messages in between. Tested on 5 hard SWE-bench Lite instances with Claude Haiku 4.5, Focus cut total token usage by 22.7% (14.9 million to 11.5 million tokens) while matching baseline accuracy exactly (3 of 5 tasks passed, both agents). The paper’s own ablation is the more interesting finding: passive prompting produced only 2.0 compressions per task and a marginal 6% savings, while an aggressive prompt that told the agent to compress every 10 to 15 tool calls pushed that to 6.0 compressions per task and the full 22.7% savings, at identical accuracy. Per-instance results ranged from 57% token savings on an exploration-heavy matplotlib bug to a 110% token increase on a pylint task that needed continuous, iterative refinement rather than distinct explore-then-implement phases.

→ Agentic memory patterns

Seven memory summarization techniques from rolling window summary to agent-driven autonomous compression
Anchored incremental summarization is the technique 3 independent production systems converge on.

Failure modes

What does summarization actually lose?

Five specific things, named by Mem0’s own engineering research, not a vague loss of detail.

  • Exact numeric values: “the retry limit is 3” becomes “retries were configured”; the 3 disappears
  • Hard constraints: a rule stated once (“no Redis in this environment”) is assumed permanent by the user and compressed out by the third summarization cycle
  • Decision reasoning: the what of a decision survives, the why does not, so an agent that knows Postgres was chosen but not why makes the wrong call at the next decision point
  • Cross-turn dependencies: a file modified at turn 12 that a tool at turn 47 depends on; compressors process each eviction span independently and miss the link
  • Implicit preferences: coding style or tone the user demonstrated but never stated explicitly

One independently-measured finding backs this up directly: a Factory Research study that probed 3 production compression methods (Factory’s own, Anthropic’s and OpenAI’s) across 36,611 real messages found artifact tracking, specifically knowing which files an agent had already read or modified, was the weakest of 6 measured dimensions for every method tested, scoring between 2.19 and 2.45 out of 5. Even the best-performing method in that study could not reliably answer “which files have we touched,” which is exactly the kind of cross-turn dependency the list above predicts will break.

There is also a timing problem underneath all 5 categories, not just a technique problem: most compressors trigger only once a session reaches 50 to 70% of the context window. By that point a real session may already hold 25 or more turns of constraints and preferences that need to survive, and extracting them at compression time is already too late, since the compressor cannot know which of those 25 turns matter unless something already flagged them before the summarizer ever ran.

Implementation

Should you compress before you summarize?

Offload large tool content to a filesystem before summarization ever triggers, so the summarizer never has to compress what a pointer could replace.

LangChain’s own Deep Agents SDK implements this as a 2-step fallback with real, disclosed thresholds. Any tool response over 20,000 tokens is offloaded to the filesystem immediately and replaced with a file-path reference plus a 10-line preview, regardless of how full the context is. Only when the session crosses 85% of the model’s context window does the SDK fall back further: first truncating old tool-call inputs (file writes and edits, whose full content is already persisted to disk), then, if that still is not enough room, running an actual summarization pass that keeps a structured in-context summary while writing the complete original messages to the filesystem as a canonical record the agent can search back into.

The reason to offload before summarizing rather than summarizing everything uniformly is cost, not just quality. A large tool response, a full file dump or a long API payload, compresses well as a pointer at zero LLM cost, while running it through a summarization model burns tokens generating a prose description of content the agent could just re-read from disk if it ever needed the detail again. Reserving the summarization step for content that genuinely cannot be reduced to a pointer, actual conversation turns and reasoning, keeps the expensive step reserved for the one place a pointer cannot substitute.

→ Memory consolidation

Two-step offload-then-summarize fallback: filesystem offload at 20000 tokens, summarization fallback at 85 percent of the context window
Offloading to disk first means the summarizer only has to compress what a pointer genuinely cannot replace.

Disambiguation

How does summarization differ from consolidation?

Summarization compresses representation; consolidation promotes to LTM and merges duplicates, often used together.

Typical pipeline: summarize session transcript, extract durable facts, merge into long-term store, evict raw transcript. Consolidation without summarization still works for fact-extraction pipelines (Engram, Mem0). Mem0’s own engineering research frames this as a 3-tier architecture worth naming explicitly: tier 1 is in-context working memory, the current session, full and lossless; tier 2 is compressed session memory, where anchored summarization keeps a single session coherent as it grows; tier 3 is the external persistent store, facts extracted before compression ever fires and retrieved at the start of the next session. Compression and consolidation are not competing techniques under this framing, they occupy different tiers, and a production agent typically needs both.

→ Memory consolidation

Frameworks

How do frameworks implement summarization today?

Every major framework ships its own summarize-before-store path, at different levels of built-in support.

  • Engram: Weaviate-backed write pipeline; summarize before store in custom middleware
  • Letta: built-in compaction tools for core/archival memory paging
  • LangChain: ConversationSummaryMemory, ConversationSummaryBufferMemory, and the newer Deep Agents SDK’s offload-then-summarize fallback
  • Mem0: extraction pipeline compresses to salient facts automatically
  • Custom: LLM summarize-before-store in your middleware

→ LangChain memory · Letta alternatives

Quality

How do you measure summarization quality?

Recall of key facts post-summary: run LOCOMO or LongMemEval before and after compression, or probe the agent directly.

LOCOMO J scores: Mem0 66.9, Zep 66.0, LangMem 58.1 (Chhikara et al., 2025 eval). If summarization drops recall below your threshold, reduce compression ratio or switch to fact extraction instead of abstractive summary.

LOCOMO and LongMemEval measure memory retrieval generally, not compression specifically, so a purpose-built probe methodology fills a real gap. Factory Research’s own study asked 4 kinds of question after each compression event: recall (“what was the original error message?”), artifact (“which files have we modified?”), continuation (“what should we do next?”), and decision (“what did we decide about the Redis issue?”), then graded the answers on 6 dimensions with an LLM judge. Across 36,611 production messages, Factory’s own anchored-summarization approach scored 3.70 overall against Anthropic’s 3.44 and OpenAI’s 3.35 on a 0 to 5 scale, a self-reported result from Factory’s own study rather than an independent benchmark. The 3 methods work differently under the hood: Anthropic’s Claude SDK regenerates a full structured summary (typically 7,000 to 12,000 characters) on every compression pass, OpenAI’s `/responses/compact` endpoint produces an opaque compressed representation with the highest compression ratio of the 3 (99.3%) but no way to read back what it kept, and Factory’s own approach merges each new eviction span into one persistent summary rather than regenerating it. The more transferable finding is methodological: compression ratio alone was the wrong metric in that study, since OpenAI’s approach compressed the most aggressively yet scored lowest on whether the agent could actually continue the task afterward; what matters is tokens spent completing the whole task, not tokens saved per individual request. LangChain’s own Deep Agents evals apply a similar idea in miniature, a needle-in-the-haystack test that embeds one fact early in a conversation, forces a summarization event, and checks whether the agent can still recover that fact, specifically to catch the goal-drift failure mode where an agent loses track of what it was doing right after a summary replaces its history.

→ AI memory metrics · LOCOMO benchmark

Probe-based evaluation: recall, artifact, continuation and decision probes graded across six dimensions
Artifact tracking scored weakest of six dimensions for every compression method a real study tested.
Three-tier memory architecture: in-context working memory, compressed session memory, and external persistent store
Summarization owns tier 2; consolidation owns tier 3; most production agents need both.

FAQ

Frequently asked questions

The techniques, trade-offs and evaluation methods above, answered as direct questions.

Should agents summarize memory every turn?

No. Summarize at session end, memory pressure thresholds or every N turns. Per-turn summarization adds latency and can lose facts before extraction runs.

Does summarization lose accuracy?

Yes. Abstractive summaries can drop fine-grained facts like exact numbers, hard constraints and decision reasoning. Validate recall@k on LOCOMO after compression. Fact extraction (Engram, Mem0) often preserves accuracy better than free-form summary.

Summarization vs retrieval for long histories?

Use both: summarize for thread overview, retrieve specific facts from LTM for precision. Mem0 uses ~1,800 vs ~26,000 tokens per LOCOMO query (Chhikara et al., 2025).

Can Claude summarize agent memory?

Yes, any LLM can summarize transcripts before store or on session end. Same pattern for Claude, GPT-4o or open models. See persist conversation memory.

How much token savings does summarization provide?

Mem0 reports ~1,800 vs ~26,000 tokens per LOCOMO query, over 90% reduction vs full-context (Chhikara et al., 2025). Your ratio depends on compression aggressiveness.

Where do you persist summaries?

Vector store as episodic memory, archival tier (Letta), or replace raw transcript in Redis/Postgres. See storage backends.

Production summarization pattern?

Session end, LLM summarize, extract facts, write to LTM, evict raw transcript. Run LOCOMO in CI to gate recall regressions. See reduce token cost.

Can an agent decide for itself when to compress its own memory?

A 2026 study (Verma, arXiv:2601.07190) tested exactly this: an agent with start_focus/complete_focus tools cut token usage 22.7% at identical accuracy on SWE-bench Lite, but only with aggressive prompting; passive prompting yielded just 6% savings.

What's the single biggest thing that breaks when you compress agent memory?

Artifact tracking, knowing which files an agent already read or modified. A Factory Research study across 36,611 messages found this was the weakest of 6 measured dimensions for every compression method tested, scoring under 2.5 out of 5.

Continue exploring

Three routes onward: promoting summaries to long-term memory, evicting what’s left, or the full persistence guide.