Guides · Cost

How Memory Reduces Token Cost in AI Agents

Token cost grows because a running agent keeps paying to re-read everything it already said, not because any single call is expensive. Memory reduces that cost by moving the work of deciding what matters from every read to a single write, so a turn re-injects a distilled fact instead of the conversation that produced it. The saving is real and it is not free: extraction discards what it judges unnecessary, and knowing what that tradeoff actually costs matters as much as knowing the saving exists.

Full context (LOCOMO)

~26,000 tokens / query

Selective memory (LOCOMO)

~1,800 tokens / query

The mechanism

Why does an agent’s token cost grow the longer it runs?

Because most agent loops send the entire accumulated history back to the model on every call, so cost compounds with each turn rather than staying flat. Nothing about any individual step is expensive. The transcript itself becomes the cost problem.

Illustrative token growth per turn in an agent loop as tool results accumulate in the transcript
Figure 1. Illustrative accumulation pattern from a published context-engineering benchmark.

A typical agent turn starts small: a system prompt and a user message. Each subsequent step appends whatever the previous step produced, a file read, a tool result, a search hit, and the entire growing list is resent as input on every following call, because that is how a model without persistent state stays aware of what already happened. By the tenth step, the model is re-reading nine steps of results it has already seen, every single time, just to answer the tenth question.

A published context-engineering benchmark shows the shape of this concretely across five iterations of a real agent run: the first turn carries roughly 900 tokens, growing to roughly 3,400 after a directory listing is appended, roughly 8,900 after a file read, roughly 14,200 after a second file read, and roughly 18,900 once a search result joins the mix. None of those individual additions is unreasonable on its own. The problem is that every one of them is still being paid for by the time the agent reaches its tenth or twentieth step, because nothing in a plain loop ever removes anything from what gets resent.

This is not a wasteful implementation choice; it follows directly from how a stateless model works. The fix is not making individual calls cheaper. It is reducing how much of the history has to be resent at all, and the two general techniques for doing that, summarizing it and extracting facts from it, trade away different things, covered further down. First it is worth seeing the actual scale of the problem in numbers rather than intuition: how much does selective memory retrieval actually save?

The published figure

How much does selective memory retrieval actually save?

On the LOCOMO benchmark, roughly 1,800 tokens per query for selective retrieval against roughly 26,000 for a full-context approach, a reduction on the order of fourteen times. This is the figure this site cites elsewhere and it is worth being precise about where it comes from rather than treating “memory saves tokens” as a vague truism.

The comparison is between two ways of giving a model what it needs to answer a question about a long-running conversation. Replaying the full conversation history means the model reads everything ever said, whether or not it is relevant to the current question. Selective retrieval means a memory system has already extracted the durable facts from that history and returns only the small number relevant to the current query. The tokens saved are not free of any cost, since retrieval itself and the extraction that made it possible both cost something, but the arithmetic remains lopsided in memory’s favour at any real conversation length.

The same shape of saving shows up outside conversational benchmarks, in ordinary multi-step agent loops that never touch a dedicated memory system at all: what does the savings look like against real model pricing?

The arithmetic

What does the savings look like against real model pricing?

Dramatic even without a dedicated memory system, purely from bounding how much of the history gets resent. Research into agent-loop context costs has worked this arithmetic directly against a specific model’s published pricing rather than quoting a rounded percentage.

Token cost comparison for a file reading agent across a naive ten step loop, a constrained context window, and a single pass approach
Figure 2. File-reading agent, 10 iterations, published Claude Sonnet 4.6 pricing.

For a file-reading agent running ten iterations, sending the full accumulated history on every call totals roughly 472,500 input tokens against published per-token pricing. Constraining the context to keep only the two most recent iterations, discarding the rest rather than resending it, cuts that to roughly 260,000 input tokens. A hypothetical single-pass version of the same task, with no loop at all, costs roughly 9,000. Every number here is reproducible against your own token counts and your own provider’s published rate; none of it depends on trusting a vendor’s summary figure.

Notice what the constrained approach still costs: a large multiple of the single-pass baseline, even though it discards most of the history. That gap is the ceiling on how far windowing alone can go, and it is exactly the gap a memory system, extracting durable facts once rather than repeatedly re-reading raw history, is built to close further. Whether the right technique for closing it is summarizing that history or extracting facts from it is not the same decision, and conflating them hides a real tradeoff: should you summarize the conversation or extract facts from it?

Two techniques

Should you summarize the conversation or extract facts from it?

They trade away different things, and the right choice depends on which kind of loss you can tolerate. Both reduce the tokens carried forward. Neither does it for free.

Summarization compresses conversation while keeping its shape, extraction keeps only durable facts and drops the rest
Figure 3. Summarization trades precision. Extraction trades completeness.

Summarization asks a model to condense a stretch of conversation into a shorter passage describing what happened. It preserves the shape of the exchange and loses precision: exact figures, direct quotes, and specific phrasing are the first casualties, because a summary by definition restates rather than repeats. It is a good fit when what matters going forward is the gist, and a poor fit when a later turn needs an exact number that got smoothed over in the retelling.

Extraction asks a model to decide, statement by statement, what is durable enough to keep as a standalone fact, and discards the rest entirely rather than compressing it. This is the mechanism a dedicated memory layer uses, described in depth on writing memories, and its loss profile is different in kind from summarization’s: instead of everything surviving in blurrier form, most of the conversation does not survive at all, and only what was judged worth keeping does. This is precisely what one independent, honestly reported test measured directly: what does an independent test of a memory layer against a raw vector store show?

One independent test

What does an independent test of a memory layer against a raw vector store show?

An 87% reduction in tokens re-injected per turn, and a real cost that came with it: roughly one in four support tickets survived extraction as a kept memory. One builder ran both storage approaches over the same 200 support tickets and the same questions, and reported both numbers rather than only the favourable one.

The setup compared an extractive memory layer, built on Engram, against a conventional vector store that returns raw retrieved chunks. Holding the data and the questions fixed and changing only what came back at read time, a distilled fact versus a raw chunk, the extractive approach re-injected 87% fewer tokens per turn for the same answer quality. The author was careful to state precisely what that figure measures: the compression of the stored unit, not a claim about retrieval speed, and not a substitute for a vector database’s job.

A vector store holding company truth compared with a memory layer holding customer truth, from one independent tester's 200 ticket comparison
Figure 4. One tester’s 200-ticket comparison, not a claim this site independently verified at scale.

The same test reported the honest cost of that saving directly: of 200 tickets, extraction kept only 52 as durable memories. For a customer’s evolving state, the author treats that as correct behaviour, since most of what is said in a support ticket is not worth remembering permanently. For a knowledge base, where any ticket might later be the one a question needs, discarding most of them would be a defect rather than a feature. The architecture the test settles on reflects that distinction directly: a vector store holding the complete, stable body of company knowledge verbatim, and a memory layer holding the personal, evolving facts about one customer, distilled rather than complete, meeting only at the moment a prompt is assembled.

Read this as one independent test at a specific, modest scale, two hundred tickets, not a benchmark this site ran or independently verified at production volume. What it demonstrates convincingly is the mechanism, not a number to quote as a universal rate: moving the work of deciding what matters from read time to write time is what produces the saving, and that mechanism does not depend on the exact scale it was measured at. The storage-layer requirements that make this practical at real volume are a separate, more mundane question: what does a cost-sensitive memory store need from its backend?

Storage requirements

What does a cost-sensitive memory store need from its backend?

Scoped isolation per user, fast mutation with read-after-write visibility, and selective retrieval that stays fast under concurrent reads and writes. These are not cost optimisations in themselves; they are the preconditions that let extraction and retrieval run often enough, on a small enough footprint, to actually produce the savings above.

Scoped ownership matters because a cost-efficient memory store is queried per user on nearly every turn, and a backend that cannot cheaply isolate one user’s memories from everyone else’s either leaks data or pays a broader query cost than it should. Fast mutation with read-after-write visibility matters because extraction runs on a write path that a subsequent turn may query almost immediately, and a backend with a long consistency lag reintroduces the staleness problem memory was supposed to solve. Caching and data-life controls matter because not every memory needs the same retention or the same tier, and a backend without those controls forces every record into the most expensive path by default.

These properties also determine whether the retrieval side of the equation stays cheap as usage grows. A store that cannot filter cheaply to one user’s records ends up scanning a wider index than it needs to on every query, which is a cost that compounds with user count in a way the token savings alone do not offset. Getting scoping right at the schema level, rather than bolting it on as an afterthought, is what keeps the per-query cost of retrieval flat as the number of users served grows, instead of climbing alongside it.

None of this requires an exotic architecture; it requires choosing a backend with these properties deliberately rather than by accident, which is the fuller subject of storage backends. One complementary technique is worth separating out clearly from all of this, since it addresses a related cost and not the same one: does prompt caching replace the need for memory?

A different lever

Does prompt caching replace the need for memory?

No. Prompt caching reduces the cost of tokens you send anyway; memory reduces how many tokens you have to send in the first place. The two are complementary, not competing, and confusing them leads to expecting caching to solve a problem it was never built for.

Provider-side prompt caching lets a static prefix, a long system prompt, a shared instruction block, be computed once and reused across calls rather than reprocessed every time. It is genuinely valuable, and it does nothing about a growing conversation history, because that history changes on every turn and therefore cannot sit in a fixed, reusable prefix. Structuring a prompt so the static, cacheable parts come first and the dynamic, per-turn content comes last is worth doing regardless of whether memory is also in use.

Memory addresses the other half of the problem: the part of the prompt that does change turn to turn, replacing an ever-growing raw history with a small set of relevant facts. A production system aiming to minimise cost typically wants both, caching what is genuinely static and letting memory keep what is genuinely dynamic small, rather than treating either as a substitute for the other.

Bringing all of this together into an order worth actually following: how do you actually reduce token cost with memory, in order?

Where to start

How do you actually reduce token cost with memory, in order?

Measure what you are actually spending before changing anything, bound the loop first, then decide between summarization and extraction based on which loss you can tolerate. Each step is cheap and each one is wasted if done out of order.

Start by counting retrieval and history tokens specifically, not just total spend, since the fix depends entirely on where the tokens are going. Bounding the context window, keeping only the last few turns rather than the full history, is the cheapest lever and worth pulling before building anything more sophisticated, since it requires no new infrastructure. From there, the choice between summarizing what remains and extracting durable facts from it should follow directly from what a later turn actually needs: exact precision on everything said, or a smaller set of durable facts with most of the rest discarded. Building or adopting a memory layer, whether hand-rolled following writing memories or a managed one such as Engram, is the step that gets extraction’s savings without hand-building the reconciliation and scoping logic that keeps a growing collection of facts from becoming its own mess, covered on memory consolidation.

FAQ

Frequently asked questions

The practical questions that follow once the mechanism is understood.

Does memory always reduce token cost?

For long-running or repeat-visit agents, usually yes, because full history replay grows without bound while extracted memory stays small. For short, single-turn interactions with no history to accumulate, there is little history to save on, and the overhead of extraction may not be worth it.

Is the 1,800 versus 26,000 token figure specific to one tool?

It comes from Mem0's published LOCOMO benchmark results (Chhikara et al., 2025) and describes selective retrieval against full-context replay generally, not a claim verified for every memory product. Treat it as evidence for the mechanism, not a number every tool will reproduce exactly.

Does extraction lose more information than summarization?

It loses more items and less precision. Summarization keeps a compressed version of everything that happened but blurs exact details; extraction keeps only what is judged durable, in full precision, and drops the rest entirely rather than compressing it. Which loss matters more depends on what a later turn actually needs.

How do I measure whether memory is actually saving tokens in my own agent?

Log input tokens per call broken down by source: system prompt, conversation history, and retrieved memory. Compare a run with full history replay against one using retrieval, on the same conversation, before assuming memory helps in your specific setup rather than trusting a general figure.

Should a knowledge base be stored the same way as per-user memory?

No. A knowledge base needs completeness, since any document might be the one a future question needs, and belongs in a store that keeps it verbatim. Per-user memory benefits from lossy extraction that keeps only durable facts. Mixing the two in one lossy pipeline risks losing knowledge-base content that should never have been dropped.

Does prompt caching make memory unnecessary?

No, they solve different parts of the cost. Prompt caching cuts the cost of resending a static prefix, such as a long system prompt. Memory cuts the size of the dynamic part of the prompt, the conversation-specific content that changes every turn and therefore cannot be cached as a fixed prefix.