Architecture · Tool Pattern
Memory as a Tool, Not a Background Process
Exposing memory as a tool means the agent explicitly calls remember, recall or forget as deliberate, visible actions, instead of an automatic pipeline silently extracting and storing facts after every turn. A concrete implementation of this pattern, memory blocks, and a genuinely different, similarly-named academic concept are both worth understanding clearly before building on either.
Memory tools
The pattern
What does exposing memory as a tool actually mean?
The agent is given a small set of callable functions, recall, remember, forget, and it decides on its own, turn by turn, whether to invoke them, instead of a background process silently extracting and writing facts after every message regardless of whether the agent asked for that.
This is a genuine architectural choice with a real tradeoff, not a stylistic preference. An automatic pipeline guarantees nothing gets missed, since extraction runs on every turn whether the agent thinks it’s relevant or not, but it also makes memory writes invisible: nothing in the agent’s own reasoning trace shows what got stored or why. A tool-based approach makes every memory operation an explicit, loggable action the agent chose to take, which is directly debuggable the same way any other tool call is, at the cost of depending on the agent to actually notice something is worth remembering and choose to act on it. Neither approach is strictly better; they trade completeness for visibility in opposite directions.
The tool surface itself is usually kept deliberately small: a handful of named operations, recall to search existing memory, remember to explicitly store something judged important, forget to mark something no longer reliable, rather than a general-purpose database interface the agent has to reason about from first principles on every call. Keeping the surface narrow matters for a practical reason: a model reasoning about whether and how to call a memory tool is spending attention on that decision instead of on the actual task, so the fewer, more specific the operations, the less overhead the pattern adds to every turn where memory isn’t actually relevant.
A concrete implementation of the tool-based side of this tradeoff gives the abstract idea real shape. What’s the difference between a memory block and a retrieval tool?
A concrete implementation
What’s the difference between a memory block and a retrieval tool?
A memory block is a labeled, character-limited region sitting directly in context that the agent rewrites through its own tool calls; a retrieval tool is a search action that fetches something from external storage, and the two solve different halves of the same problem.
A memory block carries four parts: a label naming it, a description telling the agent what belongs there, a value holding the actual tokens, and a character limit bounding how much of the context window it can consume. The agent updates a block’s value directly through a tool call when it learns something worth keeping, which is a genuinely different mechanism from calling a search tool against a vector or graph store to pull in something that already exists elsewhere. Splitting memory into several named blocks rather than one large undifferentiated region also gives the agent a way to update one specific thing, a stated preference, say, without disturbing an unrelated block holding a different kind of fact, which matters once a system is tracking more than a single category of information about a user or task.
External storage and retrieval, vector databases for semantic similarity search, graph databases for traversing relationships between stored entities, extend the same tool-based idea beyond what fits directly in context. The agent calls a search tool, gets a result back, and can choose to act on it or write something new to a block based on what it found. That distinction is worth stating precisely, since the two get treated as interchangeable constantly: retrieval is an action, it happens once and produces a result; memory is the state that persists afterward and gets rewritten as new information arrives. Calling a search tool for context is a tool for memory. It is not memory itself.
Memory blocks aren’t necessarily maintained only by the agent using them in the moment. A separate kind of process can maintain them too. Who else can modify an agent’s own memory tools?
A separate maintainer
Who else can modify an agent’s own memory tools?
A specialized agent, running asynchronously between interactions, can rewrite another agent’s memory blocks on its own, separating the work of using memory in the moment from the work of maintaining and consolidating it over time.
This division of labor is exactly what sleep-time compute is built around: a background process reasons about accumulated context while the primary agent is idle, and can rewrite memory blocks directly to improve them, consolidating redundant entries or restructuring what’s stored based on patterns only visible across many interactions rather than a single one. The primary agent still uses recall, remember and forget as its own tools during a live conversation; it simply isn’t the only process with write access to what it remembers. The mechanics of this asynchronous maintenance role are covered in full on sleep-time compute, which goes deeper into the cost tradeoff and the specific architecture this pattern requires.
Everything described so far is one coherent pattern under one name. A separate, similarly-named concept exists in recent research and deserves to be told apart clearly rather than quietly conflated. Is this the same thing as MemTool’s tool-memory management?
A naming collision
Is this the same thing as MemTool’s tool-memory management?
No, despite the similar name. MemTool is a rigorously benchmarked academic framework (arXiv:2507.21428) solving a different, adjacent problem: managing which tool definitions stay loaded in an agent’s context window across a multi-turn session, not managing facts or preferences about a user.
MemTool addresses a real, specific pain point: an agent that can discover and equip hundreds of tools or MCP servers still needs to remove the ones it no longer needs from its context window as a conversation moves on, or the tool list itself becomes the thing crowding out useful context. The paper tests three modes, autonomous (the agent manages its own tool list through function calls), workflow (a fixed, deterministic process handles it), and hybrid, across more than 13 language models and over 100 consecutive interactions. The measured results are specific and worth citing precisely: in autonomous mode, reasoning-capable models achieved 90 to 94 percent tool-removal efficiency, while medium-sized models managed only 0 to 60 percent, a sharp capability split the paper attributes to how reliably a model can judge which of its own currently-loaded tools are no longer relevant. Workflow and hybrid modes managed tool removal consistently across model sizes, at some cost to the autonomous flexibility that helped task completion in the other two modes.
This is a legitimate, well-evidenced piece of research, and it’s worth knowing about specifically because it isn’t about the same thing this page has covered so far: it never addresses facts about a user or preferences that should persist across sessions, the subject of every other section here. Citing MemTool as evidence for the memory-blocks pattern, or vice versa, would misrepresent both.
With the pattern, its concrete implementation, its maintenance model, and its naming collision all clarified, the remaining question is practical. When does this pattern actually earn its complexity?
The decision
When does this pattern actually earn its complexity?
When debuggability and explicit control over what gets remembered matter more than guaranteeing nothing gets missed, which tends to be true for agents with a narrow, well-defined set of things worth remembering rather than agents expected to passively absorb everything a user says.
A coding assistant deciding whether a specific convention is worth remembering, or a support agent explicitly noting a customer’s stated preference, both fit this pattern well: the set of things worth remembering is narrow enough that giving the agent explicit tools to manage it directly is tractable, and having a visible record of exactly what got remembered and when matters for debugging a wrong response later. An agent expected to build a broad, comprehensive profile of a user from everything they say across long, unstructured conversations is a better fit for the automatic extraction pipeline covered generally on memory management, since relying on the agent to notice and act on every worthwhile detail in real time asks it to do reliably what a background process is better suited to doing exhaustively.
The two approaches aren’t mutually exclusive either. A common production shape runs both: an automatic pipeline handles the baseline extraction so nothing important is missed by default, while the agent still has recall, remember and forget available as tools for the cases where it needs to act on memory deliberately mid-conversation, correcting something the pipeline got wrong, or explicitly confirming a preference the user just stated rather than waiting for a background process to catch it later. Choosing between the two isn’t really a single, permanent architectural decision so much as a question to revisit per category of information a system needs to track. The general step-by-step process for wiring either approach into an agent is covered on adding memory to an agent.
FAQ
Frequently asked questions
The practical decisions that follow once the pattern and its naming collision above are understood.
Is memory as a tool the same thing as MemTool?
No. Memory as a tool means exposing recall, remember and forget as callable operations for facts and preferences. MemTool (arXiv:2507.21428) is a separate research framework for managing which tool definitions stay loaded in an agent's context across a multi-turn session.
Can automatic extraction and tool-based memory be used together?
Yes, and it's a common production shape. An automatic pipeline handles baseline extraction so nothing is missed by default, while the agent still has explicit tools for cases where it needs to act on memory deliberately, like correcting a stored fact mid-conversation.
Why keep the memory tool surface small instead of a general database interface?
A model reasoning about a broad, general-purpose interface spends more attention deciding how to call it correctly. A small set of named operations, recall, remember, forget, reduces that overhead on every turn where memory isn't actually relevant.
Is retrieval the same thing as agent memory?
No. Retrieval is an action, calling a search tool that returns a result once. Memory is the state that persists afterward and gets rewritten over time. A retrieval tool is one way to access memory, but it isn't memory itself.
Can something other than the primary agent modify a memory block?
Yes. A specialized process running asynchronously between interactions, such as a sleep-time agent, can rewrite another agent's memory blocks directly, separating using memory in the moment from maintaining it over time.
Why did MemTool's autonomous mode perform so differently across models?
The paper's own benchmark found reasoning-capable models achieved 90 to 94 percent tool-removal efficiency in autonomous mode, while medium-sized models managed only 0 to 60 percent, attributed to how reliably a model can judge which currently-loaded tools are no longer relevant.