Tools · Framework Memory
LlamaIndex Memory for AI Agents
LlamaIndex’s Memory class buffers short-term chat history and flushes it into one or more long-term memory blocks once a token ratio is exceeded, all inside a single Python class you configure directly. It is a component of the LlamaIndex agent framework rather than a standalone memory service, which shapes both what it does well and where it stops.
Flush path
Definition
What is Memory in LlamaIndex, and how does an agent use it?
A component you attach to an agent that stores information with `put()` and retrieves it with `get()`, backed by a class called `Memory` that replaces the older, now-deprecated `ChatMemoryBuffer`. The concept is simple by design: as the agent runs, every turn is written in and read back out through those two calls, and everything else on this page is about what happens between them.
The framework’s own documentation is explicit that this is a component, not a product: you customise memory by using an existing `BaseMemory` class or writing your own, and by default a basic buffer keeps the last messages that fit into a token limit. That basic behaviour is what `ChatMemoryBuffer` provided and what the `Memory` class now does with more configuration available.
Because `Memory` is a class you instantiate inside your own agent code, every parameter below is something you set explicitly rather than something a managed service decides for you. That is the tradeoff running through the whole module: more control, and more that is your responsibility to configure correctly. The first parameter worth knowing is the one that decides when anything moves out of the fast, in-context buffer at all: how does short-term memory decide when to flush to long-term storage?
Short-term memory
How does short-term memory decide when to flush to long-term storage?
By a token ratio, not a message count. Once the chat history’s share of the total token budget crosses a configured threshold, the oldest messages are moved out of the short-term buffer and into whichever long-term memory blocks are configured, in a fixed-size chunk rather than all at once.
Three parameters govern this, and the framework documents specific defaults for each: a `token_limit` of 30,000, the combined ceiling for short-term and long-term memory together; a `chat_history_token_ratio` of 0.7, meaning short-term chat history is allowed to occupy up to seventy percent of that ceiling before anything moves; and a `token_flush_size` of 3,000, the amount moved into long-term memory each time the ratio is crossed rather than a full dump of everything over the line.
The practical implication is that raising the token limit alone does not change the balance between recent conversation and durable memory, since the ratio scales with it. Changing the ratio down means the agent holds less raw conversation and pushes information into long-term memory sooner, which trades a smaller in-context buffer for earlier processing by whichever blocks are attached. There is no single correct setting; it depends on how much of what is said needs to survive verbatim versus how much can be safely compressed by extraction.
One migration note worth acting on if you have existing code: the older `ChatMemoryBuffer` class is deprecated and still the framework-wide default in places, with `Memory` intended to replace it. New agent code should configure `Memory` directly rather than relying on the old default quietly continuing to work.
Flushing only matters if something useful happens to the data once it lands in long-term storage, and that is where the module’s real design choices live: what are the three long-term memory blocks, and when does each apply?
Long-term memory
What are the three long-term memory blocks, and when does each apply?
Static content, extracted facts, and vector-searchable message batches, each its own block type, combinable in one Memory instance. Flushed messages are handed to whichever blocks are configured, and each processes them differently.
`StaticMemoryBlock` holds content that was supplied once and does not change: a name, a location, an employer, anything that would otherwise be repeated in every system prompt by hand. It involves no model call and no extraction step; it is simply always present.
`FactExtractionMemoryBlock` runs an extraction prompt, overridable with a custom one, over the flushed conversation and appends the resulting facts to a running list. This is the block doing the closest thing to what this site calls writing memories: turning raw conversation into discrete, durable statements rather than keeping the conversation itself.
`VectorMemoryBlock` stores flushed message batches in a vector store and retrieves the most similar ones by embedding similarity at read time, configurable by how many batches to retrieve and how much surrounding context to include with each. This block keeps the original conversational text rather than an extracted fact, which suits recalling a specific past exchange rather than a durable preference.
Multiple blocks can run at once, so a single Memory instance might hold a static block for identity, a fact-extraction block for stated preferences, and a vector block for searchable history, merged together whenever memory is retrieved. Combining sources like that raises an obvious question the framework’s own examples answer directly: what happens when combined memory sources return the same fact twice?
Composable memory
What happens when combined memory sources return the same fact twice?
It is deduplicated rather than injected twice. `SimpleComposableMemory` lets a primary memory source pull in messages from one or more secondary sources at read time, and the framework’s own example demonstrates a specific, checkable behaviour worth knowing before combining sources.
If a message already exists in the primary memory and a `get()` call against a secondary source retrieves that same message again, the redundant copy is not added to the assembled system message. The documented example seeds a fact into every configured source and shows the resulting prompt contains it once, not once per source. This matters because it means composing several memory blocks is not automatically a way to inflate a prompt with repeats; the module actively guards against exactly that failure.
Composable memory is also how a primary memory can be reset independently of the secondary sources feeding it, which is useful when one part of an agent’s memory, such as session-scoped working state, should clear between runs while durable long-term blocks persist. The `reset()` call supports resetting only the primary source or every source at once.
All of this assumes a fairly standard chat-loop agent. Production systems are frequently not that simple, and wiring memory into something else is a real question: how do you wire LlamaIndex memory into an agent that is not a simple chat loop?
Integration
How do you wire LlamaIndex memory into an agent that is not a simple chat loop?
By retrieving before the model call and storing after it, as an explicit step in your own request handler, rather than relying on an implicit chat-session wrapper. Hosted or event-driven agent runtimes, such as an agent deployed behind a serverless entrypoint, do not always give you the same conversational session object a local chat script would.
The pattern that generalises is straightforward: on each incoming request, retrieve relevant memory and use it to enrich the prompt before the model is called, then store the resulting conversation turn back into memory after the response is generated. That separates memory from any particular runtime’s request lifecycle, since retrieval and storage become two explicit calls your own code controls rather than something bound to a specific session abstraction.
Memory can also be exposed to the agent as a callable tool rather than something invoked automatically around every turn, which gives the model itself the choice of when to search memory versus when to answer from the immediate context. That tradeoff, between memory the agent always consults and memory the agent decides to consult, mirrors the broader retrieval question covered on memory retrieval.
Sharing memory across more than one agent in the same system is a related but distinct problem, since it introduces the question of which agent owns a write and how conflicting writes from different agents get resolved; that is covered in full on multi-agent memory rather than here.
Wiring is solvable with the patterns above. The more important question for anyone choosing this module is what it cannot do no matter how it is wired, and being clear about that upfront: what does LlamaIndex Memory not do?
Architectural limits
What does LlamaIndex Memory not do?
Resolve entities across turns, reason about when a fact stopped being true, or run outside a LlamaIndex agent. These are architectural facts about the module rather than criticisms borrowed from a rival’s marketing, and they are worth stating on their own terms because they follow directly from how the module is built.
There is no entity resolution pipeline. If a conversation refers to “Alice,” “my manager,” and “the person who approved the budget,” and never states outright that these are the same person, the module has no mechanism linking them into one identity. Each mention is stored as it appeared, and any connection between them has to already be explicit in the text or be built separately.
There is no temporal reasoning. A fact extracted in January and a contradicting fact extracted later both exist as separate entries with no built-in notion of which currently applies, unlike the supersession mechanics described on conflicting memories. Beyond fact extraction, long-term retrieval is vector similarity over stored message batches, which finds what is semantically close to a query rather than what is current.
The module is also coupled to the framework. It is built to plug into a LlamaIndex agent’s request and response cycle, and using it from a different agent framework is not a supported path without building the bridge yourself. Adopting the framework purely to get its memory module is unlikely to be worth it if memory is the actual priority.
None of these are failures of the module relative to what it is designed to be: a flexible, code-level memory component for agents already built on LlamaIndex. They are the boundary of that design, and knowing them changes what you should expect it to solve on its own: is LlamaIndex Memory enough on its own?
The decision
Is LlamaIndex Memory enough on its own?
For an agent already built on LlamaIndex that needs conversational recall and a handful of stated facts, yes. For anything needing entity linking, temporal correctness, or a framework-independent memory layer, no, and that is a real gap rather than a configuration problem.
The module does the job it is built for well: buffered short-term history with a documented, configurable flush policy, three composable long-term blocks covering static facts, extracted facts, and searchable history, and a deduplication behaviour that keeps combined sources from repeating themselves. For teams already committed to the framework, building this by hand would mean reproducing what is already there.
The gap shows up specifically when a team needs the memory itself to survive independently of the agent framework, when facts need to resolve to the same entity across sessions, or when a stored fact needs to be marked stale once something newer contradicts it. Those are exactly the responsibilities a managed memory layer such as Engram takes on, running underneath a hybrid vector store rather than inside a single framework’s request cycle, which is worth weighing directly against building the extraction, entity resolution and supersession logic yourself on top of what this module already gives you.
The broader landscape of memory tools and how they differ from framework-level modules like this one is on memory tools, and the deeper mechanics behind writing, scoring and retiring memories, applicable whichever tool ends up holding them, are on writing memories and memory scoring.
FAQ
Frequently asked questions
Configuration questions that come up once the basic setup is running.
Should I still use ChatMemoryBuffer in new LlamaIndex agent code?
No. It is deprecated and the Memory class is intended to replace it, with more flexible configuration for short-term and long-term memory. New code should configure Memory directly rather than relying on the older default.
Can I combine a fact-extraction block with a vector memory block?
Yes. A single Memory instance can hold multiple long-term memory blocks at once, so a fact-extraction block for durable preferences and a vector block for searchable conversation history can run side by side, both fed by the same short-term flush.
Does LlamaIndex Memory deduplicate facts automatically?
Only across composed memory sources at read time, not within a single block over time. SimpleComposableMemory avoids injecting the same message twice when two sources both return it, but a fact-extraction block does not on its own detect that a new extracted fact contradicts or restates an older one.
How do I make LlamaIndex memory work with a hosted agent runtime?
Retrieve relevant memory explicitly before the model call and store the resulting turn explicitly after it, rather than relying on an implicit chat-session object. This makes memory two calls your own request handler controls, independent of whatever session lifecycle the hosting runtime provides.
Does LlamaIndex Memory know when a stored fact is out of date?
No. There is no built-in temporal reasoning, so a fact extracted earlier and a contradicting one extracted later both exist as separate entries with nothing marking which currently applies. Detecting and resolving that is a separate mechanism, covered on conflicting memories.
Is it worth adopting LlamaIndex just to get its memory module?
Rarely. The memory classes are built to plug into a LlamaIndex agent's request cycle, so using them from a different framework means building that integration yourself. If memory is the primary requirement rather than the LlamaIndex agent framework itself, a framework-independent memory layer is usually the better starting point.