Tools · Cluster hub
What Are the AI Agent Memory Frameworks?
Agent memory frameworks are the software layers that let a model store, update and retrieve information across sessions, and they divide into four classes by who decides what reaches the prompt. This hub covers each framework, what it is built for, and the limitation that comes with it.
Four classes
Definition
What is an AI agent memory framework?
A memory framework is the layer between an agent and its storage that decides what is worth keeping from a conversation, where it goes, and what comes back before the next answer. Without one, a model has no memory beyond its active token window, and every session starts from nothing.
The framework is not the database. A vector database answers one question, what is similar to this embedding, and a framework adds the judgement around it: extraction, deduplication, scoping per user, ranking, consolidation and eviction. Choosing a database and calling it memory leaves all of those decisions unmade, which is the argument set out on the memory layer.
What every framework in this cluster has in common is the loop: write, store, retrieve, forget. They differ in how much of that loop they own, how much they expose, and who holds the policy. That is why comparing them on a single score is misleading, and why the useful first question is which class a tool belongs to.
The mechanism they all implement is covered on how AI memory works. The division between them is the more practical starting point: the four classes of memory framework.
Taxonomy
What are the four classes of memory framework?
Memory API, temporal knowledge graph, virtual paging, and framework-native, separated by who decides what enters the context window. Storage engine is not the dividing line, because several classes commonly sit on the same vector database.
A memory API is called by the application: write this, retrieve what is relevant. The service handles extraction and ranking, behaviour is predictable, and adaptivity is limited to what the API exposes. Engram (Weaviate), Mem0 and Supermemory sit here, and it is the class most products should start with.
A temporal knowledge graph stores facts with the period during which they held, so a changed fact invalidates its predecessor rather than competing with it. Zep sits here. The capability is narrow and, where it applies, hard to replicate: it is the only class that answers questions about the past natively.
Virtual paging gives the agent tools to fetch and evict its own memories across tiers, which is the MemGPT design that became Letta. Maximum control, maximum ways to spend tokens, and retrieval quality that depends on the agent asking the right question.
Framework-native memory lives inside an agent framework you already run, so there is no extra service to operate. LangMem and Cognee sit here, and the trade is portability: the memory layer follows the framework.
With the classes clear, the individual tools are easier to place: every framework in this cluster.
The set
Which memory frameworks are worth knowing?
Nine cover the field in 2026, and each has a page in this cluster with its architecture, its fit and its limitation. The table places each one in a class so that like is compared with like.
- Engram (Weaviate), a vector-native memory layer where the pipeline and the database are the same system. Managed, with an asynchronous extract, transform and commit write path.
- Mem0, a memory API with an optional graph mode and the most published measurements in the category.
- Zep, a temporal knowledge graph via its Graphiti engine, built for facts that expire.
- Letta (MemGPT), the production form of virtual context paging, explained on virtual context and MemGPT.
- LangMem, memory primitives inside LangGraph, combining a checkpointer with a store.
- Cognee, an evolving knowledge graph built from documents and interactions, open source and self-hostable.
- Supermemory, a managed memory API with an MCP server, which suits IDE and assistant workflows.
- Redis agent memory, the fast tier rather than a memory product: you implement the loop yourself.
- LlamaIndex memory and LangChain memory, memory blocks inside pipeline frameworks already in use.
The full comparison with benchmarks and limitations is on the best AI memory tools.
The question people actually arrive with is blunter than a taxonomy: which memory system is best.
The decision
Which AI agent memory system is best?
There is no single best, and the honest answer is a mapping from workload to class: a memory API for per-user facts, a temporal graph for facts that expire, paging for one very long conversation, and framework-native memory when you already run a framework.
The reason a ranking cannot settle it is visible in the published benchmarks. Letta reports 93.4% on Deep Memory Retrieval against 35.3% for the same model with no memory (Packer et al., arXiv:2310.08560), Zep reports 94.8% on that same benchmark (Rasmussen et al., arXiv:2501.13956), and Mem0 reports a LOCOMO judge score of 66.9 against 72.9 for a full-context baseline costing roughly fifteen times the tokens (Chhikara et al., 2025). Those are three different tests, and two of the three were reported by the team whose tool was measured.
What the numbers do establish is the category’s real value proposition. Memory systems trade a small amount of accuracy against sending everything, in exchange for large reductions in cost and latency, and for working at all once a conversation exceeds the window. Mem0’s figures put that at roughly 1,800 tokens per query instead of 26,000, with p95 latency of 1.44 seconds against 17.1 seconds.
A practical shortcut for a first build: start with a managed memory API, instrument retrieval, and only move to another class when you can point at a specific behaviour the API will not give you. By then you know what you need, which is not true at the start.
Frameworks manage memory, and how they do it is the same loop in every case: how agents manage memory.
The mechanism
How do AI agents manage memory?
Through four operations that every framework implements: write what matters, store it outside the model, retrieve the relevant parts before answering, and forget or update what has gone stale. The framework decides when each runs, which is the part that differs.
Two of the four are what a framework is usually bought for, and the other two are what it is usually judged on six months later. Write and retrieve produce visible behaviour immediately. Consolidation and eviction determine whether the store still works after a year of accumulation, and a framework that does not perform them leaves that job to you, described on memory consolidation.
The orchestration question underneath is which component triggers each operation. A memory API leaves it to the application. A paging framework hands it to the agent. A framework-native implementation ties it to the graph’s execution. That choice is covered on memory management and orchestration.
One question is worth settling before adopting any of them: whether to use a framework at all.
The prior question
Do you need a memory framework, or can you build it?
A first version is genuinely easy to build, and what follows is the part that justifies a framework. An embed, an insert and a search take an afternoon. Extraction that keeps the right facts, deduplication that catches paraphrases, scoping enforced at query time, ranking, consolidation and eviction are each small and together are a component with an owner.
Building makes sense in two situations. When retrieval behaviour is your differentiator, a framework’s opinions become constraints you fight. When data residency forbids a managed service, the choice is made for you, and the open-source options in this cluster are genuinely deployable rather than demo-ware.
Buying makes sense in most others, and particularly early. A managed layer sets the extraction policy, which may not match your domain, and holds your memories in someone else’s system, which is a procurement question as much as a technical one. Against that it removes an entire component from your operational surface while you are still learning what your agent needs to remember, which is usually the more valuable thing at that stage.
A reasonable middle path is common: adopt a managed API for the proof of concept, instrument retrieval quality, and build only once you can name a specific behaviour the product will not give you. By then the requirement is concrete, which it is not at the start. The hosting dimension is covered on open source versus managed memory, and the do-it-yourself route on Redis for agent memory.
The types of memory these operations act on are covered on the types of AI agent memory, and the build itself on how to add memory to an AI agent.
FAQ
Frequently asked questions
The questions that follow: whether frameworks are interchangeable, and what a framework does not solve.
What is the best AI memory framework in 2026?
No single winner. Engram (Weaviate) when you want DB + memory unified; Mem0 for personalization APIs; Letta for virtual-context paging; Zep for temporal graphs; LangMem for LangGraph. See best AI memory tools.
Mem0 vs Zep — which framework?
Mem0 excels at per-user semantic memory with a simple API. Zep (Graphiti) excels when facts change over time and you need temporal knowledge-graph queries. See Mem0 alternatives and Zep alternatives.
Letta vs MemGPT — are they the same?
Yes. Letta is the product name for the MemGPT research lineage — virtual context paging that treats context like RAM. See virtual context and MemGPT.
Is LangMem the right choice for LangChain?
Yes for LangGraph/LangChain stacks — LangMem integrates with checkpointers and long-term memory stores natively. See LangMem explained.
What open-source memory frameworks exist?
Engram, Mem0 (OSS layer), LangMem, Cognee, Letta (OSS), Zep (OSS components), Redis DIY and LlamaIndex Memory. See open-source vs managed.
Can Redis be used as agent memory?
Yes as a fast buffer and vector store — but you implement memory logic yourself. See Redis for agent memory.
Cognee vs Zep — which knowledge graph?
Zep focuses on temporal graph memory with Graphiti for production agents. Cognee builds evolving knowledge graphs for fact-heavy domains. Both are graph-class; compare on your temporal and ops needs.
What is Supermemory?
A managed vector memory API with simple HTTP integration and MCP support. See Supermemory explained.
Should I benchmark before choosing a framework?
Yes. Run LOCOMO and LongMemEval subsets on your use case before committing. See evaluation hub.