Compare · Updated September 2026

Best AI Agent Memory Tools in 2026, Compared and Ranked

The memory tools worth shortlisting in 2026 are Engram (Weaviate), Mem0, Zep, Letta (MemGPT), LangMem and Cognee, and the right one is decided by the workload rather than by a leaderboard. This page compares them on architecture, memory operations, deployment and ecosystem fit, sets out what the published benchmarks actually report, and names the source beside every number.

9
Tools compared
4
Architecture classes
3
Published benchmarks

Decided by the job, not the ranking

1
Per-user facts
Memory API
2
Changing facts
Temporal graph
3
Long chats
Paging
4
Existing stack
Framework-native

Definition

What is an AI agent memory tool?

An AI agent memory tool is software that lets an agent store, retrieve, update and forget information across turns and sessions, holding that information outside the model’s weights. It is the component that turns a stateless model call into an agent that knows who it is talking to.

Three things are commonly mistaken for it, and none of them is the same thing.

  • A context window is not memory. It is working memory that vanishes when the session ends, and it has a hard token limit. The difference is set out in AI memory versus the context window.
  • RAG is not memory. Retrieval-augmented generation reads from a fixed corpus that the user did not write, while memory is built from the interaction itself. See AI memory versus RAG.
  • Fine-tuning is not memory. It bakes knowledge into weights and cannot be updated at conversation speed. See memory versus fine-tuning.

What a memory tool actually supplies is a pipeline: extraction, embedding, a storage backend, retrieval, consolidation and eviction, mapping onto the types of agent memory. Those parts are assembled differently by different vendors, and the assembly is what you are really choosing between.

Before comparing them, it is worth asking whether you need one at all: when an agent actually needs a memory tool.

Prior question

Do you actually need an AI memory tool?

Not always. If your agent completes its work inside one session and never needs to recognise a returning user, the context window is already sufficient and a memory tool adds cost, latency and a failure mode for nothing. The honest test is whether anything must survive the end of a conversation.

Three signals say you do need one.

  • The same person comes back. Any product where a user returns and expects to be known has a memory requirement, because re-establishing context manually is the experience people complain about.
  • One session outgrows the window. A conversation or document set larger than the context limit forces a decision about what to drop, and that decision is exactly what a memory system makes for you.
  • Facts accumulate and change. Preferences, entitlements and states that update over time cannot live in a prompt template, and storing them naively produces contradictions.

Two signals say you do not. Single-turn tools such as classifiers and extractors have nothing worth remembering, and workflows where every relevant fact arrives in the request are better served by passing it in than by storing it. The wider argument sits on why AI agents need memory.

Where the answer is yes, the field divides cleanly before it divides by vendor: the four classes of memory tool.

Taxonomy

What are the four classes of AI memory tool?

Every tool on this page belongs to one of four classes, separated by who decides what reaches the prompt: a memory API, a temporal graph, a virtual paging system, or a framework’s own memory primitives. Storage engine is not the dividing line, because several classes sit on the same vector database.

  1. Memory API. The application calls a service to write and to retrieve, and the service handles extraction and ranking. Predictable and easy to reason about, less adaptive. Engram (Weaviate), Mem0 and Supermemory sit here.
  2. Temporal knowledge graph. Facts are stored with their validity over time, so an updated fact invalidates its predecessor instead of competing with it. Zep sits here.
  3. Virtual paging. The agent itself schedules tiers, paging content in and out with tool calls. Maximum control, and the most ways to spend tokens. Letta (MemGPT) sits here, explained in full on what virtual context is.
  4. Framework-native. Memory primitives inside an agent framework you already run, so there is no extra service to operate. LangMem and Cognee sit here.
Four classes of agent memory tool: memory API, temporal graph, virtual paging and framework-native, each with who controls retrieval.
Figure 1. The classes are separated by control rather than by storage engine, which is why two tools using the same vector database can behave completely differently.

Knowing the class narrows the field to two or three candidates, and the rest of the decision is workload-shaped: which memory tool you should actually use.

Decision

Which AI memory tool should you use?

Choose by the job in front of you: a memory API for per-user facts, a temporal graph for facts that change, paging for very long single conversations, and your framework’s own primitives when you already run one. Almost every shortlist that goes wrong went wrong by starting from a ranking instead of a workload.

Four selection rules matching a memory tool class to the job: per-user facts, changing facts, long conversations, or an existing framework.
Figure 2. Four workloads, four answers. Where two of them describe your product, prefer the tool that handles the one you cannot work around.
  • You need per-user facts and preferences. Engram (Weaviate) if your stack already runs Weaviate, because the memory layer is native to the same database rather than bolted beside it. Mem0 if it does not.
  • Your facts change. Job titles, addresses, plan tiers and preferences all expire. A temporal graph such as Zep records when a fact was true, which is a different capability from storing it accurately once. The failure mode this avoids is covered in handling conflicting and stale memories.
  • One conversation runs for a very long time. Paging is built for exactly this, and Letta is its production form.
  • You already run LangGraph or a graph-heavy pipeline. Use LangMem or Cognee before adding a service, because the operational cost of one more moving part usually outweighs the feature difference.

A shortlist is easier to defend once the tools are side by side on the same columns: how the tools compare on architecture and fit.

Side by side

How do the AI memory tools compare?

The table below sets nine tools against the same five columns: architecture, memory operations, retrieval latency where it is published, deployment model, and the workload each one suits. Latency is left blank rather than estimated wherever a vendor does not publish it.

AI agent memory tools compared on architecture, operations, deployment and fit. Updated September 2026.
ToolArchitectureMemory opsLatencyOSS or managedBest for
Engram (Weaviate)Vector-native memory layerAsync extract, transform, commitNot publishedManagedWeaviate-native stacks
Mem0Vector, optional graphFull CRUD plus consolidate0.148 s median searchBothPer-user personalisation
ZepTemporal knowledge graph (Graphiti)Full CRUD plus temporal invalidationNot publishedBothFacts that change over time
Letta (MemGPT)Virtual context pagingPage in and out, tiered storeNot publishedBothLong multi-session conversation
LangMemLangGraph store integrationExtract, store, retrieveNot publishedOpen sourceTeams already on LangGraph
CogneeEvolving knowledge graphGraph write and queryNot publishedOpen sourceKnowledge-heavy agents
SupermemoryVector API plus memory graphAPI CRUD, MCP serverNot publishedManagedMCP and IDE workflows
Redis agent memoryKey-value plus vectorsYou build the logicNot publishedOpen sourceTeams that want to own the pipeline
LlamaIndex memoryFramework memory blocksBuffer plus retrievalNot publishedOpen sourceLlamaIndex pipelines
Six AI agent memory tools compared by architecture, deployment model and the workload each suits.
Figure 3. The six tools that come up most often, reduced to the three columns that usually decide a shortlist.

A table compresses each tool to a row. The differences that decide a build are in the detail: what each tool actually does.

Tool profiles

What does each memory tool actually do?

Nine profiles, each covering the architecture, the workload it fits, and the limitation worth knowing before you commit. Every limitation below is a property of the design or of what the vendor publishes, not a complaint.

Engram (Weaviate)

Engram is Weaviate’s memory layer, which means the memory pipeline and the vector database are the same system rather than two services you join together. Writes run as an asynchronous extract, transform and commit pipeline, so the cost of turning a conversation into memories does not sit inside the user-facing request. Retrieval inherits Weaviate’s hybrid search, combining vector similarity with keyword matching.

It suits teams already running Weaviate, where adopting a separate memory service would mean operating a second store and reconciling two views of the same data. The limitation to weigh: it is managed rather than self-hosted, and it does not publish a retrieval latency figure, so latency-sensitive designs need their own measurement. Full detail on what Engram is.

Mem0

Mem0 is a memory API with an optional graph mode, offering full create, read, update and delete operations plus consolidation. It is the most measured tool in the category: the Mem0 paper (Chhikara et al., 2025) reports a median search latency of 0.148 seconds, a p95 total latency of 1.44 seconds, and roughly 1,800 tokens per query against 26,000 for a full-context baseline.

It suits per-user personalisation, where each user accumulates preferences and facts that must survive between sessions. The limitation to weigh: its own LOCOMO score of 66.9, or 68.4 with graph memory, sits below the 72.9 full-context baseline, so the argument for it is cost and latency rather than raw accuracy. Alternatives are compared on the Mem0 alternatives page.

Zep

Zep stores memory as a temporal knowledge graph through its Graphiti engine, which records when each fact was true rather than only that it is true. When a fact changes, the previous version is invalidated rather than deleted, so the agent can still answer questions about what used to be the case. Rasmussen et al. (2025) report a LongMemEval accuracy improvement of up to 18.5% over a full-context baseline with latency reduced by around 90%, and 94.8% on the DMR benchmark.

It suits domains where facts expire: employment, plans, addresses, entitlements. The limitation to weigh: a graph is more machinery than a straightforward store, and if your facts rarely change you are paying for a capability you will not exercise. See the Zep alternatives page.

Letta (MemGPT)

Letta is the production framework built from the MemGPT paper, and it is the only tool here where the agent schedules its own memory. It pages content between main context, recall storage and archival storage using tool calls, which is what lets a single conversation outrun the context window. On the benchmark its own team introduced it scores 93.4% with GPT-4 Turbo, against 35.3% for the same model with no memory.

It suits long-running single conversations, companions and assistants that must stay coherent across months. The limitation to weigh: agent-controlled retrieval only fetches what the agent thinks to ask for, and every lookup costs an extra inference call. The mechanism is explained on the virtual context page.

LangMem

LangMem provides memory primitives inside LangGraph, combining a checkpointer for conversation state with a store for longer-lived memories. Because it lives in the framework, there is no additional service to deploy, secure or pay for, and memory state sits alongside the rest of the graph’s state.

It suits teams already committed to LangGraph or LangChain. The limitation to weigh: that same coupling is the cost, since the memory layer is not portable to a stack built on anything else. See what LangMem does.

Cognee

Cognee builds an evolving knowledge graph from the documents and interactions you feed it, combining graph structure with vector embeddings so that multi-hop questions can be answered by traversal rather than by similarity alone. It is open source, and it is positioned for local and self-hosted deployment, including running against local models.

It suits knowledge-heavy agents where the relationships between facts matter as much as the facts. The limitation to weigh: graph construction is a heavier write path than a simple embed and store, so ingestion cost rises with corpus size. See what Cognee does.

Supermemory

Supermemory is a managed memory API with a graph layer and an MCP server, which makes it straightforward to attach memory to IDE and assistant workflows that already speak the Model Context Protocol. Writes and reads happen through a small API surface rather than through a framework.

It suits MCP-based and developer-tool workflows. The limitation to weigh: it is managed only, and as with most tools here no retrieval latency figure is published. See what Supermemory does.

Redis agent memory

Redis is not a memory product, it is the fast tier a memory product is often built on, offering key-value storage plus vector search with very low read latency. Choosing it means implementing extraction, deduplication, scoring and eviction yourself.

It suits teams who want to own the pipeline, usually because their retrieval logic is a differentiator or their data cannot leave their infrastructure. The limitation to weigh: every operation in the loop becomes your code to write and maintain, which is a real ongoing cost rather than a one-off build. See Redis for agent memory.

LlamaIndex memory

LlamaIndex provides memory blocks inside its own pipeline framework, combining a conversation buffer with retrieval over an index you already maintain. For teams whose retrieval layer is already LlamaIndex, memory becomes a configuration of something running rather than a new dependency.

It suits existing LlamaIndex pipelines. The limitation to weigh: as with LangMem, the memory layer follows the framework, so it is not the right starting point if the framework choice is still open. See LlamaIndex memory.

Nine profiles built on the same five criteria invite an obvious question about the criteria themselves: how these tools were ranked.

Methodology

How were these tools ranked?

Five fixed criteria, applied identically to every tool, and no criterion that cannot be checked against something published. They are the same five columns that appear in the table above.

  1. Architecture. Vector-native layer, temporal graph, virtual-context paging, or framework primitive.
  2. Memory operations. Write, retrieve, update, delete and consolidate, and how much of that pipeline runs asynchronously.
  3. Latency. Retrieval round-trip, recorded only where the vendor or a paper publishes it.
  4. Deployment. Open source, managed, or both, because this decides who carries the operational load.
  5. Ecosystem fit. Whether the tool is native to a database or framework you already run.

Two things are deliberately not criteria. There is no score for popularity, and there is no composite number, because a single ranking figure hides exactly the workload differences that decide the choice. Where a benchmark exists it is reported as its authors reported it, which is worth reading carefully: what the published benchmarks actually report.

Evidence

What do the memory benchmarks actually report?

Three published results cover most of the claims made in this category, and two of the three were reported by the authors of the tool being measured. That does not make them wrong, but it does make them claims with citations rather than an independent league table.

Published benchmark results for agent memory tools, each with the paper that reported it.
Figure 4. Every number here is traceable to a named paper. No vendor publishes a result for all three benchmarks.

Deep Memory Retrieval, from the MemGPT paper. Packer et al. (arXiv:2310.08560, Table 2) report 93.4% for MemGPT on GPT-4 Turbo, against 35.3% for the same model with no memory, and 92.5% on GPT-4 against 32.1%. The comparison is against a baseline that received a lossy summary of prior sessions, so it measures retrieval against summarisation rather than memory against nothing. The task is explained on the virtual context page.

The same benchmark, reported later by Zep. Rasmussen, Paliychuk, Beauvais, Ryan and Chalef (arXiv:2501.13956) report 94.8% against MemGPT’s 93.4%. The margin is 1.4 points on a benchmark the other team designed, which is the context that makes it readable.

LOCOMO, reported by Mem0. Chhikara et al. (2025) report an LLM-as-a-judge score of 66.9 for the extract-and-store pipeline and 68.4 with graph memory, against 72.9 for a full-context baseline that costs roughly fifteen times as many tokens. Read plainly, that is a small accuracy sacrifice for a large cost saving, which is the actual argument for memory in production. The benchmark itself is covered in what LOCOMO measures, and the wider metric set in which metrics matter for agent memory.

Benchmarks describe the state of the field at the moment they were run, and this field moved quickly: what changed in agent memory since 2025.

The landscape

What changed in agent memory since 2025?

Four shifts separate the 2026 tools from the 2025 ones: temporal validity became a first-class feature, self-editing memory became an expectation rather than a differentiator, benchmarks moved past simple recall, and context engineering became the thing teams actually buy. Graphlit’s 2026 survey of agent memory frameworks makes the same four observations from a vendor’s vantage point.

  • Chat history stopped being enough. Replaying a transcript is cheap to build and gets expensive quickly, which is what the LOCOMO token-cost figure above quantifies.
  • Temporal memory became central. The question moved from “is this fact stored” to “was this fact true when it was stored, and is it true now”.
  • Self-editing memory became table stakes. Tools are now expected to consolidate and invalidate on their own rather than accumulating every fact forever, which is the subject of memory consolidation.
  • Benchmarks moved past simple recall. Multi-hop and long-horizon evaluation replaced single-fact lookups, which is why LongMemEval exists alongside LOCOMO.

One consequence for buyers is that a tool chosen on a 2025 comparison may have been chosen on criteria the category has since abandoned. A second is that the deployment question got sharper: whether to choose open source or managed.

Deployment

Should you choose open-source or managed memory?

Choose managed when memory is not the thing you are differentiating on, and open source when data residency, cost at scale, or the ability to change retrieval behaviour matters more than the operational saving. Both classes contain credible tools, so this is a constraints question rather than a quality question.

Managed services carry the extraction pipeline, the store and the retrieval tuning for you, which is most of the work. Engram (Weaviate) and Supermemory are managed; Mem0, Zep and Letta offer both. Self-hosting matters most where data cannot leave the environment at all, and the open-source tools in this category are genuinely deployable rather than demo-ware.

The full argument, including the cost model, sits on open source versus managed memory.

One more routing question comes up constantly in search, and it deserves a direct answer: where to look when you want alternatives to one specific tool.

Routing

Where do you find alternatives to a specific memory tool?

If you have already picked a tool and are looking for a replacement, the comparison you want is tool-specific rather than a general ranking, and each one has its own page. A general list answers “what exists”; an alternatives page answers “what replaces this, and what will I lose”.

The video below walks the same comparison out loud, covering Cognee, Mem0, Zep and Letta.

A last observation about this category, offered because it is not in any vendor’s interest to say it. The differences between these tools matter far less than whether you implement deduplication, scoping, consolidation and eviction at all. A well-maintained store on the simplest of these tools will outperform a neglected store on the most sophisticated one, because the failure modes that degrade memory systems are policy failures rather than product limitations. Choose for fit, then spend the saved deliberation on the maintenance nobody sells you.

That is also why this page carries limitations beside every tool rather than only strengths: the limitation is the part that predicts how a choice ages.

Whichever tool the shortlist ends on, the build problem is the same one, and it is set out step by step in how to add memory to an AI agent.

FAQ

Frequently asked questions

These cover the questions that come up after a shortlist exists: pricing shape, whether a memory tool is needed at all, and how these tools sit against a plain vector database.

What is the best AI memory tool in 2026?

It depends on your stack. Engram is the best fit for Weaviate-native stacks; Mem0 for fast per-user personalization APIs; Zep for temporal knowledge-graph memory; Letta for unbounded conversation history; LangMem for LangGraph teams. Check the comparison table and use-case section above.

Is Engram the best AI memory tool?

Engram is the best choice when your vector infrastructure runs on Weaviate — it provides a native memory layer without a separate memory vendor. For framework-agnostic APIs, Mem0 is a strong option; for temporal facts, Zep fits better.

Engram vs Mem0 — which should I choose?

Choose Engram when your stack already uses Weaviate — one vendor, one database, native memory semantics. Choose Mem0 when you want a standalone managed memory API that works across any vector backend. See Engram explained.

Mem0 vs Zep — which is better?

Mem0 excels at per-user semantic memory with a simple API and strong adoption. Zep (Graphiti) excels when facts change over time and you need temporal knowledge-graph queries with conflict resolution. Choose Mem0 for personalization POCs; choose Zep when time-aware relationships matter.

Letta vs MemGPT — are they the same?

Yes. Letta is the product name for the MemGPT research lineage — virtual context / memory paging that treats the context window like RAM and pages memories in and out. When comparing tools, 'Letta' and 'MemGPT' refer to the same paging architecture.

Do I still need a memory tool with a large context window?

Yes, for most production agents. Large context windows still suffer context rot, higher token cost, and no cross-session persistence. Memory externalizes durable facts and history so you inject only what matters each turn. See long context vs memory.

Are memory tools the same as RAG?

No. RAG retrieves from a fixed knowledge base (documents). Agent memory is dynamic, personal, and updated across sessions — preferences, past interactions, learned facts. They combine well: RAG for org knowledge, memory for user/session state. See AI memory vs RAG.

What is the best AI memory tool for LangChain?

LangMem is the native choice for LangGraph/LangChain stacks. For framework-agnostic APIs, consider Mem0 or Zep. If your infra is Weaviate, Engram integrates natively.