Compare · Cluster hub

AI Agent Memory Comparisons: Which One Do You Need?

There is no single best agent memory system, so the useful question is which comparison answers your actual decision: choosing a first tool, replacing one you already use, or deciding between two architectures. This hub routes to the right one and sets out the criteria that make any of them meaningful.

Three decisions

1
Choosing
First tool
2
Replacing
Alternatives
3
Architecting
Approach

The router

Which AI memory comparison do you need?

Three different decisions get called “comparing memory tools”, and they need different pages: choosing a first tool, replacing a specific one, or deciding between architectural approaches. Starting on the wrong one is why so many evaluations end without a conclusion.

If you are choosing for the first time, you want the full field on fixed criteria, which is the best AI memory tools. That page compares nine tools on architecture, memory operations, deployment and fit, with the limitation attached to each.

If you already use something and want to replace it, a general ranking is the wrong shape. You need to know what the closest substitute is and what you would lose, which is what the alternatives pages answer: Mem0 alternatives, Zep alternatives and Letta and MemGPT alternatives.

If the decision is architectural rather than about a product, four comparisons cover it: long context versus memory, open source versus managed, vector database versus knowledge graph, and agentic RAG versus agent memory.

A fourth kind of comparison is not on this site by design. There is no page ranking memory tools by popularity, because popularity is not a property that predicts fit, and a ranking that hides workload differences is the specific failure this cluster exists to avoid.

Before any of those, the question underneath them all deserves a direct answer: whether there is a single best memory system.

The honest answer

Is there a single best AI agent memory system?

No, and the reason is structural rather than diplomatic: these tools specialise in different jobs, so a ranking that puts them in one order is comparing systems built for different problems.

Four classes of agent memory tool: memory API, temporal graph, virtual paging and framework-native, each with who controls retrieval.
Figure 1. Comparing a temporal graph against a paging framework on a single score compares two different jobs.

The specialisations are real and fairly clean. A memory API is built for per-user facts and gives predictable behaviour. A temporal knowledge graph is built for facts that expire, and answers questions about the past that nothing else answers natively. Virtual paging is built for one conversation that outruns the window. Framework-native memory is built to avoid operating another service.

Put a product against the wrong specialisation and it looks weak for reasons that have nothing to do with quality. A paging framework evaluated on per-user fact recall will look like an expensive way to do something simpler; a memory API evaluated on multi-hop temporal questions will look like it cannot reason.

What can be compared meaningfully is fit against a stated workload, plus the published evidence read with its conditions attached. Both are covered on the best AI memory tools, and the classes on the memory frameworks hub.

Which makes the criteria the load-bearing part of any comparison: how memory tools should be compared.

The method

How should you compare AI memory tools?

On five properties that can be checked against something published: architecture, memory operations, retrieval latency, deployment model, and ecosystem fit. Anything that cannot be checked is an impression rather than a criterion.

Six AI agent memory tools compared by architecture, deployment model and the workload each suits.
Figure 2. The same columns applied to every tool. A comparison where the columns change per product is a set of adverts.

Architecture establishes the class and therefore what the tool is for. Memory operations asks which of write, retrieve, update, consolidate and evict the product actually performs, because whatever it omits becomes your work. Latency is recorded only where published, and left blank rather than estimated everywhere else.

Deployment decides who carries the operational load and whether data can leave your environment at all. Ecosystem fit asks whether the tool is native to a database or framework you already run, which frequently outweighs a feature difference, because one fewer service to operate is worth more than a marginal capability.

Two things are deliberately excluded. There is no popularity score, and there is no composite number, because a single figure hides exactly the workload differences that decide the choice. The methodology in full is on the best AI memory tools.

Criteria settle what to compare. Evidence settles how much any published number is worth: what the benchmarks let you compare.

The evidence

What do the published benchmarks let you compare?

Less than they appear to, because the three benchmarks in common use test different things and most results are reported by the team whose tool is being measured. They are still worth reading, with the conditions attached.

Published benchmark results for agent memory tools, each with the paper that reported it.
Figure 3. No vendor publishes a result for all three benchmarks, so a complete league table does not exist.

On Deep Memory Retrieval, Packer et al. (arXiv:2310.08560) report 93.4% for MemGPT on GPT-4 Turbo against 35.3% for the same model with no memory, and Rasmussen et al. (arXiv:2501.13956) report 94.8% for Zep on that same benchmark. A gap of 1.4 points, on a benchmark designed by one of the two teams, is not a basis for a decision.

On LOCOMO, Chhikara et al. (2025) report an LLM-as-a-judge score of 66.9 for Mem0’s extract-and-store pipeline, and 68.4 with graph memory, against 72.9 for a full-context baseline costing roughly fifteen times the tokens. Secondary write-ups quote figures a few tenths different from the paper’s, which reflects different configurations rather than disagreement, and is a reason to prefer the primary source when a number matters.

The pattern across all of it is the useful finding. Memory systems generally score slightly below simply sending everything, and win decisively on cost and latency: roughly 1,800 tokens per query instead of 26,000, with p95 latency of 1.44 seconds against 17.1 seconds. A vendor quoting only the accuracy figure is telling you the less interesting half. The benchmarks themselves are covered on LOCOMO and LongMemEval.

Benchmarks compare products. The more consequential comparisons are usually architectural: comparing approaches rather than tools.

Architectural choices

Which architectural comparisons actually matter?

Three, and each one is decided before any product is chosen: whether to use memory at all rather than a longer context, whether to self-host, and whether your facts need a graph. Getting these right narrows the product decision to two or three candidates.

Four selection rules matching a memory tool class to the job: per-user facts, changing facts, long conversations, or an existing framework.
Figure 4. The architectural decision comes first. The product decision is what remains after it.

Long context or memory. A larger window raises a ceiling and adds no persistence, so the question is whether anything must survive the session. Where nothing must, memory is cost without benefit. Where something must, no window size substitutes for it. See long context versus memory.

Open source or managed. This is a constraints question rather than a quality one. Managed removes an operational component; self-hosting is required where data cannot leave the environment and is preferable where retrieval behaviour is your differentiator. See open source versus managed memory.

Vector or knowledge graph. A graph earns its considerably higher write cost only when answers span several entities, when facts expire, or when history must stay queryable. Otherwise it is machinery you are not using. See vector versus knowledge graph memory.

Answering those three usually eliminates most of the field before a single product page is opened, which is the fastest route through this cluster.

The other common decision is not about choosing but about leaving: finding an alternative to a tool you already use.

Replacement

How do you find an alternative to a memory tool you already use?

By naming what specifically is not working, because the honest alternative to a tool depends entirely on which of its properties you are trying to escape. A general ranking cannot answer that, which is why each tool has its own alternatives page.

Three reasons for switching come up repeatedly, and they point in different directions. Cost or hosting pushes toward a self-hostable option in the same class, which is usually the smallest possible change. A capability gap pushes across classes, most often toward a temporal graph because facts started expiring. Operational burden pushes the other way, from something self-built toward a managed layer.

The question worth asking before any migration is what you would lose. Memory stores are not trivially portable: the extraction policy differs, the metadata schema differs, and embeddings produced by one model cannot be reused with another. A switch usually means re-extracting from source conversations where they still exist, or accepting a store that starts thinner than the one you left.

The per-tool pages set out the closest substitutes and the trade in each direction: Mem0 alternatives, Zep alternatives, Letta and MemGPT alternatives, and the client-side view on Engram.

One category of comparison is worth avoiding entirely, and naming it saves time: the comparisons that are not worth making.

Anti-patterns

Which memory comparisons are not worth making?

Four, and they consume a disproportionate share of evaluation time because they look rigorous.

  1. Ranking tools across classes on one number. A composite score across a paging framework, a memory API and a temporal graph produces an order that is arithmetically valid and practically meaningless.
  2. Comparing benchmark results reported by different teams. Different configurations, different underlying models and different scoring harnesses make the numbers non-comparable even when the benchmark name matches.
  3. Comparing on storage engine. Several of these tools run on the same vector database. What differs is the policy layer above it, which is what actually determines behaviour, as set out on the memory layer.
  4. Comparing before deciding what must persist. An evaluation without a written list of what has to survive the session has no criteria, and will conclude with whichever tool demoed best.

The pattern connecting all four is comparing artefacts instead of fit. The productive alternative is short: write down what must persist, decide the three architectural questions above, shortlist the two or three tools in the resulting class, and test them on your own queries rather than on published ones.

That last step is the one most often skipped and the one that actually settles it, and the method is on how to add memory to an AI agent.

The way to replace all of it with evidence is short and rarely done: running your own comparison.

The method

How do you run your own memory comparison?

Build a question set from your own traffic, run the shortlisted tools against it, and measure retrieval and answer quality separately. Two days of this settles a decision that weeks of reading vendor pages will not.

Build the set from real conversations. Sample fifty exchanges where the answer depends on something said earlier, and write down which stored fact each one requires. This is the artefact the whole evaluation rests on, and it is reusable: the same set becomes your regression test once a tool is chosen.

Load the same memories into each candidate. Where the tools have different extraction policies, that difference is part of what you are measuring, so feed them the same source conversations rather than the same pre-extracted facts. What each one decides to keep is a genuine differentiator and it never appears in a feature table.

Measure three things separately. Whether the required memory was retrieved, whether the answer used it, and what the round trip cost in tokens and milliseconds. Conflating them produces a single quality number that cannot tell you whether to fix extraction, ranking or the prompt.

Then run it again a month later, with a store that has accumulated real usage. This is the step nobody does and the one that catches the failure that matters most: tools that perform well on a small clean store and degrade as it fills, which is the difference between a demo and a product. The evaluation mechanics are on how to add memory to an AI agent and the metric definitions on memory metrics.

Running it well means measuring the things published comparisons leave out: what most memory comparisons omit.

The gaps

What do most memory comparisons leave out?

Four properties that decide how a choice ages, none of which appear in a feature table and all of which are discovered late.

Maintenance behaviour. Does the tool consolidate and evict on its own, or does the store grow until you build that yourself? This single property predicts whether retrieval quality holds after a year, and it is almost never compared, because it does not demo. It is covered on memory consolidation.

Scoping and deletion. Can you filter by user during the search rather than after it, list everything held about one person, and delete it completely including derived summaries? These are compliance requirements for anything holding personal data, and retrofitting them means rewriting every record.

Portability. If you leave, what comes with you? Embeddings are tied to the model that produced them, metadata schemas differ, and extraction policies are not transferable. A tool that makes export easy is worth more than a marginal benchmark advantage, and none of them advertise how hard leaving is.

Failure behaviour. What happens when the memory service is slow or unavailable? A system that blocks the response is a very different product from one that degrades to answering without memory, and that behaviour is a design decision you inherit rather than choose.

Asking those four questions of a shortlist usually separates it faster than any benchmark does, because they are the properties that differ most between products that otherwise look similar.

They also age at different speeds, which is worth knowing before trusting any comparison including this one: how quickly these comparisons go out of date.

Shelf life

How quickly do memory tool comparisons go out of date?

Faster than the underlying decisions do, which is the useful asymmetry: product details change every few months while the architectural questions have been stable for two years.

The parts that expire quickly are exactly the parts comparisons lead with. Version numbers, pricing, which tool has an integration with which framework, and benchmark scores all move, and a result published against one model generation may not reproduce against the next because the baseline itself changed.

The parts that hold are the classes and the criteria. The distinction between a memory API, a temporal graph, a paging framework and framework-native memory has been stable since these categories emerged, and it still predicts fit. So do the five criteria on this page, because they describe properties of the problem rather than of any product.

The practical consequence for reading any comparison, this one included, is to trust the framing and verify the specifics. If a page tells you a tool belongs to a class and that class suits your workload, that guidance survives. If it quotes a price or a benchmark, check it against the vendor or the paper before it decides anything, and prefer the primary source when a number matters.

That is also why the pages in this cluster carry the source beside every figure and the limitation beside every tool. A comparison you can re-verify ages into a useful document; one you cannot ages into a liability. The full field is on the best AI memory tools.

Underneath every comparison here is the same mechanism, which is worth understanding before choosing between implementations of it: what all these tools are implementations of.

The common ground

What are all these tools implementations of?

The same four operations: write what matters, store it outside the model, retrieve the relevant parts before answering, and forget or update what has gone stale. Every product in this cluster implements that loop, and the differences are in how much of it they own.

Knowing this changes how a comparison reads. When a product describes a distinctive feature, the useful question is which operation it belongs to and what it replaces. Temporal invalidation is a claim about the update operation. Automatic extraction is a claim about write. Agent-controlled paging is a claim about who triggers retrieval.

It also explains why the operations a product omits matter more than the ones it advertises. Whatever it does not do becomes your code, and the operations most often omitted, consolidation and eviction, are precisely the ones whose absence degrades a store slowly enough that nobody attributes the decline to the memory system.

The loop is walked in how AI memory works, the mechanisms in the architecture cluster, and the storage beneath them in the infrastructure cluster.

With the mechanism understood, the comparisons in this cluster become quick to read rather than a research project, which is the point of arranging them this way.

FAQ

Frequently asked questions

The questions that follow: how often these comparisons change, and whether tools are interchangeable.

What is the best AI memory tool in 2026?

See our flagship ranking at best AI memory tools — Engram (Weaviate), Mem0, Zep, Letta, LangMem and more compared on LOCOMO, LongMemEval, architecture and fit.

Mem0 vs Zep — which should I choose?

Mem0 for per-user personalization with a simple API. Zep for temporal knowledge-graph memory when facts change over time. See Mem0 alternatives and Zep alternatives.

Letta vs MemGPT — same thing?

Yes — Letta implements the MemGPT virtual-context paging architecture. See Letta alternatives.

How do I pick a memory tool by architecture class?

Match use case to class: vector API for personalization, temporal KG for changing facts, virtual paging for long conversations, LangGraph-native for LangChain stacks. See the architecture comparison table.

Open source or managed memory?

Open-source gives control and compliance flexibility; managed gives faster time-to-ship. Many tools offer both. See open-source vs managed.

Vector database or knowledge graph for memory?

Vectors for semantic similarity and personalization. Graphs for relationships, temporal facts and conflict resolution. See vector vs knowledge graph.

Do I need memory with a 1M token context window?

Yes for most production agents — context rot, cost and no cross-session persistence remain problems. See long context vs memory.

How often are tool rankings updated?

When new benchmark results publish or major framework releases ship. Rankings cite LOCOMO and LongMemEval with a last-updated date on the flagship page.

Best memory tool for LangChain?

LangMem for native LangGraph integration. Engram, Mem0 or Zep for framework-agnostic APIs. See LangMem and best tools.

Best memory tool for customer support bots?

Episodic + semantic memory for ticket history and user preferences. Engram or Mem0 are common starting points; Zep when facts change over time. See customer support use case.