Tools · Supermemory
Supermemory: A Memory API and Router for AI Agents
Supermemory is a memory platform with an unusual integration shape: alongside a conventional API it offers a router you point your model calls through, so an existing application can gain memory by changing a base URL rather than by being restructured. That is the fastest adoption path in this category, and it comes with a control trade-off worth understanding before you take it.
At a glance
The product
What is Supermemory?
A hosted memory layer that stores what an application learns about its users and retrieves it when they return, reachable either through a conventional API or by routing your existing model calls through it. It also ingests documents and connected sources, which places it closer to a combined memory and retrieval product than to memory alone.
The company was founded by Dhravya Shah and raised a $2.6 million seed round led by Susa Ventures, Browder Capital and SF1.vc, with individual investors including Google AI chief Jeff Dean, according to TechCrunch in October 2025. The code is developed in public: the `supermemoryai/supermemory` repository on GitHub is actively maintained, with over 1,900 commits and changes landing through September 2026.
What it does is the same loop every memory layer implements: extract durable statements from interactions, store them scoped to a user, retrieve the relevant ones and place them in the prompt. The mechanics are the ones described on how AI memory works, and the write path on how agents write and store memories.
What distinguishes it is not the loop but the attachment point, which is worth taking on its own: how the memory router works.
The distinctive part
How does the Supermemory router work?
You keep your existing model client and change its base URL to point at Supermemory, which prefixes the provider endpoint. Calls pass through the memory layer on the way to the model, so retrieval and injection happen in transit rather than in your code.
Concretely, an Anthropic client configured with a base URL of https://api.supermemory.ai/v3/https://api.anthropic.com/v1 continues to work as before while gaining memory behind it. Nothing else in the application changes, and the same pattern applies to other providers.
This solves a real adoption problem. Adding memory the conventional way means finding every place the prompt is composed, inserting a retrieval step, and adding a write path afterwards, which in a mature codebase is a project rather than a task. The router turns it into a configuration change.
It also means the memory layer decides what enters the prompt. That is the trade-off, taken up in full below, and it is the reason to understand the pattern before adopting it rather than after. The wider decision is on what a memory layer is and whether to build or buy one.
Whichever attachment you use, the store behind it accepts more than conversations: what Supermemory can ingest.
Inputs
What can Supermemory ingest?
Conversations, files and documents including PDFs, and data from connected applications such as mail. The company describes the ingestion surface as covering files, docs, chats, projects, emails and app data streams, which is broader than a pure conversation-memory product.
Breadth of this kind blurs a line worth keeping clear. Documents a user uploads are retrieval material and behave like a corpus; statements extracted from conversations are memory and behave differently, particularly around updating and expiry. A product that accepts both should still store and rank them separately, for the reasons on memory versus RAG.
Connected sources raise the second question, which is scope. Mail and app streams are continuous, high-volume and full of material nobody intended to be remembered, so a selection rule matters more here than in a store fed only by chat. What to keep and what to drop is covered on writing memories.
Whatever the source, the retrieved result is organised around a person: how user profiles work.
Structure
How does Supermemory structure what it knows about a user?
As a user profile combining static facts that rarely change with dynamic ones that do. The split matters more than it sounds, because the two want different retrieval and different expiry.
Static facts are the durable attributes: a name, a role, a language, a long-standing preference. They are small enough to load unconditionally at the start of a session, and they are what makes an agent open already knowing who it is talking to rather than discovering it three turns in.
Dynamic facts are current state: what the user is working on now, what they asked about last week, what changed recently. These need recency weighting and eventual expiry, since a project someone was busy with in March is not what they are doing today.
A profile shape of this kind is a useful default, and it is worth knowing what it gives up. It suits semantic memory about a person well and episodic memory less well, because a sequence of events with outcomes is not naturally a profile field. Where the ordered record matters, as in support histories, check how the product represents it rather than assuming, and see episodic versus semantic memory for what the distinction costs.
Where all of this runs is the next question, and it is usually the deciding one for anything touching customer data: whether Supermemory can be self-hosted.
Deployment
Can Supermemory be self-hosted?
Yes, on the higher plans. The company states that self-hosted deployments on bare metal or Kubernetes are available on its Scale and Enterprise tiers, with fully air-gapped deployment offered on Enterprise.
This matters more for memory than for most infrastructure, because a memory store holds statements about identifiable people, derived rather than submitted. Where the data is subject to residency requirements or where the material is a company’s most sensitive asset, as source code is, the deployment question decides the shortlist before any feature comparison starts.
Two details are worth confirming for your own case rather than taking from any summary: which components run inside your perimeter under a self-hosted plan, and whether the router in particular can be deployed there, since a proxy that must remain hosted keeps prompts flowing through a third party regardless of where the store sits.
The hosted path has the opposite profile: a free API key is available from the console, so evaluation costs nothing, and the trade-off is the usual one set out on open source versus managed memory. The controls to check either way are on memory security.
With the deployment shape clear, the comparison people usually want is with the rest of the field: how Supermemory compares with Engram, Mem0 and Zep.
Comparison
How does Supermemory compare with Engram, Mem0 and Zep?
They differ less in what they store than in how they attach and how much they decide for you. That is the axis worth comparing on, because the storage loop is broadly the same in all four.
Engram is the fit where a Weaviate stack already exists, and its hybrid search is the practical differentiator: order numbers, error codes and symbol names are lexical strings that pure vector similarity retrieves unreliably. See Engram.
Supermemory is the fastest to attach, particularly to an application already in production, and the broadest on ingestion. Mem0 is the framework-agnostic SDK route where you want the write and read paths explicit in your own code, on Mem0 and its alternatives.
Zep is the choice when facts have histories you must be able to reconstruct, since its bi-temporal graph makes validity native rather than something you maintain in metadata, on Zep and its alternatives.
The full ranking, with what each one decides on your behalf, is on the best AI memory tools. Before choosing the router path in particular, one section is worth reading first: what the router pattern costs.
The trade-off
What are the trade-offs of the router approach?
You buy adoption speed with control, and you place a third party on the critical path of every model call. Both are consequences of the same design decision, which makes this a trade-off to size rather than a drawback to avoid.
Control is the substantive one. What gets retrieved, how much of it enters the prompt and where it is placed are the decisions that shape how an agent behaves, and under a router most of them belong to the router. For a general assistant that is often fine. For a product with a specific idea of what should be remembered, it is the part you most wanted to own, as argued on the developer hub.
Position is operational. Every call traverses one more service, adding latency and a dependency, and prompts pass through infrastructure you do not run. Both are addressable, by self-hosting and by measurement, and both should be established before the pattern is load-bearing.
Portability is the quiet one. A base URL change is easy to make and easy to reverse, but the memories accumulated behind it are not automatically yours to move, so the question of how to export the store is worth asking on day one rather than on the day you want to leave.
None of which decides the matter on its own: when Supermemory is the right choice.
Fit
When should you choose Supermemory?
When you have an application in production that needs memory without a rewrite, or when documents and connected sources belong in the same store as conversations. Those two cases are where it is clearly the strongest option.
The retrofit case is the strongest. A mature codebase where prompts are composed in a dozen places is expensive to instrument by hand, and a base URL change that delivers most of the behaviour is worth a great deal. It is also a reasonable way to establish whether memory helps your product at all before committing engineering to it.
The mixed-source case is the second. If users bring documents as well as conversations and you would otherwise be running a retrieval pipeline alongside a memory store, one product covering both is a genuine simplification, provided the two are still ranked separately.
Choose differently in three situations. When the selection rule is your product, because what your agent remembers is a differentiator rather than a feature. When you already run Weaviate, where Engram keeps memory in the stack you have. And when facts have histories you must reconstruct, where a temporal graph is doing work no general store will do for you.
Whichever way it goes, evaluate it on your own traces rather than on any published comparison, using the method on how to evaluate agent memory, and start from the build order on the memory guides.
FAQ
Frequently asked questions
The practical questions about adopting it: cost, portability and how it sits beside an existing stack.
Is Supermemory open source?
The code is developed in public in the supermemoryai/supermemory repository on GitHub, which is actively maintained. Check the repository's own licence file for the terms that apply to your use, since a public repository and a permissive licence are not the same thing.
Can you try Supermemory for free?
Yes. A free API key is available from the console, which makes an evaluation against your own traces cheap to run before committing any engineering to the integration.
Does the router work with any model provider?
The pattern is a base URL prefix in front of a provider endpoint, so it applies to providers the platform supports rather than to any endpoint universally. Confirm your provider is covered before planning around it.
How do you get your memories out of Supermemory?
Ask before you adopt rather than after. A base URL change is trivial to reverse, but the memories accumulated behind it are the asset, and export is the thing that determines whether the decision is actually reversible.
Is Supermemory a replacement for a vector database?
No. It is a memory layer, which decides what to store and what to retrieve; a vector database is one of the places such a layer can keep things. See what a memory layer is.
Should documents and conversation memories go in the same store?
They can share a product but should not share a ranking. Documents behave like a corpus and memories behave like per-user facts with expiry, so keeping them separately ranked is what stops one crowding out the other. See RAG with memory.