Developers · The stack
What Is an AI Memory Layer?
An AI memory layer is the infrastructure component between an agent and its storage that decides what to remember, what to retrieve, and what to discard. It is not a database. The database stores and searches; the memory layer holds the policy that makes storing and searching useful. This page covers what it does, where it sits, and how to evaluate one.
What the layer owns
Definition
What is an AI memory layer?
An AI memory layer is an external component that lets agents and applications store, update and retrieve user context across sessions, holding the extraction, scoping, ranking and maintenance policy that a raw store does not have. The agent calls it to write a memory and calls it again to retrieve what matters before answering.
The name became common because the alternative descriptions were worse. Calling it “the database” hides the fact that most of the work is judgement rather than storage. Calling it “memory” alone is ambiguous, since the context window is memory too. “Layer” captures the position: it sits between something that reasons and something that persists.
Its scope is narrow and specific. A memory layer does not reason, generate or plan, and it does not own the conversation. It answers two questions on repeat: what from this exchange is worth keeping, and what from history is worth putting in front of the model now.
That definition raises the obvious objection, which is why an application cannot just keep the conversation and pass it back: why an AI application needs a memory layer at all.
Motivation
Why does an AI application need a memory layer?
For three reasons: models are stateless, resending history costs money and latency, and continuity is the product feature users actually notice.
- Models start from nothing. Every call begins with an empty slate, so anything not in the prompt does not exist. Without an external store, each session is the first session. See stateless versus stateful LLMs.
- Passing the whole history is expensive. Selective retrieval is the cheaper path by a wide margin: Mem0 reports roughly 1,800 tokens per query against 26,000 for a full-context baseline, with p95 latency of 1.44 seconds against 17.1 seconds (Chhikara et al., 2025). That is the economic argument, and for most teams it is the decisive one.
- Continuity changes what the product is. An assistant that recalls preferences, prior tasks and past decisions behaves like a system that knows the user, rather than a very good autocomplete that meets them each time.
None of that requires a separate service in principle. It requires the policy, and a memory layer is where that policy lives rather than being scattered across an application: where the memory layer sits in the stack.
Position
Where does the memory layer sit in the stack?
Below the agent and above the storage: the application sends a turn, the agent decides to write or read memory, the memory layer performs the extraction and search, and a vector, graph or key-value store holds the result.
Two consequences follow from that position. The layer is the only component that sees both the raw conversation and the stored history, which is why extraction and deduplication belong to it rather than to the application. And it is the natural place for scope enforcement, because it is the last point where a user identifier can be attached to a memory before storage, and the first point where it can be applied as a search filter.
Its position also explains the most common architectural mistake in this category: treating a vector database as the memory layer.
Distinction
How is a memory layer different from a vector database?
A vector database answers one question, what is similar to this embedding. A memory layer decides what deserves to be stored at all, whose it is, which of several similar memories is current, and what should be removed. The layer usually runs on the database, which is why the two are so often conflated.
The practical test is what happens over time. A system built directly on a vector database usually works well for a month and degrades after that, because nothing deduplicates on write, nothing invalidates a fact when it changes, and nothing removes what is never retrieved. Those are not database features and they were never going to appear on their own.
Deciding to build the layer yourself is legitimate, and it means owning those four policies. The storage question is covered on vector databases for AI memory, and the comparison against graphs on vector versus knowledge graph memory.
Whether bought or built, the layer performs a defined set of jobs: what a memory layer actually does.
Responsibilities
What does a memory layer actually do?
Five jobs: extraction, scoping, retrieval and ranking, consolidation, and eviction. A product that does fewer than five leaves the remainder to you, which is a fair trade as long as it is a known one.
- Extraction. Turn a turn into durable facts and discard the rest. See writing memories.
- Scoping. Attach a user or tenant identity at write time and apply it as a filter during search, not after it. The controls are covered on AI memory security and access control.
- Retrieval and ranking. Search, score by recency, importance and relevance, and inject a limited number of memories. See scoring and ranking memories.
- Consolidation. Merge duplicates and promote session memories to durable ones. See memory consolidation.
- Eviction. Remove or invalidate what has gone stale. See forgetting and eviction.
The layer also owns the memory types the agent can hold, from semantic facts to episodic events, described in the types of AI agent memory.
Before evaluating products it helps to know what a stored memory physically looks like: how AI memory is stored.
Storage
How is AI memory stored?
A stored memory is a short piece of text, its embedding, and the metadata that makes it findable and governable: whose it is, when it was written, what type it is, and whether it is still valid. The text is what reaches the model; everything else exists so the right text can be found.
Three storage shapes cover almost every system, and many run more than one.
- Vector storage. The memory’s embedding is indexed for similarity search, which is what allows a question phrased differently from the stored fact to still retrieve it. See vector databases for AI memory and embeddings.
- Graph storage. Memories become entities and relationships, which makes multi-hop questions answerable by traversal and makes temporal validity expressible. See knowledge graphs for AI memory.
- Key-value storage. Fast exact lookup for the current session’s working state, usually in front of the durable store rather than replacing it. See memory storage backends.
The metadata is where governance lives, and it is routinely under-designed. A user or tenant identifier makes filtering and deletion possible. A timestamp makes recency weighting and expiry possible. A type marker separates a durable preference from a transient event so they can be retained on different schedules. A validity flag lets a superseded fact be kept for history without competing with the current one.
Retrieval usually combines vector similarity with keyword matching, since pure semantic search misses exact identifiers and pure keyword search misses paraphrases. That combination is covered on hybrid search for memory retrieval.
With the storage shape clear, the five jobs above become a checklist to evaluate a product against: how to evaluate a memory layer.
Selection
How do you evaluate a memory layer?
Ask which of the five jobs it performs, how it scopes data, what happens when a fact changes, and what it costs per query. Benchmark scores come last, because they are measured on somebody else’s conversations.
Five questions separate the options quickly. Does it extract, or does it store whatever you send? Is scoping per user built in or your responsibility? When a stored fact changes, does the old one get invalidated, overwritten or left to compete? Can you inspect and delete an individual memory, which matters the first time a user asks? And what does a retrieval cost in tokens and milliseconds on your data?
The candidates are Engram (Weaviate), Mem0, Zep, Letta and LangMem, compared in full with their limitations on the best AI memory tools, with the hosting question on open source versus managed memory.
Which raises the decision behind the selection: whether to build the layer or buy one.
The decision
Should you build or buy a memory layer?
Buy while memory is not what makes your product different, and build when retrieval behaviour is the differentiator or when data residency forbids anything else. The question is not difficulty, since a first version is genuinely easy to write. It is the maintenance nobody scopes.
What building actually commits you to. The first version is an embed, an insert and a search, which takes an afternoon. What follows is the work: an extraction prompt that has to be tuned and re-tuned, deduplication that has to catch paraphrases, scoping enforced at query time, a ranking function, a consolidation job, and an eviction policy. Each is small; together they are a component with an owner.
What buying costs you. A managed layer sets the extraction policy, and its judgement about what is worth keeping may not match your domain. Your memories live in somebody else’s system, which is a procurement question before it is a technical one. And the retrieval behaviour is theirs to change.
A reasonable middle path exists and is common: buy for the proof of concept, instrument it, and only build once you can point at a specific retrieval behaviour that the product will not give you. By then you know what to build, which is not true at the start. The hosting dimension is covered on open source versus managed memory, and the DIY route on Redis for agent memory and memory storage backends.
One question comes up from outside engineering almost every time this subject is raised: whether ChatGPT has a persistent memory layer.
Consumer products
Does ChatGPT have a persistent memory layer?
Consumer assistants including ChatGPT do keep memories between conversations, and they work the way described on this page: a filtered set of facts is extracted from conversations, stored outside the model, and retrieved into later sessions. They are not storing your transcripts and re-reading them.
Two things follow that surprise people. The assistant remembers less than users assume, because extraction discards most of a conversation by design. And what it does remember is editable, since a memory is a record rather than a change to the model, which is why these products can show a list of what they know and let you delete an entry.
For builders the useful observation is that the consumer feature and the infrastructure component are the same architecture at different scales. The behaviour users describe as an assistant “learning” them is extraction plus retrieval, discussed from the user’s side on how AI memory works.
Which leaves the question people ask when they hear that memory is a separate layer at all: whether memory is still the bottleneck.
Outlook
Is memory still a bottleneck for AI agents?
Yes, though the bottleneck moved. It is no longer that models cannot hold enough text, since context windows now run to hundreds of thousands of tokens. It is that holding text is not the same as knowing which of it matters, and that judgement is what remains unsolved.
The evidence is direct. Liu et al. (arXiv:2307.03172) showed that models degrade on information buried in the middle of a long context even when it fits, which is why a larger window does not remove the need for retrieval. And the published benchmarks still show memory systems trading a little accuracy for a large cost saving rather than beating a full context outright, which is an honest description of an unfinished problem.
What is genuinely improving is maintenance: temporal invalidation, self-organising memory structures, and consolidation moved to background compute. Those are covered on sleep-time compute, conflicting memories and long context versus memory.
The pattern across all of it is that the hard part has moved. Storing text and searching it are solved problems with several good implementations. Deciding what deserves to be stored, whose it is, which version is current, and what should be dropped is not solved, and it is exactly what the layer exists to hold.
For a team building today, the practical consequence is that the memory layer is a component to choose deliberately rather than a solved dependency to install, and the build itself is set out in how to add memory to an AI agent.
FAQ
Frequently asked questions
The questions that follow a memory layer decision: build versus buy, where the data lives, and how it interacts with RAG.
Is a memory layer required for AI-native apps?
For any multi-session agent, yes — without it the app forgets everything between sessions. Single-shot tools may skip LTM.
Build vs buy a memory layer?
Buy (Engram, Mem0, Zep) for speed and eval-backed pipelines. Build (Redis + pgvector DIY) for full control. See open-source vs managed.
Engram vs Mem0 as a memory layer?
Engram is Weaviate-native with hybrid search and async pipelines. Mem0 is framework-agnostic API (LOCOMO J 66.9). Choose based on existing infra.
Is Weaviate required for Engram?
Yes — Engram runs on Weaviate (Cloud or self-hosted). See Engram explained.
Can Zep serve as a memory layer?
Yes — temporal knowledge graph layer with bi-temporal invalidation. LongMemEval +18.5% vs baseline (Rasmussen et al., 2025).
Can Letta serve as a memory layer?
Yes — virtual-context paging layer (core + archival). MemGPT DMR 93.4% (Packer et al., 2023). Different architecture class than vector APIs.
Open-source memory layers?
Self-hosted Weaviate+Engram, Mem0 OSS, Letta OSS, Graphiti (Zep), LangMem. See open-source vs managed.
How does MemMachine compare to Engram?
MemMachine is another memory-layer approach — compare on architecture fit, LOCOMO scores and Weaviate integration. Engram is vector-native on Weaviate; evaluate both for your stack.
How do you benchmark memory layers?
LOCOMO for in-session recall; LongMemEval for cross-session. Multiple frameworks publish scores — see memory metrics.
Memory layer in the AI-native stack diagram?
App API → Agent orchestration → Memory layer → Embeddings + vector/graph stores. See AI-native tech stack.