Memory types · Flagship guide
What Is Long-Term Memory for AI Agents?
Long-term memory for AI agents is storage outside the model that persists facts, events and learned procedures across sessions, so an agent still knows something after the context window has been cleared. This guide covers the three subtypes, the five-stage pipeline that fills them, how retrieval selects from them, and what the published benchmarks report.
The long-term memory pipeline
Definition
What is long-term memory in AI agents?
Long-term memory is the store an agent keeps outside its context window, holding facts, past events and learned procedures that must survive the end of a session. It is implemented with a vector database, a knowledge graph, or a managed memory service, and it is reached by an explicit search rather than by being present in the prompt.
The defining property is persistence across a boundary. Anything the agent knows only because it is currently in the prompt is short-term memory, and it disappears when the session does. Anything the agent can retrieve tomorrow, in a new session, from a store it wrote to yesterday, is long-term memory.
This is not the model’s training. Long-term memory sits entirely outside the weights, which is why a fact can be added, corrected or deleted in seconds, and why the same base model can serve two users who know it as two different assistants.
The clearest way to see the boundary is to put the two stores side by side: long-term memory compared with short-term memory.
The boundary
How does long-term memory differ from short-term memory?
Short-term memory is the context window: fast, bounded by a token limit, read directly by the model, and gone at the end of the session. Long-term memory is external: slower to reach, effectively unbounded, and reachable only by a search the agent runs on purpose.
The relationship is one-directional and worth stating precisely: long-term memory is useless unless something retrieves from it into the window, because the model still only reads the prompt. Every retrieval decision is therefore a decision about which scarce short-term slots to spend. The comparison in full sits on short-term versus long-term memory.
Long-term memory is not one undifferentiated store either, and the split matters for how each part is queried: the three subtypes of long-term memory.
Taxonomy
What are the types of long-term memory?
Three: semantic memory for what is true, episodic memory for what happened, and procedural memory for how to do things. The split follows the CoALA framing used across most agent memory literature, and it is functional rather than decorative, because each subtype is written and queried differently.
- Semantic memory holds facts and preferences without reference to when they were learned, such as a stated language preference. See semantic memory in AI agents.
- Episodic memory holds the ordered record of interactions and outcomes, which is what lets an agent say what it tried last time. See episodic memory in AI agents and episodic versus semantic memory.
- Procedural memory holds learned workflows and tool-use patterns, closer to a policy than to a fact. See procedural memory in AI agents.
A fourth category is sometimes counted separately, the agent’s own working notes about itself, though most implementations store those as semantic memories with a scope marker rather than as a distinct store.
Whichever subtype a memory belongs to, it arrives through the same production path: the long-term memory pipeline.
Mechanism
How does long-term memory work?
Through a five-stage pipeline: extract the durable parts of a turn, embed them, store them, retrieve them when they are relevant, and consolidate the store so it does not decay into noise. Production systems run this pipeline rather than dumping transcripts into a database, because a transcript is expensive to search and mostly irrelevant.
Extraction decides what is worth keeping, turning “I moved to Berlin and I am vegetarian” into two clean memories and discarding the rest. Its quality is measurable: the Mem0 paper (Chhikara et al., 2025) reports a LOCOMO LLM-as-a-judge score of 66.9 for its extract-and-store pipeline, 68.4 with graph memory, against 72.9 for a full-context baseline costing roughly fifteen times the tokens.
Retrieval runs before each response and has to choose which memories are worth their slot in the window. The canonical scoring comes from the Generative Agents paper (Park et al., 2023), which ranks each memory on recency, importance and relevance, with recency decaying exponentially at 0.995 per hour since last access. Selective retrieval is also why memory is cheaper than resending history: on LOCOMO, Mem0 answers with a p95 total latency of 1.44 seconds against 17.1 seconds for a 26,000-token full-context baseline, using roughly 1,800 tokens per query instead of 26,000.
Consolidation merges duplicates and promotes session memories into durable ones, usually as a background job. Each stage has its own page: writing memories, memory retrieval and memory consolidation.
The pipeline has to put its output somewhere, and the storage choice constrains what queries are possible: where long-term memory is stored.
Storage
Where is long-term memory stored?
In a vector database, a knowledge graph, or both, with a fast key-value store often sitting in front as a warm tier. The choice is not about capacity, it is about which questions the store can answer.
- A vector database indexes embeddings and answers “what is similar to this”, which suits free-text facts and preferences. See vector databases for AI memory.
- A knowledge graph stores entities and relationships and answers “what is connected to this, and when was it true”, which is what makes temporal reasoning possible. See knowledge graphs for AI memory.
- A key-value store such as Redis answers “give me this exact thing now”, and is used as the warm tier rather than the system of record. See memory storage backends.
Many production systems run a hybrid, using vectors for recall and a graph for relationships, and combining vector similarity with keyword matching at query time. That combination is covered in hybrid search for memory retrieval.
One storage decision is easy to defer and expensive to retrofit: scope. Every memory needs to record whose it is, because a store that mixes users has no way to prevent one person’s fact reaching another person’s prompt. In practice that means a user or tenant identifier on every record and a filter applied at query time rather than after retrieval, since filtering after the search returns the wrong candidates and wastes the slots. Multi-tenant products need the same discipline one level up, at the organisation boundary. Getting scope right at write time also makes deletion possible later, which matters the moment a user asks for their data to be removed.
Storage decides what is possible; a framework decides how much of it you have to build: the frameworks that implement long-term memory.
Implementations
Which frameworks implement long-term memory?
Match the need to the pattern rather than starting from a product: a vector-native layer for an existing vector stack, a managed API for the fastest start, a temporal graph for facts that expire, and paging for a single very long conversation.
The trade-offs behind each row, including the limitation attached to every tool, are set out on the best AI memory tools, and the deployment question on open source versus managed memory.
Any of these choices can be defended, and only one thing decides between them honestly: what the long-term memory benchmarks report.
Evidence
What do the long-term memory benchmarks report?
Two benchmarks carry most of the published claims, LOCOMO for long conversations and LongMemEval for long-horizon recall, and the headline results are reported by the teams whose tools are being measured. That is worth holding in mind when the margins are small.
Zep reports a LongMemEval accuracy improvement of up to 18.5% over a full-context baseline, with response latency cut by around 90% (Rasmussen et al., 2025), and 94.8% on the Deep Memory Retrieval task against 93.4% for Letta (Packer et al., arXiv:2310.08560). Mem0’s LOCOMO position is the cost argument rather than the accuracy argument: slightly below a full-context baseline on the judge score, at a fraction of the tokens.
Read together, the results say something useful and unglamorous. Long-term memory rarely beats simply putting everything in the context window on raw accuracy; it wins on cost, latency and the fact that everything does not fit. The benchmarks themselves are explained on LOCOMO and LongMemEval.
Benchmarks compare memory against a long prompt. The other comparison teams need is against the two techniques memory is most often confused with: long-term memory against fine-tuning and RAG.
Distinctions
How does long-term memory compare with fine-tuning and RAG?
Long-term memory stores what this user did and said, and updates in seconds. RAG retrieves from a corpus the organisation maintains. Fine-tuning changes the model’s weights and cannot be updated mid-conversation. They solve different problems and are routinely combined.
A practical rule: if the knowledge would be the same for every user, it is probably RAG. If it exists because of this user, it is memory. If it is about how the model behaves rather than what it knows, it may be fine-tuning. The full comparisons sit on memory versus RAG, memory versus fine-tuning and how to combine RAG and memory.
Distinctions settled, the harder operational question is how much of all this to keep: how long an agent should keep a memory.
Retention
How long should an agent keep a memory?
Long enough to be useful and no longer, which in practice means setting a policy per memory type rather than one retention rule for the whole store. The default of keeping everything forever is the option that reliably degrades, because retrieval quality falls as the store fills with near-duplicates and superseded facts.
Four policies cover most systems, and they are usually combined.
- Time to live. A fixed expiry, which suits episodic memories about transient events. A support ticket from two years ago rarely helps and frequently misleads.
- Decay. Importance falls with age rather than dropping to zero, so an old memory can still win retrieval if it is important and relevant enough. This is what the 0.995 hourly decay in the Generative Agents scoring implements.
- Invalidation. A fact is superseded rather than deleted, so the agent can still answer questions about what used to be true. This is the temporal graph approach.
- Relevance pruning. Memories that are never retrieved are removed on a schedule, on the argument that a memory nothing ever selects is costing storage and diluting search.
Semantic memories usually deserve the longest retention, because a stated preference stays true until it changes. Episodic memories age fastest. Procedural memories should be updated rather than expired, since a workflow that is wrong should be corrected, not forgotten. The mechanisms are covered on forgetting and eviction.
Retention policy is also where a memory system quietly fails, which is worth looking at directly: what goes wrong with long-term memory.
Failure modes
What goes wrong with long-term memory?
Four failures account for most of the disappointment with agent memory in production, and none of them announces itself as an error. Each is a design decision that was skipped rather than a bug that was introduced.
- The store fills with duplicates. Without deduplication at write time, the same fact accumulates in slightly different wordings, and retrieval returns three versions with equal confidence. The agent then appears indecisive about something it was told once.
- Stale facts outrank current ones. A fact stored before it changed keeps being retrieved until something invalidates it. Where the old and new versions both sit in the store with no temporal marker, recency weighting is the only thing separating them, and it is not enough.
- Retrieval never finds what is there. The agent only retrieves what it thinks to search for, so a memory phrased differently from the query may as well not exist. This is the failure hybrid search and query rewriting exist to reduce.
- Everything is remembered and nothing is prioritised. When extraction is too permissive, small talk enters the store and competes with real preferences for the limited slots retrieval can spend.
The pattern connecting all four is that they degrade slowly. A memory system is usually best on the day it launches and quietly worse each month, which is why consolidation and eviction are not optional extras. Conflicts specifically are covered on handling conflicting and stale memories.
With the failure modes visible, the remaining work is the build itself, set out step by step in how to build long-term memory for an LLM agent.
FAQ
Frequently asked questions
The questions that follow once the pipeline is clear: how much to store, how long to keep it, and what happens when memories conflict.
What is long-term memory in AI?
Long-term memory is information an AI agent persists in external storage across sessions — facts, preferences and conversation history retrievable later. See short-term vs long-term.
What is the difference between short-term and long-term memory in AI agents?
Short-term memory lives in the context window and clears when the session ends. Long-term memory persists in vector DBs, graphs or KV stores across sessions. See short-term vs long-term.
How do LLMs store long-term memory?
LLMs don't store LTM in weights — agents extract facts after each turn, embed them, and write to external stores like vector databases or knowledge graphs. See writing memories.
What is the best long-term memory framework?
Depends on stack: Engram for Weaviate, Mem0 for managed API, Zep for temporal graphs, Letta for paging, LangMem for LangGraph. See best AI memory tools.
Does Claude have long-term memory?
Claude's built-in memory is product-specific. For custom agents, implement external LTM via Engram, Mem0, Zep or your own vector store. See persist conversation memory.
How do I add long-term memory to LangGraph?
Use LangMem for native checkpointer integration, or Mem0/Zep as external APIs. See LangMem and build long-term memory.
Is a vector database the same as long-term memory?
No — a vector DB is the storage backend. LTM is the full pipeline: extract, embed, store, retrieve, consolidate and forget. See vector databases for memory.
What is Weaviate Engram?
Engram is a vector-native memory layer on Weaviate — dynamic writes and retrieval on the same platform without a separate memory store. See Engram explained.
How do I test long-term memory quality?
Run LOCOMO for long-conversation recall and LongMemEval for cross-session recall on your domain queries. See evaluation hub.