Memory types · Comparison
Episodic vs Semantic Memory in AI Agents
Semantic memory stores what is true, episodic memory stores what happened, and the difference decides how each one is written, retrieved and retired. An agent that stores both in one undifferentiated pile retrieves badly, because a fact and an event should be ranked by different signals and kept for different lengths of time.
Two questions
The distinction
What is the difference between episodic and semantic memory?
Semantic memory holds facts stripped of when they were learned; episodic memory holds events with their time and order intact. “The user is vegetarian” is semantic. “The user asked about billing in March and the issue was unresolved” is episodic, and the date is part of the content rather than metadata attached to it.
The distinction comes from human cognitive psychology, where semantic memory is knowing that Paris is the capital of France and episodic memory is remembering the trip you took there. Agent memory borrows the split because it turns out to predict something useful: the two are retrieved for different reasons.
A semantic memory surfaces because the query is about its subject. An episodic memory surfaces because the query is about a period, a sequence or a prior interaction. A question about dinner should retrieve a dietary preference; a question about what was decided last week should retrieve a session, and a ranking function tuned for the first will handle the second badly.
Both are long-term memory, meaning both live outside the model and are reached by search, as described on long-term memory for AI agents.
The consequences show up first in how each one is written: how the two are stored differently.
Storage
How are episodic and semantic memories stored differently?
A semantic memory is a short standalone statement; an episodic memory is a record with a timestamp, participants and an outcome. Storing an episode as if it were a fact strips exactly the parts that make it useful.
Semantic records should be short, atomic and deduplicated aggressively. “Prefers window seats” is one memory, and a second phrasing of the same preference is a duplicate to merge rather than a second fact. Their value is that they compress: one sentence stands in for many conversations.
Episodic records are longer and should not be deduplicated the same way. Two similar support interactions are two events, not one repeated fact, and merging them destroys the sequence that makes the history meaningful. What episodes need instead is summarisation as they age, keeping the outcome and dropping the transcript.
The write trigger differs too. A semantic memory is written when a user states something durable. An episodic memory is written at the end of an interaction, when the outcome is known, which is why episodic capture usually belongs to a background job rather than to the turn itself. Both are covered on how agents write and store memories.
Because they are stored differently, they also have to be found differently: how retrieval differs between them.
Retrieval
How does retrieval differ for episodic and semantic memory?
Semantic memory is retrieved by topic, episodic memory by time and sequence, and one ranking function serves them badly. The standard scoring of recency, importance and relevance weights those signals differently depending on which type is being searched.
For semantic memory, relevance should dominate and recency should barely matter. A dietary requirement stated eight months ago is exactly as true today, and a scoring function that decays it, such as the 0.995 per hour used in the Generative Agents work (Park et al., 2023), will eventually bury it beneath trivia mentioned yesterday.
For episodic memory, recency and order carry real information. The most recent interaction is usually the relevant one, and questions about episodes are frequently comparative: what changed since last time, what we tried before, whether this has happened previously.
The practical answer most systems land on is to search the two separately and merge the results, rather than running one query across an undifferentiated store. That requires a type marker on every record, which is a small field with a large payoff. See memory retrieval and scoring and ranking memories.
They also age at different rates, which changes what should be kept: how long to keep each type.
Retention
How long should each type be kept?
Semantic memories should be kept until superseded; episodic memories should be summarised as they age and eventually dropped. Applying one retention policy to both is what makes a store either forget things that are still true or accumulate events nobody will ever ask about.
A semantic fact has no natural expiry. It stops being true when something replaces it, which is a supersession event rather than a time-based one, and the correct handling is to mark the old version invalid rather than delete it, as covered on handling conflicting and stale memories.
An episode does have a natural decay. A support interaction from two years ago rarely helps and frequently misleads, and the useful residue is usually a one-line outcome rather than the exchange. Compressing episodes into summaries on a schedule keeps the history navigable while bounding the storage, which is the consolidation step described on memory consolidation.
One nuance is worth keeping. Episodes sometimes contain latent semantic facts: a user who mentioned a job change inside a support conversation has stated something durable inside an event. Promoting those facts out of episodes and into semantic memory, then letting the episode decay, is how a well-designed store keeps what matters while shrinking.
Which raises the question a build actually has to answer: which type your agent needs.
Selection
Which type does your agent need?
Almost every agent needs semantic memory. Episodic memory is needed when continuity across interactions matters, which is a narrower set of products than people assume.
Start with semantic. It is cheaper to build, easier to get right, and delivers the recognisable behaviour of an assistant that knows the user. Extraction is more reliable because a stated fact is unambiguous, and deduplication is straightforward.
Add episodic when the product has cases rather than conversations. A support tool where an issue spans several contacts needs it. A coding agent working across sessions on the same task needs it. A one-shot question-answering assistant does not, and building it produces a store of events nobody queries.
The clearest signal is whether users ask questions containing “last time”, “before”, “already” or “again”. Those are episodic queries, and if they do not appear in your traffic then episodic memory has no consumer. The method for checking is on why AI agents need memory.
Both types appear in the wider taxonomy alongside procedural memory on the types of AI agent memory, and the third one is worth knowing about because it is the most often missing: procedural memory in AI agents.
In practice the two are rarely used in isolation: how they work together.
Together
How do episodic and semantic memory work together?
Episodes are where semantic facts come from, and semantic facts are what make episodes interpretable. A system running both gets a compounding benefit that neither delivers alone.
The flow in one direction is extraction: an interaction happens, it is stored as an episode, and durable facts stated during it are promoted into semantic memory. That promotion step is what stops the semantic store from being limited to things the user explicitly announced.
The flow in the other direction is interpretation. An episode that says a request was escalated means little without the semantic fact that this customer is on an enterprise plan. Retrieving both together is what lets an agent answer why something happened rather than only that it did.
The promotion step deserves one more note, because it is where most of the compounding value sits and it is almost never built first. An agent that only stores what users explicitly announce will have a thin semantic store, since people state preferences far less often than they demonstrate them. Mining episodes for durable facts is what fills that gap, and it can run entirely in the background without touching the response path.
The design that supports this is unglamorous: one store, a type field on every record, separate retrieval paths, and a background job that promotes facts out of episodes and compresses the episodes afterwards. Everything else follows from those four decisions, and the implementation is on how to add memory to an AI agent.
A useful way to sanity-check a design is to ask what happens to each type after a year. Semantic memory should have grown slowly and stayed roughly current, because supersession keeps replacing rather than adding. Episodic memory should have grown quickly and then been compressed, leaving outcomes rather than transcripts. If both curves look the same, the system is almost certainly treating events as facts, and retrieval quality will already be suffering.
The wider loop these both sit inside is described on how AI memory works, and the storage beneath them on the infrastructure cluster.
FAQ
Frequently asked questions
The questions that follow: whether one store can hold both, and how procedural memory relates.
Can one store hold both episodic and semantic memories?
Yes, and most systems do. What matters is a type marker on every record so the two can be retrieved with different ranking weights and retained on different schedules. A single undifferentiated store is what produces poor retrieval, not a shared database.
Which should you build first?
Semantic memory. It is cheaper, more reliable to extract, and delivers the recognisable behaviour of an agent that knows the user. Add episodic memory when your traffic contains questions using words such as last time, before or already.
How does procedural memory relate to these two?
It is the third long-term subtype and answers a different question again: how to act, rather than what is true or what happened. It is written from corrections and outcomes rather than from statements, and it should be updated rather than expired. See procedural memory.
Do episodic memories need to be kept in full?
Rarely. The useful residue of most episodes is the outcome rather than the exchange, so compressing them into summaries as they age keeps history navigable while bounding storage. Keep the original where you can still retrieve it, since summaries lose detail by construction.
Is the human memory analogy accurate?
It is a useful borrowing rather than a claim about mechanism. The split predicts something real about how agent memories should be written and retrieved, which is why it survives, but agent memory is external storage and search rather than anything resembling human recall. See AI memory versus human memory.