Architecture · Forgetting
Forgetting and Eviction in AI Agent Memory
A memory store grows until something removes from it, and nothing will unless you decide what and when. The useful version of forgetting is rarely deletion. It is making a record unreachable in ordinary retrieval while keeping it, which is reversible, auditable, and the only version you can measure. This page covers how to choose the policy and how to tell whether it is working, which is a question almost nobody asks.
Three outcomes
The problem
Why does an agent memory system need to forget?
Because retrieval quality falls as the store grows, long before storage cost becomes an issue. Forgetting is usually presented as a capacity problem. It is a precision problem, and treating it as capacity leads to the wrong policy.
Consider what a store looks like after a year of use. A user has hundreds of records, most of them true when written and irrelevant now: a project that shipped, a preference stated in a context that no longer applies, a question they asked once. All of them are still semantically close to plausible queries, so all of them are still candidates, and every one takes a place that a current fact could have held.
The consequence is not slow search. Modern stores handle far more records than any single user will ever accumulate. The consequence is that the top few results returned to the model contain something out of date, and the model presents it as current, because nothing in the prompt says otherwise. That failure looks to a user like an agent making things up.
There is a second reason, cheaper to act on than any eviction policy: much of what needs forgetting should never have been written. A write path that stores every turn creates the problem that a forgetting policy then has to solve, and tightening extraction is more effective than any pruning rule, as set out on writing memories.
Assuming the store is already accumulating what it should, the question becomes which records to remove and on what basis: which eviction strategies exist, and what does each get wrong?
The options
Which eviction strategies exist, and what does each get wrong?
Four in common use, and each is wrong in a different direction, which is why production systems combine two. Age, use, importance and replacement. Picking one and applying it alone is the source of most forgetting complaints.
Time to live expires a record after a fixed age. It is cheap and predictable, and it is completely blind to value: a dietary restriction stated once expires on the same schedule as a remark about the weather. Used alone it produces the most common complaint about memory systems, which is an agent that has forgotten something the user considers permanent.
Least recently used removes what has not been read, which is better, because retrieval is evidence that a record was useful. Its blind spot is the fact that is rarely relevant and critical when it is. An allergy might not be retrieved for months and must not be the thing that gets dropped.
Salience keeps records by importance score, which is the closest to what anyone actually wants and is only as good as the score. If importance was assigned once at write time by a model that had seen a single turn, an eviction policy built on it is a guess with extra steps. How that score is produced and revised is on memory scoring.
Supersession retires a record when a newer one replaces it, which is the most correct of the four because it acts on evidence rather than on a proxy. Its blind spot is specific and is covered further down. The mechanics of detecting a replacement are on conflicting memories.
The pairing that works for most products is supersession plus age with a salience floor. Supersession does the accurate work whenever the user gives it something to act on, age catches the long tail of records that nobody will ever contradict, and the floor stops the age rule from reaching anything that matters. Usage-based eviction is the one to leave out unless retrieval volume per user is genuinely high, because below that it is measuring noise: a record read twice and a record read once are not meaningfully different signals.
Whichever combination you choose, the same question arrives first, and getting it wrong makes every other decision here irreversible: should you delete a memory or retire it?
The distinction
Should you delete a memory or retire it?
Retire it, in almost every case. Two very different operations get called forgetting, and conflating them is why so many memory systems have a policy nobody can tune and nobody can undo.
Retiring means the row stays and stops competing. It is excluded from ordinary retrieval by a predicate every read already carries, it can still be reached by an explicit lookup or an audit, and the decision can be undone if it turns out to have been wrong. In schema terms it is one column, which is the shape described on storage backends.
Deleting means the record is gone. That is the right operation for a narrow set of cases: a person asking for their data to be removed, a record that should never have been stored, a retention limit that requires destruction rather than concealment. Those cases are real and they are governed by obligations rather than by quality judgements, which is why they belong on memory security and privacy.
The practical argument for retiring is that it keeps the evidence. A deleted memory tells you nothing about whether deleting it was correct. A retired memory that would have been the best match for a later query is a measurable signal that the policy is too aggressive, and that signal is the entire basis of the testing approach further down this page.
Retiring on supersession is straightforward, because a newer fact justifies it. Retiring on age needs a number, and choosing that number is where most systems guess: how do you choose a time-to-live without guessing?
The number
How do you choose a time-to-live without guessing?
Derive it from how often your users come back. Every guide to memory management says to set a retention horizon and none says how to pick one, so the number is usually copied from an example and never revisited.
The reasoning is simple once stated. A memory exists so the agent knows something on the user’s next visit. If it expires before that visit, it never does its job, so the retention horizon has to be longer than the gap between sessions. That gap is measurable from data every product already collects.
Take a high percentile of the interval distribution rather than the median. A horizon that covers the typical gap but not the long one fails precisely for the user who returns after a month and most expects to be recognised, and that user is the one whose experience of memory matters most. A daily-use assistant and a quarterly-use planning tool land on very different numbers from the same reasoning, which is the point: there is no universal default.
The same logic applies to a decay constant rather than a hard expiry. The formulation this site uses elsewhere comes from published research on generative agents (Park et al., 2023), whose paper decays a memory’s recency score by 0.995 per hour, which halves it in roughly 138 hours, a little under six days. That is a sensible shape for a product people open most days and much too fast for one they open monthly, where every memory is effectively at zero by the time they return. The full treatment of decay inside the ranking function is on memory scoring.
Whatever number you land on, put a salience floor under it, so that records above an importance threshold never expire on age alone. Age should remove the forgettable, not the important.
Age handles records that stopped mattering. A harder case is the record that stopped being true and never announced it: what happens to a fact that goes stale without being contradicted?
The blind spot
What happens to a fact that goes stale without being contradicted?
It stays, and it stays confident. This is the failure mode of relying on supersession alone, and it is the reason decay is a complement to it rather than an alternative.
Supersession retires a record when a newer statement contradicts it. That works whenever the user says the contradicting thing. It does nothing at all when they simply stop mentioning the old one. Someone who moved city and never announced it, whose project ended, whose child grew out of a stage, whose job title changed: the store holds a fact that was true, has no successor, and will keep being retrieved with full confidence.
Research on human memory has treated forgetting as an active process rather than simple loss for decades, and the same framing is the useful one here: the mitigation is a gentle time-based decay applied on top of supersession, so that an old fact gradually loses weight even with nothing to replace it. It never disappears on its own, which is correct, because a fact that has been true for two years is not less true for being old. It simply stops outranking newer, better-evidenced records.
A second mitigation is worth building for high-value facts: ask. An agent that surfaces a two-year-old preference and adds a light confirmation is doing what a person would do, and a confirmed fact can have its written-at value refreshed, which resets the decay honestly rather than by assumption.
There is a third case worth separating out, because it is not staleness at all: the fact that was never true. A model inferred something from a turn and stored it as though the user had said it, and no contradiction arrives because the user never knew it was recorded. Decay eventually lowers its weight, which is a slow remedy for something that should not have been written; the faster one is marking inferred records as inferred at write time so they can be weighted differently or confirmed before they are ever asserted.
Both of those push in the direction of forgetting more. The opposite failure is more visible to users and worth guarding against explicitly: how do you avoid forgetting the things that mattered?
The visible failure
How do you avoid forgetting the things that mattered?
By putting a floor under the policy rather than by making the policy gentler. Over-eager forgetting is the failure users notice, and it almost always comes from a horizon set shorter than the interval at which people actually return.
The complaint has a distinctive shape. A user says “I told you this yesterday” or, worse, repeats a correction they already made. The second case is the damaging one, because a correction is exactly the class of memory a system should treat as most durable, and a naive age or usage policy treats it like any other sentence.
Three classes deserve exemption from any policy that acts on age or use alone. Corrections, because the user has already invested effort in them. Stated constraints, meaning the things a user expects applied without repeating them, from allergies to accessibility requirements. And anything the user explicitly asked the agent to remember, which is a direct instruction and should not be quietly overridden by a retention rule.
The implementation is a salience floor rather than an exception list: records above an importance threshold are skipped by the age and usage rules entirely, and only supersession or an explicit user request removes them. That keeps one mechanism rather than two, and it makes the exemption visible in the data rather than buried in code.
One adjacent technique looks like a way to have both, and it is worth being precise about what it does: is summarising the same as forgetting?
A related operation
Is summarising the same as forgetting?
No. It is lossy compression, and the loss is chosen by a model rather than by your policy. Summarisation is frequently offered as the humane alternative to eviction, and it is a genuinely useful technique with a different risk profile that deserves stating plainly.
Replacing twenty session records with one summary reduces volume while keeping something of everything, which is the appeal. What it also does is discard specifics without telling you which ones. The summary keeps the gist and drops the exact figure, the qualifying condition, the name. Those details are often precisely what a later query needs, and unlike an evicted record, a summarised one cannot be recovered, because the original is gone.
The safer pattern is to summarise and retire rather than summarise and delete. Write the summary as a new record, retire the originals, and keep them. Retrieval improves because the summary is a better match for general queries, and the specifics remain reachable when something needs them. This costs storage, which is the cheapest thing in the system.
The wider treatment, including when consolidation is worth running at all and how it interacts with idle time processing, is on memory consolidation. The distinction to carry away is that eviction removes records you judged unnecessary, while summarisation removes details a model judged unnecessary.
Every mechanism on this page changes what the agent knows, silently, which raises the question the sources on this subject do not: how do you test that forgetting is working?
Measurement
How do you test that forgetting is working?
Log the retrievals your retired memories would have won. Forgetting fails silently in both directions, which is why it is the least tuned part of most memory systems and the reason this is the most useful idea on the page.
Neither direction raises an error. Forgetting too much shows up as a user complaint you may never hear. Forgetting too little shows up as retrieval quality drifting downward over months, which no dashboard reports because nothing broke. Both are invisible to ordinary monitoring, so both need a signal built on purpose.
The signal is cheap if you retired rather than deleted. Run each retrieval against the full store as well as the live one, and record when a retired record would have ranked in the selected set. That count is a direct measure of how aggressive the policy is: near zero means it is safe or too gentle, and a steady stream of high-scoring retired records means it is cutting into memories that were still wanted. The comparison run is offline and sampled, so it costs almost nothing.
Two further checks are worth having. Track the age distribution of the records that actually get retrieved, because that tells you empirically how far back your agent reaches and is the same data that should be setting your horizon. And keep a small fixed set of memories that must never be forgotten, seeded deliberately, and assert on every build that a query still finds them.
Read the retired-but-wanted cases rather than only counting them. They usually cluster into a class, which turns an abstract tuning problem into a specific rule such as exempting corrections or lengthening the horizon for one memory type. The wider method is on evaluating agent memory.
All of this concerns forgetting as a quality mechanism. A separate category of removal is not a judgement call at all: what has to be deleted rather than forgotten?
The boundary
What has to be deleted rather than forgotten?
Anything a person has asked you to remove, and anything you should not have stored. These are obligations rather than tuning decisions, and retiring does not satisfy them, because the record still exists.
The distinction matters in practice because the two paths look similar in code and are not similar in consequence. A retired memory is present in the database, in its backups, and in any index that has not been rebuilt. If a person exercises a right to have their data removed, which the GDPR calls the right to erasure and several other regimes mirror, a boolean column does not discharge that, and neither does removing the row while leaving the vector in an index that was built from it.
A deletion path therefore has to cover every tier that holds a copy: the record, its vector and index entry, any summary derived from it, the event history that logged it, and the cached session that may still hold it in working state. That is easier to guarantee when the tiers share a database, which is one of the arguments on storage backends, and it is a path worth designing when the schema is designed rather than when it is first requested.
Three categories are worth deleting on sight rather than retiring: credentials or secrets that reached the store through a conversation, data belonging to a category you have no basis to hold, and anything written under an identity that turned out to be wrong. The obligations, the categories and the practical mechanics are covered on memory security and privacy, which owns that subject.
For everything else, the policy on this page holds: retire rather than delete, put a floor under age and usage rules, add decay so that stale facts fade without needing a contradiction, and keep the evidence that tells you whether any of it was set correctly. The loop these decisions sit inside is on how AI memory works, and what should reach the store at all on writing memories.
FAQ
Frequently asked questions
The policy questions that follow: defaults, exemptions, and what to do with a store that already grew too large.
What retention horizon should I start with?
Whatever comfortably exceeds a high percentile of the gap between your users' sessions, which is data you already have. Starting from someone else's default is how a memory system ends up expiring facts before the people who stated them come back. Add a salience floor so important records are exempt from age rules entirely.
Which memories should never expire on age alone?
Corrections, stated constraints such as allergies or accessibility requirements, and anything the user explicitly asked the agent to remember. Implement it as an importance threshold that age and usage rules skip, rather than as a hardcoded list, so the exemption is visible in the data.
My store already has months of accumulated memories. Where do I start?
Retire rather than prune. Add the validity column, retire by supersession first since that acts on evidence, then apply an age rule with a salience floor and watch the retired-but-wanted log before tightening it. Nothing is deleted, so every step is reversible.
Does forgetting reduce cost?
Barely, and that is the wrong reason to do it. Storage for memory records is cheap at any realistic scale. The gain is precision: fewer stale candidates competing for the small number of slots that reach the prompt, which shows up as better answers rather than a smaller bill.
Should the user be able to see and delete their own memories?
Yes, and it is the cheapest correctness mechanism available, because the person best placed to know a fact is out of date is the person it is about. It also changes the deletion path from an internal process into a product surface, which is covered on memory security.
Can I use decay instead of an eviction policy entirely?
For many products, yes. Decay lowers the weight of old records without removing them, which avoids over-eager forgetting and still keeps stale facts from outranking current ones. It leaves the store growing, so pair it with supersession so replaced records actually leave the candidate set.