Architecture · Ranking
Memory Scoring and Ranking for AI Agents
Memory scoring decides which of the memories a search returned actually reach the prompt, and in what order. Search finds candidates. Scoring is what turns a list of things that resemble the query into a small set of things the agent should know right now, and it is where most memory systems quietly go wrong: not because the fact was missing, but because it lost to four near-copies of a less useful one.
Three signals
The premise
Why is the most similar memory not the right memory?
Because similarity answers a question you did not ask. A vector search ranks memories by how close their meaning sits to the query. What you actually want is the memory most useful to answer this turn, and those two orderings diverge constantly.
The clearest case is a question about dinner. Every restaurant the user has ever mentioned is semantically close to it, so similarity happily returns five of them, including one from eight months ago that was a passing remark. The memory that matters, a stated dietary restriction, may sit lower in the list because it is phrased in language further from the question.
The reverse failure is just as common. Sort by time instead and the agent gets whatever was said most recently, which is right for conversational continuity and wrong for anything else. A user who once stated a restriction sees it overridden by a casual mention of pizza last Tuesday, because pizza is newer.
Neither signal is wrong. Each is incomplete, and a working system combines them along with a third that neither captures: how durable the memory is, independent of when it was written or what the current query happens to be. Which raises the question of what the full set is: which signals should a memory score combine?
The function
Which signals should a memory score combine?
Three, weighted and summed: relevance to the query, recency, and importance. The formulation comes from the Generative Agents paper (Park et al., 2023), which normalises each of the three and sums them with every weight set to 1.0, and it remains the default in most systems because it is simple and it works.
Where implementations differ is in the weights, and equal weights are a starting position rather than an answer. A support agent handling live tickets wants recency to dominate, because a fact from earlier in the same conversation beats a similar fact from three months ago almost every time. A personal assistant wants importance to dominate, because stated preferences should survive months of unrelated chat. A research agent wants relevance to dominate, because the query is specific and the timeline is irrelevant.
Two further signals are worth considering once the three basics are working. Access frequency, meaning how often this memory has been retrieved before, is a cheap proxy for usefulness and it is what the original formulation actually decays: recency there is time since last access, not time since creation. And source, meaning whether the fact came from the user directly or was inferred by a model, which is worth a penalty on the inferred side because inferred facts are where confident errors originate.
Resist adding more than that. Every extra signal is another weight to calibrate and another thing that can drift, and the returns fall off quickly. The three-signal function with sensible weights beats a six-signal function nobody has tuned.
Relevance comes free from the search. Recency comes free from a timestamp. Importance is the one signal that has to be produced deliberately, and it is the one most systems fumble: how do you assign importance without guessing?
Calibration
How do you assign importance without guessing?
By defining what the number means before asking anything to produce it. Importance is not how interesting a turn was. It is how long the memory stays worth surfacing, which is a different question and a much easier one to answer consistently.
The usual approach is to ask a model to rate a memory from one to ten during extraction. That works only if the prompt defines the scale with examples, because without anchors a model rates almost everything in the middle, and a signal where every value is a six contributes nothing to the ranking while still consuming a third of the weight.
Better is to score by category rather than by judgement. Corrections rank highest, because the user telling the agent it was wrong is exactly what should never be forgotten and is routinely scored like any other sentence. Stated preferences and constraints rank next, since the user expects them applied without repeating them. Session outcomes sit in the middle. Passing remarks sit at the bottom.
The bottom of the scale deserves a value of zero, and most implementations do not offer one. A memory scored zero is still stored and still retrievable by an explicit query, but it never surfaces on general recall. That distinction is what keeps a store from slowly filling its own retrieval budget with things nobody asked about, and it is milder than the alternatives on forgetting and eviction.
One more property matters: importance assigned at write time is assigned with the least context that will ever exist. A stated preference looks ordinary in the turn it appears in and looks important after the user refers back to it three times. Systems that never revise the score are stuck with a first impression, and the cheapest revision available is to raise importance on retrieval, so that memories the agent actually uses drift upward on their own.
Importance is deliberately static between revisions. The signal designed to move on its own is the other one: how fast should a memory’s recency score decay?
Time
How fast should a memory’s recency score decay?
Slowly enough that a genuinely useful memory survives a quiet week. The standard treatment is exponential: a memory’s recency score is a decay factor raised to the number of hours since it was last accessed, and the Generative Agents formulation uses 0.995 per hour.
That constant is easier to reason about as a half-life. At 0.995 per hour a memory’s recency score halves in roughly 138 hours, a little under six days, and falls to about a quarter in twelve days. Whether that is right for your agent is a question about how often your users return. For a daily-use assistant it is sensible. For a tool someone opens once a month, every memory is effectively at zero recency by the time they come back, and the signal has quietly stopped contributing anything.
Two adjustments are worth knowing. Decaying from last access rather than from creation means a memory that keeps proving useful keeps its recency, which is closer to what you want than punishing a fact for being old when it is retrieved constantly. And a floor under the decay, so that recency never reaches zero, keeps very old but highly relevant memories from being unreachable no matter how well they match.
The failure to watch for is recency swamping the other signals in a long session. Twenty turns into a conversation, everything from this session sits near the top of the recency range and everything from previous sessions sits near the bottom, so a weighted sum effectively becomes a session-only retrieval. If the agent stops remembering anything from last week during long conversations, this is why, and the fix is either a lower recency weight or a separate budget for older memories, covered in ranking across layers.
All three signals now exist, on three different scales. Combining them correctly is a smaller topic than it sounds and a bigger source of bugs: how do you combine signals measured on different scales?
The common bug
How do you combine signals measured on different scales?
Normalise each one across the candidate set before you weight it. This is the step most implementations skip, and skipping it means the weights written in the code are not the weights the system applies.
Look at the raw ranges. Cosine similarity between an embedded query and a plausible candidate rarely spans the full zero-to-one range; in practice the candidates that survive a search cluster in a narrow band, often between 0.7 and 0.9. An exponential decay factor spans nearly the whole range from one down toward zero. An importance score from a model might be an integer from one to ten. Sum those raw and the signal with the widest spread dominates entirely, regardless of what weight it was given.
Concretely: if similarity varies by 0.2 across the candidates and recency varies by 0.9, then with equal weights recency is doing more than four times the work of relevance. The system behaves like a recency-sorted list with mild semantic influence, and no amount of adjusting the weights fixes it until the scales are equalised, because the adjustment is compensating for a distortion rather than expressing an intent.
Min-max normalisation across the retrieved candidates is usually enough: for each signal, map the lowest value in this candidate set to zero and the highest to one. It is cheap, it adapts to each query, and it makes the weights mean what they say. Its one quirk is that it always produces a top score of one for each signal, even when every candidate is poor, which is a reason to apply the relevance threshold before normalising rather than after.
Once the arithmetic is honest, the next decision is how much of the ranked list to actually use: should you return a fixed number of memories or everything above a threshold?
Selection
Should you return a fixed number of memories or everything above a threshold?
A threshold, in almost every case. A fixed count returns the same number of memories whether the store holds twelve relevant facts or none, which means it pads the prompt with noise on unrelated queries and truncates the answer on rich ones.
The padding case is the expensive one. A user asks something the agent has no stored context about, the retriever returns its five nearest candidates anyway, and the model receives five irrelevant facts presented in exactly the same format as useful ones. It has no way to know they were weak matches, so it may use them. Returning nothing is a valid and often correct outcome, and a fixed count makes it impossible.
Thresholds have their own maintenance cost. The bar has to be calibrated against real queries rather than guessed, and it has to be revisited whenever the embedding model changes, because similarity values from two models are not comparable. A practical approach is to keep an upper bound as a budget guard while making the threshold the real decision, so the count varies with the query and the prompt never overflows.
There is a further reason not to trust the top of a ranked list, and it is specific to memory rather than inherited from document retrieval: do your top-scoring memories all say the same thing?
Redundancy
Do your top-scoring memories all say the same thing?
Very often, and the scoring function cannot see it. A score computed per memory says how good that memory is in isolation. It says nothing about whether the memory ranked second adds anything the memory ranked first did not.
Recent research makes the structural version of this argument. Hu et al. (2026) point out that retrieval-augmented generation was designed for large heterogeneous corpora where the main failure mode is irrelevance, while an agent’s memory is a bounded, coherent stream in which many spans are near duplicates of one another. Under those conditions, they argue, fixed similarity ranking collapses into a single dense region and returns redundant evidence, failing to separate what is needed from what is merely similar.
The practical shape is familiar to anyone who has read their own retrieval logs. A user mentioned a preference in four different conversations. All four memories are stored, all four are relevant, all four score highly, and all four occupy slots in a budget that could have held four different facts. The agent receives one fact repeated and knows less than it would have with a mixed set.
The fix is to select a covering set rather than a top slice: add a memory to the selection only if it contributes something the already-selected ones do not, measured by similarity to the selected set rather than to the query alone. It costs a comparison per candidate against a small set, which is negligible next to the model call that follows.
The cheaper mitigation is upstream. If four memories say the same thing, the write path should have merged them, which is the job described on handling conflicting memories and on memory consolidation. Deduplicating at read time is a workaround for a store that accumulated duplicates it should not have.
A good, diverse set can still be wasted by how it is assembled: where in the prompt should the highest-scoring memory go?
Ordering
Where in the prompt should the highest-scoring memory go?
At the start or the end of the retrieved block, never buried in the middle. Ranking that sorts candidates and then concatenates them in arbitrary order hands most of its own benefit back, because position inside a context affects whether the model uses the information at all.
Liu et al. (2023) documented the effect in a multi-document question answering task: accuracy is highest when the answer sits in the first document, nearly as high when it sits in the last, and drops sharply when it sits in the middle of twenty. The swing is roughly twenty points and it is driven by position alone, with the same documents and the same question.
For a handful of memories the effect is mild, which is why it goes unnoticed in small systems. It becomes serious once the retrieved block is long or once memories are injected alongside a substantial conversation history, at which point the best-scoring memory can end up exactly where the model attends least.
Two habits follow. Put the top-scoring memory adjacent to the instruction that will use it, which usually means immediately before the current user turn rather than in a block at the top of the system prompt. And keep the block short, because the effect is a function of length: five well-chosen memories in a compact block do not have a middle worth worrying about.
Ordering assumes one block of memories. Real agents assemble a prompt from several sources at once: how do you rank across several memory layers into one budget?
Budgeting
How do you rank across several memory layers into one budget?
By allocating the budget per layer first, then ranking within each. Scores from different stores are not comparable, so a single global ranking across them is arithmetic that looks reasonable and means nothing.
A production agent typically assembles a prompt from working state for the current session, durable facts about the user, and retrieved documents or knowledge. Each comes from a different store with a different scoring scheme, and a similarity of 0.82 from one is not the same quantity as a similarity of 0.82 from another, because the corpora and often the embedding models differ.
Fixed allocation is the workable pattern: decide in advance what share of the token budget each layer gets, fill each share by ranking within that layer, and let a layer return less than its share when nothing clears its threshold. Session state usually gets the largest share because it is almost always relevant, durable facts get a small guaranteed floor so they are never crowded out, and retrieved knowledge takes the remainder.
The floor for durable facts matters more than it seems. Without it, a long conversation fills the entire budget with session context, and the agent stops applying the preferences it learned months ago at exactly the point in a conversation where a user expects it to know them. The tiers themselves are set out on types of AI agent memory, and where each lives on storage backends.
Every decision on this page is a hypothesis about what helps the agent answer better, which means none of it is finished until it is measured: how do you tell whether your scoring is actually working?
Measurement
How do you tell whether your scoring is actually working?
By logging what was retrieved with its component scores, then reading the cases where the agent got it wrong. Scoring cannot be tuned from aggregate accuracy, because a wrong answer does not say whether the fact was missing, outranked, or present and ignored.
Log the candidate set with each signal’s value before and after normalisation, the combined score, and which candidates were selected. That single record separates the three failure modes immediately: the fact was not in the candidate set, which is a retrieval problem; it was in the set and scored too low, which is a weighting problem; or it was selected and the model did not use it, which is a prompt assembly problem and often the position effect above.
Tune one weight at a time against a fixed set of real queries with known correct answers. Twenty to fifty queries collected from actual usage is enough to see a weight change move the ranking, and it is far more informative than a benchmark score, because the failures are yours. The wider method is on evaluating agent memory and the standard datasets on memory benchmarks.
Expect diminishing returns and stop when you reach them. Published sweeps of retrieval budget on conversational benchmarks show accuracy climbing steeply from a minimal budget and then flattening, with the last increments costing substantially more tokens for a fraction of a point. The same shape applies to weight tuning: the move from untuned to roughly right is large, and everything after it is small.
Be careful about what most frameworks give you by default, which is usually similarity-ranked top-k with no recency, no importance and no normalisation. That default is a reasonable place to start and a poor place to stay, and the gap between it and a tuned three-signal function is most of the difference between an agent that seems to remember and one that seems not to. The read path this feeds is set out on memory retrieval, and what should have been written in the first place on writing memories.
FAQ
Frequently asked questions
The implementation questions that come up once the function is in place: defaults, drift and what to log.
What weights should I start with for recency, importance and relevance?
Equal weights, as in the original Generative Agents formulation, then move one at a time against real queries. Which weight to raise first depends on the agent: recency for anything conversational and live, importance for an assistant that must apply months-old preferences, relevance for research and lookup work.
Should importance be assigned by a model or by rules?
Rules where the category is knowable, a model where it is not. Corrections, stated preferences and explicit instructions can be detected structurally and scored by category, which is more consistent than any rating prompt. Reserve the model call for turns that do not fall into a known class.
What should I log to debug a bad retrieval?
The full candidate set with each signal's raw and normalised value, the combined score, and the selection outcome. Without the per-signal values you cannot tell whether a fact was missing from the candidates, outranked inside them, or selected and then ignored by the model, and those need three different fixes.
Does scoring need to change when I switch embedding models?
Yes, and specifically the relevance threshold. Similarity values are not comparable across models, so a bar tuned for one model can admit almost everything or almost nothing under another. Re-calibrate the threshold against the same query set before and after the switch.
How many memories should reach the prompt on a typical turn?
Few enough that a wrong one is unlikely to be included, which for most agents means a handful rather than a page. Precision matters more than recall here, because every extra memory is another chance for a stale fact to be asserted and another slot the position effect can bury.
Is reranking a replacement for a scoring function?
No. A reranker reorders candidates by relevance to the query and knows nothing about recency, importance or redundancy, so it improves one of the three signals rather than replacing the combination. Details on hybrid search.