Advanced · Memory and reinforcement learning
Memory and Reinforcement Learning in AI Agents
An agent that records what it tried and what happened has the raw material for learning, and the interesting recent work applies that signal to the memory store rather than to the weights. The lesson lands on the next turn instead of the next training run, nothing already known is disturbed, and every update stays inspectable.
The loop
The connection
What does reinforcement learning have to do with agent memory?
Reinforcement learning needs an action, an outcome and a way of adjusting behaviour. An agent with episodic memory already produces the first two, and the memory store is a place to put the third. That is the whole idea, and its appeal is that the adjustment costs a write rather than a training run.
Ordinary reinforcement learning adjusts model parameters. The variant this page is about adjusts what is stored and how it is ranked, leaving the model untouched, which is why the research describes it as non-parametric learning: the agent gets better while its weights stay exactly where they were.
The signal is unusually easy to come by in agent settings. A test suite passes or fails, a build succeeds, a task completes, a user accepts or rejects a suggestion. None of that needs a human labeller, and all of it is already flowing past systems that discard it.
That discarding is the actual gap. Most agents write what happened and never connect it to what they did next, so the store accumulates episodes without ever using them to decide anything. The write mechanics are on how agents write and store memories, and the type in question on episodic memory.
The obvious alternative is to train on the experience instead, which runs into a wall the field has now named: why continuous fine-tuning does not work for agents.
The constraint
Why does continuous fine-tuning not work for agents?
Because each lesson costs a training run, and training on new experience degrades what the model already knew. The first makes the loop too slow to close inside a task; the second, catastrophic forgetting, makes repeated small updates actively risky.
Timing is the plainest objection. An agent discovers at 10am that an approach fails, and should act on that at 10:01. No training pipeline closes in a minute, so a fine-tuned agent is learning on a schedule that has nothing to do with the work it is doing.
Forgetting is the deeper one. Training on a stream of recent experience shifts the model toward it and away from everything learned before, so an agent that fine-tunes continuously trades yesterday’s competence for today’s. Guarding against that costs replay data and evaluation, which returns you to a slow, expensive loop.
A memory update has neither property. It applies immediately, disturbs nothing else, and can be inspected and reversed, which is the same argument set out on memory versus fine-tuning and taken further on continual learning for agents.
Moving learning into the store does not remove the underlying tension, it relocates it: the stability and plasticity problem.
The tension
What is the stability and plasticity problem in agent memory?
A system that never changes cannot learn, and a system that changes on every signal cannot keep anything. The useful designs decouple the two: a stable model doing the reasoning, and a plastic memory layer absorbing what changes.
Too stable is the ordinary failure and looks like an agent proposing a fix that failed last week. The episode is stored; nothing consults it when choosing what to do, so the store is a record rather than a policy.
Too plastic is rarer and worse. If one bad outcome retires a strategy that works nine times in ten, the agent oscillates and never settles, and because the discarded strategy was correct the failures accumulate quietly. This is why a single outcome should adjust a memory’s standing rather than delete it.
The decoupled arrangement is what the recent work formalises, and it is worth noticing that it describes what most memory systems already do. The gap is the credit step: without it, a store grows without ever being informed by whether using it worked.
Which is the mechanism itself: how an agent learns what is worth remembering.
The mechanism
How does an agent learn what is worth remembering?
By recording which memory it used, observing what happened next, and adjusting that memory’s standing accordingly. Four steps, and only the third is unusual: the credit goes to the retrieved memory rather than to the model.
The first requirement is that retrieval is recorded. If you cannot say which memories were in the prompt when an action succeeded, no credit can be assigned, which is one more reason to log the retrieval decision as described on the developer hub.
The second is a usable outcome signal. Test results, task completion, an escalation, an explicit correction from the user, all work. Ambiguous signals are the hard part: a conversation that simply ends says very little, and treating silence as approval teaches the wrong lesson.
Credit assignment is where this gets genuinely difficult. If four memories were in the prompt and the task failed, which one was responsible, and was it any of them? Coarse approaches spread credit across everything retrieved, which is noisy but workable at volume; finer approaches need the agent to say which memory it acted on, which is a design decision in the prompt rather than in the store.
The reward, once assigned, has to change something at read time: how utility changes retrieval.
Retrieval
How does learning change memory retrieval?
It adds a utility term, so relevance stops being sufficient on its own. Similarity says a memory is about the right subject. Utility says acting on it has previously worked, and only the second distinguishes a good strategy from a well-phrased bad one.
The two-phase shape is the practical one. Similarity narrows the field to memories about the right thing, then utility ranks what survives, which keeps the retrieval cheap and puts the learned signal where it changes the decision. It also degrades gracefully: with no outcome data yet, ranking falls back to similarity and nothing breaks.
This extends the standard scoring of recency, importance and relevance with a fourth term derived from experience rather than assigned at write time, and it fits the existing function directly, as described on memory scoring.
One practical caution. Utility weighting narrows exploration: a strategy that scored badly once may never be retrieved again and so never gets a second data point. Some exploration has to be preserved deliberately, which is the same problem reinforcement learning has always had, arriving in a new place.
How well any of this works outside a paper is the fair question: what the research actually shows.
Evidence
What does the research on memory and reinforcement learning show?
That runtime learning on episodic memory beats passive retrieval on agent benchmarks, without any weight updates. The evidence is early, benchmark-shaped, and pointed in a consistent direction.
The clearest recent example is MemRL (Zhang et al., 2026), which applies reinforcement learning to episodic memory at runtime. Its argument is the one this page has been making: fine-tuning is expensive and prone to catastrophic forgetting, while memory methods that rely on passive semantic matching retrieve noise. Its answer is a two-phase retrieval that filters that noise and identifies high-utility strategies from environmental feedback, evaluated on HLE, BigCodeBench, ALFWorld and Lifelong Agent Bench.
A parallel line of work frames the same thing as learning in token space: the agent improves by changing what is written and retrieved rather than what the weights encode, which makes the learned material readable, editable and portable between models in a way parameters are not.
Two honest caveats. These results are on benchmarks with clean success signals, and production environments rarely offer anything so unambiguous. And the approach inherits every weakness of the underlying memory system: a store with poor write hygiene learns to prefer its own bad memories, which is worse than not learning at all. Evaluating that is on how to evaluate agent memory.
None of which prevents borrowing the useful half now: what you can build today.
Practical
What can you build today without a reinforcement learning stack?
Record outcomes against memories and use them at retrieval time. That is most of the benefit and none of the machinery. No policy gradients, no training infrastructure, just two extra columns and a term in the ranking function.
Three steps get you there. Store the outcome: when an episodic memory records an attempt, record what happened, so “tried X” becomes “tried X, it failed with this error”. That one change makes the store useful for decisions rather than only for recall.
Count use and success: keep how often a memory was retrieved and how often the turns it appeared in went well. Crude, and enough to separate strategies that work from ones that merely match.
Rank with it: add the counts as a term alongside relevance and recency, weighted low at first. A memory recording a failed approach should still be retrievable, since knowing what failed is exactly what stops the agent repeating it. Failure memories are not low-utility, they are high-utility with a negative sign.
The single highest-value version of this is the simplest: never propose a fix that a memory records as already attempted and failed. It needs no learning theory, and it removes the behaviour users find most exasperating in support and coding agents alike, as covered on memory for coding agents.
Before letting a system adjust its own store automatically, one more thing is worth thinking through: the risks of an agent that reshapes its own memory.
Risks
What are the risks of an agent that rewrites its own memory?
Reward hacking, self-reinforcing errors, and a store that no longer explains itself. A system that adjusts its own memory has a feedback loop, and feedback loops fail in ways that open-loop systems do not.
Reward hacking follows from a proxy signal. If the reward is task completion, strategies that complete tasks by avoiding the hard part score well, and the agent learns to take the easy route rather than the correct one. Proxy signals are all that is usually available, so this is a design constraint rather than an avoidable bug.
Self-reinforcement is the compounding one. A memory extracted from the agent’s own unverified output gets retrieved as evidence, contributes to an outcome scored as good, and gains standing on that basis. Restricting extraction to user turns and verified tool results prevents it, and no amount of ranking fixes it afterwards.
Opacity is the operational cost. Once utility scores are moving on their own, explaining why an agent behaved as it did means reconstructing a history of adjustments, not reading a row. Keeping the score history and its inputs is what makes that recoverable, and it is the same argument as the provenance fields on writing memories.
The practical safeguard is a rate limit on how far any single outcome may move a memory, plus a review path for the memories that reach the top of the store. A system that learns slowly and can be inspected is worth a great deal more than one that learns quickly and cannot, which is also the theme of memory security.
FAQ
Frequently asked questions
The questions this area attracts: whether it is production-ready, what signal to use, and how it differs from ordinary scoring.
Is reinforcement learning on agent memory production-ready?
The full research versions are not. The useful half is: record outcomes against the memories that were retrieved and add them as a term in ranking. That needs no training infrastructure and delivers most of the benefit.
What can be used as a reward signal for agent memory?
Anything the environment already produces: a passing test, a completed task, an accepted suggestion, an avoided escalation. Silence is the signal to avoid, since treating an ended conversation as approval teaches the wrong lesson.
How is this different from memory scoring?
Standard scoring uses recency, importance and relevance, all assigned at write time or derived from the query. This adds a term derived from what happened after the memory was used, which is the only one that reflects whether acting on it worked.
Does this replace fine-tuning for agents?
For learning from experience during operation, largely yes, because the loop closes in seconds and disturbs nothing else. Fine-tuning keeps its place for behaviour, format and domain vocabulary. See memory versus fine-tuning.
Should memories that record failures be retrieved?
Yes, and prominently. Knowing that an approach already failed is exactly what stops an agent repeating it. Treat failure memories as high utility with a negative sign rather than as low-value records to bury.
What stops an agent learning the wrong lesson from its own output?
Restricting extraction to user turns and verified tool results. A memory drawn from the agent's own unverified text can be retrieved as evidence, contribute to a good outcome and gain standing on that basis, and no ranking change repairs it afterwards.