Advanced · Idle-time inference
Sleep-Time Compute for AI Agent Memory
Sleep-time compute is a specific technique: reasoning over an agent’s existing context while it is idle, then writing the result back to memory so the next query answers from it instead of re-deriving it. It only works for agents with somewhere persistent to write that result, which is why it is a memory-architecture technique rather than a generic term for background jobs. The paper that introduced it measured real gains, and real limits worth knowing before you build it.
Four steps
Definition
What is sleep-time compute, and what problem does it solve?
Reasoning that happens between interactions rather than during them, so the model has already thought about the context before the next question arrives. The name and the technique come from a specific 2025 paper, and it is worth using it precisely rather than as a synonym for any background job.
The problem it addresses is the cost of test-time scaling. Letting a model think longer at the moment of answering, the technique behind reasoning models like o1 and R1, improves accuracy on hard problems, and it does so at a real cost: minutes of latency and, for the most demanding configurations, tens of dollars per query. That cost is paid every time, even when several queries share the same underlying context and are asking closely related things about it.
Sleep-time compute observes that many agent applications are inherently stateful: a coding agent works against the same repository across many turns, a document assistant answers repeated questions about the same file, a conversational agent carries the same dialogue history forward. In all of these, the context is already available before the next question is asked. Nothing stops the model from thinking about that context in advance, off the clock, and sleep-time compute is the name for doing exactly that.
It is worth distinguishing sleep-time compute from a technique it resembles on the surface: speculative decoding, which also reduces latency by generating tokens ahead of when they are confirmed to be needed. The difference is what happens to the speculative work when the guess is wrong. Speculative decoding discards tokens that turn out not to match the eventual output, so a wrong guess costs nothing beyond the wasted computation. Sleep-time compute’s generated reasoning is used regardless of what the user actually asks, on the working assumption that the query is anticipatable from the context. That assumption is testable, and it is the subject of one of the paper’s own experiments further down this page.
Two things have to be true for it to apply at all: the workload has to be stateful, meaning there is persistent context to reason over, and there has to be a place to write the result so a later query can read it. Neither condition is automatic, which is why the mechanism is worth looking at directly: how does an agent turn raw context into learned context while idle?
The mechanism
How does an agent turn raw context into learned context while idle?
By prompting the model to generate a new, re-represented version of the existing context, then persisting that version for later use. The paper calls the input “raw context” and the output “learned context,” and the distinction is the entire idea in two words.
Concretely, a coding agent working on a repository might use idle time to identify architectural patterns, note likely places a bug would hide, or pre-derive an optimisation, none of which was asked for yet. A conversational agent might use the same window to reconcile something the user said against what is already stored, which overlaps with what memory consolidation does, though sleep-time compute is one specific way of producing that consolidation rather than the whole subject.
The result has to go somewhere that survives past the inference call, or the exercise is wasted. This is the reason sleep-time compute is fundamentally a memory-architecture technique: it presupposes a stateful agent with persistent context, and an agent that resets between turns has nowhere for the learned context to live. Where that learned context is written, and how it differs from other tiers of memory, is covered on types of AI agent memory.
The technique is not free. It spends compute during idle time that would otherwise go unused, on the bet that the result will be read later. Whether that bet pays off depends on how the write is actually organised, which is the implementation question: what does a sleep-time agent architecture actually look like?
Implementation
What does a sleep-time agent architecture actually look like?
Two agents split by latency budget: one that talks to the user and one that rewrites memory. This is the pattern used in the reference implementation that popularised the technique, and it is a clean way to reason about the split even if you build it differently.
The primary agent handles the conversation, calls tools, and reads memory freely, but it does not edit its own persistent memory directly. That job belongs to a separate sleep-time agent, which runs between turns or on a schedule and has the ability to rewrite the memory the primary agent reads from. Splitting the two avoids a specific failure of single-agent designs: an agent that has to manage its own memory and hold a conversation at the same time is slower on every turn and more likely to call the wrong kind of operation at the wrong moment.
Because the split is by latency rather than by capability, the two agents can run different models entirely. The primary agent benefits from a fast, cheap model, since every millisecond it spends is felt by a person waiting for a reply. The sleep-time agent has no such constraint and can run a slower, stronger model, since its only cost is compute that would otherwise sit idle.
The remaining implementation choice is frequency: how often the sleep-time agent runs. A higher frequency means more tokens spent and a learned context that is fresher when the next query arrives; a lower frequency spends less but risks the learned context lagging behind what actually happened in the most recent turns. There is no universal answer, and it is worth treating as a tunable parameter rather than a fixed default, revisited against the workload the same way a retention horizon is on forgetting and eviction.
An architecture is only worth building for a measurable return, and the paper that introduced this pattern measured one directly: what does sleep-time compute actually save, measured?
Measured results
What does sleep-time compute actually save, measured?
Roughly five times less test-time compute for equal accuracy, and up to eighteen points of accuracy gain when scaled further, according to the paper’s own reported figures. These are the numbers to cite, because the paper that introduced the technique ranks in its own search results, so there is no need to pass along a vendor’s paraphrase of it.
Two modified mathematical reasoning benchmarks, built by splitting existing problems into a context and a question, are where the headline figures come from. Sleep-time compute produced a Pareto improvement in the accuracy-versus-test-time-compute curve, reducing the compute needed to reach the same accuracy by roughly five times. Scaling sleep-time compute further shifted accuracy up by 13 percentage points on one benchmark and 18 on the other, at a fixed test-time budget.
The second result is about cost rather than accuracy: when several queries share the same underlying context, the sleep-time reasoning only has to be paid for once and can be reused across all of them, cutting the average cost per query by up to 2.5 times at ten queries per context in the paper’s own analysis. That is the amortisation effect, and it is the reason the technique suits repeated-use workloads far better than one-off interactions.
All of these gains assume the workload has a property the paper studied directly, and knowing whether yours has it is more useful than any of the headline numbers on their own: how do you know if your workload will benefit at all?
The condition
How do you know if your workload will benefit at all?
Check whether the question is predictable from the context, because the paper shows the benefit scales directly with how predictable it is. This is the single most useful finding in the research for deciding whether to build the architecture at all, rather than assuming it helps everywhere.
The paper quantifies this rather than gesturing at it: it scores how likely a given question is given its context, under a reference model, then bins examples into groups by that score and compares sleep-time compute against standard test-time compute within each group. The result is not subtle. The accuracy gap between the two widens as questions become more predictable from their context, and narrows toward nothing as questions become harder to anticipate.
Translated into an engineering check: does your context, a codebase, a document, a conversation history, tend to constrain what will be asked about it. A support agent answering from a known account and order history is a strong case, because the space of likely questions is narrow. An open-ended research assistant fielding genuinely novel questions about a new context every time is a weak case, because there is little for idle-time reasoning to usefully anticipate.
Predictability answers whether the technique can work. It does not answer whether it is the better choice in every situation where it can, which is a separate and more surprising finding: does sleep-time compute always win?
The honest limit
Does sleep-time compute always win?
No, and the paper’s own case study shows exactly where it loses. Every source that discusses the agentic coding evaluation in the paper cites the win. Fewer state the condition under which the paper’s own numbers favour the alternative instead.
In a software-engineering case study built around multi-file pull requests, sleep-time compute beats standard test-time scaling at lower test-time budgets: when the agent has little time to reason at the moment of answering, having already thought about the repository in advance wins clearly. At higher test-time budgets, the comparison flips, and standard test-time scaling overtakes sleep-time compute. When the agent is allowed to reason for a long time at query time anyway, pre-computed context stops being the advantage it was.
That single result reframes the whole decision. Sleep-time compute is not a strictly better version of test-time scaling. It is a way to buy back a low-latency budget by spending idle time instead, and its advantage shrinks as you give the agent more room to think in the moment regardless. A system that can already afford a generous test-time budget has less to gain from it than one under tight latency constraints.
Cost and accuracy are not the only things at stake when a model reasons unsupervised for long stretches: what can go wrong when an agent thinks without supervision?
Risks
What can go wrong when an agent thinks without supervision?
An error made during sleep-time reasoning gets written into memory as though it were settled, and the primary agent has no way to tell the difference. This is the sharpest practical risk of the technique and it follows directly from the mechanism: the whole point is that the primary agent trusts the learned context rather than re-deriving it.
If the sleep-time agent misreads the codebase, draws the wrong inference from a conversation, or hallucinates a connection that is not there, that mistake is now sitting in memory with the same authority as anything correctly derived. The primary agent will read it back on the next turn and act on it with no signal that it was never verified. This is a variant of the same failure this site covers on conflicting memories, with one difference: the bad write here did not come from the user, it came from the system’s own idle-time reasoning, which makes it easy to trust by default.
The mitigation follows the same shape as everywhere else memory can be wrong. Treat sleep-time writes as provisional rather than authoritative, mark them with their source the way any inferred fact should be marked per storage backends, and where the stakes justify it, have the sleep-time agent flag lower-confidence inferences for review rather than writing everything with equal weight. A second, more mundane risk is operational rather than architectural: running a second agent means a second thing to monitor, schedule and debug, and multi-agent orchestration overhead is a real cost even when the split works exactly as intended.
A third risk sits underneath both of the others and is easy to miss because it looks like a storage decision rather than a safety one. How the memory store retrieves what the sleep-time agent wrote determines how much of a bad inference actually reaches the primary agent on any given turn. A retrieval step that is too permissive surfaces a hallucinated inference alongside the sound ones with no way to tell them apart at read time, which is the same relevance problem covered on memory retrieval, applied here to writes the user never saw and therefore cannot catch by simply disagreeing with the agent.
With the mechanism, the measured gains, the condition for benefit, the limit, and the risk all on the table, the decision comes down to a short checklist: should your agent use sleep-time compute?
The decision
Should your agent use sleep-time compute?
Only if the agent is stateful, the workload is predictable from its context, and either latency is tightly constrained or many queries share one context. All three conditions have to hold, because each addresses a different way the technique can fail to pay for itself.
Statefulness is the precondition: without persistent memory to write the learned context into, there is nothing for this technique to do, and it collapses back into ordinary test-time reasoning. Predictability is the condition the paper measured directly, and it is worth checking against real query logs rather than assuming. And the latency-or-volume test separates the two cases where the investment clearly pays off: a product where users feel every millisecond, or a context read by enough queries that amortisation covers the cost of thinking about it once.
Where all three hold, sleep-time compute is a genuine architectural option, not a marketing term for background processing, and it sits alongside the broader work of memory consolidation as one specific way of doing it well. Where the workload is unpredictable, where the agent is stateless, or where the test-time budget is already generous, the honest answer from the paper’s own results is that standard reasoning at query time remains the better choice, and no amount of idle-time cleverness changes that.
FAQ
Frequently asked questions
The implementation questions that follow: scheduling, model choice, and what happens when it goes wrong.
How often should a sleep-time agent run?
There is no universal default. Higher frequency spends more tokens and keeps the learned context fresher; lower frequency spends less but risks the learned context lagging behind recent turns. Treat it as a tunable parameter and revisit it against real usage, the same way a memory retention horizon is tuned on forgetting and eviction.
Does the sleep-time agent need a stronger model than the primary agent?
It benefits from one, because it has no latency budget to respect. A common pattern pairs a fast, cheap model for the primary agent that talks to the user with a slower, more capable model for the sleep-time agent that has time to reason more thoroughly between turns.
Is sleep-time compute the same thing as memory consolidation?
No. Memory consolidation is the general operation of merging, summarising and cleaning stored memories. Sleep-time compute is one specific technique for producing that outcome, defined by reasoning over existing context during idle time and writing the result back for reuse. See memory consolidation for the broader operation.
Can sleep-time compute be used with a stateless agent?
No. The technique requires somewhere persistent to write the learned context so a later query can read it. A stateless agent that discards everything between calls has nothing for the idle-time reasoning to feed into, which collapses the approach back into ordinary test-time reasoning.
What is the cheapest way to test whether sleep-time compute would help our workload?
Sample real queries against their contexts and check how predictable each question is given what the agent already had available. The paper's own results show the benefit tracks predictability directly, so a quick manual review of a few dozen real query-context pairs will indicate whether the investment is likely to pay off before any architecture is built.
How do you catch a bad inference the sleep-time agent wrote to memory?
Mark sleep-time writes with their source so they are distinguishable from user-stated facts, and where the stakes justify it, route lower-confidence inferences through a review step rather than writing everything with equal authority. The same supersession mechanics used for any conflicting memory apply, covered on conflicting memories.