Advanced · Degradation

What Is Context Rot?

Context rot is the performance degradation LLMs show as input length grows, and it is worst on exactly the tasks that matter most: aggregation, multi-hop reasoning, and long conversations, while the benchmark most vendors cite, single-fact retrieval, is where it shows up least. A model can score near-perfectly on that benchmark and still lose track of a conversation, a summary, or a subtle detail buried in a long agent transcript.

Least to most severe

1
Single fact
2
Conversation
3
Aggregation

The premise

What is context rot, exactly?

The tendency for a language model’s effective performance to decline as the length of its input grows, even when the relevant information is technically present in the context window and the task itself doesn’t get any harder. The information isn’t lost or deleted; the model simply becomes less reliable at using it correctly as more tokens surround it.

This directly contradicts a common assumption. Widely cited long-context benchmarks like Needle in a Haystack, placing a known fact in a long document and asking a model to retrieve it, often show near-perfect scores even at very large context lengths, which has fed the perception that long context is a solved problem. A July 2025 research report from Chroma, evaluating 18 leading models including GPT-4.1, Claude, Gemini and Qwen, found otherwise: performance degrades consistently across a range of controlled tasks as input length grows, even on tasks as simple as retrieving a single fact or exactly reproducing a short block of repeated text.

Why a benchmark that shows near-perfect scores and a real phenomenon of measurable degradation can both be true at once comes down to what the benchmark is actually testing, and what the model is actually doing internally as more tokens are added. Why does it happen mechanically?

Mechanism

Why does it happen mechanically?

Attention is a finite resource distributed across every token in the context, so every token added dilutes the share available to every other token, and the computational cost of relating tokens to each other grows with context length rather than staying flat. This isn’t a training gap that a bigger model fixes; it’s a structural property of how a transformer’s attention mechanism works.

Anthropic’s own engineering team describes this directly: “context must be treated as a finite resource with diminishing marginal returns. Like humans, who have limited working memory capacity, LLMs have an ‘attention budget’ that they draw on when parsing large volumes of context. Every new token introduced depletes this budget by some amount.” Because the transformer architecture lets every token attend to every other token across the full context, a model’s ability to capture the relationships that matter gets stretched thinner as more tokens compete for the same fixed attention budget, and the per-token computational cost of that attention grows as context length grows, not just the total cost.

The computational side of this is worth stating precisely, because it explains why the problem gets worse faster than input length grows. Before generating each new token, a model compares that token to every previous token in its context: with 100 prior tokens, that’s roughly 100 attention comparisons per new token; with 1,000 prior tokens, roughly 1,000. Since this cost applies per generated token, a session with ten times more tokens takes on the order of a hundred times more compute, not ten. No architecture in wide production use today has fully escaped this scaling, which is why frontier labs have pursued engineering workarounds, like FlashAttention’s more efficient attention computation, rather than a fundamentally different mechanism.

This mechanism alone would predict uniform degradation as tokens increase. What actually happens is more specific, and more useful to know before choosing which tasks to trust a long context with. Why do benchmark scores hide the real severity?

Task dependence

Why do benchmark scores hide the real severity?

Because context rot is not uniform across task types. It is least severe on exactly the task most long-context benchmarks measure, single-fact retrieval, and most severe on aggregation, multi-hop reasoning, and long conversations, the tasks that actually resemble what a production agent does. A model that looks solved on Needle in a Haystack can still fail badly the moment a task requires connecting two facts instead of finding one.

Context rot is least severe on single fact retrieval and most severe on aggregation and multi-hop reasoning tasks
Figure 1. Needle in a Haystack scores look strong precisely because that task sits where context rot is least severe.

NoLiMa, an independently published benchmark (arXiv:2502.05167) that requires a model to connect two related facts rather than match one directly, makes this concrete with numbers from a separate research team. At 32,000 tokens, requiring one additional reasoning hop drops GPT-4o from 99% to 70% accuracy, Claude 3.5 Sonnet from 88% to 30%, Gemini 2.5 Flash from 94% to 48%, and Llama 4 Scout from 82% to 22%.

NoLiMa two hop reasoning accuracy drop at 32000 tokens for GPT-4o, Claude 3.5 Sonnet, Gemini 2.5 Flash and Llama 4 Scout
Figure 2. NoLiMa (arXiv 2502.05167) found every tested model’s two-hop reasoning accuracy dropped sharply at 32,000 tokens.

Those numbers come from a task built specifically to require an extra reasoning step. Chroma’s own controlled experiments, run across the same 18 models under simpler conditions, isolate exactly which factors make an ordinary retrieval task behave this way too. What does Chroma’s own research actually show?

The primary source

What does Chroma’s own research actually show?

Distractors hurt more as context grows, and models handle them differently by family, and, most counterintuitively, models perform better on a randomly shuffled, incoherent haystack than on one with a natural logical flow. These findings come from controlled experiments designed specifically to isolate input length as the only variable, holding task difficulty constant.

Even a single distractor, a piece of information that is topically related to the correct answer but doesn’t actually answer the question, measurably reduces accuracy relative to no distractor at all, and adding several compounds the effect further as input length grows. The models don’t fail the same way, either: Claude models tend to abstain when uncertain, explicitly stating that no answer can be found, while GPT models show the highest rates of confidently stating an incorrect answer pulled from one of the distractors. Whether an agent that fails silently or one that fails loudly is the bigger production risk depends entirely on how failures downstream get handled, which makes this a genuine design consideration rather than a minor detail.

The haystack-structure finding is the most surprising result in the report. The intuitive expectation is that a needle placed in a coherent, logically flowing document would stand out and be easier to find than the same needle buried in randomly shuffled, unrelated sentences. Chroma found the opposite, consistently, across all 18 models tested: shuffling the haystack and removing its logical coherence improved retrieval performance. The mechanism isn’t fully explained, but it suggests attention is influenced by structural patterns in the input in ways that don’t match a simple, human-intuitive model of “obvious anomalies are easy to spot.”

A separate experiment in the same report tests something more basic than retrieval: exact replication. Given a long sequence of a repeated word with one unique word inserted somewhere in it, and asked to simply reproduce the text exactly, models still degrade as the sequence grows, even though the task requires no reasoning or search at all. Failure modes vary by family in ways that matter for production reliability rather than raw accuracy. GPT-4.1 refused the task outright in 2.55% of attempts once sequences passed roughly 2,500 words, typically responding with a flat refusal rather than an attempt. Claude Opus 4 showed the slowest overall degradation in its family but was the only Claude model to refuse outright, in 2.89% of attempts, sometimes citing a concern about reproducing copyrighted material even for an obviously non-copyrighted repeated word. Gemini and Qwen models were more likely to generate words never present in the input at all, effectively hallucinating content into a task that should have had zero room for invention. The whole evaluation covered 18 models across 8 input lengths and 11 needle positions per configuration, 194,480 individual model calls in total, which is the scale it took to establish that these patterns are consistent rather than noise.

Every experiment described so far uses a synthetic document as the haystack. The report includes one experiment that uses something closer to what this site’s readers actually build: a long conversation history. What does this mean for a conversation with a long history?

Conversational memory

What does this mean for a conversation with a long history?

Every model tested performs significantly better when given only the relevant part of a conversation than when given the same facts inside a full, ~113,000-token history, and Claude models show the largest gap of all, largely because they abstain under the added ambiguity rather than guess. This is the one experiment in Chroma’s report that speaks directly to conversational memory rather than generic document QA.

Using LongMemEval, a benchmark built specifically for testing long-term conversational memory, Chroma compared a focused input, roughly 300 tokens containing only what’s needed to answer a question, against a full input containing the entire ~113,000-token conversation history the question was drawn from. Every model family performed measurably worse on the full input, because it now has to do two things at once, find the relevant part of the history and reason over it, where the focused version only requires the second. Related information also matters independently of the raw fact: this compounds directly with the “lost in the middle” finding from a 2023 Stanford study (arXiv:2307.03172), which found that a model retrieves information placed at the very start or end of a context far more reliably than the same information placed in the middle, a positional bias that gets worse, not better, as the context grows longer.

This is also exactly the argument for external memory over stuffing a full history into every prompt: retrieving only the relevant slice of a conversation, rather than the entire transcript, is functionally the same move as Chroma’s focused-versus-full comparison, and it produces the same measured benefit. The general case for that architecture is made on memory versus the context window; this page’s job is narrower, showing the specific degradation that architecture is built to avoid.

Question type matters within the full-context condition too. Non-thinking models generally handled knowledge-update questions best, then multi-session questions, then temporal-reasoning questions worst, in both the focused and full-input conditions. Enabling extended thinking narrowed the gap between focused and full performance for models that support it, but did not close it: even with full reasoning capability switched on, the same models still performed measurably worse on the full 113,000-token input than on the 300-token focused version of the identical question. Thinking helps a model use a long context more effectively; it does not make the underlying attention-budget constraint go away.

Everything so far treats context rot as an accuracy problem: a worse answer, a missed detail. In at least one documented case, the consequence is considerably more serious than that. Why does this matter beyond a slightly worse answer?

The stakes

Why does this matter beyond a slightly worse answer?

Because the same mechanism that makes a model less accurate at finding a fact also makes an AI safety monitor less reliable at catching a dangerous action, and the measured effect is not small: 2x to 30x more misses once the transcript being reviewed contains hundreds of thousands of tokens of benign activity. This is a documented, higher-stakes version of the exact same underlying phenomenon.

AI safety monitors miss dangerous agent actions 2x to 30x more often after 800000 tokens of benign activity
Figure 3. A 2026 arXiv paper found safety monitors reviewing long agent transcripts miss dangerous actions far more often once preceded by hundreds of thousands of tokens of benign activity.

A 2026 arXiv paper studied language models used as automated monitors, reading a full agent transcript and flagging actions that weaken security, sabotage safety work, or exfiltrate data, the same kind of monitoring frontier labs use internally to oversee autonomous agents. Testing frontier models in that monitor role, the researchers found recall on a dataset of subtle attacks dropped from 98.6% to 88% once 800,000 tokens of benign activity were prepended to the transcript, and for a blatantly obvious dangerous action inserted into an otherwise long transcript, recall fell from 99.7% at 100K tokens to 69% at 800K tokens. The paper’s own conclusion is blunt: monitor evaluations that don’t account for context length are likely overestimating how much safety the monitoring system actually provides, and the longest sessions, exactly where an agent has the most autonomy and the most opportunity to cause harm, are the ones where this failure mode is worst.

A less dramatic but far more common version of the same failure shows up in ordinary agent-building work, and it’s worth seeing what it actually looks like in practice before deciding how to avoid it. How do you actually avoid it?

Mitigation

How do you actually avoid it?

Retrieve relevant context into a bounded window instead of accumulating everything, and scope sessions deliberately rather than letting one conversation run indefinitely. Both levers attack the same root cause from different angles: the first keeps the token count a model actually has to reason over small, and the second keeps unrelated history from ever entering the context in the first place.

Three ways to avoid context rot: retrieve instead of accumulate, scope sessions deliberately, watch aggregation tasks closely
Figure 4. Retrieval based memory and deliberate session scoping are the two practical levers for avoiding context rot.

One developer’s account of building with a coding agent illustrates what this looks like when it goes wrong: three hours into a session building platform-specific RSS feeds, automatic context compaction summarized the accumulated conversation, flattened old and new decisions together, and the agent came back proposing to build an RSS feed system, the exact system already finished earlier in the same session. The fix that developer settled on afterward is the same principle in practical form: restart sessions deliberately rather than letting one run indefinitely, keep durable facts in files or a memory store rather than relying on conversation history to hold them, and avoid routing unrelated tangents through a session that’s meant to stay focused on one task, since every off-topic exchange spends tokens that later crowd out what actually matters.

At an architectural level, this is exactly the argument for retrieval-based memory over full-history context stuffing, discussed generally on the context window problem and applied to runtime design on memory management. For teams that would rather not build and tune that retrieval layer themselves, a managed option such as Engram handles the extraction and retrieval decisions as a service, which sidesteps the manual session-scoping discipline described above by keeping the context an agent actually sees small on every turn, regardless of how long the underlying interaction history grows.

Video

What is LLM context rot, in four minutes?

A short walkthrough of the phenomenon covered above, for anyone who’d rather watch the explanation than read it.

FAQ

Frequently asked questions

The practical questions that follow once the mechanism above is understood.

Do million-token context windows fix context rot?

No. Chroma's research tested models with context windows well into the hundreds of thousands of tokens and still found consistent degradation as input length grew, on tasks as simple as retrieving one fact. A bigger window means more room, not more reliable use of that room.

Is context rot the same thing as the lost-in-the-middle problem?

They're related but not identical. Lost-in-the-middle refers specifically to positional bias, information in the middle of a context being retrieved less reliably than information at the start or end. Context rot is the broader phenomenon: performance degrading as input length grows, including on tasks where position isn't the main factor, like exact text replication.

Does extended thinking or reasoning mode fix context rot?

It narrows the gap but doesn't close it. Chroma's LongMemEval results showed that models with thinking enabled still performed worse on a full, long input than on a focused version of the same question, even though thinking improved both conditions.

Why do Claude and GPT models fail differently under context rot?

Claude models tend to abstain under ambiguity, stating that an answer can't be determined, while GPT models are more likely to generate a confident but incorrect answer pulled from a distractor. Neither is unconditionally safer; which failure mode is worse depends on how your system handles a wrong answer versus a refusal.

Does summarizing a long conversation prevent context rot?

Partially, and it introduces its own risk. A joshowens.dev account of a coding-agent session found that automatic summarization can flatten distinct, even contradictory, pieces of history into one summary that loses which decision was actually current, causing the agent to propose redoing work already finished.

Is context rot only a chatbot problem?

No. A 2026 paper found the same mechanism degrades AI safety monitors reviewing long autonomous-agent transcripts for dangerous actions, with recall dropping substantially once hundreds of thousands of tokens of benign activity precede the action being checked for.

Continue exploring

Three routes onward: the general context-window argument, the runtime architecture that acts on it, and how to test whether it’s actually happening to you.