Compare · Context windows

Long Context vs Memory: Do You Still Need Memory?

Million-token context windows do not replace AI agent memory: they cost more, suffer context rot and forget across sessions; external memory gives durable, selective recall at sustainable cost.

Long context

~26,000 tokens/query, session-bound

External memory

~1,800 tokens/query, cross-session

Long context

What do long context windows actually provide?

Entire conversation or document corpus in the prompt for one inference, no retrieval engineering for single-session tasks.

Providers offer 128K to 1M+ token windows. Strengths: full text available in one call. Limits: per-call cost scales with all history, latency grows, and it stays session-bound unless you re-send everything every turn. The deeper technical reasons this doesn’t scale, the quadratic attention cost and the real gap between a model’s stated and effective context size, are covered in full on the context window problem rather than repeated here.

→ Memory vs context window

Beyond the math

Why doesn’t a bigger window solve the engineering problem?

Even a free, unlimited context window wouldn’t answer 5 questions only an application’s memory policy can answer.

Oracle’s own engineering research frames this well: more context means the model can see more text at once, which helps for long documents and extended single sessions, but it never answers the harder questions a real deployment actually faces. Which facts should survive across sessions? Which older details are still relevant? Which decisions are authoritative when 2 memories disagree? Which prior failed attempts shouldn’t be repeated? Which memory belongs to this user, this project, or this task, rather than someone else’s? A bigger window gives you more room. It does not give you a policy for what to store, what to summarize, what to retrieve, what to trust, and what to actually pass into the model for a specific turn. That policy has to come from the application architecture, not the context window size.

Five policy questions a bigger context window never answers, regardless of size
A bigger window gives you more room; it doesn’t give you a memory policy.

External memory

What does external AI memory provide instead?

A selective store outside the window, retrieving only relevant slices each turn.

Cross-session persistence, lower per-turn tokens, structured update and forget, and benchmark-proven recall. Engram, Mem0 (LOCOMO J 66.9), Zep (66.0) and Letta (DMR 93.4%) implement external memory patterns. No single memory layer does everything, though: Oracle’s own framework maps 6 distinct approaches to 6 distinct continuity needs, each with an honest weakness. Sliding-window memory keeps recent turns coherent but drops older context entirely. Summarization compresses older dialogue but can lose detail or drift. Vector retrieval recalls semantically related history even when the user paraphrases, but similarity isn’t the same as correctness, a retrieved chunk can be related yet outdated. Structured memory holds precise facts, preferences and decisions that the application can query and govern directly. Episodic memory tracks what happened and what was already tried, so an agent doesn’t repeat a failed approach. A memory manager coordinates all of the above, deciding per turn which pieces actually belong in front of the model; none of the 5 storage layers work well without one.

→ How AI memory works

Comparison

How do long context and memory compare, dimension by dimension?

Use long context for single-session depth; memory for persistence and cost.

DimensionLong context onlyExternal memoryCombined
Cross-sessionRe-send all historyNative persistenceMemory + recent context
Cost at scaleO(all history) tokensO(retrieved chunks)Lowest at scale
Recall precisionDegrades (context rot)Selective top-kBest of both
Context rot riskHigh on long threadsLow (curated injection)Managed via engineering
Setup complexityLowMediumMedium to high
Best forSingle-session depthMulti-session agentsProduction assistants

Beyond facts

Is remembering facts and preferences actually enough?

Not for repeated operational work: treating a context window like RAM means every session re-pays the cost of re-teaching an agent things it already handled.

One engineering essay names this the “Context Tax”: a recurring overhead, paid in both latency and tokens, of re-explaining infrastructure, tools and prior successful approaches at the start of every single run, because the agent’s context resets like volatile memory rather than persisting like storage. Ordinary fact-based memory (preferences, identity, past conversations) doesn’t fully solve this either, since the problem here isn’t forgetting who the user is, it’s re-planning a workflow from scratch every time even after solving it successfully before. The proposed answer is treating a proven successful procedure as its own kind of memory: once a multi-step workflow has been solved and verified, crystallize it into a versioned, reusable procedure instead of re-deriving it, so a repeated task stops paying the full reasoning cost on every run. The source itself reports this cutting token consumption by over 90% after the first run pays the one-time “exploration cost,” a claim worth treating as this blog’s own, not an independently verified benchmark, but the underlying idea, procedural memory as a third answer alongside long context and fact memory, holds regardless of the exact number. See procedural memory for how this pattern works in more depth.

The Context Tax: re-paying the cost of re-explaining infrastructure and workflows every session when context resets like volatile memory
Neither a bigger window nor fact-based memory fixes repeated re-planning; procedural memory does.

Context rot

What is the context rot problem?

Accuracy degrades as context grows: needle-in-haystack failures and attention dilution.

Even with million-token windows, models miss facts buried in long histories. LOCOMO benchmarks long-context agents against selective memory retrieval; external memory consistently wins on multi-session recall at lower token cost (Mem0 ~1,800 vs ~26,000 tokens, Chhikara et al., 2025).

→ Context rot

Cost

How does cost actually compare, tokens vs retrieval?

Long context: O(all history) tokens per turn. Memory: O(retrieved chunks) tokens.

Mem0 p95 total latency is 1.44 s vs 17.1 s full-context (Chhikara et al., 2025). Letta paging (MemGPT) offers a third path: virtual context without sending everything.

→ Reduce token cost with memory · Virtual context and MemGPT

Long context enough

When is long context alone enough?

Single-session tasks; corpus fits budget; no cross-session personalization; prototype phase. Examples: one-shot doc Q&A, single coding session. A useful litmus test either way: would a human expert be better at this specific task if they remembered previous interactions with this user? If yes, you likely need memory. If no, keep it simple, a stateless agent is easier to build, test and debug, and that simplicity has real value.

Need memory

When do you need external memory?

Five signals, any one of which is usually enough to justify the added engineering.

  • Multi-session users
  • Personalization and CRM history
  • Cost control at scale
  • Selective recall requirements
  • LOCOMO/LongMemEval-proven LTM

→ Why agents need memory

Combined

How do you combine long context and memory in practice?

Memory retrieves relevant history, injects it into the context window, and the model reasons on a curated subset.

Best practice: Engram, Mem0 or Zep retrieve top-k facts; recent messages fill the remaining budget; RAG adds org docs. Letta paging is the advanced form (MemGPT DMR 93.4%, Packer et al., 2023). When 2 layers disagree, a simple preference order avoids most conflicts: trust a structured decision over a summary describing the same fact, trust newer memory over older when both carry equal authority, and trust scoped memory (project-specific, user-specific) over generic memory. Tagging each memory record with a state, active, superseded, rejected or archived, rather than deleting it outright, keeps the system debuggable when an agent gives a wrong answer and a developer needs to trace which layer supplied the evidence. See conflicting memories for the deeper treatment of resolving contradictions once they occur.

→ RAG with memory · Letta alternatives

Six memory layers mapped to six continuity needs: sliding window, summarization, vector retrieval, structured memory, episodic memory, and a memory manager
No single layer covers every continuity need; a memory manager decides which pieces belong in front of the model each turn.
Conflict resolution order when memory layers disagree: structured over summary, newer over older, scoped over generic
A simple preference order, plus tagging records as active, superseded, rejected or archived, keeps a combined system debuggable.

FAQ

Frequently asked questions

The policy questions bigger context doesn’t answer, and how to combine both layers.

Do 1M token context windows replace agent memory?

No. They are session-bound, expensive at scale and suffer context rot. External memory provides cross-session persistence at ~1,800 vs ~26,000 tokens per query (Mem0, Chhikara et al., 2025).

What is context rot?

Accuracy degradation as context grows: models miss facts buried in long histories. External memory injects only relevant slices. See context rot.

Claude memory vs long context?

Claude's window is short-term memory for one session. Long-term memory requires external tools (Engram, Mem0). See persist conversation memory.

MemGPT vs long context?

MemGPT pages archival memory in and out, virtual context without sending the full history. DMR 93.4% (Packer et al., 2023). See virtual context.

Cost comparison: long context vs memory?

Mem0 uses ~1,800 vs ~26,000 tokens per LOCOMO query; p95 latency 1.44s vs 17.1s full-context (Chhikara et al., 2025).

What does LOCOMO say about long context vs memory?

LOCOMO benchmarks multi-session recall. Mem0 J 66.9, Zep 66.0, LangMem 58.1: selective memory outperforms full-context baselines on long-horizon tasks.

RAG or memory or both?

Both, for production agents. RAG handles org docs; memory handles per-user facts. See memory vs RAG.

Engram vs long context?

Engram retrieves selective Weaviate memories at low token cost instead of resending the full history. Pair with recent context in the window. See Engram explained.

When to use Letta over long context?

When you need MemGPT-style paging: core memory in-context, archival memory paged on demand. See Letta alternatives.

Best hybrid long context + memory architecture?

Retrieve top-k from Engram, Mem0 or Zep, inject with recent messages, add RAG for org docs, apply context engineering for the token budget. See context engineering.

Is fact-based memory enough, or do agents need something more?

For repeated operational tasks, often not. One engineering framing calls repeatedly re-explaining solved workflows a "Context Tax." Procedural memory, crystallizing a proven workflow into a reusable procedure, addresses that gap; ordinary preference and fact memory doesn't. See procedural memory.

What happens when memory layers disagree with each other?

A simple order avoids most conflicts: trust structured decisions over summaries, newer memory over older, and scoped memory over generic. Tagging records as active, superseded, rejected or archived keeps the system debuggable. See conflicting memories.

Continue exploring

Three routes onward: the full memory-vs-context comparison, the best tools, and building long-term memory.