Compare · Memory vs context window

AI Memory vs Context Window: What Is the Difference?

A context window is a model’s temporary working memory for one session, bounded by a token limit and gone when the session ends. Memory is external storage that persists across sessions and is reached by an explicit search. They are not competitors: memory exists to decide what goes into the window.

Two different things

1
Window
This session
2
Memory
Across sessions
3
Together
Always

The distinction

What is the difference between memory and the context window?

The context window is the finite amount of text a model can consider in a single call, measured in tokens. Memory is a store outside the model that survives beyond the call and has to be searched to be used. One is a capacity; the other is a system.

Short-term context window memory compared with the long-term external store, by speed, size and whether it survives the session.
Figure 1. The model reads only the left-hand side. Everything on the right has to be retrieved into it before it can be used.

Three properties separate them and all three matter in a build. Persistence: the window is discarded when the session ends, while memory is written to a database and is still there next month. Access: the model reads the window directly on every forward pass, while memory only exists for the model if something searched for it and placed it in the prompt. Cost shape: filling the window costs tokens on every single call, while a memory store costs storage plus one search.

The window is sometimes described as short-term memory, and that is a fair analogy as long as one thing is clear: it is not a store the system writes to deliberately. Text enters the window because someone assembled a prompt containing it, which means the interesting engineering is always in the assembly rather than in the window itself.

That framing also explains the relationship people most often get backwards: how the two work together.

The relationship

How do memory and the context window work together?

Memory’s entire job is to put the right things into the context window at the right moment. They are not alternatives to choose between; one is the delivery mechanism for the other.

The read path in five stages: embed the query, filter by scope, search, rank, and inject into the prompt.
Figure 2. Every retrieval pipeline ends in the same place: a decision about what occupies the window.

Follow one turn through. The user asks a question. Before the model is called, the system searches memory, scoped to that user, ranks what comes back, and selects a handful of memories. Those are written into the prompt alongside the system instructions and the recent conversation. The model then reads the window, which now contains a small amount of carefully chosen history, and answers.

Two design consequences follow immediately. Every memory injected is a slot the rest of the prompt does not get, so retrieval is a budgeting exercise rather than a search problem. And the model cannot tell the difference between a memory retrieved from three months ago and something the user said thirty seconds ago, because both arrive as text in the same window.

That second point is the source of a subtle failure. A retrieved memory presented without its age can be treated as current, which is why stored facts need timestamps and why superseded facts have to be filtered out before injection rather than after, covered on handling conflicting and stale memories.

Given the window is the bottleneck, the obvious question is whether making it larger removes the need for the whole apparatus: whether a bigger window replaces memory.

The obvious question

Does a bigger context window replace memory?

No, for three reasons: the window still ends with the session, filling it costs tokens on every call, and models attend badly to content buried in the middle of a very long prompt. A larger window raises a ceiling; it does not add persistence.

The persistence point is the decisive one and the most often overlooked. A window of a million tokens is still empty at the start of the next session. Whatever the user said last week has to be resupplied by something, and that something is a memory system regardless of how much room there is to put it.

The attention point is empirical rather than theoretical. “Lost in the Middle” by Liu et al. (arXiv:2307.03172) shows that models degrade on information positioned in the middle of a long context even when it fits comfortably inside the window. Fitting is not the same as attending, so a longer prompt is not reliably a better-informed one.

The cost point is arithmetic. Resending a conversation means paying for the whole transcript on every turn, and that grows without limit as the conversation does. Selective retrieval sends a roughly constant amount regardless of relationship length: Mem0 reports about 1,800 tokens per query against 26,000 for a full-context baseline, with p95 latency of 1.44 seconds against 17.1 seconds (Chhikara et al., 2025).

The full argument, including where long context genuinely wins, is on long context versus memory.

Which raises the practical question of what happens when the window does fill: what happens when a context window is full.

The limit

What happens when the context window fills up?

Something has to leave, and the system’s choice about what leaves is the whole design. Three strategies exist and they lose different things.

Four stages of context pressure: the queue fills, a warning fires at 70 percent, the agent writes memories out, then eviction and recursive summary at 100 percent.
Figure 3. MemGPT’s approach: warn at 70% so the agent can write anything important to external memory before the window is full.

Truncation drops the oldest turns. It is the simplest and it loses the earliest part of a conversation silently, which is usually where the user explained what they actually wanted.

Summarisation replaces old turns with a shorter summary. It keeps the gist and loses detail by construction, and the loss compounds when summaries are themselves summarised later.

Paging writes evicted content to external storage and retrieves it when needed. It is the only approach that can recover the original, and it is the design MemGPT introduced, described on virtual context and MemGPT.

The important observation is that the first two are what happens by default when nobody chooses, and both degrade the conversation invisibly. There is no error when a window overflows; there is only an agent that has quietly stopped knowing something. That is why eviction is a policy to design rather than a limit to hit, covered on forgetting and eviction.

People reasonably ask how large windows have become, and the answer is less decisive than it sounds: which model has the longest context window.

The specification

Which model has the longest context window?

The answer changes every few months, which is itself the useful finding: context length is a specification vendors compete on and a poor basis for an architecture decision. Windows have gone from a few thousand tokens to hundreds of thousands, and the leading figure has changed hands repeatedly.

Three things are worth knowing instead of a number. The advertised maximum is not the usable maximum. Filling a very large window costs proportionally more on every call, and attention quality degrades over long spans regardless of the stated limit. Effective recall is measured, not specified. That is exactly what benchmarks like LongMemEval and LOCOMO exist to test, covered on LongMemEval and LOCOMO. The window is per call, not per relationship. However large it is, it starts empty at the next session.

The practical consequence for a build is that a system designed around one model’s window becomes fragile when the model changes, and models change often. A system designed around retrieval works across model generations, because it sends a small, curated prompt whatever the ceiling happens to be. That portability is an underrated argument for memory and it does not appear on any specification sheet.

Where a very large window genuinely helps is a bounded, one-off task: analysing a single long document, reviewing a codebase in one pass, summarising a contract. Those are cases where everything relevant genuinely exists at once and no continuity is required, and paying to fill the window is the right call. The comparison in full is on long context versus memory.

One recurring confusion is worth clearing up, because it sends people to entirely the wrong answers: memory here does not mean RAM.

Disambiguation

Is AI memory the same as RAM?

No. Memory in this sense is a software system for storing and retrieving facts, not the physical RAM a machine needs to run a model. The two get conflated because both are called memory and because the field borrows RAM as an analogy for the context window.

The hardware question, how much RAM or VRAM is needed to run a given model, is about serving infrastructure: model weights, batch size and quantisation. It matters if you are self-hosting and is irrelevant to whether your agent remembers a user’s preferences. An agent calling a hosted API uses no local RAM for the model at all and can still have excellent memory.

The analogy that causes the overlap is a good one and worth keeping, as long as its limits are clear. Treating the context window as RAM and external storage as disk is exactly how virtual context management is framed, and it accurately describes the speed and scarcity relationship. What it does not mean is that adding physical memory to a server enlarges a model’s context window, which is set by the model architecture rather than by the machine.

A rough test: if the question is about gigabytes, it is hardware. If it is about what the assistant remembers about a person, it is the subject of this site. The types involved are covered on the types of AI agent memory, and the human comparison on AI memory versus human memory.

With the terms separated, the practical question becomes which one your problem needs: when you need memory rather than a bigger window.

The decision

When do you need memory rather than a bigger window?

Whenever anything has to survive the end of a session, and whenever the cost of resending history has started to matter. Those two tests settle almost every case.

  • The user comes back. Any product where a returning person expects to be known needs persistence, and no window size provides it.
  • Conversations run long. Once a session regularly approaches the window, you are already choosing what to drop, and choosing deliberately beats truncating silently.
  • Token cost is visible. When resending the transcript is a noticeable line item, selective retrieval pays for the engineering quickly.
  • Facts change. A window holds whatever was pasted into it. A memory system can mark a fact superseded, which a prompt cannot do.

The inverse is equally worth stating. If your agent completes its work inside one session and never needs to recognise a returning user, the context window is already sufficient, and adding memory adds cost, latency and a new class of bug. That case is made on why AI agents need memory, including where a stateless design is the right answer.

A sensible progression for most teams is to start inside the window, resending history while conversations are short, and add memory when one of the four tests above starts firing. Building it before the need exists is how simple products acquire a store nobody maintains.

Once memory is needed, the build is short and set out step by step in how to add memory to an AI agent.

The discipline

What is context engineering, and how does it relate?

Context engineering is the practice of deciding what occupies the window on each call, and memory is one of its inputs. The name emerged because prompt engineering stopped describing the job once prompts became assembled rather than written.

A production prompt is now a composition: system instructions, retrieved memories, retrieved documents, tool definitions, recent conversation, and the user’s actual question. Each competes for the same finite space, and the assembly logic that arbitrates between them is where a great deal of quality lives. That logic is not a prompt; it is code with a budget.

Memory’s contribution to that budget is unusual because it is unbounded on the input side. A document corpus has a natural retrieval limit and a conversation has a natural length, but a memory store keeps growing for as long as the relationship lasts. Deciding how many memories are worth their slots, on every turn, is the part that has to be measured rather than assumed.

The related failure is worth naming: filling the window with everything available produces worse answers than filling it selectively, which is the same “lost in the middle” result cited above. More context is not better context. The discipline is covered on what context engineering is and the degradation on context rot.

Which turns the abstract question into a concrete one every build has to answer: how much of the window memory should occupy.

The budget

How much of the context window should memory use?

Less than teams expect, and the number should be set deliberately rather than left to a library default. Every memory injected competes with the system instructions, the retrieved documents, the tool definitions and the actual conversation, all of which have a better claim on the space than a marginally relevant fact from last spring.

A workable way to think about it is as fixed allocations rather than a single limit. The system instructions and tool definitions are effectively fixed overhead. The recent conversation is whatever the current exchange requires. What remains is discretionary, and memory competes with document retrieval for it. Setting explicit caps on each, rather than letting whichever pipeline runs first consume the space, is what keeps behaviour predictable as a product grows.

The counter-intuitive part is that a smaller memory allocation frequently produces better answers. Injecting three well-ranked memories beats injecting fifteen, because the fifteen include twelve that are merely topical and each one is a chance for the model to anchor on the wrong detail. This is the same effect Liu et al. documented for long contexts generally, and it applies just as much to a prompt padded with retrieved memories as to one padded with documents.

The practical method is to measure rather than reason about it. Take a set of questions whose answers depend on stored facts, vary the number of injected memories, and look at where answer quality stops improving. That number is usually smaller than the default, and it is specific to your data. The ranking that decides which memories survive the cut is covered on scoring and ranking memories.

Budgeting gets harder in agentic loops, where the window is consumed by something other than conversation: what happens to the window during multi-step work.

Agentic work

What happens to the context window during multi-step work?

It fills with tool output, which is the fastest way to exhaust a window and the case most memory discussions ignore. A conversational assistant adds a few hundred tokens per turn. An agent that reads files, queries APIs and runs searches can add thousands per step, and it takes many steps.

The result is that agentic systems hit the window limit during a single task rather than across a relationship. Halfway through a long piece of work, the early steps, including the original instruction, are the oldest content in the window and the first candidates for eviction. An agent that truncates naively will forget what it was asked to do while still busy doing it.

Three responses are used in practice. Summarise tool output before it enters the window, keeping the result and discarding the raw payload. Write intermediate findings to memory and retrieve them when needed rather than carrying everything forward. Page deliberately, which is the MemGPT approach of moving content to external storage under the agent’s own control, described on virtual context and MemGPT.

All three are the same insight in different clothing: in a long task the window is a working surface rather than a record, and anything that must survive the task belongs in a store. That is why agentic systems tend to need memory earlier and more urgently than chat products do, covered on why AI agents need memory and agentic application architecture.

The underlying constraint that all of this exists to manage is set out on the context window problem.

The four operations of the AI memory loop: write and extract, store, retrieve and rank, then update or evict.
Figure 5. Memory is not a bigger window. It is the loop that decides, on every turn, what the window should contain.

One last practical note for anyone reading vendor material on this topic. Claims about a model’s memory almost always mean one of the two things on this page, and they are rarely labelled. A larger context window is a capability of the model. Persistent recall across sessions is a capability of the system built around it. A product announcement that mixes the two is describing a bigger window and calling it memory, and the test is simple: ask whether it still knows anything after the session ends.

The same test applies to your own build. If a feature is described internally as memory but disappears when a session ends, it is a prompt-assembly feature and should be named as one.

A closing way to hold the distinction. The context window is a property of the model you are calling, fixed by whoever trained it, identical for every user, and reset on every request. Memory is a property of the system you are building, decided by you, different for every user, and persistent by design. Confusing the two leads teams to solve a system problem by shopping for a model, which is the most expensive way to not fix it.

FAQ

Frequently asked questions

The questions that follow: which model has the longest window, and whether memory reduces token cost.

Is the context window the same as memory?

No. The context window is volatile working memory for the current inference — it clears when the session ends. AI memory persists in external stores across sessions. See comparison table above.

What is working memory in AI agents?

Working memory is the information actively in the context window for the current task — recent turns, tool outputs and retrieved memories. It maps to short-term memory in the agent taxonomy. See short-term memory.

What does context memory mean?

Often confused with AI memory. 'Context memory' usually means information in the context window (working memory), not persistent cross-session storage. For durable memory, use an external store. See how AI memory works.

What is role context memory?

In agent frameworks, role context is the system prompt and persona instructions in the context window — not persistent user memory. Durable facts about users belong in external memory stores.

Does a 1M-token context window replace memory?

No. Even 1M-token windows are session-bound, costly per token and suffer context rot. Memory provides cross-session persistence, selective retrieval and structured update/forget. See long context vs memory.

How do I persist conversation beyond the context window?

Extract salient facts after each turn to an external memory store; retrieve top-k memories before each response and inject into the window. See persist conversation memory.

What is the difference between context engineering and memory?

Context engineering curates what goes into the window each turn. Memory is the external system that stores and retrieves facts across sessions. They work together. See context engineering.

Is Claude's memory the same as the context window?

Claude's product memory is a managed feature for that app. The API context window is still the working-memory limit per call. Custom agents need their own memory layer. See add memory guide.

How does Letta relate to the context window?

Letta (MemGPT) treats the context window like RAM and pages memories to a deep store — effectively unbounded history without stuffing every token into the window. See virtual context and MemGPT.

What is the best production pattern?

External memory store + selective retrieval into the context window. Validate with LOCOMO/LongMemEval on your domain. See add memory to an agent and evaluation hub.