Architecture · MemGPT
What Is Virtual Context? MemGPT and Memory Paging for LLMs
Virtual context management lets an LLM agent page information in and out of its context window the way an operating system pages between RAM and disk. MemGPT introduced the technique in 2023, raised deep memory retrieval accuracy from 32.1% to 92.5% on GPT-4, and became the Letta framework. This page explains the mechanism, the benchmark numbers with the models pinned, and what the design still does not solve.
The three memory tiers
Definition
What is MemGPT?
MemGPT is a system that gives an LLM agent effectively unbounded context by letting the agent manage several memory tiers itself, moving information between the context window and external storage through function calls. It was introduced in the 2023 research paper “MemGPT: Towards LLMs as Operating Systems” by Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica and Joseph E. Gonzalez at UC Berkeley, published as arXiv:2310.08560.
What MemGPT does, stated plainly: it decides what the model sees. A conversation that outgrows the context window normally loses its earliest turns, and MemGPT keeps them reachable by writing them to external storage and retrieving them on demand.
The name now carries two meanings, and the distinction matters when you search for it. MemGPT is the research system; Letta is the open-source framework the same team built from it. Documentation, releases and installation instructions live under Letta, while the paper, the benchmark numbers and the original design vocabulary remain under MemGPT.
The technique that makes the decision possible has its own name, and it is the term the paper leads with: virtual context management.
Mechanism
What is virtual context management?
Virtual context management is a technique that borrows from hierarchical memory in operating systems, where moving data between fast and slow memory creates the appearance of a much larger memory resource than physically exists. The MemGPT paper states the analogy directly: the system provides extended context within the LLM’s limited context window by intelligently managing different memory tiers, and uses interrupts to manage control flow between itself and the user.
The mapping is exact enough to be useful. The context window plays the role of RAM: fast, scarce, and the only tier the model reads directly. The external stores play the role of disk: slower to reach, effectively unbounded, and reachable only through an explicit request. The agent plays the role of the operating system, deciding what to load and what to evict.
The consequence is the part worth remembering. A model using virtual context does not get a bigger window, it gets a policy for what occupies the window. That is why the technique keeps working as context windows grow, and why it belongs to the context window problem rather than to any one model generation.
The operating-system comparison is not decoration, and the paper’s title states it as the thesis. It is worth taking literally for a moment, because it answers a question people ask directly: what an LLM operating system actually is.
The analogy
What is an LLM operating system?
An LLM operating system is a layer that manages a model’s limited resources on its behalf, exactly as an operating system manages a computer’s: it decides what occupies the scarce fast memory, moves the rest to slower storage, and interrupts execution when something needs attention. The model becomes the processor rather than the whole machine.
Three properties of a conventional operating system carry over, and the paper uses all three.
- Virtual memory. The prompt is treated as physical memory and external stores as disk, so the agent appears to hold far more than it can physically fit.
- Interrupts. The agent can pause to run a memory operation before it answers, instead of being forced to respond in a single pass. This is what the paper means by control flow between the system and the user.
- System calls. Reading and writing memory is a function call the model issues, which is why the pattern generalises to any model with reliable tool use.
Calling the model an operating system is a claim about responsibility rather than about capability: the agent now owns a scheduling problem it did not have before. Seeing how it solves that problem means looking at the three memory tiers it schedules across.
Architecture
What are MemGPT’s three memory tiers?
MemGPT divides memory into three tiers: main context, recall storage and archival storage. The three differ in what they hold, how the agent reaches them, and whether the model can see them without asking. The definitions below use the paper’s own names.
- Main context
- The prompt itself, and the only tier the model sees directly on any given inference call. It holds the system instructions, the working context the agent maintains about the user and task, and the most recent messages. It is bounded by the model’s context window.
- Recall storage
- The message database. It holds the full conversation history, including messages that have been evicted from main context, and the agent searches it with paginated queries to bring past exchanges back into the prompt.
- Archival storage
- A read and write database of arbitrary-length text objects, used for facts, documents and anything the agent chooses to keep beyond the conversation. In the paper’s experiments it is a vector store, implemented with the pgvector extension for PostgreSQL.
The agent moves data between tiers through ordinary function calls, which is why this design is often described as memory as a tool rather than memory as a database. That framing is covered separately in how memory works as a tool for agents, and the tiering idea generalises beyond MemGPT in what a tiered memory hierarchy is.
Tiers only matter at the moment the fast one runs out of room, so the design’s real test is what happens when the context window fills up.
Eviction
What happens when the context window fills up?
When the message queue reaches 70% of the context window, MemGPT inserts a system message warning the agent of memory pressure, so the agent can store anything important in external memory before it is evicted. The threshold is the most concrete number in the design: it converts “the window is nearly full” from an error state into a scheduled event the agent can act on.
Once the limit is reached, evicted messages are replaced by a recursive summary that keeps the gist of what was flushed. The full messages remain in recall storage and can be searched back into the prompt, so eviction removes them from the prompt rather than from the system.
The honest cost sits in the same mechanism. Recursive summarisation preserves the gist and loses detail by construction, and the paper acknowledges the tradeoff without quantifying how much is lost. Any agent that relies on summaries rather than on retrieval inherits that loss, which is why the eviction policy matters as much as the storage. The general version of this problem is covered in how agents decide what to forget.
That machinery exists to serve two workloads in particular, which is where MemGPT is actually used.
Applications
What is MemGPT used for?
MemGPT was evaluated on two workloads, and both are ones a fixed context window handles badly: perpetual conversation, and document analysis over a corpus larger than the window. The paper calls these conversational agents and document analysis, and they remain the two clearest reasons to adopt paging.
- Perpetual chatbots and companions. An assistant that runs for months has to stay consistent about what the user told it in week one. MemGPT keeps that history in recall storage and retrieves it when a later question depends on it, which is exactly what the benchmark below measures. The same requirement is described in memory for personal assistants.
- Large document analysis. Where a corpus exceeds the context window, the agent pages through archival storage instead of trying to hold everything at once, and can collate an answer across several retrievals. The paper introduces a nested key-value retrieval task specifically to test that multi-hop behaviour.
Support and coding assistants sit on the same requirement from a different angle, covered in memory for customer support agents and memory for coding agents.
Both workloads make the same claim, that the agent recalls what a fixed window would have dropped. The paper backs that claim with a benchmark it built for the purpose, and the numbers repay reading closely: what Deep Memory Retrieval measured, and what MemGPT scored.
Evaluation
What is the Deep Memory Retrieval benchmark, and how did MemGPT score?
Deep Memory Retrieval, usually shortened to DMR, is the task the MemGPT team introduced to test conversational consistency: the agent is asked a question that can only be answered from an earlier session, and its reply is scored against a gold answer. It runs on the Multi-Session Chat dataset from Xu et al. (2021), in which human labellers held five sessions of roughly a dozen messages each while playing a consistent persona. The MemGPT team added a sixth session containing a single question and answer pair, and scored responses with ROUGE-L recall alongside an LLM judge.
The table below gives every result from Table 2 of the paper. Each pair compares a fixed-context baseline against MemGPT running on the same underlying model.
| Underlying model | Fixed-context baseline | With MemGPT | ROUGE-L (baseline to MemGPT) |
|---|---|---|---|
| GPT-3.5 Turbo | 38.7% | 66.9% | 0.394 to 0.629 |
| GPT-4 | 32.1% | 92.5% | 0.296 to 0.814 |
| GPT-4 Turbo | 35.3% | 93.4% | 0.359 to 0.827 |
Two details decide how to read those numbers, and most summaries omit both. First, the model names are pinned to specific endpoints: GPT-4 is gpt-4-0613 with an 8,192 token window, GPT-4 Turbo is gpt-4-1106-preview with 128,000 tokens, and GPT-3.5 Turbo is gpt-3.5-turbo-1106 with 16,385 tokens. A 92.5% result against an 8k baseline is a different claim from the same result against a 128k baseline. Second, the baselines were not blind: they received a lossy recursive summary of the five prior sessions, while MemGPT had to retrieve the full history through paginated search. The comparison is between summarisation and retrieval, not between memory and nothing.
Later systems report higher figures on the same task. Zep, a temporal knowledge graph architecture for agent memory (Rasmussen, Paliychuk, Beauvais, Ryan and Chalef, arXiv:2501.13956), reports 94.8% against MemGPT’s 93.4%. The gap is real but narrow, and it is measured on a benchmark the MemGPT team designed, which is worth stating whenever the two are compared. Benchmarks for agent memory are covered in what the LOCOMO benchmark measures and which metrics matter for agent memory.
Every figure above sits in Table 2 of the paper rather than in its abstract, which matters if you go looking for them yourself: where to read the MemGPT paper.
Source
Where can you read the MemGPT paper?
The paper is “MemGPT: Towards LLMs as Operating Systems”, catalogued as arXiv:2310.08560, submitted on 12 October 2023 and revised to version 2 on 12 February 2024. Its seven authors are Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica and Joseph E. Gonzalez, working at UC Berkeley. The paper also released the augmented Multi-Session Chat dataset, the nested key-value retrieval dataset, and embeddings for 20 million Wikipedia articles.
One thing to know before you open it. The abstract contains no benchmark numbers at all. Every figure quoted on this page comes from Table 2 in the body of the paper, which is why searches for the DMR results so often land on pages that do not carry them.
Reading the paper answers what MemGPT is. Running it is a separate question, because the software people install today carries a different name: Letta.
Disambiguation
Is MemGPT the same thing as Letta?
No. MemGPT is the research system described in the paper, and Letta is the open-source agent framework that the same team developed from it. If you are looking for the concept, the tier vocabulary or the benchmark results, you want MemGPT. If you are looking for something to install and run, you want Letta.
The practical consequence shows up in searches for alternatives. People who ask for MemGPT alternatives almost always want a framework they can deploy, which makes the honest comparison set Engram (Weaviate), Mem0, Zep and LangMem rather than other research prototypes. That comparison is made in full on which Letta alternatives are worth using.
Before choosing any of them, most teams ask a blunter question first: whether a longer context window removes the need for paging at all.
Trade-offs
Does virtual context beat a longer context window?
A longer window and virtual context solve different halves of the problem, and a longer window alone does not remove the need for memory management. The empirical reason is documented in “Lost in the Middle” by Liu et al. (arXiv:2307.03172), which shows that models degrade on information buried in the middle of a long context even when that information fits inside the window. Fitting is not the same as attending.
Virtual context sidesteps that by retrieving only what is needed for the current step, which keeps the occupied portion of the window small and relevant. It carries its own cost: every archival lookup adds an inference call to generate the query plus a search against the store, and the paper reports no wall-clock latency comparison. On a long task those round trips accumulate.
The choice between the two is set out in whether long context replaces memory. If a longer window is not the answer, the next comparison is against the other way of putting memory into a prompt: a vector memory API.
Architecture class
How is virtual context different from a vector memory API?
A vector memory API retrieves into a prompt that someone else assembles, while virtual context makes the agent responsible for assembling the prompt. The distinction is about who holds the policy, not about which storage engine sits underneath, and both classes commonly use the same vector store.
The three architecture classes below cover most production systems in 2026.
| Class | Examples | Who decides what is in context | Best for |
|---|---|---|---|
| Virtual paging | Letta (MemGPT) | The agent, through tool calls | Unbounded multi-session conversation |
| Memory API | Engram (Weaviate), Mem0 | The application, through a retrieve call | Per-user facts and preferences |
| Temporal graph | Zep | The graph, through time-aware queries | Facts that change and relate to each other |
None of the three is strictly better. Paging gives the agent the most control and the most ways to waste tokens; a memory API gives the application predictable behaviour and less adaptivity. See how the memory APIs compare. Choosing between them is easier once the weaknesses of the paging design are on the table, and the paper is unusually clear about where its own limits sit.
Limitations
What are the limits of MemGPT’s design?
Four limits are visible in the paper itself, and they matter more than the headline accuracy figure when you are deciding whether to build on this design.
- The evaluation is narrow. Multi-Session Chat is built from human labellers playing consistent personas across five short sessions. Production conversations are messier, longer and less consistent, and the paper does not test them.
- Retrieval quality depends on the agent’s own query. Archival storage only returns what the agent thinks to ask for, so an agent that does not know what it is missing will never issue the lookup that would have found it.
- Latency is unreported. Each archival lookup costs an extra inference call plus a search, and the paper publishes no wall-clock comparison against a fixed-context baseline.
- Summarisation is lossy by design. Evicted content is compressed into a recursive summary, and the paper acknowledges without quantifying the detail lost in that step.
None of these invalidates the result. They set the conditions under which the 92.5% figure was earned, and those conditions are what a team should check its own workload against before adopting the pattern. The broader question of how to keep memory accurate over time is covered in how agents handle conflicting and stale memories.
Those same four limits are the agenda for everything published since, which is what came after MemGPT.
What came next
What came after MemGPT?
Three lines of work extend or replace the fixed-tier design: temporal knowledge graphs, agentic memory that organises itself, and background consolidation that runs while the agent is idle. Each one targets a limit named in the section above.
- Temporal knowledge graphs. Zep (Rasmussen, Paliychuk, Beauvais, Ryan and Chalef, arXiv:2501.13956) stores facts and their validity over time rather than paging message history, and reports 94.8% on the same DMR benchmark against MemGPT’s 93.4%. It addresses stale facts, which paging does not handle.
- Agentic memory. A-MEM (Xu, Liang, Mei, Gao, Tan and Zhang, arXiv:2502.12110, February 2025) lets the memory structure organise itself instead of relying on fixed tiers, and reports improvement over prior systems across six foundation models. The paper’s abstract states the improvement in general terms rather than as a single number.
- Background consolidation. Sleep-time compute moves summarisation and reorganisation off the critical path, so the latency cost the MemGPT paper never measured is paid while the user is not waiting. See how sleep-time compute consolidates memory.
Virtual context remains the reference design that the rest are measured against, which is why its vocabulary of main context, recall and archival storage still turns up in systems that do not page at all. If you are moving from reading to building, the next step is choosing where memory writes happen in your own agent loop, set out step by step in how to add memory to an AI agent and in how to build long-term memory for an LLM agent.
FAQ
Frequently asked questions
These questions cover what people ask once the mechanism is clear: licensing, how MemGPT sits against the memory APIs, and what the technique needs from the underlying model.
Is MemGPT open source?
Yes. The MemGPT research code was released with the paper, and the Letta framework built from it ships open-source components alongside managed hosting options. See open-source vs managed memory.
How does MemGPT compare with Mem0?
They belong to different architecture classes. MemGPT and Letta page content across memory tiers under the agent's control, while Mem0 is an extraction API that writes facts to a vector store the application queries. Paging suits unbounded conversation; an extraction API suits per-user facts. See Mem0 alternatives.
How does MemGPT compare with Zep?
MemGPT pages conversation history between tiers. Zep stores facts in a temporal knowledge graph so relationships and changes over time are queryable. Zep reports 94.8% on the DMR benchmark against MemGPT's 93.4% (Rasmussen et al., arXiv:2501.13956). See Zep alternatives.
Does Engram use MemGPT paging?
No. Engram (Weaviate) is a vector-native memory layer rather than a paging framework, so the agent does not manage tiers itself. Choose Letta when agent-controlled tier paging is a requirement. See Engram explained.
Can you build virtual context without Letta?
Yes. The pattern is a hot tier such as Redis, a cold vector store, and a summarisation step that runs on a context-pressure threshold. Letta supplies those pieces already wired together and tested against the paper's benchmarks. See memory storage backends.
Does MemGPT work with open-weight models?
The paper reports results on OpenAI endpoints, and the technique itself requires only reliable function calling, so it transfers to any model that supports tool use well enough to issue and interpret memory operations. Weaker function calling shows up as missed retrievals rather than as an error.
Continue exploring
What is a tiered memory hierarchy?
How hot, warm and cold tiers are organised beyond MemGPT.
Read →How does an agent decide what to forget?
Eviction policies, decay and the cost of recursive summaries.
Read →Which memory framework should you use?
Engram, Mem0, Letta and Zep compared on benchmarks.
Compare →