Architecture · Hierarchy
What Is a Memory Hierarchy?
“Memory hierarchy” names three genuinely different design decisions that happen to share a name: how fast and expensive each tier is to access, who is allowed to see a given memory, and which technology, an event log, a graph, or a vector store, a memory actually lives in. Most explanations pick one and imply it’s the whole picture.
Three axes
Disambiguation
Why does “memory hierarchy” mean three different things?
Because the word “hierarchy” only describes a shape, tiers ordered by some priority, and three completely different priorities produce three completely different hierarchies. Reading an explanation of one and assuming it covers all three is where most of the confusion around this term comes from.
A storage-tier hierarchy orders memory by speed and cost: what stays in the model’s context window versus what gets paged in from somewhere slower when needed, the sense borrowed directly from how a computer’s own memory is organized. An organizational-scope hierarchy orders memory by who is allowed to see it: global facts every agent on a team can read, group-scoped facts for a subset of agents, private facts for one agent alone, which is a question about multi-agent systems specifically and is covered in full on multi-agent memory rather than here. A technology-layer hierarchy orders memory by what kind of store it lives in: an event log, a knowledge graph, a vector index, each suited to a different kind of question.
This page covers the first sense in depth, including a second implementation of it that’s genuinely different from the one most explanations default to, and the third sense briefly, with pointers to where each underlying technology gets its own full treatment. Start where the term originated: a direct borrow from computer architecture. What is the canonical hot, warm, cold storage hierarchy?
The canonical model
What is the canonical hot, warm, cold storage hierarchy?
Three tiers, directly modeled on a computer’s own memory hierarchy: core memory (hot, always in context, like RAM), recall memory (warm, searchable recent history, like a disk cache), and archival memory (cold, a long-term vector store, like cold storage). This is MemGPT’s design, and it remains the most widely implemented version of a storage-tier hierarchy for agent memory.
The analogy isn’t decorative; it’s borrowed from a real, decades-old engineering problem. Early personal computers had brutally limited memory, a Commodore 64 shipped with 64 kilobytes total, and entire programs had to fit inside that ceiling. The Intel 80486, in the late 1980s, popularized on-chip caching alongside virtual memory: caching made frequently accessed data feel fast, virtual memory made the total available space feel effectively infinite, and the two together let a program behave as though it had far more memory than actually existed close to the processor. MemGPT’s core/recall/archival split does the same job for an agent: the context window is the scarce, fast resource; recall and archival memory are progressively slower, cheaper stores an agent pages information into and out of on purpose, through explicit function calls, rather than everything sitting in context all the time.
What makes this an active hierarchy rather than a passive cache is that the agent itself decides what moves between tiers, choosing what to retain, summarize, or archive, instead of simply receiving whatever gets injected. The full mechanics of how that paging actually works are covered on virtual context and MemGPT; this page’s job is narrower, establishing the tier concept and where it comes from.
MemGPT’s tiers are fixed in advance: three named stores, each with a defined role. That’s one way to build a storage hierarchy. It isn’t the only way a real production system does it today. Is a fixed tier the only way to build a storage hierarchy?
A newer pattern
Is a fixed tier the only way to build a storage hierarchy?
No. A second, genuinely different pattern computes its index at query time instead of building one in advance: composable Unix tools, grep, ls, glob, running directly over a filesystem, used by Claude Code, Cursor, and Arize’s own production agent Alyx. There is no pre-built vector index or fixed tier at all; the index is generated fresh, on demand, by the tool call itself.
An index, in the traditional sense, answers “given this query, where is the data” using a structure computed ahead of time. A command like grep answers the same question differently: it scans at the moment it’s called and streams back a result that carries the same semantic role, a map from query to location, without ever being stored anywhere. Files never need pre-indexing, and the result is only ever as stale as the last time the command ran, which for a filesystem is never. This is what lets a fixed context window feel, functionally, like it has access to an unbounded amount of data: the filesystem plus a handful of composable tools acts as the slow, cold tier, and only what a specific query actually needs gets pulled into the fast tier that is the model’s context.
Arize’s own evaluation of this pattern is concrete rather than theoretical: given 10,000 files, only a few of which contained the target data, with no filename hints about which ones, both Claude Code and Cursor used grep, cut, sort and uniq together to progressively narrow the search down to the relevant rows. The execution wasn’t flawless either, which is itself informative: one model recognized mid-task that a tool’s output was too large to fit its remaining context, and backtracked to a different approach that would actually fit, exactly the kind of self-correction a fixed, pre-built tier has no mechanism to perform on its own.
Both patterns so far organize memory by how fast and cheap it is to reach. A third, different way to organize a hierarchy is by what kind of store the data actually lives in. How does a hierarchy by storage technology differ from one by speed?
A third axis
How does a hierarchy by storage technology differ from one by speed?
Instead of ordering stores from fast to slow, this hierarchy splits memory by data model: an immutable event log as ground truth, a graph layer for relationships and temporal validity, and a vector layer for fast semantic search, typically composed as a pipeline rather than three interchangeable stores. Each layer answers a structurally different question, not the same question at different speeds.
The event layer is the audit trail: every action, decision and outcome logged immutably, giving perfect replay and the ability to reconstruct exactly what happened and when, the concern this site covers in depth on storage backends. The graph layer sits above it, modeling entities and how they relate to each other, and is what makes a multi-hop question like “which project depends on the vendor whose contract just changed” answerable at all, covered fully on knowledge graphs. The vector layer is usually the entry point: fast semantic search that surfaces what’s topically relevant before the graph layer traverses relationships within that narrowed set, the mechanics of which are covered on vector databases. A query commonly moves through all three in sequence: semantic search narrows the field, the graph expands and validates relationships within it, and the event log confirms the result against what’s actually recorded as having happened.
Three different hierarchies, three different organizing questions. The genuinely useful next question isn’t which one is correct, since they aren’t competing answers to the same problem. Do these layers compete, or do they compose?
Composition
Do these layers compete, or do they compose?
They compose. A production agent typically draws on working memory, episodic or experiential memory, and semantic or relational memory simultaneously, each answering a different question about what the agent needs to know right now, not choosing one pattern and discarding the others.
Working memory is what the agent is processing this turn, the context window itself. Episodic or experiential memory is what this agent or this user has actually experienced, whether that’s a flat vector store or MemGPT-style tiers. Semantic or relational memory is the structured layer of entities and how they connect, when a task genuinely needs multi-hop reasoning rather than similarity search alone. Treating these as a menu to pick one item from, rather than as layers that each cover different ground, is a common source of under-built memory systems: a team that adopts only a vector store, for instance, has working and episodic memory but nothing that handles relational questions well, and later has to bolt on a graph layer rather than having planned for it.
Even understood this way, composing multiple layers correctly is not a fully solved problem, and it’s worth being honest about exactly where the difficulty still is. What’s still genuinely unsolved here?
The open question
What’s still genuinely unsolved here?
Cross-layer consistency: keeping what’s true in one layer from silently contradicting what’s stored in another. Academic work proposing a formal hierarchy for multi-agent memory identifies exactly this as the field’s most pressing open challenge, not a detail left for implementation.
The failure mode is concrete: an event log records that a user cancelled a subscription, but the graph layer’s cached relationship between that user and the subscription hasn’t been updated yet, and a vector-indexed summary from before the cancellation still ranks highly for a related query. Nothing in any individual layer is wrong; each one is internally consistent with what it was told. What’s missing is a mechanism that keeps the layers from drifting apart as facts change, and no widely adopted standard for that currently exists. This is a genuinely different problem from choosing which layers to use in the first place, and teams building a multi-layer system should expect to design an explicit reconciliation strategy rather than assume the layers will stay in sync on their own.
With all three senses of hierarchy on the table, storage tier, technology layer, and the open problem of keeping them consistent, the practical question is which of these decisions actually applies to what’s being built. Which hierarchy does your agent actually need?
The decision
Which hierarchy does your agent actually need?
Start with a storage-tier hierarchy the moment context length becomes a real constraint, add a technology-layer split only once a single store genuinely can’t answer the questions being asked of it, and treat the organizational-scope hierarchy as a separate decision entirely, made when more than one agent needs to share memory. Building all three from day one, before any of them is a measured bottleneck, is over-engineering a single-agent prototype.
A team just past the prototype stage usually needs the first: some form of tiering, whether MemGPT’s fixed core/recall/archival split or a dynamic filesystem-plus-tools approach, to keep the context window from being the whole memory system. A technology-layer split becomes worth the added complexity specifically when relational or temporal questions start showing up that a flat vector store answers poorly, not as a default starting architecture. The organizational-scope hierarchy, global, group, and private visibility across multiple agents, is a different problem again, covered on multi-agent memory, and applies only once there’s more than one agent to scope memory between in the first place.
For teams that would rather not design and maintain the storage-tier and technology-layer decisions themselves, a managed option such as Engram handles the extraction, tiering and retrieval logic as a service, which sidesteps building and operating the hierarchy directly while still getting its practical benefit: a small, fast context on every turn, backed by something larger and slower underneath it.
FAQ
Frequently asked questions
The practical questions that follow once the three senses of hierarchy above are understood.
Is the MemGPT tiered model the only correct memory hierarchy?
No. It's the canonical, most widely implemented storage-tier hierarchy, but a dynamic, runtime-computed index over a filesystem using composable tools is a documented, current alternative used by Claude Code, Cursor and production agents like Alyx, with no fixed tiers at all.
Does a memory hierarchy replace a vector database?
No. A storage-tier hierarchy decides what sits in context versus what gets paged in from elsewhere; a vector database is often the technology that slower tier is built on. They're different layers of the same problem, not alternatives to each other.
Do I need a graph layer if I already have a vector store?
Only once relational or temporal questions start showing up that similarity search answers poorly, such as multi-hop relationships between entities. Adding a graph layer before that need is measured is added complexity without a corresponding benefit.
Is organizational-scope hierarchy the same as storage-tier hierarchy?
No, they're separate design questions that happen to share the word hierarchy. Organizational scope (global, group, private visibility) only applies once more than one agent needs to share memory; storage tiering applies even to a single agent.
What actually breaks when memory layers get out of sync?
A fact updates in one layer (an event log records a cancellation, for instance) while a graph relationship or a vector-indexed summary elsewhere still reflects the old state. Each layer stays internally consistent with what it was told; nothing keeps them consistent with each other automatically.
Should a new agent project start with a full multi-layer memory hierarchy?
No. Start with basic tiering once context length becomes a real constraint, and add a technology-layer split only when a single store demonstrably can't answer the questions being asked of it. Building all three axes before any is a measured bottleneck is over-engineering.