Architecture · Cluster hub

What Is AI Memory Architecture?

AI memory architecture is the arrangement of stores, operations and policies that turns a stateless model into an agent with continuity: tiered storage underneath, a write and read loop on top, and maintenance running in the background. This hub covers each mechanism and how they fit together.

Three layers

1
Context
Hot
2
Buffer
Warm
3
Store
Cold

The layers

What are the layers of an AI memory architecture?

Three tiers, ordered by how fast the model can reach them: the context window as the hot tier, a fast buffer such as Redis as the warm tier, and a vector or graph store as the cold tier. Production systems organise memory this way for the same reason operating systems do, because the fastest storage is always the scarcest.

Where the memory layer sits: application, then agent, then the memory layer, then the vector, graph or key-value storage beneath it.
Figure 1. The tiers sit below the memory layer, which holds the policy deciding what moves between them.

The hot tier is the context window itself. It is the only memory the model reads directly, it is bounded by a token limit, and it disappears when the session ends. Everything else in the architecture exists to decide what occupies it, which is covered on short-term memory in AI agents.

The warm tier holds the current session’s working state. A key-value store gives sub-millisecond exact lookup for things like the active task, recent tool results and the conversation buffer. It is not the system of record; it is what keeps the current turn fast.

The cold tier is the durable store, a vector database, a knowledge graph, or both, holding everything that must survive the session. It is reached only by an explicit search, which is why retrieval quality matters more here than raw capacity.

Some architectures add a fourth tier by paging between them under the agent’s own control, which is the MemGPT design covered on virtual context and MemGPT, and the general form on what a tiered memory hierarchy is.

Layers describe where memory lives. The operations describe what happens to it: the mechanisms that run on the layers.

The operations

What operations make up a memory architecture?

Six: write, retrieve, score, consolidate, forget, and orchestrate. Every framework implements the same six, however it packages them, and a system missing one of them fails in a way that is characteristic of the missing piece.

The four operations of the AI memory loop: write and extract, store, retrieve and rank, then update or evict.
Figure 2. The four-operation loop is the shape most systems ship. Scoring sits inside retrieval, and orchestration decides when each operation runs.
  • Writing extracts durable facts from a turn and deduplicates them before storage. Skipping deduplication is the root of most later problems. See how agents write and store memories.
  • Retrieval searches, filters by scope and injects the top candidates before the model answers. See memory retrieval.
  • Scoring ranks candidates on recency, importance and relevance, deciding which memories are worth their prompt slots. See scoring and ranking memories.
  • Consolidation merges duplicates and promotes session memories to durable ones, in the background. See memory consolidation.
  • Forgetting evicts what is stale so the store does not grow without limit. See forgetting and eviction.
  • Orchestration decides when each of the others runs, which is the part that is easiest to leave implicit and hardest to debug later. See memory management and orchestration.

Two more mechanisms sit alongside these and matter once a system is live: resolving conflicting memories and summarising and compressing memory.

The same six operations get arranged differently depending on what the product needs: the common architecture patterns.

Patterns

What are the common memory architecture patterns?

Four arrangements cover most production systems, and they differ in who decides what enters the context window. That single question separates them more cleanly than storage technology does.

Four classes of agent memory tool: memory API, temporal graph, virtual paging and framework-native, each with who controls retrieval.
Figure 3. Storage engine is not the dividing line. Two systems on the same vector database can behave completely differently.

Application-controlled retrieval is the most common: the application calls a memory service before invoking the model, and decides itself how many memories to include. Predictable and easy to reason about, less adaptive when a turn needs something unusual.

Agent-controlled paging gives the model tools to fetch and evict its own memories, which is the MemGPT and Letta design. Maximum flexibility, more inference calls, and retrieval quality that depends on the agent asking the right question.

Graph-backed memory stores entities and relationships with validity intervals, which is what makes temporal questions answerable. Higher write cost, described on knowledge graphs for AI memory.

Framework-native memory uses the primitives of an agent framework already in use, trading portability for one less service to operate.

Most real systems mix these. A common shape is application-controlled retrieval over a vector store for per-user facts, plus a small graph for the handful of relationships that need time, plus background consolidation. The tool implementations are compared on the best AI memory tools.

Whichever pattern is chosen, the same design mistakes recur: what goes wrong in memory architectures.

Failure modes

What goes wrong in memory architectures?

Four design errors account for most systems that work at launch and disappoint six months later, and every one of them is an operation that was left out rather than implemented badly.

  1. Write without deduplication. The store accumulates near-identical memories, and retrieval spends several slots on one fact. The fix belongs at write time, not at read time.
  2. Retrieval without scope. Memories with no user or tenant identifier cannot be filtered safely, cannot be deleted on request, and eventually leak across boundaries. Retrofitting scope means rewriting every record.
  3. No consolidation. Nothing merges, nothing supersedes, and the store’s quality declines slowly enough that the memory system is rarely blamed.
  4. No eviction policy. Keeping everything forever is a decision, usually an unexamined one, and it is the option that reliably degrades retrieval and raises cost.

The common thread is that a memory architecture is usually built as a write path and a read path, with the maintenance operations deferred. They are cheap to add on day one, when the store is small, and expensive later, when adding them means processing a year of unfiltered writes.

A useful sanity check on any architecture is to ask which component would notice a problem first. If the answer is “a user”, the system has no instrumentation and every one of the four failures above will be discovered late and expensively. If the answer is a retrieval metric running on a fixed question set, the architecture has a feedback loop, and that matters more than which database sits underneath it.

The way to catch all four early is to measure retrieval rather than assume it: a fixed set of questions whose answers depend on stored facts, checked regularly for whether the right memory was retrieved and whether the answer used it. That evaluation is described in how to add memory to an AI agent, and the metrics on which metrics matter for agent memory.

How much of this a system needs varies more than most guides admit: how architecture differs by product.

By product

How does memory architecture differ by product?

The product decides which memory types dominate, and that in turn decides which operations need the most engineering. A support agent and a coding agent both run the same six operations, and they weight them completely differently.

Six types of AI agent memory laid out: working, semantic, episodic, procedural, sensory and buffer, and shared memory.
Figure 4. Architecture follows from which of these types your product actually needs, not from a reference diagram.

A support agent leans on semantic memory of the account and episodic memory of the case. Its hardest operation is conflict resolution, because customer facts change constantly and a stale entitlement produces a wrong answer with real consequences. Retention is short for episodes and long for account facts.

A personal assistant is almost entirely semantic, accumulating preferences about one person over a long period. Its hardest operation is consolidation, because a store that runs for years without merging becomes a slow, contradictory pile. Users also expect to inspect and delete memories, which makes scope and provenance product features rather than implementation details.

A coding agent is the outlier: its most valuable memory is procedural, meaning the conventions of a codebase rather than facts about it. Its hardest operation is capture, since the useful signal arrives as corrections and successful sequences rather than as statements. See memory for coding agents.

A multi-agent system adds shared memory and inherits a consistency problem none of the others have, described on shared memory in multi-agent systems.

Team size belongs in this decision too, and rarely appears in reference architectures. A two-person team shipping a first version should implement write, retrieve and deduplicate, and defer everything else with a note about when to revisit. A team operating a system with a year of accumulated memories has the opposite problem: the maintenance operations are the ones carrying the quality, and the write path has long since stopped being interesting. Architectures published by large vendors describe the second situation and are read by teams in the first.

The practical implication is that copying an architecture from a different product class is how teams end up over-engineering one operation and under-engineering the one that matters. The infrastructure these architectures run on is covered in the infrastructure cluster, and the full operating loop in how AI memory works.

FAQ

Frequently asked questions

The questions that follow: how many tiers a system needs, and whether to build or buy the layer.

What are the main components of AI memory architecture?

The core components are: write (extraction and persistence), retrieve (search and ranking), consolidate (merge and promote memories), forget (eviction and decay), plus supporting layers for hierarchy (tiered storage), scoring (relevance ranking) and orchestration (lifecycle management). See the architecture hub for deep dives on each.

Should agents write or retrieve memory first?

In production, retrieval usually runs every turn — the agent needs context before responding. Writing happens after a turn when the system extracts what is worth remembering. Most frameworks run retrieve → generate → write in sequence. See writing memories and memory retrieval.

When does memory consolidation happen?

Consolidation runs when short-term memories should be promoted to durable long-term storage — after a session ends, on a schedule, or during background 'sleep-time' processing. It merges, summarizes and deduplicates memories. See memory consolidation and sleep-time compute.

Is forgetting necessary in AI agent memory?

Yes. Without eviction, memory stores grow unbounded, retrieval quality degrades, and token costs rise. Agents need TTLs, decay policies and relevance pruning to keep memory useful. See forgetting and eviction.

What is the MemGPT memory architecture?

MemGPT treats the context window like RAM and pages memories between a fast tier (context) and a deep store (database) — giving agents effectively unbounded conversation history. Letta implements this virtual-context paging pattern. See virtual context and MemGPT.

Graph vs vector memory architecture — which is better?

Vector stores excel at semantic similarity search and fast retrieval. Knowledge graphs excel at relationships, temporal facts and conflict resolution. Many production systems use hybrids (Zep's Graphiti, Mem0's optional graph). See vector vs knowledge graph memory.

What is memory-as-a-tool architecture?

Instead of automatic background writes, memory operations (store, search, update, delete) are exposed as tools the agent calls explicitly — giving the model control over what to remember. Grounded in research like AgeMem. See memory as a tool.

How do you orchestrate memory in production?

Production orchestration coordinates when to write, retrieve, consolidate and forget — with monitoring, conflict handling and cost controls. Frameworks like Engram, Mem0, Zep and LangMem each implement different orchestration patterns. See memory management and orchestration.

How should agents handle conflicting memories?

When new information contradicts old facts, agents must update, version or invalidate stale memories. Strategies include temporal graph edges (Zep), explicit versioning, or overwrite-with-confidence scoring. See handling conflicting memories.

How do you benchmark memory architecture choices?

Use public benchmarks like LOCOMO (long-conversation recall) and LongMemEval (cross-session recall) to compare architectures on your use case before committing. See the evaluation hub.