Infrastructure · Cluster hub

What Infrastructure Does AI Memory Run On?

AI memory runs on three storage shapes and one retrieval technique: vector databases for similarity, knowledge graphs for relationships, key-value stores for speed, and hybrid search to combine meaning with exact matching. This hub covers each one, what it answers, and how to decide between them.

Three shapes

1
Vector
Similarity
2
Graph
Relationships
3
Key-value
Speed

The storage shapes

What storage does AI memory use?

Three shapes, each answering a different question: a vector database answers “what is similar to this”, a knowledge graph answers “what is connected to this and when was it true”, and a key-value store answers “give me this exact thing now”. Most production systems use at least two.

A vector database stores and searches embeddings, while a memory layer adds extraction, scoping, ranking, consolidation and eviction on top.
Figure 1. Storage is the smaller half of a memory system. Choosing a database and calling it memory leaves every policy decision unmade.

Vector databases index embeddings so a question phrased differently from a stored fact can still retrieve it. This is the default for free-text memories and the reason semantic search works at all. Covered on vector databases for AI memory, with the embeddings themselves on embeddings explained.

Knowledge graphs store entities and typed relationships, which makes multi-hop questions answerable by traversal and makes validity over time expressible. The cost is a much harder write path. Covered on knowledge graphs for AI memory.

Key-value stores such as Redis provide exact, very fast lookup for the current session’s working state. They sit in front of the durable store rather than replacing it. Covered on memory storage backends.

Underneath the shapes, what actually gets written is more than the text: how a memory is stored.

The record

How is a memory actually stored?

As a short piece of text, its embedding, and the metadata that makes it findable and governable: whose it is, when it was written, what type it is, and whether it is still valid. The text is the only part that reaches the model; everything else exists so the right text can be found and managed.

Short-term context window memory compared with the long-term external store, by speed, size and whether it survives the session.
Figure 2. All of this infrastructure serves the right-hand store. The model still only reads what is placed in the left one.

The metadata is where governance lives, and it is routinely under-designed. A user or tenant identifier makes scoped search and deletion possible, and adding it later means rewriting every record. A timestamp makes recency weighting and expiry possible. A type marker separates a durable preference from a transient event so the two can be retained on different schedules. A validity flag lets a superseded fact be kept for history without competing with the current one, as described on handling conflicting memories.

One infrastructure decision has consequences people discover late. The embedding model is part of the storage format. Queries and memories must be embedded by the same model, so changing it means re-embedding the entire store, which is a migration rather than a configuration change. That constraint is worth knowing before a store is large.

With the record shape settled, retrieval technique is the next choice: how the search itself works.

Selection

How do you choose memory infrastructure?

Start from the questions your product has to answer, not from the storage technology. The shape of your questions determines the shape of the store, and that ordering prevents most expensive mistakes in this category.

  • Single-hop questions about facts. A vector database is enough, and adding a graph is machinery you will not use.
  • Questions spanning several entities. A graph earns its write cost, because similarity search cannot join facts that were stored separately.
  • Questions about the past. You need validity intervals, which a graph gives natively and a vector store can approximate with a validity field and disciplined filtering.
  • Sub-second working state. A key-value tier in front, holding the current session rather than the system of record.

Two constraints override the question shape when they apply. Data residency can rule out managed services entirely, which pushes toward self-hosted stores. An existing stack is a strong argument on its own: a team already running Weaviate or Postgres gets a better outcome extending what they operate than adding a second database for a marginal capability gain.

One constraint sits above all the others and is worth stating separately, because it is the one people discover after the fact. Whatever you choose has to support filtering during the search rather than after it. Scoped retrieval, validity filtering and type filtering all depend on it, and a store that can only filter results post-hoc will return other users’ candidates, discard them, and leave you with fewer useful memories than you asked for. It is an unglamorous capability and it constrains the choice more than raw performance does.

A second constraint worth checking before committing is whether the store supports updating a record in place. Memory needs supersession and merging, both of which are updates rather than inserts, and a store optimised purely for append-and-search makes those awkward enough that teams skip them.

The most common shape in production is unglamorous and works: a vector store for the bulk of memories, a small key-value tier for session state, hybrid search on top, and a graph added later only if multi-hop questions turn out to be real rather than hypothetical. The managed options that hide these choices are compared on the best AI memory tools and open source versus managed memory.

Two questions decide most builds before any of that matters: whether you need a dedicated vector database.

The common question

Do you need a dedicated vector database?

Not at first. A Postgres database with the pgvector extension handles memory stores well into the hundreds of thousands of records, and it removes an entire service from your architecture. The MemGPT paper’s own experiments used exactly that arrangement.

Where the memory layer sits: application, then agent, then the memory layer, then the vector, graph or key-value storage beneath it.
Figure 3. The storage tier is the most replaceable part of the stack, which is an argument for starting with whatever you already run.

The case for starting with what you have is strong. Your team already knows how to back it up, monitor it, and reason about its failure modes. Memories can be joined against application data in one query. And there is no second consistency model to think about when a user is deleted.

A dedicated vector database earns its place at scale and on features. Purpose-built engines handle far larger indexes efficiently, offer richer filtering alongside similarity, and ship capabilities such as native hybrid search that would otherwise be yours to build. Weaviate, Qdrant, Milvus and Pinecone all sit here, and the choice between them is a normal infrastructure decision rather than a memory-specific one.

The honest trigger for migrating is measurable rather than architectural: search latency climbing past your budget, filtering that has become awkward to express, or an index that no longer fits comfortably in memory. Migrating on any of those is a good decision, and migrating in anticipation of them usually is not.

Four classes of agent memory tool: memory API, temporal graph, virtual paging and framework-native, each with who controls retrieval.
Figure 4. Managed memory layers make the storage question disappear, which is either the main attraction or the main objection.

There is a third option that neither builds nor migrates: let a managed memory layer own the storage decision. Engram (Weaviate), Mem0 and Supermemory all hide the database entirely, and for teams whose differentiation is elsewhere this is usually the right trade. The cost is that retrieval behaviour becomes someone else’s to change and your memories live in their system, which is a procurement question as much as a technical one.

Whichever route you take, one operational detail is worth settling early because it is painful later: backups and deletion. A memory store holds personal data by construction, so it needs the same retention and erasure handling as any other user data. A store that cannot answer “show me everything you hold about this person” and “delete it” is a compliance problem waiting rather than a technical shortcut, and the scope field that makes both possible has to be written from the first record.

The one thing that does not change with storage is the policy layer above it. Extraction, scoping, ranking, consolidation and eviction have to exist regardless of which database holds the vectors, which is the argument made on the memory layer. Infrastructure is the layer below the mechanisms in the architecture cluster.

FAQ

Frequently asked questions

The questions that follow: whether Postgres is enough, and when to add a second store.

What is the best vector database for AI agent memory?

Depends on your stack and scale. Pinecone, Qdrant, Weaviate and pgvector are common choices — match latency, ops burden and whether you need a separate memory framework on top. See vector databases for memory.

Do I need a knowledge graph for agent memory?

Not always. Vector stores suffice for semantic recall and personalization. Knowledge graphs add value when facts have relationships, change over time, or need conflict resolution — as with Zep's Graphiti. See vector vs knowledge graph.

Which embedding model should I use for memory?

Choose based on your language, dimension budget and latency needs. OpenAI, Cohere and open-source models (e.g. sentence-transformers) are common. The embedding model must stay consistent between write and retrieve. See embeddings for memory.

How does RAG infrastructure relate to agent memory?

RAG retrieves from a fixed document corpus; agent memory is dynamic and personal. They combine well: RAG for org knowledge, memory for user/session state. See RAG explained and memory vs RAG.

Is Redis enough for agent memory?

Redis excels as a fast buffer and semantic cache, but you typically implement memory logic yourself. For full CRUD memory APIs, use a framework like Engram, Mem0 or LangMem. See Redis for agent memory.

How do embeddings connect to long-term memory?

Text is embedded into vectors at write time; retrieval embeds the query and finds nearest neighbors in the store. The same model must be used for both. See embeddings for memory.

When should I use hybrid search for memory?

When pure vector search misses exact keyword matches (IDs, product names, codes). Hybrid combines vector similarity with keyword/BM25 search. See hybrid search.

Is pgvector production-ready for agent memory?

Yes for many workloads — especially when you already run Postgres and want SQL + vectors in one store. Scale and latency requirements may push you to dedicated vector DBs. See storage backends.