Infrastructure · Storage
Storage Backends for AI Agent Memory
Agent memory is four different workloads, and the question is not which database is best but which store each tier needs. Working state wants speed and expiry. Event history wants ordering. Durable facts want similarity search. Instructions want to be read exactly. Almost every comparison of this decision is written by a database vendor whose own product turns out to be the answer.
Four tiers
Requirements
What does agent memory actually need from a store?
Durability, scoped reads, ordering by time, and the ability to change a record without losing what it replaced. Notice that similarity search is not on that list. It is a requirement of one tier, and treating it as the requirement is how teams end up running a vector database for data that would have been better off in a table.
Durability is the one people assume and occasionally do not have. A memory written during a conversation has to survive a process restart, which rules out anything held only in application memory once the product has more than one instance. It also has to survive a partial failure: a write that stored the record but not its metadata is worse than no write, because it produces a memory nobody can scope or expire. That is the atomicity guarantee in the ACID properties, the standard every relational database is measured against, and it is the property most often given up by accident rather than on purpose.
Scoped reads are the second requirement and the one with consequences. Every read of a memory store is filtered to an identity, which means the store has to make that filter cheap and, ideally, make it hard to omit. A store where filtering is an afterthought pushes the boundary into application code, where it eventually gets forgotten in one code path.
Ordering by time is third, and it is why a pure vector store on its own is awkward. Half the useful questions about a memory store are temporal: what happened in this thread, what did the user say most recently, which of these two statements came later. Those are index scans, not similarity searches.
Fourth is revision. Memories change, and the record of what they replaced has value, so a store needs a way to mark something superseded rather than only to delete it. What that machinery does is on handling conflicting memories.
Those four requirements are not evenly distributed across the memory a system holds, which is the actual structure of the decision: which store fits which memory tier?
Tier by tier
Which store fits which memory tier?
Cache for working state, a relational table for history, a vector store for durable facts, and files for instructions. That mapping is the default, and every deviation from it should be a decision rather than an accident.
The tiers themselves are covered on types of AI agent memory. What matters here is that their access patterns differ so sharply. Working state is read on every single turn and thrown away within hours. Event history is written constantly and read rarely, usually filtered by thread. Durable facts are few, read often, and revised. Instructions are read exactly and change on a human timescale.
The semantic tier is the one with a dedicated product category, and it is covered in full on vector databases for agent memory. The short version for this page is that it is the only tier where the storage choice carries an embedding, and that single property is what makes it the expensive tier to change later.
Two of the four assignments deserve more scrutiny than they usually get, because both have a cheaper answer than the default suggests. The first is the relational tier, which frequently turns out to be capable of more than the tier it was assigned: is a plain SQL database enough for agent memory?
Relational
Is a plain SQL database enough for agent memory?
For many products, yes, and it is the cheapest adequate answer. No vendor comparison will tell you this, because a comparison written by a database company exists to move you onto that company’s database.
Start with what the queries actually are. Give me this user’s memories, ordered by when they were written. Give me everything from this thread. Give me the facts that are still valid. Give me the memory that superseded this one. Every one of those is a filtered, ordered read on an indexed table, and a relational database has been extremely good at exactly that for decades.
Only one common query is not: find the memories that mean something similar to this. That single requirement is what pulls teams toward a vector store, and it is worth noticing that most relational databases can now answer it too, through a vector column and an index in the same database holding everything else. The tradeoff against a dedicated store is covered on vector databases.
The cost argument is real and rarely stated. A database the team already runs, already backs up, already monitors and already knows how to debug carries no additional operational cost. A separate managed vector service carries a bill, an integration, another set of credentials, another failure mode and another thing to be paged about. For a memory store of a few million rows that is a large amount of overhead for one query shape.
The honest limits: at very large vector counts a dedicated store wins on latency and on index options, and if hybrid search with proper fusion is central to your retrieval, the specialised stores support it better. Neither limit binds until you are well past the scale most products reach.
If the relational tier can absorb more than expected, so can the tier below it, and this one is stranger: can the filesystem be a memory backend?
Files
Can the filesystem be a memory backend?
Yes, and it is the one most coding agents actually use. Instructions in a file, notes in a directory, tool output written to disk instead of into the prompt. It is a real architecture with real advantages and a boundary that is easy to describe.
The advantages are not trivial. Files are readable by a person without a client, diffable, version controllable, and trivially inspectable when something goes wrong, which is more than can be said for a row of floats in a managed service. An agent can also read part of a file rather than all of it, which keeps large outputs out of the context window entirely, and that pattern is a large part of why the approach works so well for coding agents.
The boundary is concurrency and structure, and it is set out carefully in Oracle’s published comparison of filesystems and databases for agent memory, which is one of the few sources to treat files as a serious option rather than a stopgap. A filesystem gives no atomic multi-step update, so a failure part way through leaves memory in an inconsistent state. Two writers without correct locking interleave or overwrite each other, and locking semantics differ across platforms and network filesystems. Permissions are per file rather than per record, so one agent reading some of another’s memory means duplicating data or building an authorisation layer.
Retrieval is the other half. Search over files is keyword matching unless you build an index, and the moment you build an index with metadata, ranking and recency weighting, you have written a database with worse guarantees than the one you did not use.
So the rule is workable: files are a good backend for a single agent operating in its own workspace, and stop being one as soon as several writers share the memory or several users are isolated within it. The tier above them has a similar test: do you need a cache tier at all?
Working state
Do you need a cache tier at all?
Only when more than one process has to see the same session. That is the whole test, and it removes a service from a large number of architectures that added one because a diagram said to.
Working state is the conversation so far plus whatever the agent is holding for this task. If your application runs as a single process per session, that state can live in the process, and adding a cache buys network latency in exchange for nothing. The moment requests for one session can land on different instances, or a background worker needs to see what the foreground turn wrote, shared state becomes a requirement.
When it is a requirement, the properties that matter are expiry and speed. A session store that never expires becomes a memory leak with a monthly bill, and a time-to-live on every key is what makes the tier self-maintaining. Speed matters because this tier is read on every turn, and it is one of the few places where single-digit millisecond access is genuinely worth engineering for.
The mistake worth avoiding is promoting the cache into the durable tier. Working state is disposable by design, and a system that keeps important facts only in a store with a time-to-live loses them silently when a session expires. What should be promoted into durable memory before that happens is on writing memories, and the boundary between the two is on short-term versus long-term memory.
Once each tier has a candidate store, a structural question arrives that vendors answer for you: should everything live in one database or several?
Shape
Should everything live in one database or several?
One, until a tier proves it needs its own. The split stack is drawn in more architecture diagrams, and the unified stack is right for more teams than those diagrams suggest.
A split stack runs a cache, a relational database and a vector store side by side, each doing what it is best at. It is the shape every reference architecture shows and it is genuinely better at the extremes. It also means three systems to operate, three sets of credentials, three failure modes, and no way to write across two of them atomically.
A unified stack keeps every tier in one database that supports vectors alongside tables. It gives up some specialisation and gains one system, one backup, one monitoring setup and, importantly, one transaction. Both of the vendors arguing loudest for this shape sell such a database, which is a reason to check the argument rather than to dismiss it.
The honest version is that this is decided by what you already run. A team on Postgres with moderate memory volume should add a vector column and stop. A team with an existing large-scale vector deployment should not migrate it into a relational database to satisfy a diagram. The cost of being wrong here is low, for reasons covered further down.
Whichever shape you pick, the same table has to be designed, and it is the most actionable thing on this page: what does a memory schema actually look like?
Schema
What does a memory schema actually look like?
Three groups of fields: identity, provenance and lifecycle. Most writing about memory architecture stops at diagrams, and the difference between a memory system that works and one that does not is usually visible in the column list.
Identity is the user, and the tenant if the product has them. It is indexed, and it appears in every read. The reason to treat it as a boundary rather than a filter is that omitting it does not produce a slow query, it produces one user’s agent stating another user’s facts.
Provenance is which thread produced the record, whether it came from a person or was inferred by a model, and when it was written. The timestamp is the field teams most often leave out and most often need, because without it nothing downstream can prefer a recent statement to an old one, which is most of what memory scoring does.
Lifecycle is whether the record still applies, what replaced it if something did, and the exact text that was embedded. Keeping the text is not optional in any store where the source conversation might be deleted, because it is the only thing that makes changing the embedding model possible later, as set out on embeddings for agent memory.
Two habits are worth adopting with the schema. Retire rather than delete, so that a superseded fact remains auditable and a mistaken supersession is reversible. And index the pair of identity and timestamp together rather than separately, because nearly every read filters on the first and orders by the second.
A schema is straightforward inside one database. It gets harder the moment the memory lives somewhere the application data does not: what breaks when memories and application data live apart?
Consistency
What breaks when memories and application data live apart?
Dual writes, and there is no clean solution to them. Writing a record to your application database and a memory to a separate store is two writes with no transaction across them, so any failure between the two leaves the system inconsistent.
The concrete case is ordinary. A user updates a setting, the application row is written, the memory write fails, and the agent now confidently tells the user the old value. The reverse is worse: the memory is written and the application write rolls back, so the agent remembers something that never happened. Neither produces an error a user sees at the time.
Three mitigations exist and each costs something. Put the memory in the same database as the application data, which removes the problem entirely and is the strongest argument for the unified shape. Write to a queue and let a worker apply the memory write with retries, which makes the memory eventually consistent and means a read immediately after a write may miss it. Or accept the inconsistency and reconcile on a schedule, which is honest and slow.
Eventual consistency deserves a specific warning in memory systems, because it is invisible. If the agent writes a fact and then reads it two seconds later during the same conversation, a store that has not committed yet returns nothing and the agent asks the user something they just answered. Any managed memory pipeline that processes writes asynchronously has this property, and it is worth testing for directly rather than discovering in a session log.
All of these decisions look weightier than they are, because of an asymmetry nobody mentions: how hard is it to move to a different backend later?
Reversibility
How hard is it to move to a different backend later?
Cheap for three tiers out of four. This is the point missing from every comparison in this category, and it changes how much the initial decision is worth agonising over.
Working state needs no migration at all, because it expires on its own. Switch the store, let the old sessions age out, and the problem solves itself within hours. Event history is rows with a stable shape, so moving it is an export and an import. Instructions in files move the way files move.
Only the semantic tier is expensive, and only because of the vectors. Moving between vector stores means re-embedding every record if the embedding model changes, and even when it does not, the vectors have to be rewritten and every index rebuilt. That is the tier worth thinking hard about, and it is also the tier that has a specific prerequisite: the source text has to still exist.
Two practices keep the door open. Keep the memory access behind a narrow interface in your own code rather than calling a vendor SDK from twelve places, so the store can be swapped without touching the application. And keep the original text of every memory, which is the single field that makes the semantic tier portable at all.
Given that reversibility, the sensible strategy is not to optimise the first choice: which backend should you start with?
Starting point
Which backend should you start with?
The database you already run, with a vector column added, and nothing else. That is not a compromise position. For most products it is where the architecture should end up, and it is certainly where it should begin.
The reasoning follows from everything above. Three of the four tiers are cheap to move later. The one that is not is portable as long as you keep the source text. The dual-write problem disappears when memory and application data share a transaction. And the operational cost of one system you already understand is far lower than that of three you do not.
Add a dedicated store when a specific constraint appears, not in anticipation of one. Vector counts large enough that index options matter. Hybrid search central enough that native fusion is worth a service, as covered on hybrid search. Session traffic heavy enough that a real cache earns its keep. Each of those is a measurable trigger, and none of them arrives quietly.
If you would rather not build the layer above the store at all, managed memory products supply extraction, reconciliation and retrieval on top of a backend they operate. Engram runs on Weaviate, and Mem0 and Zep sit on backends of their own; the tradeoff between buying that layer and writing it is on memory tools and on Engram.
The order of work matters more than the choice. Get the schema right, because that is what determines whether memories can be scoped, aged and superseded at all. Keep the store behind your own interface. Then pick the simplest thing that satisfies today’s queries, and let the triggers above tell you when to change it. The whole loop is on how AI memory works, and the build sequence on adding memory to an agent.
FAQ
Frequently asked questions
The practical follow-ups: what to index, what to keep, and when a second store is justified.
What should I index on a memory table?
The identity and timestamp together as a composite index, because nearly every read filters on the first and orders by the second. Add an index on thread for history queries and one on the validity marker if most reads exclude superseded records. Index the vector column separately, since it uses a different index type entirely.
When is a dedicated vector database worth adding?
When one of three specific triggers appears: vector counts large enough that index family and memory footprint start to matter, hybrid search central enough to your retrieval that native fusion is worth a service, or latency requirements a vector column in your existing database cannot meet. Anticipating those triggers is not the same as hitting them.
Can I keep memories in the same tables as my application data?
Yes, and it removes the dual-write problem entirely, which is its main advantage. Keep memories in their own tables within that database rather than as columns on existing rows, since memories have their own lifecycle, their own retention rules and their own access patterns.
What happens to memory when a user deletes their account?
Every tier has to be covered, which is easier to guarantee when they share a database. Memories scoped by identity delete with a single predicate, but a superseded record, an event log row and a cached session are three separate places the same data can persist. Design the deletion path when you design the schema, not afterwards.
Is a document database a reasonable memory backend?
Yes, particularly one with vector search built in, since a memory record is naturally a document with variable metadata. The considerations are the same as for a relational store: how cheap the scoped read is, whether it can order by time efficiently, and whether it is already part of your stack.
Should event history and durable facts be separate tables?
Yes. They have different volumes, different retention and different read patterns: history is written constantly and read rarely, facts are few and read on most turns. Keeping them in one table means the fact reads pay for the history volume, and it makes retention policy harder to express.