Infrastructure · Vector DB
Vector Databases for AI Agent Memory
A vector database stores memory records as embeddings and returns the ones nearest to a query, filtered by whatever metadata you wrote alongside them. That is the whole of its job. It does not decide what was worth remembering, it does not notice that a new fact contradicts an old one, and it does not forget. Knowing exactly where its responsibility ends is what separates a memory system that works from one that confidently repeats something the user corrected last week.
Vector role
Definition
What does a vector database actually store for agent memory?
Three things per record: a vector, the original text, and metadata. The vector is the embedding of that text, the text is what eventually enters the prompt, and the metadata is every attribute you will later want to filter or sort by.
Most explanations of vector databases are written for document retrieval, where the record is a passage cut out of a larger file. A memory record is a different object. It is usually one statement about one person, it was produced by a model reading a conversation rather than by a splitter reading a file, and it has a lifecycle: it becomes true at some point, it can be superseded, and it may need to stop being retrievable without being deleted.
That difference shows up in the metadata, which is where memory systems live or die. A document store often needs nothing more than a source path. A memory store needs, at minimum, the identity the record belongs to, the session or thread it came from, when it was written, and some marker of whether it is still current. None of those fields is inferred for you. They exist only if the write path put them there.
The vector database treats all of it uniformly: it indexes the vector for similarity and the metadata for filtering, and it has no opinion about what the fields mean. That neutrality is a feature at retrieval time and a hazard at design time, and it starts with what happens on the way in: how does a memory become a searchable vector?
Write path
How does a memory become a searchable vector?
Four steps, and only the last two involve the database. Something decides the turn is worth keeping, something reconciles it against what is already stored, an embedding model turns the surviving text into a vector, and the store writes that vector with its metadata.
The first step is the fork that separates a memory pipeline from a retrieval pipeline. A document pipeline chunks everything it is given, because every passage might be asked about later. A memory pipeline extracts, because most of a conversation is not worth remembering. Storing the raw turn gets you a searchable record of the sentence a preference happened to appear in, filler included, rather than the preference itself.
The practical consequence is a model call on the write path. Extraction is a judgement about what counts as durable, which no embedding can make, and it is the reason a memory system costs more per turn than a RAG ingest does. What that judgement should include is covered on writing memories.
The second step is the one most often skipped. Before writing, the pipeline should ask whether a semantically similar record already exists, and if it does, whether the new statement agrees with it, refines it, or contradicts it. A store that never asks accumulates both halves of every contradiction as independent rows with similar embeddings, which means both will be retrieved together, at nearly the same score, forever. Handling conflicting memories is the page for that problem.
The third step is where a choice you cannot easily reverse gets made. The embedding model fixes the dimension of every vector in the collection, and changing model later means re-embedding the entire store, because vectors from two models are not comparable. Choosing an embedding model covers the tradeoffs; the point here is that the vector database inherits that decision rather than making it.
Only the fourth step is the database’s own work, and once the row is written, the interesting question is how it is found again: how does an ANN index find a memory, and what does it give up?
Indexing
How does an ANN index find a memory, and what does it give up?
It gives up certainty, deliberately, in exchange for speed. Approximate nearest neighbour search does not examine every vector in the collection. It examines a small fraction chosen by the index structure, and it returns the best it found there, which is usually but not always the best that exists.
Comparing a query vector against every stored vector is exact and linear in the size of the collection, which is fine at ten thousand records and unusable at ten million. Every production index therefore approximates, and the parameter you set is how hard it looks before answering. Recall is not a property you receive from the product. It is a number you configure, and configuring it lower is how the same store gets faster.
Three families cover almost all deployments. Graph indexes, of which HNSW is the standard, connect vectors into a navigable network and walk it toward the query. They give the best unfiltered latency and hold the graph in memory, so the footprint follows the size of the collection rather than the traffic. Clustered indexes, the IVF family, partition the space and search only the nearest partitions. Quantised variants, IVF with product quantisation, compress each vector into a short code, which cuts memory by roughly an order of magnitude and cuts recall along with it.
For memory workloads the clustered family deserves more attention than it usually gets, because almost every memory read is filtered. A graph index has to walk edges that mostly lead to other users’ records and discard them, while a clustered index can skip whole partitions that the filter has already excluded. That is a structural difference, not a tuning difference, and it grows with the number of tenants.
The failure mode of a badly configured index is quiet in a way that matters here. A missed document in a search result looks like a search result with one fewer item. A missed memory looks like an agent that does not know something it was told, which the user reads as the system having no memory at all. There is no error, no warning, and nothing in a log to indicate that the record was sitting in the store the whole time.
The index decides which candidates are considered. What decides which of them may be returned at all is a different mechanism entirely: why is metadata filtering a correctness boundary and not a speed trick?
Scoping
Why is metadata filtering a correctness boundary and not a speed trick?
Because in a memory store the user predicate is the only thing standing between one person’s facts and another person’s agent. Every guide that covers metadata filtering covers it as a way to narrow a search so it runs faster. In a memory system that framing is wrong, and it leads teams to test the filter for performance rather than for isolation.
Consider what each failure looks like. Drop the filter in a document search and the query returns passages from a collection the user did not intend to search: a worse answer, recoverable, invisible in the worst case. Drop the same filter in a memory search and the retrieved rows are somebody else’s stated preferences, medical details or account facts, which the model then presents in its own voice as things it remembers about the person it is talking to. The mechanism is identical. The consequence is a different category of problem.
Three rules follow from treating it as a boundary. The predicate belongs in the query construction that every read path shares, not in each caller, because a boundary enforced in twelve places is enforced in eleven. It should be tested with an assertion that a query scoped to one identity never returns a record belonging to another, which is a test you can write once and run on every build. And it should hold at the storage layer where the product supports it, through per-tenant collections or namespaces, so that a missing predicate returns nothing rather than everything.
Beyond identity, the fields worth writing are the ones that answer questions the agent will actually ask: which thread produced this, when was it written, does it still apply, and where did it come from. Time is the one teams most often omit and most often need, because without a written-at value there is no way to prefer a recent statement over an old one, and preferring the recent statement is most of what memory scoring does.
High cardinality filters have a cost, which is where the performance framing does apply. Filtering on a field with millions of distinct values forces the index to consider and reject many candidates, and some stores handle that far better than others. It belongs on the evaluation list in which vector database should you choose, but never as a reason to relax the boundary.
Scoping controls which records may be returned. It says nothing about whether similarity was the right way to rank the ones that survive: does vector search alone retrieve the right memory?
Retrieval quality
Does vector search alone retrieve the right memory?
Often, but it misses exact strings, and memories are full of them. Similarity search finds records whose meaning is close to the query, which is the right behaviour for a question phrased differently from the stored fact and the wrong behaviour for a question naming an order number, an error code, a ticket reference or an internal product name.
An embedding places a rare token somewhere unhelpful, because the model has little evidence about it, so two different ticket references end up near each other and neither ends up near the query that names one of them. A keyword search has the opposite profile: it is exact where the vector is vague and useless where the wording differs. Running both over the same collection and fusing the rankings is standard practice, and most established stores support it natively.
The full treatment, including how the two rankings are combined, how to weight them, and whether reranking is worth its latency, is on hybrid search for agent memory. What matters when choosing a database is only whether it can do it in one query rather than requiring you to run two systems and merge the results in application code.
There is also a class of question that neither retriever answers, no matter how they are fused: anything asking how two facts relate, or what was true at a particular moment. Those need traversal rather than ranking, which is the subject of vector versus knowledge graph memory. Recognising that a query is of that kind is more useful than tuning retrieval that was never going to answer it.
Retrieval quality is a solvable problem. The harder issue is the set of jobs the store never does at all: what does a vector database not do for memory?
Boundaries
What does a vector database not do for memory?
It does not decide what to store, it does not reconcile contradictions, and it does not forget. Those three gaps are the entire reason the memory tooling category exists, and every product in it will tell you about them, because each is selling the layer that fills them.
The first gap is selection. Hand a raw conversation turn to any vector store and you get an embedded, searchable copy of that turn. You do not get the fact inside it as a separate retrievable statement, and the difference shows at read time, when the retrieved record carries the surrounding conversational filler into a context window you were trying to keep small.
The second gap is supersession. If a stored record says the user prefers one thing and a later statement says they now prefer another, both live in the index as independent vectors with no relationship between them. Similarity search will happily return both, because both are relevant to the query, and nothing in the store marks one as no longer true. Metadata filters narrow which records a query considers; they do not decide which of two candidates is currently correct.
The third gap is forgetting. A vector store grows until something deletes from it, and nothing will, because deletion is a policy question rather than a database feature. Left alone, the oldest and least useful records eventually outnumber the useful ones and start crowding retrieval results. Forgetting and eviction covers the policies; the store will implement whichever one you choose and none by default.
These are not shortcomings. A vector database is doing exactly what it was built to do, and the three jobs above are a different layer of the problem. That layer is either a product, such as Engram, Mem0 or Zep, or it is code you write and own. Both are legitimate. What is not legitimate is assuming the database is doing it, which is the most common way a memory feature ships broken.
Every source that names these gaps is selling the fix, so nobody writes the other half of the analysis: when is a plain vector store enough on its own?
The counterweight
When is a plain vector store enough on its own?
When the records are appended rather than revised. That single property decides it. If nothing you store can later become false, the extraction and reconciliation layers are solving a problem your application does not have, and adding them buys complexity and a per-write model call in exchange for nothing.
Three shapes qualify. The first is a knowledge base an agent reads from: documentation, policies, past tickets, transcripts kept for reference. Records are added and occasionally removed wholesale, never edited into a contradiction, and the read is a retrieval the agent cites rather than an assertion it makes in its own voice.
The second is short-lived context. Session memory that expires when the session does never has time to accumulate contradictions, and a store scoped to one thread with a time-to-live is complete as it stands. The distinction is set out on short-term versus long-term memory.
The third is a small store per user. With a few dozen records per identity, retrieval quality is not the binding constraint, because almost everything relevant comes back regardless of method. Effort spent on retrieval sophistication at that size would be better spent on what gets written in the first place.
The honest test is to ask what happens when a user says something that contradicts what the agent already knows. If the correct behaviour is to keep both records and let recency sort it out at read time, a plain store with a written-at field handles it. If the correct behaviour is to mark the old record superseded so it never surfaces again, you need the layer, and no amount of retrieval tuning substitutes for it.
Once you know which of those two you are building, the product question becomes answerable: which vector database should you choose for agent memory?
Selection
Which vector database should you choose for agent memory?
Four decisions separate them, and none of them is recall on a benchmark. Whether you want a memory layer included or only a store, whether the data must live inside a database you already run, how filtering behaves under multi-tenancy, and whether hybrid search happens in one query.
| Option | What it is in a memory stack | Fits when |
|---|---|---|
| Engram on Weaviate | Managed memory pipeline over a hybrid vector store, so extraction and reconciliation are not yours to build | You want the memory layer and the store from one place |
| Weaviate | Open source hybrid store used directly, with the memory layer written by you | You want native hybrid search and control of the pipeline |
| Pinecone | Serverless store with high write throughput and many namespaces | Tenant count is large and you are pairing it with a memory framework |
| pgvector | A vector column in the Postgres you already operate | Memories must join to relational data or stay in one database |
| Qdrant | Store built around filtered search | Nearly every read is scoped by user or by attribute |
| Milvus | Distributed store for very large collections | Vector count is in the hundreds of millions |
| Chroma | Embedded store for local development | You are prototyping and will migrate before production |
The decision that matters most is the first one, and it is not really a database decision. Buying a store means you own extraction, reconciliation and eviction. Buying a memory product means you accept its opinions about all three. Teams that need those rules inspectable per write, or that need a memory readable in the same transaction that wrote it, are better served by the store plus their own code, which is the tradeoff set out on Engram.
The second decision is quieter and often decisive. If your agent has to answer questions that combine a memory with a row in your application database, a separate vector service means two round trips and no way to join. A vector column in the database you already run removes an entire class of consistency problem, and it is why pgvector wins arguments that have nothing to do with vector search quality. The wider tier question is on storage backends for agent memory.
Benchmarks published by vendors about their own products are not useful input here. Every store in the table is fast enough for a memory workload, which involves a handful of queries per turn against a collection scoped to one user. What differs is operational shape, and that is where the next question leads: what breaks first at scale, and what does it cost?
Scale and cost
What breaks first at scale, and what does it cost?
Memory footprint breaks first, and the bill is dominated by embedding calls rather than by storage. Both are arithmetic you can do before committing to anything.
Take the footprint. A vector of 768 dimensions stored as 32 bit floats occupies 3,072 bytes. A billion of them is 3.07 terabytes before any index overhead, and a graph index adds its own edges on top. That number is why quantisation exists and why compressed indexes are the norm at that size. It is also why the number rarely applies to memory: a memory store holding a few hundred records per user reaches a billion vectors only at a scale most products never see. Run the multiplication for your own record count before assuming you have a scale problem.
What does bite earlier is filtered search under many tenants. A store that is fast on unfiltered queries can degrade sharply when every query carries a predicate that excludes most of the index, and the degradation depends on index family, as set out in Figure 2. This is the load test worth running before launch, and it has to be run with realistic tenant counts, because the behaviour at ten tenants tells you nothing about the behaviour at ten thousand.
On cost, the bill has three lines. Storage and query charges from the store itself, which are the visible ones and usually the smallest. Embedding API calls, one per write and one per read, which scale with conversation volume rather than with corpus size. And model inference on the extraction step if you run a memory layer, which is the largest of the three per turn.
The saving that offsets all of it sits at the other end. Retrieving a small set of relevant memories rather than replaying the conversation cuts the tokens sent to the model on every turn, and the published research on Mem0 measures roughly 1,800 tokens per query for selective retrieval against roughly 26,000 for a full context approach on the LOCOMO benchmark (Chhikara et al., 2025). Storage grows, inference shrinks, and inference is the expensive half. The full evidence is on agent memory benchmarks.
Costs and load are predictable. The failures are less so, because the familiar ones behave differently here: which failures look different when the records are memories?
Failure modes
Which failures look different when the records are memories?
All the standard retrieval failures, with worse consequences. Context pollution, bad chunking and stale records are the familiar list from document retrieval. Each one changes character when the record is a statement about a person that the agent will repeat as its own knowledge.
Context pollution in document retrieval means a few weak passages dilute the context and the answer gets vaguer. In memory it means a superseded fact arrives alongside the current one at a similar score, and the model has no way to tell them apart, so it may assert the old one. The fix is not retrieving fewer records. It is marking supersession on the write path so the old record never competes.
Chunking in document retrieval splits a sentence awkwardly and costs a little relevance. In memory it can split a fact from the condition that qualifies it, so that a stored preference loses the circumstance under which it applies and gets retrieved for situations where it is wrong. Storing whole extracted statements rather than fixed size windows avoids the problem entirely, which is another reason the extraction step earns its cost.
Staleness in document retrieval means an outdated page is served and someone notices the date. In memory there is no date to notice, because the model presents the record as something it simply knows. Without a written-at field and a read path that prefers recent records, staleness is undetectable from the outside, which is what makes it the most damaging of the three.
A fourth failure has no document analogue at all: silent scope leakage, covered in Figure 3. It produces no error, degrades no metric, and appears only as an agent that knows something it should not. It is worth an explicit test rather than trusting that it cannot happen.
Diagnosing any of these means being able to see what was retrieved and why, which is what evaluating agent memory exists for. Once the failure modes are understood, the last question is architectural: where does the vector store sit in the rest of the memory stack?
Placement
Where does the vector store sit in the rest of the memory stack?
It is the semantic tier, and it is one of several. A production agent keeps working state for the current turn, a record of what happened, durable facts it has learned, and the instructions that shape how it behaves. Only the third of those is naturally a vector problem.
Session state is small, hot and read on every turn, which is a cache workload. Event history is append only and queried by time or by thread, which is a relational workload. Learned facts are queried by meaning, which is the vector workload. Procedural instructions are configuration that has to be read exactly rather than approximately. The four are set out properly on types of AI agent memory.
Teams split into two camps about how to host them. A split stack runs a cache, a relational database and a vector store side by side, each doing what it is best at, at the cost of three systems to operate and no way to write across them atomically. A unified stack puts every tier in one database that supports vectors alongside tables, trading some specialisation for one system, one backup and one transaction.
Both camps have vendors arguing their case, and the argument usually reduces to what the team already runs. If Postgres is already in production and the memory volume is moderate, adding a vector column is less work and fewer failure modes than adding a service. If the vector workload is large or the hybrid search requirements are specific, a dedicated store earns its keep. Neither choice is a mistake, and neither is worth a rewrite once the other is running.
What is worth stating plainly is that a vector database is a component, not an architecture. It answers one question well: which stored records are most similar to this one. The system around it decides what got stored, whether it is still true, and what deserves a place in the prompt. The whole loop is set out on how AI memory works, and the build sequence on adding memory to an agent.
FAQ
Frequently asked questions
The decisions that come up once the architecture is settled: schema, migration, and what to test.
What metadata fields should a memory collection have from day one?
Identity, thread, written-at, source and a validity marker. Identity is the isolation boundary, written-at is what lets a read prefer recent statements, and the validity marker is what a supersession step writes to. Fields you skip cannot be backfilled for records already stored, so add them before the first write rather than after.
Can I change the embedding model later without re-indexing?
No. Vectors from two models are not comparable, so a model change means re-embedding every record in the collection and rebuilding the index. Plan for it as a migration with a dual-write window rather than a configuration change. See embeddings for AI memory.
How do I test that memory scoping actually holds?
Write an assertion that a query scoped to one identity never returns a record belonging to another, seed the store with records from at least two identities, and run it on every build. Scope leakage produces no error and degrades no metric, so it is only ever caught by a test that looks for it explicitly.
Should each tenant get its own collection or a shared collection with a filter?
Separate collections or namespaces make a missing predicate return nothing instead of everything, which is the safer failure. A shared filtered collection is simpler to operate and scales better past a few thousand tenants. Pick separation when the data is sensitive, and a filter when tenant count is high.
How many memories should a retrieval return per turn?
Few enough that a wrong one is unlikely to be included. Retrieval size is a precision decision, not a recall decision, because every extra record is another chance for a superseded fact to reach the prompt. Tune it against real queries as described on evaluating agent memory.
Do I need a vector database if my agent only has a handful of facts per user?
Usually not. At a few dozen records per identity you can load them all and let the model sort it out, which avoids an entire service and a per-read embedding call. Add the store when the record count per user makes that impractical, not before.