Advanced · Security and privacy
AI Memory Security: Access Control, PII and Multi-Tenancy
A memory store holds personal data by construction, so it inherits obligations a document index does not: scoped retrieval, deletion on request, and a boundary that prevents one user’s memories reaching another user’s prompt. Most of those are cheap on day one and expensive to retrofit, which is why they belong in the first design rather than the first audit.
Four controls
The exposure
What makes agent memory a security problem?
Memory turns transient conversation into stored personal data, and it does so automatically, without anyone deciding record by record what was worth keeping. That combination is what makes it different from the systems teams are used to securing.
Three properties create the exposure. It is written by inference rather than by a form, so what ends up stored is whatever an extraction step judged durable, which nobody reviewed. It is retrieved into a prompt, which means a retrieval failure is a disclosure rather than a wrong answer. And it accumulates, so a store that held little of consequence at launch may hold a great deal a year later without any change to the code.
The failure that matters most is cross-user leakage: one person’s memory retrieved into another person’s context. It is not exotic, it is the default outcome when scope is not enforced inside the search, and the symptom is an assistant that mentions something the user never said.
A second category is quieter. Memory can capture things a user shared in one context and did not expect to persist, which is a privacy problem even when nothing leaks to anyone else. The mechanism for both is the same write path described on how agents write and store memories.
Almost every control follows from one field being present on every record: scoping memories correctly.
The core control
How do you scope memories to the right user?
Attach a user or tenant identifier at write time, and apply it as a condition inside the search rather than as a filter on the results. The distinction between those two sounds like an implementation detail and is the whole control.
Filtering after the search fails in two ways. It wastes the candidate budget, since results discarded as belonging to someone else were still retrieved. And it makes correctness depend on application code running correctly every time, rather than on the store being incapable of returning the wrong rows.
Multi-tenant products need the same discipline one level up. A tenant identifier alongside the user identifier means a support agent acting on behalf of an organisation can see that organisation’s memories and no others, and it makes an accidental cross-tenant query impossible rather than merely unlikely.
The practical requirement this places on infrastructure is filtering during vector search. Not every store does this well, and it constrains the choice more than raw performance does, as noted on the infrastructure cluster and vector databases for AI memory.
Scope decides who can read a memory. The next control decides whether it should have existed: what not to store.
Minimisation
What should an agent never store in memory?
Credentials, payment details, government identifiers and anything in a special category such as health, and in general anything whose value is lower than the cost of holding it. The safest memory is the one that was never written.
Extraction is the control point, and it is usually treated purely as a quality filter rather than also as a privacy one. The same step that decides whether a statement is durable can decide whether it is storable, and adding an exclusion list there is cheaper than any downstream remediation.
Three categories are worth excluding explicitly. Secrets, because a credential mentioned in passing should never become a retrievable memory. Identifiers, since a stored national ID or card number converts a memory store into a regulated database. Sensitive personal categories, which carry heavier obligations and are frequently mentioned incidentally rather than deliberately.
Minimisation also has a quality benefit, which makes it easier to justify. A store holding less holds a higher proportion of useful material, retrieves faster, and produces fewer irrelevant candidates competing for prompt slots. The privacy argument and the retrieval argument point the same way, which is unusual and worth using.
Even a well-minimised store will hold something a user later wants removed: how deletion actually works.
Erasure
How do you delete a memory properly?
Remove the record, its embedding, and anything derived from it, including summaries and consolidated memories that absorbed its content. Deleting only the original leaves its shadow influencing answers, which is both a privacy failure and a confusing one to debug.
Derived data is the part teams miss. Consolidation merges memories and compresses episodes into summaries, so a fact deleted from the store may survive inside a summary written last month. A deletion path that does not account for that will pass a spot check and fail an audit, and it will produce the unsettling behaviour of an assistant repeating something a user asked it to forget.
Two design choices make this tractable. Record provenance on derived memories, so a summary knows which records it came from and can be regenerated when one is removed. And prefer invalidation over merging where the source matters, since a superseded record that is still individually addressable is far easier to erase than one that was absorbed.
There is a real tension here with the supersession approach described on handling conflicting memories, which deliberately keeps outdated versions so history stays answerable. Both can be true: keep superseded versions for continuity, and delete completely when a user asks. The difference is who is asking and why, and that distinction should be explicit in the data model rather than decided per incident.
Deletion is one obligation among several that arrive with a personal-data store: what else memory brings with it.
Obligations
What obligations come with storing agent memory?
The same ones that attach to any store of personal data: knowing what you hold, being able to show it, being able to delete it, and keeping it only as long as it is needed. The novelty is not the obligations, it is that memory acquires them silently.
A team that would never build a user database without a retention policy can add a memory layer and end up with one by accident, because the store filled itself. The four capabilities worth building deliberately are an inventory of what a store holds about one person, an export, a delete, and a retention rule per memory type.
Retention deserves the most thought, because memory types age differently. Semantic facts stay true until superseded, while episodes have a natural decay, which is the argument set out on long-term memory for AI agents. A single retention period applied to both either discards facts that are still valid or keeps events long past their usefulness.
Transparency is the control that does double duty. A product that shows users what it remembers, and lets them edit or delete it, satisfies a large part of the obligation and is also the feature that makes users forgiving of occasional errors. That is covered from the product side on memory for personal assistants.
Multi-agent systems add one more surface worth naming separately: memory security across several agents.
Multi-agent
How does memory security change with multiple agents?
More writers means more paths to a scope mistake, and shared memory means a single bad write is visible to every agent rather than to one. The controls are the same; the blast radius is not.
Two additional controls matter here. Provenance per agent, so a memory records which agent wrote it and whether it was observed or inferred, which makes a bad batch removable rather than merely regrettable. And least privilege between agents, so an agent that only needs to read a subset is not granted the whole store by default.
The failure specific to this setting is inference laundering: one agent writes a guess, a second reads it as established, and a third treats the second’s conclusion as corroboration. Nothing in the store distinguishes the original inference from the facts built on it, so confidence compounds without new evidence. Provenance is what prevents it, and it is covered on shared memory in multi-agent systems.
The build-side view of all of this is on how to add memory to an AI agent, and the component that should own these controls on the memory layer.
FAQ
Frequently asked questions
The questions that follow: whether memory should be encrypted, and who can see stored memories.
Should agent memories be encrypted?
At rest and in transit, on the same basis as any other store of personal data. Encryption does not address the failure that matters most here, which is cross-user retrieval, so it complements scoping rather than substituting for it.
Who can see stored memories?
Whoever the scope filter allows, plus anyone with database access. That second group is easy to forget: a memory store is queryable by your own team like any other table, so the same access controls and audit expectations apply.
Does a managed memory service change the obligations?
It changes who operates the store, not who is accountable for the data in it. Check that the provider supports per-user deletion, export, and filtering during search, because those are the capabilities the obligations rest on.
What is the most common memory security mistake?
Filtering by user after retrieval instead of during it. The store returns other users' memories and application code drops them, which makes correctness depend on that code running correctly every time rather than on the query being incapable of returning the wrong rows.
How does deletion interact with keeping superseded facts?
They serve different purposes and should be modelled separately. Superseded versions are kept so the agent can answer questions about the past. A deletion request removes a record and everything derived from it, including summaries that absorbed its content.