Developers · Cluster hub
Building Applications Where Memory Is a First-Class Component
The retrieval code is the small part. Adding memory introduces a new class of user data into your schema, a listing and deletion surface you now owe people, tests that stop being deterministic, and a store somebody has to operate. This hub covers what changes in the application around the model, and routes to the eight developer guides on this site.
What memory touches
Scope of work
What does a developer actually build when adding memory?
Not the retrieval, which any library provides. You build the policy: what counts as worth remembering in this product, whose memory it is, how long each kind lasts, and what the user can see and change. Those four are product decisions wearing engineering clothes.
This is worth being clear about early, because tool selection absorbs a disproportionate share of the planning. Extraction prompts, store adapters and similarity search are genuinely interchangeable, and swapping one library for another is an afternoon. The selection rule that decides what your agent remembers is not portable, and getting it wrong is not visible until the store is full.
The third column is the one nobody scopes. A memory store means deletion requests you must be able to satisfy, backups that now contain statements about users, migrations run against live data, and an extraction cost that scales with usage rather than with headcount.
The decisions themselves, and the order to take them in, are on the memory build guides. What is worth setting out first is how much of your application, outside the AI parts entirely, this touches: what changes in the application around the model.
The blast radius
What changes in an application when you add memory?
Four things, none of which are model work: the data model, the privacy surface, the test suite and the operational footprint. Teams estimate the retrieval code and are surprised by these, which is the usual reason a two-week memory feature takes two months.
The data model gains a class of record that is unlike anything else in it: statements about a person, produced by inference rather than by that person filling in a field. It needs a scope key, a source, a timestamp, a confidence and an expiry, and retrofitting those is close to impossible once rows exist without them.
The privacy surface follows immediately. If your product stores what a user said, that user can reasonably ask what is held and require its removal, and both have to be supported operations covering the store, its backups and any derived index. The controls are on memory security.
Testing changes most and is discussed least, so it has its own section below. Operations gains a store to back up, an extraction queue that can fail without erroring on any request, and a cost line that grows with conversations rather than with users.
Before any of that is committed to, one architectural decision constrains the rest: choosing a memory architecture for your application.
Architecture
How do you choose a memory architecture for your application?
From the shape of what you need to remember, not from the shape of your existing infrastructure. Three questions settle it in most cases, and none of them is about which database you already run.
Do the facts change? If entitlements, policies or account states have histories you must be able to reconstruct, you need temporal invalidation, and a graph that models validity natively will save you building it. If the facts are stable preferences, a simple store is enough.
Do questions span several facts? A question needing three linked facts from different sessions is answered badly by top-k similarity, which returns three plausible fragments and no connection between them. That is the case graph representations exist for, covered on knowledge graphs for AI memory.
Do exact strings matter? Order numbers, error codes, SKUs and symbol names are lexical, and pure vector search retrieves them unreliably. Where they matter, hybrid retrieval is not an optimisation, it is a requirement, as set out on hybrid search.
Answer all three with no and a relational table with text search will carry you a long way, which is the right place to start. The patterns themselves, and how the layers fit together, are on the memory architecture hub, and where the layer sits in a stack on what a memory layer is.
Whichever architecture you pick, it will meet the same four constraints once traffic arrives: why memory becomes the bottleneck.
Scale
Why does memory become the bottleneck at scale?
Because it adds work at three different frequencies, and only one of them is visible in per-turn monitoring. Retrieval runs per turn, extraction runs per conversation, and store growth acts over a user’s whole tenure.
Retrieval latency sits on the critical path of every reply and should be watched at the p99 rather than the average, because a slow tail on a search is felt as an agent that hesitates. Prompt weight is the tokens each injected memory costs on every turn, and past a handful of memories it costs accuracy too, an effect measured in research on long-context degradation and described on context rot.
Extraction is usually the largest cost and the least monitored, because it happens after the conversation and appears in no request trace. One model call per closed thread across a million conversations is a real number that arrives on an invoice rather than on a dashboard.
Store growth is the structural one. Precision falls as a user’s store grows, so the experience degrades for long-tenured users while looking fine in testing, where every account is new. The remedies are a selection rule at write time and a consolidation pass, on writing memories and memory consolidation, with cost control on reducing token cost with memory.
All four are discovered late for the same reason, which is that the test suite could not have caught them: testing an application that has memory.
Testing
How do you test an application that has memory?
By treating the memory store as a database fixture: seeded before each case, reset after, and asserted on directly rather than only through the agent’s reply. Memory makes an application stateful between runs, which is the assumption most existing suites are built to avoid.
Seed and reset. A test that runs against whatever the previous test wrote will pass alone and fail in a suite, or worse, pass in both and for the wrong reason. The store gets the same treatment as any other database in your fixtures.
Assert on retrieval separately. Check which memories came back and in what order, not only the final text. A wrong answer can mean the memory was never stored, was stored and not retrieved, was retrieved and ranked below the cutoff, or was retrieved and ignored by the model, and only separate assertions distinguish those four.
Pin the clock. Any scoring function with recency decay is time-dependent, so a suite that passes today fails in three months for no reason anyone will find quickly. Freeze time in tests as you would for any other temporal logic.
Test the absences. Nothing stored, a memory belonging to another user, and a memory that has been superseded. The second is a privacy test and belongs in the suite that gates deploys, because a scope leak found in production is an incident rather than a bug.
Tests tell you the code is right in the cases you thought of. Production tells you the rest, and only if you logged enough to reconstruct it: what to log when memory is in the loop.
Observability
What should you log when memory is in the loop?
The retrieval decision, not just the outcome: which memories were considered, what they scored, which were injected, and which were dropped by the cutoff. Without that record a memory failure cannot be diagnosed after the fact, and memory failures are almost always reported after the fact.
The reason is that four distinct faults produce one identical symptom. A user reports that the agent forgot something. It could have been never written, written and superseded, retrieved below the cutoff, or retrieved and ignored by the model. A log holding only the prompt and the reply cannot separate them, so every such report becomes an investigation from scratch.
Four fields cover most of it: the memories returned with their scores, the subset actually injected, the store and scope the search ran against, and the retrieval latency. Together they turn a vague complaint into a specific line to look at, usually in under a minute.
Two things should be logged on the write side as well, because they fail silently. The number of candidates extracted from each closed conversation, which goes to zero when a background queue breaks without erroring, and the proportion of writes that were merges rather than appends, which is the earliest signal that deduplication has stopped working.
One caution: these logs contain user statements, so they inherit the same retention and deletion obligations as the memory store itself, as covered on memory security. A deletion request that clears the store and leaves the traces has not been satisfied.
Instrumented this way, the remaining failures are design rather than diagnosis: what developers get wrong first.
Pitfalls
What do developers get wrong when adding memory first time?
Choosing the store before the policy, keying memory to the session, running extraction on the turn, and shipping without a way to inspect what was stored. The fourth is the one that makes the other three expensive.
Store before policy is the most common. A vector database is chosen in week one, and the question of what deserves to be in it is deferred until the store is full of conversation fragments. The volume decision determines whether the storage decision matters at all, so it comes first.
Session keying demos perfectly, since a demo is one conversation, and fails at exactly the moment a returning user would have noticed the benefit. Inline extraction doubles per-turn latency for a result nothing reads until the next session.
No inspection tool is the multiplier. Without a way to list what is stored for a user, every one of the above is diagnosed by guesswork, and the support question “why did it say that” has no answer. An internal page that lists a user’s memories with their sources is a day of work and pays for itself in the first week.
A fifth is architectural rather than technical: treating memory as a feature to be added rather than as a property of the application. Memory changes what your product knows about people, which is a design decision before it is an engineering one. The application shapes that follow are on AI-native applications.
Which is where the rest of this cluster picks up: which developer path to take next.
Routing
Which developer path should you take next?
Pick by what you are deciding right now, not by what you are curious about. The eight guides in this cluster are ordered here as a project encounters them.
- Deciding whether the thing you are building is AI-native at all: what separates AI-native from AI-enabled applications, with examples on AI-native app examples.
- Designing the application shape: what an AI-native application is and its tech stack.
- Choosing an agent pattern: agentic architecture patterns, and what each one needs from memory.
- Building the thing: how to build an AI agent, and what an agentic application is.
- Deciding on the memory component itself: what a memory layer is and whether to build or buy one.
Two adjacent clusters matter more than the rest of this one. The build order and the minimum system worth shipping are on the memory guides, and the internals of the write and read paths on the memory architecture hub.
If you read one thing next and are building rather than surveying, make it adding memory to an agent, which is the end-to-end version of everything above.
FAQ
Frequently asked questions
The questions that come up in planning: effort, ownership and what memory does to an existing codebase.
How long does it take to add memory to an existing application?
The retrieval and write code is days. The schema decisions, the listing and deletion surface, and the test changes are what set the real timeline, and they are the parts that cannot be retrofitted cheaply.
Does adding memory require a vector database?
No. A relational table with text search handles a few thousand facts per user. A vector index earns its place when retrieval starts missing memories phrased differently from the query. See storage backends.
Who owns agent memory in a team, product or engineering?
The policy is product: what is remembered, for how long, and what the user can see. The store and its operations are engineering. Projects go wrong when the selection rule is treated as an implementation detail.
Can you add memory without changing the database schema?
Not usefully. Memories need a scope key, a source, a timestamp and an expiry to be correctable and deletable later, and adding those to rows that already exist means guessing values that are no longer recoverable.
How do you handle a user asking to delete their memories?
Make it a supported operation across the store, its backups and any derived index, and keep memories as discrete rows rather than one summary blob so a deletion is a row rather than a rewrite. See memory security.
What should be monitored once memory ships?
Retrieval latency at the p99, tokens added per turn, extraction calls per closed conversation, and store growth per user. The last two are the ones that appear on an invoice before they appear on a dashboard.