Architecture · The write path

How Agents Write and Store Memories

Writing a memory is four decisions in sequence: what triggered the write, whether the claim is worth keeping, whether it already exists, and what metadata travels with it. Retrieval gets most of the attention, but a store full of restatements, unsourced guesses and facts that stopped being true cannot be rescued by better search.

The write path

1
Trigger
Something happened
2
Extract
Claim proposed
3
Reconcile
Skip or merge
4
Index
Findable

Triggers

What triggers an agent to write a memory?

Four things: an explicit instruction to remember, an inference that something said is durable, the completion of a task with an outcome, and a scheduled reflection pass over memories already stored. They arrive at very different rates and with very different reliability.

Four triggers for a memory write: an explicit user instruction, a model inferring a durable fact, a task outcome, and a scheduled reflection pass.
Figure 1. Explicit writes are rare and trustworthy. Inferred writes are the bulk of the store and the source of most of its problems.

An explicit write is a user saying remember this. It is the only trigger where intent is unambiguous, so it should never be silently dropped by a salience filter, and it deserves a higher confidence value than anything a model proposes on its own.

An inferred write comes from an extractor reading the conversation and deciding a statement is durable. This produces most memories in most systems, which is why the selection rule in the next section matters more than any other single choice in the pipeline.

An outcome write happens when a task ends and the result is known, which is what makes episodic memory useful rather than merely voluminous. A reflection write is a background pass that reads recent memories and writes a more general one, the pattern introduced in the Generative Agents research (Park et al., 2023). It is also the one trigger that can feed on its own output, a risk taken up under the failure modes below.

Whatever fires the write, the same question follows immediately: how an agent decides what is worth remembering.

Selection

How does an agent decide what is worth remembering?

Two tests, both answerable at write time: would this change a future answer, and will it still be true in a month. Advice to be selective is everywhere and is not implementable. These two questions are, because an extraction prompt can ask them and a threshold can act on the result.

The first test removes the bulk of a conversation. Pleasantries, restatements of what the agent just said, and the mechanics of getting to an answer change nothing later. A user’s constraint, a decision reached, a correction they made, and an outcome all change later answers, which is precisely why they are worth the storage.

The second test sorts what survives into where it belongs. A durable preference goes to the user scope. A detail true only inside the current task belongs to the thread and should expire with it. Writing the transient one into permanent memory is how an agent ends up insisting in June on something that was true for an afternoon in March.

Scoring the survivors is the refinement. An importance value assigned at write time, alongside recency and relevance, is what lets retrieval rank a stated dietary requirement above a passing remark, and the scoring functions are set out on memory scoring. The alternative, storing everything and hoping the ranker sorts it out later, is the first failure mode in this pipeline.

Once a claim has passed both tests, it has to be written down in a shape that survives being read months later: what a stored memory record contains.

The record

What does a stored memory record contain?

One atomic claim, plus the provenance that says where it came from and the lifecycle fields that say when it stops applying. The text alone is the part people design; the other two groups are what make a memory correctable rather than permanent.

A stored memory record split into three groups of fields: the claim, its provenance, and its lifecycle metadata.
Figure 2. The claim decides what the memory can answer. Provenance decides whether it can be trusted, and the lifecycle fields decide when it should stop being returned.

The claim should be one statement, not a paragraph. Atomic records deduplicate cleanly, retrieve precisely and can be individually retired. A record holding three facts cannot be corrected without rewriting the two that were fine.

The provenance is the source turn, the timestamp and whether a person stated the claim or a model inferred it. This is the group most implementations skip, and it cannot be added afterwards, because by then the conversation that produced the memory is gone. Without it, a user asking why the agent believes something has no answer, and a suspect memory cannot be traced back to whatever misread produced it.

The lifecycle fields are confidence, the period the claim covers, and which earlier memory this one supersedes. They are what let a store distinguish a fact that has been replaced from one that is merely old, and they make the difference between an audit that is a query and one that is a manual reconstruction. Deletion and retention obligations sit on the same fields, as covered on memory security.

With the shape settled, the pipeline’s real work begins, because most candidates are not new: stopping the same fact being stored twice.

Reconciliation

How do you stop the same fact being stored twice?

Search the store for near matches before every write, then skip, merge or supersede rather than appending. Append-only writing is the default in most first builds and it is why memory stores degrade steadily instead of failing visibly.

A candidate memory passing through near-match search, then a decision to skip, merge or supersede, before being committed and indexed.
Figure 3. The search before the write is the same search retrieval uses, which is why deduplication costs so little to add.

The three outcomes are distinct. Skip when the candidate says what an existing memory already says, adding at most a fresh timestamp. Merge when the candidate is a fuller version of the same claim, replacing the text and keeping the older provenance so the original source is not lost. Supersede when the two cannot both be true, which ends the old memory rather than deleting it.

Similarity thresholds need care in both directions. Set too loose, and two genuinely different preferences collapse into one wrong memory. Set too tight, and a preference stated in five different phrasings becomes five records that all retrieve together and crowd out everything else. Paraphrase is the hard case, because “I avoid dairy” and “I am lactose intolerant” are near-identical in embedding space while differing in a way that matters.

Deduplication also has to run per scope. Two users may hold contradictory preferences with no conflict at all, so a near-match search that crosses the scope boundary is both a correctness bug and a privacy one.

The supersede case deserves its own treatment, because deciding which of two incompatible claims wins is not a similarity question: what happens when a new memory contradicts an old one.

Contradiction

What happens when a new memory contradicts an old one?

The write path should end the old memory rather than delete it, recording when it stopped applying and which memory replaced it. Deleting loses the history, and keeping both without a relationship between them leaves retrieval to pick, which it will do on similarity rather than on truth.

Not every contradiction is a change of mind. A user who says one thing on Monday and the opposite on Friday may have changed their preference, or may have been describing a different context, or the earlier extraction may simply have been wrong. Those three cases want different resolutions, and the provenance fields from the record above are what let the pipeline tell them apart at all.

The general rule is that a claim stated by a person beats a claim inferred by a model, a more recent statement beats an older one at equal confidence, and anything the user has explicitly corrected is treated as authoritative from then on. Where a memory system offers bi-temporal storage, the invalidation is native: the fact keeps both the period it was true and the period the system believed it.

Contradiction is where the write path and the retrieval path meet, and the full resolution rules, including what to show a user when both versions matter, are set out on handling conflicting memory updates. Expiry for memories that nothing contradicts, but that have simply stopped being useful, is on forgetting and eviction.

A committed memory is still inert until it can be found, which is the last mechanical step: how a memory is embedded and indexed.

Indexing

How is a memory embedded and indexed after it is written?

The claim is embedded as its own unit, stored with its metadata as filterable fields, and added to an index that supports both vector similarity and exact matching. What gets embedded is a design decision, and embedding the wrong span is a common quiet failure.

Embed the claim, not the surrounding conversation. A memory embedded together with the pleasantries around it carries their meaning into the vector, so it retrieves for queries about nothing in particular and misses the query it was written for. This is the write-side reason atomic records matter.

Metadata belongs in filterable fields rather than in the embedded text. Scope, type, timestamp and confidence should constrain the search before similarity is computed, since no embedding reliably encodes a user ID, and a scope enforced by similarity alone is not enforced.

Pure vector search is not enough on its own for memory. Order references, error codes, product names and dates are lexical, and similarity retrieves them unreliably, which is why hybrid retrieval combining vector and keyword search is the practical default, as covered on hybrid search. The embedding choice itself is on embeddings for AI memory, and where the records physically live on storage backends.

All four steps have a cost, and where they run decides who pays it: whether writes belong on the turn or in the background.

Timing

Should writes run on the turn or in the background?

Run the cheap writes on the turn and the model-dependent ones in the background. The distinction is not about importance. It is about whether the next turn needs the result, and extraction almost never does.

Appending the turn and updating a rolling summary are inexpensive and immediately relevant, so they belong on the hot path. Extraction, deduplication search, embedding and reflection each involve a model call or an index write whose result will not be read until a later session, and running them inline adds latency to every turn in exchange for nothing the user experiences.

LangChain on the same split, which the LangMem SDK names as conscious formation on the hot path and subconscious formation in the background.

Background writing has its own obligation: the job has to be observable. A silently failing extraction queue produces an agent that appears to work for a session and remembers nothing afterwards, and because nothing errors on the turn, the symptom shows up days later as a vague complaint that memory is not working.

There is one exception worth honouring. An explicit remember-this instruction should be visible immediately, because a user who says it and then asks whether it was recorded expects a truthful yes, and a queue that runs a minute later cannot give one.

Timing errors are one of a small set of recurring write-path defects: what goes wrong in a memory write pipeline.

Failure modes

What goes wrong in a memory write pipeline?

Four defects account for most of it, and all four present as poor retrieval, which is why they are usually diagnosed in the wrong half of the system. Each has a different fix, and better search fixes none of them.

Four write-pipeline failures: storing everything, weighting all memories equally, extracting from generated text, and never expiring anything.
Figure 4. The middle two are the dangerous pair, because both make the store confidently wrong rather than merely noisy.

Storing everything defers the selection decision to retrieval, where it cannot be made well: the ranker sees a hundred plausible candidates and no signal about which were ever worth keeping. Retrieval quality then falls as usage grows, which reads as a scaling problem and is a write policy.

Weighting every memory equally lets a model’s guess rank beside a fact the user stated outright, so whichever is phrased closer to the query wins. A confidence field, populated from the trigger type, costs almost nothing and removes the whole class.

Extracting from generated text is the failure that compounds. If the extractor reads the agent’s own replies as well as the user’s, an invented detail becomes a stored memory, is retrieved later as evidence, and gets reinforced by the next reflection pass. Restricting extraction to user turns and to verified tool output prevents it, and no amount of retrieval tuning undoes it once it has started.

Never expiring anything leaves stale facts retrievable indefinitely while the index slows down. The remedy is the lifecycle fields above plus a maintenance pass, described on memory consolidation, and measured through the write-side metrics on memory evaluation metrics.

Most teams meet these defects through whichever library they adopted, so it is worth knowing what each one does on your behalf: how the frameworks implement memory writes.

Implementations

How do the frameworks implement memory writes?

They differ mainly in how much of the selection and reconciliation they do for you, and in whether writes are synchronous. Nearly every library will store a memory; the useful comparison is what each decides on your behalf before it does.

  • Engram writes into Weaviate, so the deduplication search and the retrieval search run over the same hybrid index. Lexical matching on the write side is what keeps order numbers and error codes from being merged by similarity alone. See Engram.
  • Mem0 runs extraction and a conflict decision automatically on each write, choosing between add, update and delete rather than appending. See the memory tools hub.
  • Zep writes into a bi-temporal graph, so superseding a fact records when it stopped being true as a first-class property rather than as metadata you maintain. See Zep and its alternatives.
  • LangMem exposes the hot-path and background split directly, which makes the timing decision explicit instead of implied. See LangMem.

Whichever you use, the same audit applies: find out what it stores when a user states something durable, whether it searches before writing, and whether it records where a memory came from. A library that appends without reconciling has moved the problem rather than solved it. The full comparison is on the best AI memory tools, and the read half of this loop on how agents retrieve memories.

FAQ

Frequently asked questions

The questions that follow a first write pipeline: who does the extracting, what it costs, and how to tell whether it is working.

What actually performs the extraction when a memory is written?

A model call with a prompt that asks which statements in the turn are durable and worth keeping. It can be a smaller and cheaper model than the one answering the user, since classifying a sentence is an easier task than generating the reply.

How much does a memory write pipeline cost to run?

The dominant cost is the extraction model call, which is why it belongs in a background job batched over a closed thread rather than run per turn. Embedding and storage are close to negligible next to it. See reducing token cost with memory.

Should an agent write memories from its own replies?

No, unless the content came from a verified tool result. Extracting from generated text lets an invented detail become a stored memory that is later retrieved as evidence, and the error compounds with every reflection pass.

How do you tell whether the write path is working?

Measure the store, not just the answers: how many candidates were written versus discarded, the share of writes that were merges rather than appends, and how many memories were never retrieved. See memory evaluation metrics.

Can memories be written directly by a user rather than extracted?

Yes, and they should be. A user-stated memory carries higher confidence than an inferred one, and an interface that lets people correct or remove a memory is also the cheapest fix for a bad extraction.

What is the difference between writing a memory and consolidating it?

A write records one claim from one interaction. Consolidation is a later pass over many stored memories that merges, generalises and retires them. See memory consolidation.