Guides · Implementation

How to Build Long-Term Memory That Actually Improves

Long-term memory built well is a closed loop, not a one-time write to storage: capture what happened, analyze it for a recurring pattern worth fixing, update memory with a specific correction, and confirm the next run actually reads it. Most of the visible improvement in an agent’s behavior comes from getting that loop right, not from accumulating more facts.

The loop

1
Capture
2
Analyze
3
Update
4
Verify

The core idea

Is building long-term memory a one-time job or an ongoing loop?

An ongoing loop. Treating it as a one-time pipeline, extract facts, write them to a store, done, is what leaves most of the value on the table. The full loop has three working parts plus a step that’s easy to skip: capture traces, analyze them for recurring signal, update memory, then confirm the update is actually used, a structure LangChain’s own engineering research on agent memory documents directly.

Building long term memory as a loop: capture traces, analyze for signal, update memory, verify the update actually loads
Figure 1. Long-term memory built well is an ongoing loop, not a one-time write to storage.

Capture is the evidence layer: a well-instrumented trace records the user’s input, the model’s calls, the tools it used and what they returned, and the outcome, since with an agent, unlike deterministic software, you often can’t tell how it actually behaved without inspecting the trajectory it took. Analysis is where captured evidence turns into a decision, looking for a pattern that recurs rather than reacting to a single one-off mistake, since a single strange output is more often noise than signal. Update is where a specific, scoped change actually gets written, a corrected instruction, a new preference, a reusable example, not a vague note that something went wrong. The step teams most often skip is the fourth: confirming the update is actually read on the next run. A memory system that writes correctly but whose runtime caches the old prompt or the old tool schema has produced a change that looks committed but has zero effect.

Analysis is also where a subtle failure mode shows up if it’s skipped or rushed: the same visible symptom can have several different underlying causes, and treating the first plausible explanation as the correct one produces a fix that addresses the wrong thing. An agent that ignores a tone instruction, for example, might be doing so because the instruction is too vague to act on, because it’s technically present but buried somewhere the agent doesn’t reliably read, because it’s missing from the specific skill or workflow where it actually applies, or because a separate instruction elsewhere quietly contradicts it. Each of those four causes calls for a different update, and writing the wrong one doesn’t just fail to fix the problem, it adds a memory that now has to be reconciled with whatever eventually does get corrected.

Knowing this is a loop still leaves the harder design question open: a loop that captures and updates the wrong things is busywork. What should actually be worth remembering in the first place? What should you actually try to remember?

Design first

What should you actually try to remember?

Name the specific facts the system actually needs before writing any extraction code, a fixed schema it’s explicitly hunting for, rather than accumulating an open-ended pile of free-text memories and hoping the useful parts surface later. The storage technology matters far less than getting this decision right.

Schema-first long term memory design: name the specific fields the system needs instead of storing open-ended free text
Figure 2. Naming the specific facts a system needs makes it possible to know what’s still missing.

A schema-first approach names the specific fields a system is trying to fill for a given task, a user’s stated goal, a stable preference, a fact that shapes future behavior, rather than treating “remember more” as an unqualified goal. This has a direct practical payoff: when a field is still empty, the agent knows exactly what it still needs to ask for, and when a field is filled, it can skip small talk and act on what it already knows. This is a different concern from choosing between unstructured, free-flowing conversational memory and structured, stable facts; most working systems need both, but only the structured half benefits from a schema, since the whole point of the unstructured stream is that it isn’t reducible to fixed fields. The general architecture decisions behind choosing a memory system, including how to weigh structured versus unstructured storage for a specific use case, are covered in full on adding memory to an agent; this page’s concern is narrower, the design habit of naming what you’re hunting for before you start extracting.

Naming what to remember only solves half the noise problem. The other half is deciding what to actually let into memory once a conversation happens, since not everything that gets said is worth keeping. Should every conversation turn become a memory?

Filtering, not accumulating

Should every conversation turn become a memory?

No. A large share of any raw conversation log is small talk, repetition, or transient reasoning that doesn’t belong in durable storage at all, and a system that stores it verbatim degrades its own retrieval quality as that noise accumulates. Filtering has to happen actively; it doesn’t happen on its own.

The practical failure mode is specific: as low-value content piles up alongside genuinely useful facts, retrieval has to work harder to separate signal from clutter, storage and embedding costs rise for no corresponding benefit, and the system as a whole gets slower and less precise even as it technically “remembers more.” The fix isn’t a single filter step bolted onto extraction; it’s treating every candidate memory as something that has to earn its place, worth revisiting alongside the extraction and validation discipline covered in depth on writing memories, which goes further into exactly how that judgment gets made at the point of extraction.

None of this means capturing less at the trace level. The trace itself, the full record of what happened, stays valuable as raw evidence for debugging, even when most of it never gets promoted into durable memory. The filtering decision is specifically about what earns a place in the store an agent actually consults during future reasoning, not about how much gets logged in the first place, since a trace kept purely for observability carries a very different cost and access pattern than a memory the agent reads on every relevant turn.

Filtering decides what enters memory. A separate, and in practice more consequential, question is which kind of memory actually changes an agent’s behavior once it’s stored. Why does procedural memory matter more than most teams expect?

What actually moves behavior

Why does procedural memory matter more than most teams expect?

Because most visible improvements in an agent’s behavior come from fixing how it acts, instructions, workflows, tool-use rules, not from adding another stored fact about what it knows. A team that only builds semantic memory (facts and preferences) has built half a memory system.

Procedural memory, how the agent should behave, drives more visible improvement than semantic memory, what the agent knows
Figure 3. Instructions and workflow rules often fix a repeated mistake more directly than adding another stored fact.

When an agent repeatedly formats output incorrectly, calls a tool in the wrong order, routes work to the wrong place, or ignores a stated tone rule, the fix is rarely “remember one more fact.” It’s procedural: clarify the rule, reorder the steps, or move the behavior into a more specific instruction that actually owns that case. Diagnosis matters here more than it first appears, because the same visible symptom can trace back to several different causes: a rule that’s too vague, a rule that exists but sits in the wrong place to be found, a rule that’s missing entirely, or a rule that’s being silently contradicted by another instruction elsewhere. Writing a fix before correctly diagnosing which of these is actually happening tends to produce a memory update that doesn’t actually change anything, or worse, one that fixes the symptom in one place while leaving the underlying contradiction intact somewhere else.

Getting this diagnosis right, and getting the resulting update to actually stick, is where the loop described earlier either holds together or quietly falls apart. What keeps a self-improving memory loop from quietly breaking?

Engineering discipline

What keeps a self-improving memory loop from quietly breaking?

Four disciplines, each addressing a specific way the loop can fail without anyone noticing: diagnose the real cause before writing a fix, don’t promote every trace to memory, confirm updates actually get reloaded, and protect anything important with an eval that can catch it silently regressing later.

Four disciplines that keep a self-improving memory loop reliable: diagnose first, not every trace is memory, confirm reload, protect with an eval
Figure 4. A self-improving memory loop degrades quietly without these four disciplines in place.

Most trace data should remain history rather than becoming a memory update; only a small, deliberately filtered subset earns durable status, echoing the noise-filtering point above but applied specifically to the analysis step of the loop rather than raw ingestion. A runtime that caches prompts, tool definitions, or skills needs an explicit refresh path when memory changes, or the system can end up storing the correct update while continuing to operate on stale context indefinitely, a failure mode that’s invisible from the outside because the write appears to have succeeded. And any memory update significant enough to actually shape future behavior deserves a way to detect if that behavior later regresses, since a memory system with no evaluation attached has no way to tell the difference between an update that’s working and one that silently stopped mattering months ago.

All four disciplines exist to answer one practical question a team eventually has to ask about whatever they’ve built. How do you know the long-term memory you built is actually working?

The check

How do you know the long-term memory you built is actually working?

The behavior that prompted a memory update should visibly not recur, and that has to be checked deliberately, not assumed the moment a fact is written to a store. A memory system that can’t demonstrate this is passing accuracy checks on the write side while leaving the actual goal, changed behavior, unverified.

This closes the loop this page opened with: capture the trace, analyze it for the real cause, write a scoped update, confirm the update actually loads, and then specifically test that the originally observed failure doesn’t happen again under similar conditions. Session-level persistence, making sure memory survives an actual restart rather than living only in the current process, is a distinct and equally important concern covered in depth on persisting conversation memory. Teams building this loop from scratch are also building an evaluation and iteration practice alongside it, and for teams that would rather adopt this as a managed capability instead of building and operating the loop themselves, an option such as Engram handles the extraction, filtering and retrieval pipeline as a service, which is worth weighing against the engineering discipline described above before committing to build it in-house.

Video

How do you add persistent memory to an agent, shown directly?

A practical walkthrough covering much of the same ground from a hands-on angle.

FAQ

Frequently asked questions

The practical decisions that follow once the loop above is understood.

Is a bigger vector store the goal when building long-term memory?

No. A store full of unfiltered, low-value memories degrades retrieval precision as noise accumulates. The goal is a schema that names what's actually needed and a filter that keeps low-value content out, not raw volume.

Why did a memory update not change the agent's behavior?

The most common cause is that the update was written correctly but never actually reloaded, because the runtime cached the previous prompt, tool definitions, or skill set. Confirm the next run actually reads the new memory before assuming the fix failed.

Should you fix a repeated agent mistake by adding a fact to memory?

Usually not directly. Most repeated behavioral mistakes trace back to an instruction that's too vague, missing, misplaced, or contradicted elsewhere, which calls for a procedural fix rather than another stored fact.

How do you avoid writing a memory that fixes the wrong root cause?

Diagnose before writing. The same symptom, like an ignored tone rule, can stem from several different causes, and each calls for a different fix. Writing a correction before confirming which cause is actually responsible risks adding a memory that later conflicts with the real fix.

Does every conversation trace need to be captured for this loop to work?

Capturing the trace and promoting it to memory are different decisions. Traces are worth keeping broadly as raw evidence for debugging; only a small, deliberately filtered subset should ever become a durable memory an agent reads on future turns.

How do you know a memory-driven fix hasn't quietly stopped working?

Attach an eval to anything important enough to shape future behavior. Without one, there's no way to detect a regression, since a memory update that silently stops mattering looks identical from the outside to one that's still working.