Tools · Framework Memory

LangChain Memory for AI Agents

LangChain’s memory architecture separates two concrete objects: thread-scoped state persisted through a checkpointer for short-term memory, and namespace-scoped JSON documents in a Store for long-term memory. Both are configured directly in your own agent code, and the design questions that matter, how to shape a long-term memory and when to write it, are yours to answer rather than the framework’s to decide for you.

Two scopes

1
Thread
Checkpointer
2
Namespace
Store
3
Search
Filter or vector

Definition

How does LangChain separate short-term memory from long-term memory?

By recall scope, backed by two different mechanisms: a thread for short-term memory, a namespace for long-term memory. The distinction is not just conceptual; it maps to two different objects a developer actually configures.

LangChain short-term memory scoped to a thread via a checkpointer compared with long-term memory scoped to a namespace via a Store
Figure 1. Thread-scoped state for the current conversation, namespace-scoped documents for everything else.

Short-term memory is thread-scoped: a thread groups the messages and other stateful data belonging to one ongoing conversation, similar to how an email client groups replies into one conversation view. This state is managed as part of the agent’s graph state and persisted to a database via a checkpointer, which is what lets a conversation resume exactly where it left off. It is read at the start of every step and updated as the graph runs.

Long-term memory has no thread. It is saved within a custom namespace, which can include a user id, an organisation id, or any other label useful for organising memories hierarchically, and it is recalled across sessions and threads rather than being tied to one conversation. That single design choice, namespace instead of thread, is what makes a fact learned in one conversation available in a completely different one later.

Both mechanisms exist to answer the same underlying question at different scopes: what does the agent need available right now versus what should it be able to recall regardless of which conversation it is currently in. Deciding what kind of thing belongs in the long-term store at all is a separate question the framework borrows an answer to from elsewhere: where does LangChain’s memory taxonomy come from?

Taxonomy

Where does LangChain’s memory taxonomy come from?

A published research framework for cognitive language agents, cited directly in the framework’s own documentation. LangChain’s docs map three categories, facts, experiences and instructions, onto the same semantic, episodic and procedural distinction psychology uses for human memory, and they attribute this mapping to the CoALA paper (Sumers et al., 2023, arXiv:2309.02427) rather than presenting it as an original framework decision.

Semantic memory holds facts and concepts about a user, an organisation, or the agent itself, and is what most personalization features are built on. Episodic memory holds past experiences, meaning what the agent actually did in prior interactions, useful for an agent that needs to recall its own prior actions rather than only facts about the user. Procedural memory holds rules and instructions, the agent equivalent of the system prompt, and is the category most often left implicit rather than treated as memory that can be inspected and updated.

The taxonomy is a helpful organising tool and not a strict schema the framework enforces; nothing prevents storing all three kinds of information in the same namespace. What it does provide is vocabulary for a design conversation that otherwise defaults to “just save whatever seems useful,” and the full treatment of these three categories, independent of any one framework, is on types of AI agent memory.

Knowing what kind of thing to store is only half the decision. The other half is a design fork the framework’s own conceptual guide states plainly and does not resolve for you: should long-term memory be one profile or a growing collection?

Shape

Should long-term memory be one profile or a growing collection?

A profile is simpler to read and riskier to update; a collection is safer to write and harder to search comprehensively. This is a real design fork with named failure modes on each side, not a stylistic preference.

Profile memory as one continuously updated document compared with collection memory as many narrow documents in LangChain
Figure 2. Each shape trades a different kind of error: lost fields against duplicate or conflicting entries.

A profile is a single, continuously updated JSON document representing well-scoped facts about a user or entity. Updating it means passing in the previous version and asking a model to produce a new one, or a patch to apply to it. As the profile grows, this becomes more error-prone: a model reconciling many fields at once can silently drop or overwrite something that should have survived, which is why large profiles often benefit from being split into several smaller documents.

A collection is many narrower documents, each added as new information arrives rather than merged into an existing whole. Each individual write is simpler and less likely to lose information, since a model generating one new small document has an easier job than reconciling a complex one. The complexity moves elsewhere: the model must now decide whether an incoming fact should become a new entry or update an existing one, and different models default to different mistakes here, some over-inserting duplicates and others over-updating and merging things that should have stayed separate. Searching a collection is also a harder problem than reading one profile, since relevant context may now be spread across several documents rather than sitting together in one place.

Neither shape is a mistake. A profile suits a small, stable set of facts a system needs as a whole on every turn; a collection suits an open-ended stream of experiences or preferences where completeness matters more than a single coherent snapshot. Whichever shape is chosen, the tooling for finding what is relevant in the store matters as much as how it is written: how is long-term memory actually searched?

Timing

Should memory be written during the response or in the background?

Both are real options with opposite tradeoffs, and the framework’s own documentation states both sides rather than recommending one default. This is a genuine engineering decision, not a settled question.

Writing LangChain agent memory on the hot path during the response compared with writing it as a background task
Figure 3. Immediate availability against added latency, on one side; no latency against staleness risk, on the other.

Writing on the hot path means the agent decides what to remember as part of producing its response, the same pattern ChatGPT’s own memory feature uses with a save-memory tool the model calls mid-conversation. The advantage is immediacy: a memory written this way is available on the very next turn, and the write can be made visible to the user as it happens, which is a real transparency benefit. The cost is added latency on every response and split attention, since the agent is now reasoning about what to store as well as what to answer.

Writing in the background means memory formation runs as a separate task, triggered on a schedule, after a period of inactivity, or manually, rather than inline with the response. This removes memory-related latency from the user-facing path entirely and lets writes be batched or deduplicated more easily. Its cost is a scheduling problem the docs name directly: infrequent background runs leave other threads working from stale context in the meantime, so choosing the trigger and frequency is itself a real design decision, not a default to accept unexamined.

Both approaches assume the conversation history feeding either path does not grow without bound, which raises a related and more immediate problem: how do you manage a conversation history that keeps growing?

Short-term management

How do you manage a conversation history that keeps growing?

Trim it, delete specific messages from it, or summarize it, and these are three distinct operations rather than one blurred idea of “compression.” Long conversations threaten to exceed a model’s context window, and even within the window, models perform worse and cost more as history grows.

Three LangChain techniques for managing growing short-term memory: trim messages, delete messages, and summarize messages
Figure 4. Three distinct operations, not one blurred idea of compressing the history.

Trimming counts tokens in the message history and cuts to a limit, typically keeping the most recent messages. It is the cheapest option and the most destructive: whatever falls outside the kept window is simply gone, with nothing preserved.

Deleting removes specific messages from state directly, which is precise rather than positional. This suits a case where a particular message should be removed, such as a tool call whose output is no longer relevant, rather than trimming from one end of the whole history.

Summarizing runs a model over the older portion of the history and replaces it with a condensed summary message, available as built-in middleware in the framework. This preserves the gist of what trimming or deletion would discard outright, at the cost of a model call and the loss of exact wording, figures, or phrasing that a summary does not carry forward.

All three operate on the thread-scoped conversation, not on the long-term store, and none of them decides what should be promoted into long-term memory before the short-term history is cut. That promotion decision, and how well or poorly the framework makes it for you, is the honest limit worth naming directly: what does LangChain’s memory not decide for you?

Honest limits

What does LangChain’s memory not decide for you?

What is worth remembering in the first place, whether two stored facts contradict each other, and when a memory has gone stale. These are architectural gaps rather than missing features, and the framework’s own writing concedes as much: long-term memory is described directly as “a complex challenge without a one-size-fits-all solution,” with a framework of questions offered rather than a built-in answer.

Extraction, the decision that a given statement is worth a memory write at all, is left to whatever prompt or tool schema the developer builds. There is no built-in judgement about durability versus small talk; a poorly specified extraction step over-writes trivial facts just as readily as important ones, and the reference templates are starting points to adapt rather than tuned defaults.

Contradiction handling is likewise not automatic. Storing memories as a growing collection, the shape most likely to accumulate contradictions over time, provides no built-in mechanism to detect that a new fact conflicts with an old one; that logic has to be written on top, following the mechanics on conflicting memories.

And staleness has no native handling either. A memory written six months ago carries the same weight at retrieval time as one written yesterday unless the developer adds a timestamp field and a scoring function that uses it, which is the subject of memory scoring rather than anything this framework provides out of the box.

None of this makes the framework’s memory system worse than alternatives; it makes it a well-documented set of primitives rather than a finished pipeline. Whether that tradeoff is the right one depends on what you are building: is LangChain’s memory enough on its own?

The decision

Is LangChain’s memory enough on its own?

For an agent already built on LangGraph that needs configurable short-term state and a searchable long-term store, yes, provided you build the extraction, contradiction handling and staleness logic yourself. For a team that wants those solved rather than built, no.

The primitives here are genuinely well designed: a clear thread-versus-namespace split, an explicit profile-versus-collection choice with documented tradeoffs, a real hot-path-versus-background decision rather than a hidden default, and three distinct short-term management techniques. Teams already invested in LangGraph get a coherent, well-documented foundation to build on rather than a black box.

The gap is everything the framework explicitly declines to solve: what counts as durable enough to store, whether a new fact contradicts an old one, and when a stored fact should stop being trusted. Those are exactly the responsibilities a managed memory layer such as Engram takes on as a service, worth weighing directly against writing that logic yourself on top of these primitives.

The broader landscape of memory tools, including how framework-level modules like this one compare to managed layers, is on memory tools, and the deeper, framework-independent mechanics of writing, scoring and retiring memories are on writing memories and memory scoring.

FAQ

Frequently asked questions

Configuration questions that come up once the basic architecture is in place.

Do I need a database for LangChain memory, or does the in-memory default work?

The default in-memory store and checkpointer are fine for development and testing but do not survive a process restart. Production use requires a database-backed checkpointer and store, such as Postgres or Redis, covered on storage backends.

Can the same fact live in both short-term and long-term memory?

Yes, and it often should during the transition. A fact mentioned in the current conversation exists in thread-scoped short-term state until something promotes it into the namespace-scoped long-term store, at which point it becomes available to future threads as well.

How do I stop the same fact from being stored twice in a collection?

This is not handled automatically. The write path, whether on the hot path or in the background, has to check the new fact against existing entries and decide to update or skip rather than blindly append, which is the deduplication logic the framework's own docs point to reference packages for rather than providing natively.

Should I use semantic search or content filtering for long-term memory?

Content filtering when memories carry structured fields worth matching on directly, such as a category. Semantic search when the query's exact wording is unknown in advance, which is the more common case for facts extracted from natural conversation. Many production setups combine both.

What happens if the summarization middleware drops an important detail?

It is gone from the summary and only recoverable if the original messages are still retained somewhere, since summarization is lossy by design. For facts that must survive exactly, extract them into long-term memory before the short-term history is summarized rather than relying on the summary to preserve them.

Is LangChain memory tied to LangGraph specifically?

The checkpointer and Store mechanisms described here are LangGraph primitives. Using them means building on LangGraph's execution model; a different orchestration framework would need its own equivalent mechanisms or a framework-independent memory layer instead.