Guides · Persistence

How to Persist Conversation Memory Across Sessions

Persisting conversation memory means writing what the conversation established into a store outside the request, then reading it back the next time the same person appears. The model never held the history in the first place, so the work is not making it remember. It is deciding what to keep, where it belongs, and what to load back when the next session opens.

The loop

1
Rehydrate
Session start
2
Retrieve
Each turn
3
Append
Turn end
4
Extract
Async

Definition

What does persisting conversation memory mean?

It means maintaining an external, editable memory state that survives between interaction sessions and that the agent can write to, update and query while it runs. That definition is worth being precise about, because it separates persistence from two things it is often confused with.

It is not the same as saving a chat log. A log is a record for people to read later; persisted memory is a store the agent queries before it answers. The same rows can serve both purposes, but only one of them affects the reply.

It is also not fine-tuning. Persisted memory sits outside the weights and can be corrected, contradicted or deleted in a single write, which is exactly what a conversation demands, since what a user says today routinely overrides what they said in March. That distinction is worked through on memory versus fine-tuning.

The research literature treats persistence as a lifecycle rather than a storage format: write durable information out of the interaction, maintain it as new observations arrive, and read it back with enough surrounding context to be useful. Infini Memory (Ji et al., 2026) sets out that three-part framing, and this guide follows the same order.

Before any of it is worth building, it is worth being exact about what is actually lost and when: why a chatbot forgets at the session boundary.

The cause

Why does a chatbot forget when the session ends?

Because the API call is stateless: the conversation lived in the request, not in the model, and when the request ends there is nothing left holding it. Within a session the agent looks like it remembers only because your code replays every prior turn on each call.

A conversation across two sessions: history is replayed inside one session, lost at the session boundary, and restored by reading an external store.
Figure 1. The loss is not gradual and it is not a limitation of the model. It happens at one specific point, which is also the point where a store can be read.

This surprises people because the failure looks like forgetting rather than like absence. An agent that has just discussed your project for an hour and greets you as a stranger the next morning appears to have lost something, when in fact nothing was ever written down. The behaviour is covered in more depth on stateless versus stateful LLMs.

The second surprise is that the boundary is not always the one you expect. A session ends when your application decides it does, and a great many systems open a fresh thread on every page load, every reconnect or every new tab. Users experience that as an agent with a much shorter memory than the code appears to give it.

Enlarging the context window does not move the boundary, it only lets more of one session fit inside it, as set out on memory versus the context window.

Given that something has to be written down, the first real decision is what: the transcript, a summary or extracted facts.

The decision

What should you persist: the transcript, a summary or extracted facts?

Persist facts for anything that must outlive the thread, a rolling summary for continuity inside a long thread, and raw turns only when you may need the exact wording again. Most guides list these three and then demonstrate whichever one the tool they are selling implements.

Three persistence options compared: raw turns kept verbatim, a rolling summary that is lossy, and extracted facts stored individually.
Figure 2. The three are not competing designs. They answer different questions, and a production system usually runs all three at once.

The test that resolves it is whether you will ever need to reconstruct what was actually said. Support, legal, healthcare and anything with a dispute process will, so raw turns are kept regardless of cost. A consumer assistant almost never will, and keeping every turn there buys storage and retrieval noise in exchange for nothing.

Raw turns lose nothing and decide nothing, which pushes the whole cost to retrieval: search over a year of verbatim chat returns plausible fragments from conversations that ended long ago. Rolling summaries are cheap to load and irreversibly lossy, so what the summariser drops in March cannot be recovered in June. Extracted facts are small, individually retrievable and correctable, which is why they are the only one of the three that reliably survives the thread that produced them.

The practical default is turns for the current thread, a summary for the thread once it grows past what you want to reload, and facts promoted to the user. That promotion step needs somewhere to promote to: separating one user’s threads from another’s.

Scoping

How do you separate one user’s threads from another’s?

By keying memory in two scopes at once: a thread scope holding this conversation, and a user scope holding what is true of the person across all of their conversations. Systems that model only one of the two produce the two most common complaints about agent memory.

Thread-scoped memory keyed by thread ID holding messages and a summary, beside user-scoped memory keyed by user ID holding durable facts.
Figure 3. Both scopes are read before a reply. Only the user scope survives the thread that wrote it.

With only a thread scope, every new conversation starts empty and the user re-explains themselves weekly. With only a user scope, everything said anywhere is retrievable everywhere, so an aside from a finished conversation surfaces in an unrelated one and the agent looks like it is not listening.

The keys should be independent: a thread ID that means one conversation and a user ID that means one person, with threads belonging to users rather than standing in for them. Frameworks express this as namespaces or as a users-and-threads model, and the shape is the same underneath. Getting it wrong is not merely untidy, because in a multi-tenant product a memory key that is too coarse is how one customer’s context reaches another, as covered on memory security.

A third scope appears once several agents share one user: what one agent writes, another can read. That case has its own rules, set out on shared memory in multi-agent systems.

With the keys settled, the writes can be placed on the turn: writing conversation memory at the end of a turn.

The write path

How do you write conversation memory at the end of a turn?

Append the turn and update the thread summary synchronously, and push fact extraction into a background job. Splitting the write this way is what keeps persistence off the critical path, because the cheap writes are the ones the next turn depends on and the expensive one is not.

The write and read paths of conversation memory: rehydrate at session start, retrieve each turn, append at turn end, extract asynchronously.
Figure 4. Four operations, two of which run while the user waits and two of which do not.

The synchronous half is a database append and, past a threshold, a summary update. Neither needs a model call of its own beyond what the turn already made, so neither is felt. The asynchronous half is the extraction step that reads the closed thread and decides which statements are durable enough to promote to the user scope, and it takes a model call plus a deduplication pass against what is already stored.

Running extraction inline is the most common performance mistake in a first build. It adds a second model call to every turn, doubles the latency the user experiences, and does it for a result that will not be read until the next session opens, by which point a background job would long since have finished.

What gets extracted, how duplicates are merged and how contradictions are settled belong to the write pipeline itself, covered on how agents write and store memories. Whichever store sits underneath, the read path is what the user actually experiences: reloading the right context when a user returns.

The read path

How do you reload the right context when a user returns?

Load the user-scope facts and the thread summary unconditionally at session start, then retrieve additional memories per turn based on what the user actually asks. The two reads answer different questions and running only the second is why some agents feel cold for the first few turns.

The session-start read is small, fixed and predictable: the handful of facts that apply to any conversation with this person, plus the summary of the thread being resumed. It goes into the system prompt, so the agent opens already knowing who it is talking to rather than discovering it three turns in.

The per-turn read is a search, and it should be narrow. Retrieving twenty memories on the chance that one is relevant fills the window with noise and measurably degrades the answer, which is the effect described on context rot. Four to six well-scored memories is the working range, and how they are scored is set out on memory scoring.

Google Cloud Tech walking the same two reads, with the session-start load and the per-turn retrieval shown separately.

One detail is easy to miss: label what you inject. Memories dropped into the prompt without a marker are indistinguishable from things the user said in this conversation, and an agent that cannot tell the difference will claim you told it something today that you told it in March. The retrieval mechanics are covered on how agents retrieve memories.

Threads that keep growing eventually force the question the summary was standing in for: when to summarise a long conversation.

Compression

When should a long conversation be summarised instead of stored whole?

Summarise when reloading the thread would cost more context than the conversation is worth, and keep the raw turns underneath the summary rather than replacing them with it. Summarisation is a read-time convenience, not a storage strategy.

The trigger is a budget rather than a turn count. Decide how much of the window a resumed thread may occupy, and summarise when the transcript exceeds it. A support thread might reach that in fifteen turns and a casual assistant thread not for a hundred, so a fixed number chosen once will be wrong for one of them.

What a summary must preserve is what a later question will need: decisions taken, commitments made, unresolved items and anything the user corrected. What it can drop is the diagnostic back-and-forth. Summarisers left unguided invert this and produce a fluent paragraph about the topic of the conversation that answers nothing anyone will ask.

Summaries also degrade when rewritten repeatedly. Each pass compresses the previous compression, and after enough rounds the specifics are gone while the confident tone remains. Summarising from the raw turns each time rather than from the last summary avoids this, and costs one more model call in a background job where nobody is waiting. Techniques and their trade-offs are on memory summarisation, and the difference from consolidation on memory consolidation.

Compression loss is one of four documented ways persisted memory goes wrong: what breaks when conversation memory is persisted badly.

Failure modes

What goes wrong when conversation memory is persisted badly?

Four failures recur, and the research names them: fragmentation, conflict, compression loss and insufficient retrieval. The Infini Memory paper (Ji et al., 2026) sets out this taxonomy, and having names for them is what turns a vague complaint that memory is not working into something you can locate.

Fragmentation is evidence about one person or one task scattered across many small records, so no single retrieval returns enough to answer with. It is the standard cost of storing every observation as an isolated row.

Conflict is old and new versions of the same fact both sitting in the store, because append-only writing never retires anything. The agent then retrieves whichever scores higher, which is not reliably the true one. Resolution rules are on handling conflicting memories.

Compression loss is the summarisation failure above: temporal order, source and specificity stripped out while the text still reads well. Insufficient retrieval is a single-shot search returning fragments that are individually relevant but collectively inadequate for a question needing several linked facts.

The paper’s own answer is to group related evidence into topic documents and let the agent read memory through several tool calls rather than one search, an approach that reports 64.7% overall on MemoryAgentBench. Whether or not you adopt that design, the diagnosis is the useful part: each failure has a different fix, and treating all four as a retrieval-quality problem fixes none of them. How to tell which you have is covered on memory evaluation metrics.

Which leaves the question most readers arrive with, having already used memory somewhere else: whether built-in platform memory removes the job.

Build or inherit

Do built-in platform memory features remove the need to build this?

Not for an agent you build yourself. Platform memory belongs to the platform’s own chat product and does not travel to your application through the API. The memory that makes ChatGPT or Claude feel continuous is a product feature layered on the model, not a property of the model you are calling.

What the APIs do give you is progressively more of the plumbing: conversation and thread objects that persist message history server-side, so you are not maintaining the transcript yourself. That removes real work, and it removes the least interesting part. The decisions in this guide, what is durable enough to promote, how memory is scoped across threads, what gets injected and what is left out, remain yours because they are product decisions rather than infrastructure.

There are two reasons to keep the memory state on your own side even where a platform will hold it. The first is portability: memory tied to one provider’s thread objects moves only as far as that provider does, and model switching is common enough to matter. The second is inspection, since a support or compliance question about what an agent knew is answerable only if you can read the store.

The practical shape for most teams is a thin memory layer of their own over whatever the platform provides, or a dedicated library that already implements the write and read paths. Those options are compared on the best AI memory tools, with the hosting trade-off on open source versus managed memory. The end-to-end build, once these decisions are made, is on adding memory to an agent.

FAQ

Frequently asked questions

The questions that follow a first persistence build: where the data goes, how long to keep it, and what to do about cost.

Where is persisted conversation memory actually stored?

In a store you control: a relational table for turns and summaries, and a vector store or knowledge graph for retrievable facts. Many teams start with Postgres for both and add a vector index when retrieval quality stops being good enough. See storage backends for agent memory.

How long should conversation memory be kept?

Thread memory can expire when the thread does. User-scope facts should live until they stop being true rather than until a fixed date, which makes an update and expiry policy more useful than a retention window. See forgetting and eviction.

Does persisting conversation memory increase token cost?

It usually lowers it. Replaying an entire transcript on every turn pays for the whole history repeatedly, while retrieving a few relevant memories pays for a fixed small addition. See reducing token cost with memory.

Should conversation memory be written on every turn or at the end?

Both, but different things. Append the turn and update the summary each turn, since the next turn depends on them. Extract durable facts after the thread closes, in a background job where the extra model call costs the user no latency.

Can you persist conversation memory without a vector database?

Yes. Threads, summaries and a modest set of user facts work in a relational database with text search. A vector index earns its place when the fact set grows past what keyword matching retrieves reliably. See vector databases for AI memory.

How do you let a user see or delete what was persisted?

Store memories as discrete records with a source and a timestamp rather than as one opaque blob, so a listing is a query and a deletion is a row. Retrofitting this onto a rolling summary is close to impossible. See memory security.