Advanced · Frontier research

Latent Memory and Token-Level Memory in LLMs

Latent memory is memory the model carries in its own internal representations, hidden states, cached attention values and learned vectors, instead of in readable text. One kind of it already runs in production on almost every serious deployment. The rest is research, and the reason it is still research has less to do with whether it works than with what you lose when memory stops being something a person can read.

Three forms

1
Token
Text outside
2
Parametric
In the weights
3
Latent
In the activations

Definition

What is latent memory in AI agents?

Memory carried implicitly in a model’s internal representations rather than stored as explicit, human-readable tokens or as dedicated parameters. That is the definition given in the survey Memory in the Age of AI Agents (Hu et al., 2025, arXiv:2512.13564), which is the most complete map of the field currently available and the source this page follows.

Three forms agent memory can take: token-level memory held as text outside the model, parametric memory folded into weights, and latent memory carried in the model's own activations.
Figure 1. The three forms differ in what the memory is physically made of, which decides everything else about how it behaves.

That survey organises agent memory by form, function and dynamics, and the form axis has exactly three positions. Token-level memory is text: statements, summaries and transcripts held outside the model and inserted into the prompt when they are needed. This is what nearly every memory product on the market is made of. Parametric memory is weights: knowledge folded into the model by training, or into an adapter loaded beside it. Latent memory is everything the model computes and carries internally, including the key-value cache, hidden states and vectors learned specifically to stand in for text.

The word “latent” is doing precise work here. It does not mean hidden in the sense of private, and it does not mean compressed. It means the memory exists in the model’s own representational space and is never rendered into language on its way to being used. There is no sentence anywhere in the system that says what the memory holds.

That is also the property that produces the benefit and the problem at the same time, and it is the reason the term gets applied so loosely. Before going further it is worth removing the thing most readers are already thinking of, because it is not this: whether latent memory is the same as storing embeddings in a vector database.

The confusion

Is latent memory the same as storing embeddings in a vector database?

No, and this is the single most common misreading of the term. An embedding in a vector store is an index. Latent memory is the content. Both are vectors, which is why the two get collapsed, but they occupy different positions in the system.

Embeddings in a vector store index text that is retrieved and read as language, while latent memory is a representation the model consumes directly and never decodes into words.
Figure 2. Both involve vectors. Only one of them stores the memory as a vector rather than storing text and finding it with one.

Follow what happens to a memory in an ordinary retrieval-backed system. A statement is written as text. It is embedded so that it can be found later. A query is embedded, the nearest statements are located, and those statements are pulled out and put into the prompt. The model reads words. The vector was a filing system, and at no point did the model consume the vector itself. That is token-level memory with a vector index in front of it, which is what nearly every product described as an AI memory layer is doing.

In a latent memory system, the vector is what the model consumes. There is no decoding step, no text version, and nothing to print. If you asked the system to show you the memory, the honest answer would be a tensor.

The distinction has a practical consequence that goes beyond terminology. Everything you can do with a text memory, reading it in a support ticket, correcting a wrong fact, showing a user what is stored about them, moving it to a different model next quarter, follows from the memory being text. None of it follows automatically when the memory is an activation. The infrastructure behind the ordinary case is on vector databases and embeddings.

With the boundary drawn, the research becomes much easier to read, because it separates cleanly into three groups: the three kinds of latent memory.

Taxonomy

What are the three kinds of latent memory?

Generated, reused and transformed, distinguished by where the latent state came from. The taxonomy is the survey’s, and it is the most useful thing in the literature for a practitioner, because the three groups have completely different maturity.

Three origins of latent memory after the Memory in the Age of AI Agents survey: generated internal states, reused key-value cache, and transformed or compressed latent state.
Figure 3. The taxonomy comes from the survey Memory in the Age of AI Agents. One of the three is already running in production everywhere.

Generate covers work where a model or an auxiliary module deliberately produces compact internal states to stand in for something longer. The survey groups a decade of methods here: gist tokens trained so that a long prompt collapses into a handful of internal tokens, summary vectors that act as soft prompts in place of an entire document, and learned memory tokens that hold facts across a long conversation. MemGen (Zhang et al., ICLR 2026) sits in this group and is the clearest recent example, generating a latent token sequence during reasoning rather than retrieving text.

Reuse covers latent state that already existed and is carried forward instead of recomputed. The main instance is the attention key-value cache, kept between turns or across requests rather than rebuilt from the prompt each time. Nothing new is learned and nothing is compressed; the model simply does not throw away work it already did.

The three are not variations on one technique. Generate involves training something; reuse involves training nothing and simply retaining what already exists; transform involves discarding most of what exists on a rule. Reading a paper is much easier once you have decided which of the three it belongs to, because the claims each group can make are different: generate claims capability, reuse claims cost, transform claims both at some loss of fidelity.

Transform covers taking existing latent state and cutting it down: pruning the cache to the entries that carry the most attention weight, pooling across layers, or compressing so that the essentials survive in a fraction of the space. This is the group with the most immediate engineering value, because it addresses a cost that every long-context deployment already pays.

Those three groups are not equally speculative, and confusing their maturity is why this topic reads as either futuristic or as something you have been doing all along: whether any form of latent memory is used in production today.

Maturity

Is any form of latent memory used in production today?

Yes: the reuse kind is everywhere, under names that do not mention memory. Prompt caching, prefix caching and key-value cache reuse are all the same idea, which is holding onto internal state so a later request does not recompute it, and they are standard features of modern inference stacks and model APIs.

It is worth sitting with that, because it changes how speculative the topic feels. If your application sends a long system prompt on every request and your provider bills the repeated prefix differently from the new tokens, you are already running on latent state carried between requests. Nobody calls it memory in the product documentation, and by the survey’s definition it is exactly that.

What has not shipped is the generate kind. Learned latent memory that carries a user’s history between sessions, in a form no one ever decodes, is confined to research systems and evaluations at the time of writing. The transform kind sits between the two, with cache compression techniques appearing in serving frameworks as performance features rather than as memory features.

This split explains an odd property of the discourse. Practitioners read a paper about latent memory and conclude it is years away; infrastructure engineers read the same paper and recognise their cache. Both are right about different thirds of the taxonomy.

Reuse ships easily for a reason that is worth understanding before relying on it, because unlike a row in a database this kind of memory is expensive to keep: where latent memory lives between turns.

Lifecycle

Where does latent memory live between turns?

In memory attached to the machine that produced it, for as long as something is willing to pay to keep it there. That single sentence contains most of the operational difficulty, because text memory has none of these properties.

A row in a database is small, cheap, durable and reachable from anywhere. Latent state is none of those. Attention cache for a long conversation is measured in gigabytes rather than kilobytes, it lives in accelerator memory where space is the scarcest resource in the system, and it is meaningful only to the process that created it. Keeping it is a decision to hold expensive capacity idle between one turn and the next.

Three consequences follow, and they are the reason this form is treated as infrastructure rather than as a feature. It has to be evicted, because capacity is finite and something must decide which conversations lose their state first, which is the same problem as forgetting and eviction with a much shorter clock. It creates affinity, because a request that wants to reuse state has to reach the machine holding it, which constrains how work is routed. It can be offloaded, moved to slower memory or to disk and brought back, which is cheaper than recomputing it and slower than having kept it.

This is why the reuse kind is usually bounded to minutes or hours rather than being the durable store of what a system knows. It is a cache with an unusually high value per entry, and treating it as long-term memory means paying to keep accelerator capacity warm indefinitely for a user who may never come back.

The generated kind behaves differently and better in this respect: a learned memory vector is small enough to sit in an ordinary database, which is what makes that line of research a candidate for durable memory rather than only for caching. It still cannot be read, which is the constraint the last two sections of this page are about, and none of that changes what the form is good for: what latent memory buys you.

The case for it

What does latent memory buy you?

Three things, and each of them is a direct consequence of skipping text. Latency, context budget, and the preservation of signal that text cannot carry.

Latency. A text-based memory system does work before the model can start: search the store, rank the results, assemble them into a prompt, and let the model read them as tokens. Latent state that is already in the right form skips all of it. Reused attention state skips the prefill for everything it covers, which is why cache reuse is a performance feature before it is anything else.

Context budget. Every token of retrieved memory is a token unavailable to the conversation, and the competition for that space is the real constraint in most agent systems, covered on memory versus the context window. Latent representations are dramatically more compact than the text they stand in for, which is precisely the claim the compression work is built on. A long document reduced to a handful of vectors costs a fraction of what the document costs.

Signal that text loses. This is the subtlest of the three and the most interesting. Writing a memory as a sentence forces a lossy commitment: the extraction step decides what mattered and everything else is discarded. A latent representation can preserve fine-grained contextual detail that no summary sentence would have kept, including things nobody thought to look for. The survey makes this argument explicitly, and it is the reason the research continues despite the practical difficulties.

The three compound rather than adding up. A representation that is smaller also arrives faster, and a system that skips the retrieve-rank-assemble sequence removes a set of failure modes along with the latency, since there is no ranking to get wrong and no budget to overspend. That is the version of the argument worth taking seriously, and it is why the idea keeps returning despite everything in the next section.

A fourth benefit is sometimes claimed and deserves care. Latent memory is not plaintext, so it is often described as more private. It is more opaque, which is not the same thing, and opacity that stops you from auditing your own system is not a security property. That argument is treated on memory security and privacy.

Which is the doorway into the honest half of this page: what you give up by keeping memory in latent form.

The case against it

What do you give up by keeping memory in latent form?

Four capabilities that text gives you for nothing, and every one of them is something a production system eventually needs. The surveys tend to file this under interpretability. In practice it is four separate engineering problems.

Four costs of keeping agent memory in latent form: it cannot be audited, edited, moved to another model, or shown to a user who asks what is stored about them.
Figure 4. Each of these is an engineering or compliance problem rather than a theoretical objection, which is why the frontier work stays in research.

You cannot read it. When an agent says something strange about a user, the first debugging step in a text system is to look at what it remembered. In a latent system there is nothing to look at, so the same investigation becomes an experiment rather than a query.

You cannot edit it. A fact that changes is one row to update when memory is text. Facts change constantly, which is the whole subject of conflicting memories. A superseded fact inside a compressed activation has no address to update, so correction means regenerating the state that contains it, if you can identify which one that is.

You cannot move it. Latent states belong to the representation that produced them. Change model version, quantisation or provider and the stored states may no longer mean what they meant, which turns a routine upgrade into a migration with no clear procedure. Text memories survive every model change, and that portability is worth more than it appears until the first time you need it.

You cannot easily show or delete it. A user entitled to know what is stored about them, or to have it removed, needs an answer. “It is distributed across a compressed representation” is not one. This is the constraint most likely to keep latent memory out of consumer products regardless of how well it performs.

None of this makes the research wrong. It explains the shape of what has shipped: the forms that keep no user-specific content, like cache reuse, went to production quickly, while the forms that would hold a person’s history in an unreadable state have not. There is a nearby form with a similar profile, and comparing the two clarifies both: how latent memory relates to parametric memory and fine-tuning.

Neighbours

How does latent memory relate to parametric memory and fine-tuning?

They share the property of being unreadable and differ in where the information sits and how long it lasts. Parametric memory is in the weights and persists until the weights change; latent memory is in the activations and lasts as long as those states are kept.

The survey draws the line on form rather than on learning mechanism, which resolves a case that otherwise causes confusion. A method that trains an auxiliary model to produce memory vectors is still latent memory, because the memory itself is instantiated as reusable representations rather than absorbed into the model’s parameters. Training was involved; the product of training is not what carries the memory.

The practical difference is scope and lifetime. Parametric memory is shared by every caller of the model and changes only when you retrain, which makes it the wrong place for anything about an individual user and the right place for durable skills and domain behaviour. Latent memory can be per user and per session, held for as long as you choose to hold it, which makes it a candidate for exactly the case parametric memory cannot serve. The comparison with training is on memory versus fine-tuning and the axis itself on parametric versus non-parametric memory.

There is a live research direction that blurs the two, where an agent’s accumulated experience is written into small adapters loaded alongside the base model, giving something that behaves like memory and is stored like weights. It belongs to the same family of questions as learning from experience, covered on continual learning and memory and reinforcement learning.

All of which leaves the question a practitioner actually has, which is what to do about any of this on Monday: whether to build on latent memory now.

The recommendation

Should you build on latent memory now?

Use the reuse and transform kinds, which are performance features you should already be taking advantage of, and build your actual memory in text. That is not a hedge; the two halves of the recommendation are about different problems.

Cache reuse and cache compression are cost and latency work. If your agent sends a long stable prefix on every request, or holds long conversations, the savings are real and available today through your serving stack or model provider, with no change to how your memory is designed. Treat it as infrastructure, and measure it as latency and cost rather than as recall.

What an agent knows about a user should stay text, and the reasons are the four in the previous section rather than any doubt about the research. You will need to read it during an incident, correct it when a fact changes, keep it across a model upgrade, and produce it when someone asks. A store of statements gives you all four for free.

The thing genuinely worth watching is the boundary case, where a system keeps text as the source of truth and a latent representation as a cache in front of it. That gets the latency benefit without giving up the properties above, since anything unreadable can be regenerated from the text that produced it. Nothing in the current tooling makes this easy, and it is where a practical version of this research would land first.

One practical note for teams evaluating vendors. A product describing itself as using latent memory is, in almost every case, describing vector search over text, and the question that settles it in one sentence is whether anyone can print a memory. If the answer is yes, the memory is text and the vector is an index, which is the right architecture for nearly everyone and not what this page is about.

For now, the work that decides whether an agent seems to remember is upstream of all of it, in what gets written and how it is retrieved, covered on writing memories and memory retrieval. The wider frontier, including the directions this page borders on, is surveyed on advanced AI memory research.

FAQ

Frequently asked questions

The boundary questions that come up once latent memory is separated from the vector search it is usually confused with.

Is prompt caching the same thing as latent memory?

It is one kind of it. Prompt caching, prefix caching and key-value cache reuse all keep internal state so a later request does not recompute it, which matches the definition of latent memory in the reuse category. It is sold as a performance feature, which is why almost nobody calls it memory.

Can latent memory be stored in a database?

Generated latent memory can, because a learned memory vector is small. Reused attention state usually cannot in any practical sense, because it is large and lives in accelerator memory. That difference is why one line of this research is a candidate for durable memory and the other is a cache.

Does latent memory make an agent's memory private?

It makes it unreadable, which is not the same as private. Nobody outside can interpret it easily, and neither can you, so incident response and user data requests both become harder. Treat opacity as a cost to manage rather than as a security control.

What happens to latent memory when you change model?

It may stop meaning anything. Latent states are expressed in the representation of the model that produced them, so a version change, a quantisation change or a provider change can invalidate the store. Text memories survive all three, which is a significant practical advantage.

Is latent memory the same as parametric memory?

No. Parametric memory sits in the weights and is shared by every caller until the weights change. Latent memory sits in activations and can be specific to one user or one session. The survey classifies by the form the memory takes, so a latent state produced by a trained module is still latent memory.

Which papers should you read first on latent memory?

The survey Memory in the Age of AI Agents (arXiv:2512.13564) for the taxonomy the field is converging on, then MemGen (ICLR 2026) as a worked example of generated latent memory during reasoning. Both are more useful than the secondary write-ups, which use the term inconsistently.