Glossary · 50+ terms

AI Memory Glossary

The vocabulary of AI agent memory, grouped by what each term describes: the kinds of memory, the write path, the read path, maintenance, evaluation, and the research the field borrowed its words from. Where a term is used inconsistently across sources, the entry says so and gives the reading this knowledge base uses.

Core

What do the core AI memory terms mean?

Six terms carry most of the weight, and three of them are routinely used for two different things. Everything else in this glossary is a refinement of one of these.

AI memory vocabulary grouped into four areas: memory types, the write path, the read path, and maintenance.
Figure 1. The four groups below the core terms. A term that seems ambiguous in isolation is usually clear once you know which group it belongs to.

AI memory

Storage and retrieval outside the model that lets an agent use information from earlier turns and earlier sessions. It is application code rather than a property of the model, and it is distinct from the context window, from retrieval over documents and from fine-tuning. The full guide.

Agent memory

The same thing as AI memory, named from the agent’s side. Used interchangeably here. In papers it often implies the full loop of write, store, retrieve, consolidate and forget rather than storage alone.

Memory layer

The component that sits between the agent and its stores, deciding what is written and what is retrieved. Also used commercially to mean a hosted product doing that job. What a memory layer is.

Context window

The bounded span of tokens a model can attend to in one forward pass. It holds the conversation while a session runs and holds nothing afterwards, which is why it is working memory rather than a store. Memory versus the context window.

Statelessness

The property that each model call is independent, carrying no state from the last. Every appearance of memory in an agent is state your application supplied. Stateless versus stateful LLMs.

Context engineering

Composing what goes into the prompt: instructions, retrieved documents, tool output and memories, under a token budget. Memory retrieval is one input to it rather than a synonym for it. Context engineering.

The first of the four groups is the one most sources open with, and the one where counts disagree: what the types of agent memory are called.

Types

What are the types of agent memory called?

Four names recur across almost every source, borrowed from cognitive psychology, with two more appearing once systems run at scale or in groups. Sources that give three types are counting long-term memory only.

Working memory

What the agent is attending to right now: the current instructions, recent turns and fresh tool output. In practice it is the context window itself. Called short-term memory in most engineering writing. Short-term memory.

Semantic memory

Durable facts held without reference to when they were learned: a preference, a plan tier, a specification. The type most systems implement first, because it produces the recognisable behaviour of an agent that knows you. Semantic memory.

Procedural memory

Learned skills, tool-use patterns and behavioural rules, closer to a policy than to a fact. The least mature of the four in production systems. Procedural memory.

Sensory or buffer memory

The very short staging step where raw input sits before anything is decided about it. Folded into working memory by shorter taxonomies. Sensory memory.

Shared memory

A store several agents read and write, so what one learns the others can use. Introduces write conflicts that single-agent memory never has. Shared memory.

Parametric memory

Knowledge held in the model weights, acquired in training or fine-tuning. Not editable at runtime, which is the whole reason external memory exists. Parametric versus non-parametric.

Non-parametric memory

Anything stored outside the weights and reached by retrieval: vectors, graphs, rows, documents. Nearly all agent memory is non-parametric.

Elastic on the short-term and long-term split, which is the division the six types above sort into.

Whatever type a memory is, it arrives through the same sequence of steps: the terms that describe the write path.

Write path

What terms describe how memories are written?

The write path turns a turn of conversation into a stored record, and each of its steps has a name that is worth keeping straight. Most memory-quality problems live here rather than in retrieval.

The life of one memory in four stages, with the glossary terms that apply at each: created, stored, used and retired.
Figure 2. The same term list, ordered by when it applies. Sections three to six of this glossary follow these four stages.

Extraction

The model call that reads an interaction and proposes which statements are durable enough to store. The source of most memories in most systems, and of most bad ones. How memories are written.

Salience

Whether a candidate is worth keeping at all. The implementable form is two questions asked at write time: would this change a future answer, and will it still be true next month.

Deduplication

Searching for near matches before writing, then skipping, merging or superseding rather than appending. Its absence is why append-only stores degrade steadily instead of failing visibly.

Provenance

The fields recording where a memory came from: the source turn, the timestamp, and whether a person stated it or a model inferred it. Cannot be added later, because by then the conversation is gone.

Hot path writing

Writing memory during the turn, while the user waits. Appropriate for cheap appends and for an explicit instruction to remember something. Also called conscious formation.

Background formation

Extraction and consolidation run after the interaction, in a job the user never waits for. The correct place for anything involving a model call. Also called subconscious formation.

Reflection

A scheduled pass that reads recent memories and writes a more general one. Powerful and the one write that can feed on its own output if extraction is not fenced to user turns and verified tool results.

Once written, a memory is inert until something can find it: the terms that describe retrieval.

Read path

What terms describe how memories are retrieved?

Retrieval is a search, a ranking and an insertion, and the vocabulary separates the three more carefully than everyday use does. A memory that is stored, found, and then badly inserted still fails.

Memory retrieval

Finding the stored memories that bear on the current turn. Distinct from document retrieval mainly in what is indexed: one person’s facts rather than a shared corpus. How agents retrieve memories.

Memory scoring

Ranking candidates by a combination of recency, importance and relevance, the formulation introduced in the Generative Agents research. The weights differ by memory type. Memory scoring.

Recency decay

Reducing a memory’s score as it ages. Suits episodes and damages semantic facts, since a preference stated a year ago is exactly as true today.

Reranking

A second pass that reorders an initial candidate set with a more expensive model. Worth its latency when the first pass returns roughly the right memories in roughly the wrong order.

Context injection

Placing retrieved memories into the prompt. Unlabelled injection is a common defect: memories that are indistinguishable from what the user just said produce an agent claiming you told it something today that it stored in March.

Context rot

The measured decline in answer quality as irrelevant material fills the window. The reason retrieval should return four to six memories rather than twenty. Context rot.

Stores that are only written to and read from get worse over time, which is what the fourth group is for: the terms that describe memory maintenance.

Maintenance

What terms describe memory maintenance?

Maintenance is everything that happens to a memory after it is stored and before it stops being retrieved. These terms are used loosely in commercial writing, and the distinctions are real.

Consolidation

A pass over many stored memories that merges duplicates, generalises patterns and retires what is superseded. Distinct from writing, which handles one claim from one interaction. Memory consolidation.

Summarisation

Compressing a conversation or a set of memories into shorter text. Lossy and irreversible, which is why it belongs over the raw record rather than in place of it. Memory summarisation.

Forgetting

Deliberately removing or demoting memories so the store stays useful. A design feature rather than a failure, and the counterpart to salience on the write side. Forgetting and eviction.

Eviction

Removing memories to stay within a budget, by age, by score or by a cap per user. Named for the cache operation it resembles.

Invalidation

Marking a fact as no longer applying from a given date, rather than deleting it. Preserves the ability to answer what was true earlier, which matters wherever a policy or a plan has changed.

Bi-temporal validity

Storing both the period a fact was true in the world and the period the system believed it. Makes invalidation precise and makes a fact’s history reconstructable.

Memory conflict

Two stored memories that cannot both be true, usually created by append-only writing. Retrieval then returns whichever scores higher, which is not reliably the correct one. Conflicting memories.

Everything above lives somewhere physical, which is the next group of terms: storage and infrastructure vocabulary.

Infrastructure

What terms describe memory storage and infrastructure?

These are borrowed from search and database engineering, and they mean the same here as they do there. The only memory-specific twist is that partitioning is a correctness requirement rather than a performance one.

Embedding

A vector representation of text that places similar meanings near each other. What makes similarity search possible, and what makes exact identifiers hard to match. Embeddings.

Vector database

A store indexed for nearest-neighbour search over embeddings. One of several backends for memory, and not required for a first version. Vector databases.

Knowledge graph

Memory stored as entities and the relationships between them, which makes multi-hop questions answerable and makes fact invalidation explicit. Knowledge graphs.

Namespace

The key that scopes a set of memories, typically a user or a thread. The framework term for what a schema calls a partition key.

Partition

A hard boundary between one scope’s memories and another’s. In a multi-tenant product it is enforced at the store rather than by a filter applied after the search. Memory security.

Memory hierarchy

Running several tiers at once: the context window as the hot tier, a fast buffer as the warm tier, and a durable store as the cold tier. Memory hierarchy.

Memory as a tool

Exposing memory reads and writes as tool calls the model chooses to make, rather than as steps your code runs unconditionally. Memory as a tool.

Once a system is running, the vocabulary shifts to measurement: the evaluation and benchmark terms.

Evaluation

What are the memory evaluation and benchmark terms?

Two benchmarks have become the standard references, and a handful of measures describe what they test. Vendor claims almost always cite one of them, so knowing what each measures is how those claims get read.

LOCOMO

A benchmark for recall within very long conversations, testing whether a system can answer questions about material from far earlier in the same dialogue. LOCOMO.

LongMemEval

A benchmark for recall across separate sessions, which is the harder case and the one closer to how products are used. LongMemEval.

MemoryAgentBench

A benchmark evaluating memory as an agentic capability rather than as retrieval accuracy alone, covering maintenance and multi-step evidence gathering.

Retrieval precision

The share of retrieved memories that were actually relevant. The measure that matters most in production, because the cost of a wrong memory is higher than the cost of a missing one. Memory metrics.

Ablation

Running the same workload with memory retrieval disabled for a held-out slice, to attribute a change to memory rather than to everything else that moved.

Several of the terms above were borrowed rather than invented, and knowing the source usually settles what they were meant to denote: the terms that come from the research literature.

From the literature

What memory terms come from the research literature?

Four papers supplied most of the field’s vocabulary, and each term is clearer read against the problem its authors were solving.

Four research sources and the terms they introduced: CoALA, Generative Agents, MemGPT and Graphiti.
Figure 3. Terms drift once they reach product documentation. The source is usually the fastest way to recover the original distinction.

CoALA

A framework for language agents (Sumers et al., 2024) that formalised the working, semantic, episodic and procedural split. The reason those four names recur across otherwise unrelated sources.

Virtual context

Paging memories between a bounded in-prompt tier and an unbounded external one, borrowed from operating-system virtual memory by MemGPT (Packer et al., 2023). Virtual context.

Main and external context

MemGPT’s names for what is currently in the prompt and what is held outside it. Roughly working memory and long-term memory, expressed as tiers of one address space.

Generative Agents

The research (Park et al., 2023) that introduced reflection as a scheduled write and the recency, importance and relevance scoring function most memory systems still use in some form.

Graphiti

The temporal knowledge-graph engine behind Zep (Rasmussen et al., 2025), which made bi-temporal validity a first-class property of a memory record. Zep and its alternatives.

Topic documents

Memory organised as maintainable per-topic documents rather than isolated records, proposed by Infini Memory (Ji et al., 2026) to address fragmentation and compression loss.

Terms drift most where two of them share a word, which is the last group here: the terms most often confused.

Disambiguation

Which memory terms are most often confused?

Five pairs account for nearly all of the confusion, and in each case the two terms share a word without sharing a referent. This is the section to check when a source seems to contradict another one.

Three commonly confused pairs: memory versus RAG, semantic memory versus semantic search, and agent memory versus agent types.
Figure 4. Three of the five. In each, the shared word is doing different work on either side.

Memory versus RAG

Both are retrieval. RAG searches a corpus shared by every user; memory searches what is true of one user or one agent. A support answer usually needs both. Memory versus RAG.

Semantic memory versus semantic search

Semantic memory is a category of what is stored. Semantic search is a method for finding anything stored, episodes included. The shared adjective refers to meaning in one case and to a memory type in the other.

Short-term versus working memory

Used interchangeably for agents, and both mean the context window. Cognitive psychology distinguishes them; agent engineering does not, so a source drawing the distinction is usually importing it from the human literature.

Agent memory versus types of agents

Nearly identical queries, unrelated answers. Memory taxonomies count four to six kinds of memory. Agent taxonomies count reflex, goal-based and utility-based agents, which is about how an agent decides rather than what it keeps.

Memory versus fine-tuning

Fine-tuning changes the weights and is expensive, slow and hard to reverse. Memory changes a row and is immediate and correctable. Anything a user might contradict tomorrow belongs in memory. Memory versus fine-tuning.

For the concepts behind the vocabulary rather than the definitions, the place to start is how AI memory works, and for building with them, the memory build guides.

FAQ

Frequently asked questions

The questions readers arrive at this glossary with, about which terms are standard and which are vendor coinages.

Which AI memory terms are standard and which are vendor coinages?

The four memory types, embedding, retrieval, consolidation and forgetting are standard across the literature. Memory layer, context engineering and most product-named tiers are commercial coinages, useful but not defined the same way by any two vendors.

Is short-term memory the same as the context window?

For agents, yes in practice: short-term memory is what the model can see this turn, which is the context window. Sources that separate them are importing a distinction from human cognitive psychology. See short-term versus long-term memory.

Why do sources give different numbers of memory types?

Because they count different things. Three means long-term memory only, four adds working memory, six adds sensory or buffer memory and shared memory. See the types of AI agent memory.

Is agent memory the same as an agent's knowledge base?

No. A knowledge base is shared across users and is authored deliberately. Memory is produced by interactions and scoped to one user or agent, which is why it needs partitioning and a deletion path.

What does memory drift mean?

A store gradually diverging from what is true, usually through append-only writing that leaves superseded facts retrievable. It presents as an agent that is confidently out of date. See conflicting memories.

Where should a beginner start with this vocabulary?

The six core terms above, then the four memory types. Everything else is a refinement of one of those and is easier to place once the loop is clear. See how AI memory works.