Memory types · Working memory

Short-Term Memory in AI Agents

Short-term memory in AI agents is the information held in the current task and recent turns: primarily inside the context window, before it is consolidated into long-term storage or discarded.

Short-term lifecycle

1
Accumulate
In context window
2
Use
Current reasoning
3
Promote
Or discard
4
Clear
Session ends

Definition

What is short-term (working) memory in agents?

Working memory is the active scratchpad for the current goal, recent messages and tool outputs, living in the context window plus small buffers.

It is volatile: cleared when the session ends unless promoted to long-term storage. Also called “context memory” in some frameworks, but that term often confuses working memory with persistent AI memory; see memory vs context window. Working memory isn’t one undifferentiated blob either. It splits into 3 practical sub-layers: the working memory proper (the exact tokens sent into the next LLM call), a session buffer (recent turns and tool calls keyed by thread or session ID, so a conversation can resume mid-task), and an ephemeral cache (fast, short-lived entries, often in Redis, that expire quickly and were never meant to survive the session). A 2025 research paper proposing a taxonomy of agent memory (arXiv:2512.13564) breaks working memory down further by processing stage rather than storage location: single-turn processing (condensing and abstracting the immediate input and observation) versus multi-turn processing (consolidating session state, folding a growing history into a smaller representation, and using that folded state for step-by-step planning).

→ Types of AI agent memory

Working memory's own processing stages: single-turn input condensation and observation abstraction, versus multi-turn state consolidation and context folding
Working memory isn’t just a location; it has its own processing stages, single-turn and multi-turn.
Not LSTM: “Long short-term memory” (LSTM) is a recurrent neural network architecture (Hochreiter & Schmidhuber, 1997), not agent working memory. Agent short-term memory means the context window and buffers, not a neural net layer type. For the cognitive analogy, see AI memory vs human memory.

Mechanism

How does short-term memory work in practice?

Messages accumulate in the prompt; when the window fills, agents summarize, truncate or page out oldest turns.

Framework buffers include LangChain ConversationBuffer and LangGraph checkpoint state for recent turns. Tool results inject directly into working memory for the current reasoning step. A concrete, minimal pattern used in production: a session-keyed buffer (a Redis hash or in-memory store) holds messages and tool results; at inference time, the context sent to the model is built as a rolling summary plus the last N turns plus the key tool outputs still relevant; a hard token budget caps the total, with low-value content, chit-chat, verbose logs, trimmed first when the budget is tight.

→ Memory vs context window

Failure modes

What goes wrong when short-term memory design is sloppy?

Three specific failures, not a vague warning about running out of tokens.

Context pollution: keeping irrelevant turns around drowns out the details that actually matter, degrading reasoning quality even while token count still fits the budget. Token and latency blow-ups: naively dumping entire logs or full histories into the prompt means paying for, and waiting on, the full record every single call, not just the parts the current turn needs. Hidden coupling: when every agent in a codebase implements its own ad-hoc short-term memory instead of sharing one component, debugging across a fleet of agents becomes painful, since there’s no single place to inspect what any given agent actually saw. A related pattern shows up even in simpler chat systems: state stays in the prompt until it silently falls off the edge of the window, nothing gets promoted to durable memory at the right moment, and the system ends up recalling too much because it never learned to distinguish what’s still active from what was only ever meant to be temporary. A short checklist catches most of this before it ships: is there a hard cap on turns or tokens passed to the model, is there an explicit priority for what survives when trimming, and is short-term memory handled by one shared component rather than reinvented per agent?

→ The context window problem

Three sub-layers inside short-term memory: working memory, session buffer, and ephemeral cache
Short-term memory isn’t one undifferentiated blob; each sub-layer serves a different purpose.
Three failure modes in short-term memory design: context pollution, token and latency blow-ups, and hidden coupling
Each failure mode compounds quietly until debugging across a fleet of agents becomes painful.

Constraints

What are the limits of short-term memory?

Token caps, truncation of oldest turns, per-token cost and no cross-session persistence by default.

GPT-4: 128K tokens; Claude 3.7 Sonnet: 200K (as cited in Chhikara et al., 2025). When the limit hits: summarize, drop or promote to long-term memory.

→ The context window problem

Transition

How does short-term memory get promoted to long-term memory?

Extraction, then consolidation, then external store.

Promote when a fact is durable (preferences, identity, recurring context); discard when ephemeral (a greeting, a one-off clarification).

→ Memory consolidation · Long-term memory

Promotion pipeline from short-term to long-term memory: extraction, consolidation, external store
A durable fact gets promoted; an ephemeral turn gets discarded when the session ends.

Comparison

How does short-term memory compare to long-term memory?

Same underlying question as the definition above, laid out side by side.

DimensionShort-termLong-term
LocationContext windowExternal store
PersistenceSessionCross-session
CapacityModel token limitUnbounded (store)
CostPer token in windowPer embed + retrieve

→ Short-term vs long-term (full guide)

Patterns

What frameworks and patterns exist for working memory?

Context window only; buffer plus summarization; Letta paging tiers.

  • Context window only: simplest; no persistence across sessions
  • Buffer plus summarization: LangChain buffers compress old turns
  • Letta paging: core memory in the window, the rest paged to a deep store (the MemGPT pattern)
  • External memory plus window: Engram, Mem0 and Zep retrieve into the window each turn; working memory stays in-context

→ Virtual context & MemGPT · Letta alternatives

FAQ

Frequently asked questions

The sub-layers, the failure modes, and how promotion to long-term memory actually works.

What is working memory in AI agents?

Working memory is the active information in the context window for the current task: recent messages, tool outputs and retrieved memories. It is volatile and session-bound. See definition above.

Is short-term memory the same as the context window?

Mostly yes for agents. The context window is the primary short-term/working memory tier, often plus small buffers. Persistent memory lives outside the window. See memory vs context window.

Should I use LSTM for agent memory?

No. LSTM is a neural network architecture for sequence modeling (Hochreiter & Schmidhuber, 1997), not agent working memory. Agents use context windows and external memory stores, not LSTM layers for conversation memory.

What does context memory mean?

Usually information in the context window (working memory), not persistent cross-session storage. For durable memory, use an external store. See memory vs context window.

What is role context memory?

Role context is the system prompt and persona in the window: instructions for how the agent behaves, not facts about users. User facts belong in external long-term memory.

How long does short-term memory last?

For the current session only, until the context window is cleared, the session ends, or facts are promoted to long-term storage. No cross-session persistence by default.

When should I promote short-term to long-term memory?

When a fact is durable: preferences, identity, recurring context. Discard ephemeral turns (greetings, clarifications). See memory consolidation.

Does Claude have short-term memory?

Claude's API context window holds short-term/working memory per call. Cross-session persistence requires external memory (Engram, Mem0, Zep, etc.). See persist conversation memory.

How does n8n handle agent short-term memory?

n8n passes conversation history in the prompt (working memory). For persistence across sessions, wire Engram, Mem0 or Zep APIs. See add memory to an agent.

Buffer memory vs vector memory?

Buffer memory keeps recent turns in the prompt (short-term). Vector memory stores embeddings in an external index (long-term). Production agents use both. See long-term memory.

What are the sub-layers inside short-term memory?

Working memory proper (the exact tokens sent to the model), a session buffer (recent turns keyed by thread ID), and an ephemeral cache (fast-expiring entries, often Redis, never meant to persist).

What's the most common short-term memory design mistake?

Context pollution: keeping irrelevant turns around until they drown out the details that actually matter. A close second is dumping entire logs into every prompt instead of trimming low-value content first.