Memory types · Working memory
Short-Term Memory in AI Agents
Short-term memory in AI agents is the information held in the current task and recent turns: primarily inside the context window, before it is consolidated into long-term storage or discarded.
Short-term lifecycle
Definition
What is short-term (working) memory in agents?
Working memory is the active scratchpad for the current goal, recent messages and tool outputs, living in the context window plus small buffers.
It is volatile: cleared when the session ends unless promoted to long-term storage. Also called “context memory” in some frameworks, but that term often confuses working memory with persistent AI memory; see memory vs context window. Working memory isn’t one undifferentiated blob either. It splits into 3 practical sub-layers: the working memory proper (the exact tokens sent into the next LLM call), a session buffer (recent turns and tool calls keyed by thread or session ID, so a conversation can resume mid-task), and an ephemeral cache (fast, short-lived entries, often in Redis, that expire quickly and were never meant to survive the session). A 2025 research paper proposing a taxonomy of agent memory (arXiv:2512.13564) breaks working memory down further by processing stage rather than storage location: single-turn processing (condensing and abstracting the immediate input and observation) versus multi-turn processing (consolidating session state, folding a growing history into a smaller representation, and using that folded state for step-by-step planning).
Mechanism
How does short-term memory work in practice?
Messages accumulate in the prompt; when the window fills, agents summarize, truncate or page out oldest turns.
Framework buffers include LangChain ConversationBuffer and LangGraph checkpoint state for recent turns. Tool results inject directly into working memory for the current reasoning step. A concrete, minimal pattern used in production: a session-keyed buffer (a Redis hash or in-memory store) holds messages and tool results; at inference time, the context sent to the model is built as a rolling summary plus the last N turns plus the key tool outputs still relevant; a hard token budget caps the total, with low-value content, chit-chat, verbose logs, trimmed first when the budget is tight.
Failure modes
What goes wrong when short-term memory design is sloppy?
Three specific failures, not a vague warning about running out of tokens.
Context pollution: keeping irrelevant turns around drowns out the details that actually matter, degrading reasoning quality even while token count still fits the budget. Token and latency blow-ups: naively dumping entire logs or full histories into the prompt means paying for, and waiting on, the full record every single call, not just the parts the current turn needs. Hidden coupling: when every agent in a codebase implements its own ad-hoc short-term memory instead of sharing one component, debugging across a fleet of agents becomes painful, since there’s no single place to inspect what any given agent actually saw. A related pattern shows up even in simpler chat systems: state stays in the prompt until it silently falls off the edge of the window, nothing gets promoted to durable memory at the right moment, and the system ends up recalling too much because it never learned to distinguish what’s still active from what was only ever meant to be temporary. A short checklist catches most of this before it ships: is there a hard cap on turns or tokens passed to the model, is there an explicit priority for what survives when trimming, and is short-term memory handled by one shared component rather than reinvented per agent?
Constraints
What are the limits of short-term memory?
Token caps, truncation of oldest turns, per-token cost and no cross-session persistence by default.
GPT-4: 128K tokens; Claude 3.7 Sonnet: 200K (as cited in Chhikara et al., 2025). When the limit hits: summarize, drop or promote to long-term memory.
Transition
How does short-term memory get promoted to long-term memory?
Extraction, then consolidation, then external store.
Promote when a fact is durable (preferences, identity, recurring context); discard when ephemeral (a greeting, a one-off clarification).
Comparison
How does short-term memory compare to long-term memory?
Same underlying question as the definition above, laid out side by side.
| Dimension | Short-term | Long-term |
|---|---|---|
| Location | Context window | External store |
| Persistence | Session | Cross-session |
| Capacity | Model token limit | Unbounded (store) |
| Cost | Per token in window | Per embed + retrieve |
Patterns
What frameworks and patterns exist for working memory?
Context window only; buffer plus summarization; Letta paging tiers.
- Context window only: simplest; no persistence across sessions
- Buffer plus summarization: LangChain buffers compress old turns
- Letta paging: core memory in the window, the rest paged to a deep store (the MemGPT pattern)
- External memory plus window: Engram, Mem0 and Zep retrieve into the window each turn; working memory stays in-context
FAQ
Frequently asked questions
The sub-layers, the failure modes, and how promotion to long-term memory actually works.
What is working memory in AI agents?
Working memory is the active information in the context window for the current task: recent messages, tool outputs and retrieved memories. It is volatile and session-bound. See definition above.
Is short-term memory the same as the context window?
Mostly yes for agents. The context window is the primary short-term/working memory tier, often plus small buffers. Persistent memory lives outside the window. See memory vs context window.
Should I use LSTM for agent memory?
No. LSTM is a neural network architecture for sequence modeling (Hochreiter & Schmidhuber, 1997), not agent working memory. Agents use context windows and external memory stores, not LSTM layers for conversation memory.
What does context memory mean?
Usually information in the context window (working memory), not persistent cross-session storage. For durable memory, use an external store. See memory vs context window.
What is role context memory?
Role context is the system prompt and persona in the window: instructions for how the agent behaves, not facts about users. User facts belong in external long-term memory.
How long does short-term memory last?
For the current session only, until the context window is cleared, the session ends, or facts are promoted to long-term storage. No cross-session persistence by default.
When should I promote short-term to long-term memory?
When a fact is durable: preferences, identity, recurring context. Discard ephemeral turns (greetings, clarifications). See memory consolidation.
Does Claude have short-term memory?
Claude's API context window holds short-term/working memory per call. Cross-session persistence requires external memory (Engram, Mem0, Zep, etc.). See persist conversation memory.
How does n8n handle agent short-term memory?
n8n passes conversation history in the prompt (working memory). For persistence across sessions, wire Engram, Mem0 or Zep APIs. See add memory to an agent.
Buffer memory vs vector memory?
Buffer memory keeps recent turns in the prompt (short-term). Vector memory stores embeddings in an external index (long-term). Production agents use both. See long-term memory.
What are the sub-layers inside short-term memory?
Working memory proper (the exact tokens sent to the model), a session buffer (recent turns keyed by thread ID), and an ephemeral cache (fast-expiring entries, often Redis, never meant to persist).
What's the most common short-term memory design mistake?
Context pollution: keeping irrelevant turns around until they drown out the details that actually matter. A close second is dumping entire logs into every prompt instead of trimming low-value content first.