Fundamentals · Statelessness

Stateless vs Stateful LLMs: What Is the Difference?

Every large language model is stateless: each call is an isolated event, and the model keeps nothing once it has answered. Stateful behaviour is built on top, either by resending the conversation or by storing facts and retrieving them. This page explains the distinction using the analogy most engineers already have, the stateless HTTP request.

One model call

1
Prompt in
Everything sent
2
Forward pass
No stored state
3
Answer out
Returned
4
Cleared
Nothing kept

The base case

Why are LLMs stateless?

Because a model call is a pure function of its input: the prompt goes in, a forward pass runs, an answer comes out, and the working state is discarded. Nothing from the previous call is consulted, and nothing from this one is kept. Two identical prompts sent an hour apart are, to the model, the same event happening twice.

Four stages of a stateless model call: request arrives, forward pass, response returned, working state cleared.
Figure 1. The model is stateless in the same sense an HTTP request is. Anything that persists across calls is held by something else.

This is a design property rather than a limitation someone forgot to fix. Statelessness is what makes model serving scalable: any request can go to any machine, nothing has to be kept in sync, and a failed call can be retried anywhere. The same reasoning made stateless web services the default two decades ago.

The confusion arises because the experience of using an assistant is obviously stateful. It answers follow-up questions, refers back to earlier turns, and remembers preferences across weeks. All of that is produced outside the model, which is the single most useful thing to understand about building on one.

Engineers who have built web services already hold the right mental model, so it is worth making the comparison explicit: how this compares with stateless and stateful APIs.

The analogy

How does this compare with stateless and stateful APIs?

A REST API is stateless in exactly the sense a model is: each request carries everything the server needs, and the server keeps nothing about the client between calls. If you have built or consumed one, you already understand LLM statelessness, and the analogy holds further than most analogies do.

A stateless request carries its own context. That is why REST requests send an authentication token every time rather than relying on the server to remember a login, and it is why an LLM prompt carries the conversation history every time rather than relying on the model to remember the last turn. In both cases the burden of continuity moves to the caller.

A stateful protocol does the opposite. An old-style session on the server, or a long-lived database connection, keeps context between messages, so later messages can be shorter, but the server has to hold state per client and any machine that dies takes its sessions with it. The trade is the same one in both worlds: statelessness costs bandwidth and buys scalability.

The analogy has one important limit. A stateless API can be given a session store and become stateful for practical purposes. A model cannot, because its weights are fixed at inference: no amount of surrounding infrastructure makes the forward pass remember. What surrounds it can only change what goes into the prompt, which is why the entire field is about prompt assembly rather than about model state.

The general architecture distinction, covered outside the AI context. The trade it describes, bandwidth against scalability, is the same one a model call makes.

That constraint is what makes the next question the only real question: how stateful behaviour gets built on a stateless model.

The construction

How do you make a stateless model behave statefully?

Two ways, and only two: resend the conversation on every call, or store durable facts and retrieve the relevant ones before answering. Everything marketed as agent memory is a more sophisticated version of the second.

Two ways to make a stateless model appear stateful: resending the whole history, or storing facts and retrieving them.
Figure 2. Both produce a stateful experience. Only the second still works in month six.

Resending the history is what every chat interface does at first, and it is the correct starting point. It requires no infrastructure and it is exactly right until the conversation grows. Then two things break: the cost rises with every turn, since the whole transcript is re-billed each time, and eventually the conversation exceeds the context window and the earliest turns are silently dropped.

Storing and retrieving replaces the transcript with a small set of extracted facts and a search. The cost per turn stops growing with the conversation, because what is sent is a handful of relevant memories rather than everything that was ever said. The published figures are stark: Mem0 reports roughly 1,800 tokens per query against 26,000 for a full-context baseline, with p95 latency of 1.44 seconds against 17.1 seconds (Chhikara et al., 2025).

The four operations of the AI memory loop: write and extract, store, retrieve and rank, then update or evict.
Figure 3. This loop is what statefulness actually is in practice. The model stays stateless throughout.

The mechanism is described in how AI memory works, and the build in how to add memory to an AI agent.

The obvious objection is that context windows keep growing, so perhaps the first option simply becomes viable: whether a bigger window makes a model stateful.

The objection

Does a longer context window make a model stateful?

No. A larger window raises the ceiling on how much can be resent, and changes nothing about the model keeping state. Every call is still isolated, and the caller still supplies everything.

There is also a quality argument against relying on size alone. Liu et al. (arXiv:2307.03172) showed that models degrade on information buried in the middle of a long context even when it fits comfortably inside the window. Fitting is not the same as attending, so a longer prompt is not reliably a better-informed one.

And the economics run the wrong way. A window ten times larger makes the expensive option ten times more expensive to fill, while a retrieval system sends roughly the same amount regardless of how long the relationship has been running. That is why memory systems keep being adopted as windows grow rather than being made obsolete by them.

The full argument is on long context versus memory and the context window problem.

Statelessness is not a defect to be engineered away, and it is worth being clear about when it is exactly what you want: when a stateless agent is the right choice.

The trade

When is a stateless agent the right choice?

Whenever the task is complete inside one request and nothing needs to survive it. Classification, extraction, translation, summarisation of a supplied document, and most single-shot tool calls are all better stateless, and adding memory to them adds cost and a new class of bug for nothing.

Statelessness also buys real properties that are easy to undervalue. A stateless agent is trivially horizontally scalable, is reproducible because the same input gives the same behaviour, is easy to test since there is no accumulated history to set up, and carries no data-retention obligation because it stores nothing about anyone.

Stateful behaviour is worth its cost when the same person returns and expects to be known, when a task runs longer than one window, or when the agent should improve at something rather than merely recall facts about it. Those cases are set out on why AI agents need memory.

Short-term context window memory compared with the long-term external store, by speed, size and whether it survives the session.
Figure 4. A stateless agent runs on the left-hand store alone. Adding the right-hand one is the decision this page is about.

The honest framing is that statefulness is a feature you add deliberately, with a store, a policy and a maintenance cost, rather than a property models are missing. What that feature involves is covered across the types of AI agent memory and the memory architecture cluster.

FAQ

Frequently asked questions

The questions that follow: whether any model is truly stateful, and what the API providers actually store.

Are all LLMs stateless?

At the model/API layer, yes — weights don't update per conversation and APIs don't retain state between calls unless you resend context or add an external memory layer. Product features (ChatGPT memory, Claude memory) are application-level exceptions.

Is ChatGPT stateful?

ChatGPT the product can be stateful via built-in memory features. The underlying OpenAI API you call from your app is stateless by default — you build state with external memory or by resending history. See why agents need memory.

How do I make GPT stateful?

Add an external memory layer: write facts after each turn, retrieve before the next, inject into the prompt. Use Engram, Mem0, Zep or a DIY vector store. See add memory to an AI agent.

Stateful vs memory — same thing?

Stateful is the outcome; memory is the mechanism. An agent is stateful when it persists context across sessions — usually via a memory store, checkpointer or session DB. Memory is the most scalable path for personalization.

Which LLM APIs are stateless?

OpenAI Chat Completions, Anthropic Messages and Google Gemini APIs are stateless at the protocol level — each request is independent. State is your application's responsibility.

How does Letta make agents stateful?

Letta (MemGPT) pages memories between context (RAM) and a deep store — effectively unbounded history without stuffing every token into the window. See Letta alternatives and virtual context and MemGPT.

How do I fix stateless n8n AI agents?

Wire HTTP nodes to Engram, Mem0 or Zep APIs — retrieve before the LLM node, write after the response. Same five-step pattern as any framework. See add memory guide.

Can a stateless LLM power a CRM agent?

Yes with external memory. The LLM stays stateless; Engram or Zep stores account history, preferences and changing facts. Zep fits when relationships and validity windows change over time. See customer support use case.

Stateless vs stateful — which for a POC?

Stateless (prompt-only) ships fastest for demos. Stateful (memory layer) is required for multi-session personalization — Engram or Mem0 Cloud are common POC paths. See best AI memory tools.

Does memory replace the context window?

No — memory complements the context window. Working memory stays in-context; long-term facts live in external storage and are retrieved selectively. See memory vs context window.