Developers · Definition

What Makes an App Agentic, and What Changes About Memory?

An agentic app is not a chatbot with a bigger buffer. The dividing line is who decides what gets remembered: a chatbot’s developer configures a fixed window, while an agentic app’s own model calls memory tools mid-task and decides what to keep, update, or discard. That single shift, from a passive buffer to an agent-controlled store, is what the term actually adds on top of “chatbot.”

Three context types

1
Session
2
User
3
Organisation

Definition

What makes an app agentic rather than a single-turn chatbot?

The ability to act on a goal across multiple steps, using tools, without a person directing each step, and to carry state forward while doing it. That definition is common ground across the literature: perceive the situation, reason about what to do, take an action, observe the result, and repeat until the goal is met, rather than answering one prompt and stopping.

Where accounts diverge is on what has to be true architecturally for that loop to count as more than a single isolated agent. A recent taxonomy of the field draws the line at persistence and coordination: a basic AI agent runs the perceive-reason-act loop with a simple memory buffer for recalling recent turns, while what the field is now calling agentic AI adds persistent memory, orchestration across multiple agents, and shared context as architectural components in their own right, not just a bigger version of the same buffer.

That is a more useful test than autonomy alone, because autonomy is a spectrum every vendor claims to be further along than the last, whereas persistent, shared, agent-controlled memory is a specific thing you can check for in an architecture diagram. It is also the reason this page exists on a site about agent memory rather than being folded into a general definition of agentic AI: the interesting distinction, the one that actually determines what you have to build, runs straight through the memory layer.

Underneath the loop, the same taxonomy breaks a basic agent into four subsystems worth naming, because each one is where a design choice actually gets made: a perception module that takes in a prompt or an API response and puts it into a shape the rest of the system can use; a reasoning module that applies rules, a plan, or a chain of intermediate steps to decide what to do next; an action module that turns that decision into a real effect, a message sent, a database updated, a tool invoked; and a learning module, which in the simplest agents is nothing more than a memory buffer recalling recent turns. The fourth subsystem is the one this page is about, and it is also the one that grows the most when a single agent becomes an agentic app: a buffer that only remembers the current conversation becomes a store that persists, that other parts of the system can read, and that the agent itself decides how to use.

The sharpest version of that distinction is not about whether memory exists but about what kind: how is agentic memory different from a chatbot’s context window?

The core distinction

How is agentic memory different from a chatbot’s context window?

A chatbot’s memory is a fixed buffer injected whole into every prompt. Agentic memory is a store the agent itself queries selectively and writes to deliberately. That reframing, more than any list of capabilities, is what separates the two in practice.

Chatbot memory as a fixed injected buffer compared with agentic memory as a store the agent queries and edits itself
Figure 1. The difference is agency over what is kept, not just whether memory persists.

A useful direct comparison, drawn from engineering practice rather than theory, breaks it into four questions. What gets stored: raw conversation history in a chatbot, versus structured facts, preferences and events extracted from it in an agentic app. How long it lasts: one session in a chatbot, indefinitely in an agentic app, so the agent recalls something from weeks ago without the user repeating it. How it is recalled: the entire recent history injected regardless of relevance in a chatbot, versus selective retrieval based on the current task in an agentic app. And the question underneath all three: who decides what an agentic app remembers?

Agency

Who decides what an agentic app remembers?

The agent itself, through explicit tool calls, rather than a developer’s fixed configuration set once at build time. This is the detail that actually separates the two architectures, and it is easy to miss because it lives in a code path rather than in a diagram.

In a chatbot the developer sets a fixed memory buffer; in an agentic app the agent itself calls memory tools to decide what to keep
Figure 2. A tool call the agent chooses to make is what separates the two.

In a chatbot, a developer sets a buffer size or a summarisation trigger, and the model has no say in what survives; the mechanism is entirely outside the model’s control. In an agentic app, the model is given tools such as add memory, update memory, or delete memory, and it calls them, or does not, based on its own judgement about what is worth keeping during that turn. The write path becomes a decision the agent makes rather than a policy applied to it.

That shift has real consequences worth building for deliberately rather than discovering by accident. It means the extraction step, deciding whether a given statement is worth a memory call at all, is doing real work and deserves real design attention, which is covered in full on writing memories. It also means the agent can get this wrong in both directions: forgetting something it should have kept, or writing something trivial that crowds out better memories later, which is the same failure covered on forgetting and eviction.

The practical upshot for anyone building one of these apps is that the memory tools themselves become part of the surface you design and test, not an implementation detail hidden inside a framework default. A poorly specified add_memory tool, one with no guidance on what counts as worth storing, tends to over-write: the agent calls it defensively on nearly every turn, producing a store full of trivial, half-relevant entries that then have to be filtered back out at read time. A well-specified one, with clear criteria for what is durable versus what is small talk, produces a store worth querying later. That specification work is invisible in a demo and decisive in production.

Deciding what to remember assumes you already know what an agentic app needs to be tracking in the first place, and that turns out to be more specific than “everything the user has ever said”: what does an agentic app actually need to keep track of?

What to track

What does an agentic app actually need to keep track of?

Three distinct kinds of context, each answering a different question about the current moment. This is a more useful breakdown for deciding what to build than the standard short-term-versus-long-term split, because it is organised by where the information comes from rather than by how long it survives.

Three kinds of context an agentic app tracks: session context, user context, and organisational context
Figure 3. Organised by where the information comes from, not by how long it persists.

Session context is what is happening right now: the task in progress, which thread this is, what tools have already been called. It is what lets a vague instruction like “create a follow-up” resolve correctly, producing a ticket in the tracker rather than a calendar invite, because the app knows a bug triage meeting is the frame the instruction was made in.

User context is who this specific person is: their role, their stated preferences, facts from prior interactions. It is what lets an app explain something in plain language to a product manager and in technical detail to an engineer, without being told which each time. This is the tier most agent-memory writing focuses on, and it maps closely to what this site covers on types of AI agent memory.

Organisational context is the constraint layer: policy, escalation rules, access permissions specific to the team or company the app operates inside. It is the least discussed of the three and often the one that determines whether an agentic app is safe to give real authority to, since it is what stops the app from taking an action a person in that role should not be able to take alone.

Tracking all three is a design decision, not something a framework provides by default. Where these different kinds of context should physically live, and how they interact with orchestration across more than one agent, is the architectural question: where does memory sit in an agentic app’s architecture?

Placement

Where does memory sit in an agentic app’s architecture?

As a component alongside orchestration, not bolted onto one agent’s loop. A single AI agent’s memory buffer is private to that loop: it recalls what happened in this conversation and nothing a different part of the system produced. An agentic app’s persistent memory is written to and read from by more than one part of the system, which is what makes it a shared resource rather than a personal notebook.

A basic AI agent with a simple memory buffer compared with an agentic app where persistent memory sits alongside orchestration
Figure 4. Adding memory alone does not make an app agentic; sharing it across a system does.

Concretely, when a system decomposes a goal across several agents, one planning, one retrieving information, one synthesising a result, those agents typically communicate through a shared memory or context that an orchestrator reads and writes as it monitors progress. A frequently cited multi-agent pattern makes the roles explicit: a planner agent decides what needs to happen next, a retrieval agent goes and finds the information, and a synthesis agent turns what came back into a final answer, with an orchestrator agent watching dependencies between the three and deciding when the task is actually done. None of that coordination is possible if each of the three agents holds its own private, disconnected memory buffer with no way to see what the others have already found or decided.

That is a genuinely different problem from single-agent memory: it introduces questions about who is allowed to write what, and what happens when two agents’ writes disagree, which is covered in full on multi-agent memory rather than here.

The reasoning loop itself, ReAct, planner-executor, reflection, and how each pattern uses memory differently, is its own deep subject and is covered on agentic architecture. What matters for this page is narrower: memory has to be treated as an architectural component with its own interface, not an implementation detail hidden inside whichever agent happens to need it first, because the moment a second agent or a second session needs the same fact, a private buffer cannot serve it.

With the distinction, the mechanism, and the architecture in view, the practical question is how to tell whether an app you are building, or evaluating, actually clears the bar: what should you check before calling your app agentic?

The checklist

What should you check before calling your app agentic?

Whether the model itself decides what to remember, whether that memory survives past one session, and whether more than one part of the system can read and write it. An app that fails all three is a chatbot with a longer buffer, whatever the marketing calls it.

The first check is agency over the write: does the model call a memory tool and decide what is worth keeping, or does a fixed buffer size decide for it. The second is persistence: does the app recall something from a previous session without the user repeating it, or does closing the window discard everything. The third is sharing: if the system has more than one agent or more than one entry point, can they see the same memory, or does each hold its own private, disconnected copy.

Each check has a concrete test rather than a judgement call. For agency, look at the actual tool schema exposed to the model: if there is no add_memory, update_memory, or delete_memory equivalent, the model has no lever to pull regardless of how the marketing describes the app. For persistence, end a session, start a fresh one, and ask the app something that depends on the first session; if it has no answer, the memory was scoped to the conversation rather than to the user. For sharing, trace what happens when two entry points, two agents, or two sessions from the same user touch the same fact; if one cannot see what the other wrote, the memory is private rather than architectural.

None of the three requires exotic infrastructure. They require treating memory as a first-class component with its own store, its own write policy, and its own retrieval logic, which is the whole subject of this site. The build sequence for actually wiring this into an agent is on adding memory to an agent and building an AI agent; the broader distinction between an AI-native application and one that merely calls an LLM API is on AI-enabled versus AI-native apps.

FAQ

Frequently asked questions

The practical distinctions builders ask about once the definition is settled.

Is retrieval-augmented generation enough to make an app agentic?

No. RAG retrieves documents to answer a question; it does not give the model a tool to decide what to remember about the user or the task going forward. An app can use RAG and still have no persistent, agent-controlled memory, which is the property that actually defines agentic memory.

Does an agentic app need more than one agent to count as agentic?

No. A single agent with persistent, self-directed memory and tool use over multiple steps already qualifies. Multiple agents sharing memory is a further step, common in production systems tackling complex goals, but it is not the defining property.

What is the simplest way to add agentic memory to an existing chatbot?

Give the model a small set of memory tools, such as add_memory and get_memory, backed by a store scoped to the user, and let the model decide when to call them rather than injecting a fixed history buffer on every turn. The build sequence is on adding memory to an agent.

How is organisational context different from user context?

User context describes the person the agent is talking to: their role, preferences, and history. Organisational context describes constraints that apply regardless of who is asking, such as escalation rules or access permissions. An agent can know a user well and still need organisational context to know what it is allowed to do for them.

Can an agentic app have memory without being multi-step or tool-using?

In principle yes, but it would not usually be called agentic. The term describes systems that combine persistent, agent-controlled memory with autonomous multi-step action; memory alone, without the action loop, is closer to a personalization feature on a conventional application.

What is the most common mistake teams make when adding memory to make an app agentic?

Giving the model a memory tool with no clear criteria for when to use it, which leads to over-writing: the agent calls add_memory defensively on nearly every turn, filling the store with trivial entries that then have to be filtered back out at read time. Specifying what counts as durable is as important as building the store itself.