Developers · Architecture
Agentic Architecture Patterns, Read for Memory
ReAct, Plan-and-Execute, ReWOO and Reflexion aren’t just four different control-flow structures, they imply four genuinely different memory shapes, from ReAct’s implicit, ever-growing context history to Reflexion’s explicit episodic critique buffer, which is the clearest instance of “agent memory” among all four. Multi-agent topologies make the same choice at a different scale: scoped per worker, or shared across a whole peer group.
Four patterns
The default
What does ReAct actually demand of memory?
Everything, by default: the full conversation history, every thought, action and observation, accumulates in context on every step, with no structure and no forgetting built in. ReAct, Reason, Act, Observe, repeat, is the pattern most people mean when they say “agent” without qualification, and its memory footprint is the direct cause of its best-known limitation.
Because a 10-step task means 10 full LLM calls, each one carrying a longer history than the last, ReAct’s cost scales with exactly the same accumulation this site covers generally as the context-window problem: nothing here is forgotten deliberately, everything just piles up until the context itself becomes the bottleneck. This is also what makes ReAct naturally adaptive: since the full history is present at every step, the agent re-reasons on each turn and can course-correct when a tool fails or returns something unexpected. The tradeoff is structural, not incidental: adaptability and unbounded accumulation are the same property, viewed from two angles.
ReAct also has a well-documented practical advantage that follows directly from this same accumulation: every thought is logged in order, so when something goes wrong, the full reasoning trail is available to trace exactly where it happened. Teams operating in regulated settings, finance, healthcare, tend to specifically favor ReAct over less transparent architectures for this reason, since an unstructured loop with no step-by-step record leaves nothing to audit when a decision needs to be explained after the fact. The same accumulated history that makes ReAct expensive at scale is what makes it debuggable at any scale.
Plan-and-Execute exists specifically to answer ReAct’s memory problem differently, by replacing implicit accumulation with an explicit, inspectable object. How is Plan-and-Execute’s memory different from ReAct’s?
An explicit object
How is Plan-and-Execute’s memory different from ReAct’s?
Plan-and-Execute holds one explicit, structured plan rather than an ever-growing history: a stronger model generates the full step sequence upfront, a cheaper model executes each step against it, and the plan itself, not the accumulated conversation, is what gets logged, validated and referenced.
This single design choice is what cuts the LLM call count so sharply: the planner makes one call, the executor runs tools directly without re-reasoning at every step, and a replanner call only fires on failure, which is a fundamentally different memory shape than ReAct’s turn-by-turn accumulation. A 10-step task needing 10 calls under ReAct needs one or two under Plan-and-Execute, precisely because the plan itself is doing the job that ReAct’s growing history did implicitly. The cost of that structure shows up when reality doesn’t match the plan: if a step returns something the planner didn’t anticipate, the plan can derail, and replanning adds back some of the adaptive overhead Plan-and-Execute was built to avoid in the first place.
ReWOO pushes this same idea, separating planning from memory-heavy execution, even further than Plan-and-Execute does. Why does ReWOO need almost no working memory at all?
Minimal footprint
Why does ReWOO need almost no working memory at all?
Because the entire plan, including placeholders for results the planner hasn’t seen yet, gets written out in a single pass, tools then run in parallel against those placeholders, and one final call synthesizes everything, which means the pattern never has to hold a growing, evolving working state the way ReAct or even Plan-and-Execute does.
The original ReWOO paper reports 5x token efficiency and a 4% accuracy improvement over ReAct on HotpotQA, gains that follow directly from needing only two LLM calls total rather than one per step. The tradeoff is exactly what you’d expect from a pattern with almost no working memory: zero mid-execution adaptation. If an early result means a later step should change, ReWOO has no mechanism to notice, since the plan was locked in during the very first call, and a placeholder chain built on an unexpected early result can pass garbage downstream to everything that depends on it.
None of the three patterns so far actually remembers anything between separate attempts at a task. The fourth pattern is built specifically around that gap. Is Reflexion’s self-critique actually a form of agent memory?
The clearest case
Is Reflexion’s self-critique actually a form of agent memory?
Yes, directly: after a failed attempt, an evaluator scores the result, a self-reflection step writes a verbal critique of what went wrong, and that critique goes into an episodic memory buffer the agent reads on its next retry. Of the four patterns covered here, this is the one where “memory” isn’t a useful lens to apply, it’s the literal mechanism.
The original paper reports real gains from this mechanism: GPT-4’s HumanEval coding pass rate improved from 80% to 91% with Reflexion added, and ReAct combined with Reflexion completed 130 of 134 AlfWorld decision-making tasks. What makes Reflexion distinct from the other three patterns is that it’s the only one that gets better on the same task across retries, since ReAct, Plan-and-Execute and ReWOO each run once and return whatever result they get, while Reflexion’s episodic buffer specifically steers the next attempt away from a previously failed approach.
The mechanism has a documented limit worth taking seriously before relying on it: a 2025 replication study found that single-agent Reflexion can consistently repeat its own earlier misconceptions across retries, because the same model both generates the output and writes the critique meant to correct it, which reinforces rather than corrects its own blind spots. This is a genuinely useful caution for anyone treating a self-critique buffer as a reliable substitute for actual verification: the memory mechanism only helps when the evaluator judging success is meaningfully independent of the model producing the output, covered more generally on this site’s approach to evaluating agent memory.
Every pattern covered so far is single-agent. The same memory-shape question, scoped narrowly or shared broadly, shows up again, at a different scale, the moment more than one agent is involved. How does memory get scoped across a multi-agent system?
A different scale
How does memory get scoped across a multi-agent system?
Two dominant topologies answer this differently: supervisor-worker scopes each worker’s memory to its own sub-task and aggregates outcomes at the top, while swarm or blackboard topologies share one memory surface across every peer agent with no supervisor-level scoping at all. This is the multi-agent version of the same tradeoff every single-agent pattern above makes at a smaller scale.
In a supervisor-worker topology, a supervisor agent decomposes a task and routes sub-tasks to specialized workers, each running its own loop with memory scoped to what it needs for its own piece, and the supervisor aggregates results and decides whether the overall task is complete. This trades some parallelism efficiency for isolation: a worker’s memory can’t leak into another worker’s task by construction. It also has a well-known failure mode worth naming plainly: coordination overhead between the supervisor and its workers can dominate the total cost of a task that was simple enough not to need decomposition in the first place, which is why this topology earns its complexity specifically on tasks that genuinely split into independent sub-problems, research with topic-specific subagents, multi-domain reasoning, rather than being reached for by default.
Setting explicit interface contracts between the supervisor and its workers, and requiring structured rather than free-text output from each worker, is what keeps this scoping from breaking down in practice: a supervisor that has to parse an unstructured summary from each worker to decide what happened is reintroducing exactly the kind of implicit, unbounded context ReAct relies on, at the coordination layer instead of the single-agent layer.
A blackboard or swarm topology inverts that choice entirely: peer agents post to and read from one shared surface with no coordinator scoping who sees what, which is the same shared-memory tradeoff this site covers generally on shared memory in AI agents, applied here specifically to how agents coordinate rather than to how a single agent stores facts. The general design questions around which of these scoping choices fits a given system, and what actually goes wrong when isolation is skipped, are covered in full on multi-agent memory.Four single-agent memory shapes and two multi-agent scoping answers are a lot of options. What actually decides which one fits a specific task. Which pattern, and which memory shape, actually fits your task?
The decision
Which pattern, and which memory shape, actually fits your task?
Start with ReAct’s implicit history unless a specific failure mode you’re already hitting points elsewhere: token cost at scale points toward ReWOO, unpredictable multi-step workflows point toward Plan-and-Execute, and a task with clear pass or fail criteria where getting it right matters more than speed points toward Reflexion’s episodic buffer. Most production systems don’t run one pure pattern; they combine them.
The most common hybrid pairs ReAct’s per-step adaptation with a Reflexion retry cycle triggered specifically when a result fails validation, giving step-by-step flexibility for the common case and self-correction for the harder one. A second common combination runs ReWOO for the fast path and falls back to Plan-and-Execute’s replanning when a tool returns something unexpected. Neither hybrid is worth building on a task simple enough for a plain ReAct loop to finish in three or four steps: adding planning or reflection overhead to a task that doesn’t need it makes the system slower and harder to debug for no corresponding benefit, which is the same over-engineering trap the supervisor-worker topology above falls into when applied to a task that didn’t need decomposing.
The pattern chosen matters less, in practice, than two things most teams get wrong before they even pick one: whether the tools an agent calls are actually reliable, since a well-architected loop calling a flaky API fails regardless of which of these four patterns wraps it, and whether success can actually be measured, since Reflexion’s self-critique has nothing to correct against if pass or fail was never defined precisely in the first place.
In both cases, the memory question and the architecture question are really the same question asked twice: what does this task actually need remembered, in what shape, and for how long. Teams building any of these patterns and wanting the underlying episodic-memory or context-management pipeline handled as a managed service rather than assembled by hand can look at Engram directly; the fuller path from pattern choice to a running agent is covered on building an AI agent.Video
Five agentic design patterns, explained directly
A walkthrough covering ReAct and multi-agent patterns with a Microsoft AutoGen implementation.
FAQ
Frequently asked questions
The practical decisions that follow once the memory shape of each pattern is understood.
Which agentic pattern needs the least memory infrastructure?
ReWOO, by design. Its two-call structure with placeholder results means it never has to hold a growing working state, unlike ReAct's accumulating history or Reflexion's episodic buffer. The tradeoff is zero mid-execution adaptation if a result surprises the plan.
Is Reflexion's episodic memory the same as long-term agent memory?
No. Reflexion's buffer is short-term and scoped to a single task's retries, discarded once the task ends. Long-term memory persists across separate sessions and tasks entirely, which is a different mechanism covered on this site's memory-types pages.
Does ReAct need an external memory store to work?
Not to function, but its accumulating in-context history is exactly the scaling problem external memory is built to solve. Without it, a long task simply grows its context until cost and context rot become the limiting factor.
Should a supervisor-worker system share one memory store across all workers?
Generally no. Scoping each worker's memory to its own sub-task, with the supervisor aggregating outcomes, is what keeps one worker's context from leaking into another's task. A shared store without that scoping reintroduces the isolation problems supervisor-worker topologies exist to avoid.
Can Reflexion be combined with a multi-agent system?
Yes. Reflexion's self-critique loop works on top of any underlying pattern, including a single worker inside a supervisor-worker topology, and doesn't require a single-agent architecture to function.
Why did a Reflexion-based agent stop improving after a few retries?
This is a documented limitation: since the same model generates both the output and the self-critique, it can repeat its own earlier misconceptions instead of correcting them, with most genuine improvement landing in the first one or two retries.