Memory types · Procedural
What Is Procedural Memory in AI Agents?
Procedural memory is the store of rules, workflows and tool-use patterns that tells an agent how to do something, as opposed to semantic memory, which tells it what is true. It is the memory type that makes an agent better at a task over time rather than merely better informed about it, and it is the one most systems never implement.
What it holds
Definition
What is procedural memory?
Procedural memory is the encoded set of instructions, rules and workflows that determines how an agent performs a task and uses its tools. Where semantic memory answers “what is true”, procedural memory answers “how do we do this here”, and the two are retrieved for different reasons.
The human analogy is the one everybody reaches for and it holds up well: procedural memory is closer to habit or muscle memory than to recall. A person who can ride a bicycle does not retrieve facts about balance, they execute a learned procedure. An agent that has learned to check inventory before quoting a delivery date is doing the same thing, and the useful consequence is that neither has to reason it out again each time.
In an implementation the distinction is concrete. A semantic memory is retrieved because it matches the topic of the question. A procedural memory is retrieved because it matches the situation: the agent is about to use a tool, or handle a class of request, or produce a particular kind of output. That difference in retrieval trigger is what makes procedural memory awkward to bolt onto a store designed for facts.
It is also the type least often implemented deliberately. Most systems encode procedure in the system prompt, which works until the procedure needs to differ per user, per team or per project, at which point a static prompt cannot carry it.
The category is clearest with concrete cases: examples of procedural memory.
Examples
What are examples of procedural memory in agents?
Three kinds appear repeatedly: tool-use patterns, behavioural rules, and format or style conventions. Each one is something an agent should apply without being told again.
Tool-use patterns are learned sequences. An agent that discovers it should check stock before promising a delivery date, or read the test file before editing an implementation, has learned an ordering. Storing that ordering means it stops rediscovering it, which is both faster and more consistent.
Behavioural rules are conditions paired with responses, usually accumulated from corrections. “Escalate refunds over five hundred to a human” is a rule an agent is told once and should apply forever. When a user corrects an agent’s behaviour, the correction is procedural: it changes what the agent does, not what it believes.
Format and style conventions are the third and the most commonly needed. A codebase that writes tests before implementations, a team that wants bullet summaries rather than prose, a customer who prefers short replies. None of these are facts about the world and all of them change output.
A useful test when classifying a memory: if it changes what the agent says is true, it is semantic. If it changes how the agent acts, it is procedural. The distinction against the other types is on the types of AI agent memory.
Classification matters because the two types are maintained differently: how procedural memory differs from semantic memory.
The distinction
How does procedural memory differ from semantic and episodic memory?
The three differ in what they store, what triggers retrieval, and how they should be maintained over time. Storing all three in one undifferentiated pile is the shortcut that produces bad retrieval.
Retrieval trigger. A semantic memory surfaces because the query is about its subject. An episodic memory surfaces because the query is about a period or an event. A procedural memory should surface because the agent is entering a situation it applies to, which is often before the user has finished asking.
Maintenance. Semantic memories are superseded when the fact changes. Episodic memories age and are eventually summarised. Procedural memories should be updated rather than expired: a workflow that has become wrong needs correcting, and deleting it just means the agent reverts to guessing. That is a genuinely different lifecycle, covered on forgetting and eviction.
Failure signature. A missing semantic memory looks like an agent that forgot a fact. A missing procedural memory looks like an agent that is inconsistent: it did the right thing last week and something else today. Users describe the second as unreliability rather than as forgetting, which is why the cause is often misdiagnosed.
The neighbouring types are covered on semantic memory and episodic memory.
With the distinction clear, the practical question is how it gets built: how to implement procedural memory.
Implementation
How do you implement procedural memory?
Capture corrections and successful sequences as explicit rules, store them with the situation they apply to, and retrieve them by situation rather than by topic. The hard part is the trigger, not the storage.
- Capture from corrections. When a user corrects behaviour rather than a fact, that is a procedural memory arriving. “No, always check the invoice first” should be written as a rule, not stored as a message.
- Capture from outcomes. When a sequence of tool calls succeeds where a previous attempt failed, the difference is worth recording. This is where an agent genuinely improves rather than merely accumulates.
- Store the condition with the action. A rule without its trigger condition cannot be retrieved at the right moment, and a rule retrieved at the wrong moment is worse than none.
- Retrieve before acting, not before answering. Procedural memory should be consulted when the agent selects a tool or plans a sequence, which is earlier in the turn than semantic retrieval usually runs.
A practical shortcut many teams use is to keep a compact, editable set of rules and inject all of them rather than retrieving selectively. That works while the rule set is small and it has the advantage of being inspectable by the user, which matters because procedural memory changes behaviour and people reasonably want to see why.
A failure worth anticipating is the rule that was right once and is wrong now. Procedures encode assumptions about tools, data and process, and all three change. A rule that says to check a system that has since been replaced will keep being retrieved and applied, and because it produces plausible behaviour rather than an error, nothing surfaces it. Reviewing procedural memories on a schedule is cheap insurance, and it is a different job from the eviction that semantic memories need.
Frameworks differ in how much of this they support: some provide only a fact store, so procedure has to be modelled as facts with a convention. The options are compared on the best AI memory tools, and the wider build on how to add memory to an AI agent.
That capture step raises the question people ask about this memory type more than any other: whether procedural memory means the agent is learning.
The distinction
Does procedural memory mean the agent is learning?
It means the system improves, not that the model does. The weights are unchanged; what changed is a store the agent reads before acting. The distinction sounds pedantic and has entirely practical consequences.
Because the improvement lives in data rather than in weights, it is inspectable: you can list the rules an agent has accumulated. It is reversible: a rule learned from a misread correction can be deleted in a second, where an equivalent fine-tune would need retraining. And it is portable: the same rule set can be attached to a different base model, which is not true of anything trained in.
Those properties also set the ceiling. Procedural memory cannot teach a model a capability it lacks. If the model cannot reliably call tools in sequence, storing a rule about the correct sequence does not fix it, because the failure is in execution rather than in knowledge of the procedure. That is the boundary where fine-tuning or a different model becomes the right answer instead, discussed on memory versus fine-tuning.
There is a genuine risk worth naming. An agent that accumulates rules from its own inferences, rather than from user corrections and verified outcomes, will happily learn a wrong procedure and then apply it consistently. Consistent wrongness is harder to spot than random wrongness, which is why provenance on procedural memories matters more than on semantic ones.
One practical consequence is worth planning for early. Because procedural memories change behaviour, they need a review path that semantic memories do not. A wrong fact produces one wrong answer and is usually noticed; a wrong rule silently changes every future response of that kind. Teams that ship this well tend to keep procedural memories visible to the user, editable, and small enough to read in one screen, which also happens to make the agent easier to trust.
The wider question of systems that improve over time is covered on continual learning versus memory and memory and reinforcement learning.
FAQ
Frequently asked questions
The questions that follow: whether procedural memory is just a system prompt, and how it relates to fine-tuning.
What is an example of procedural memory in AI?
"Always run linter before commit", "escalate to tier-2 after two failed resets", or a stored 5-step deploy workflow. Memory of how, not what.
Procedural vs semantic memory?
Procedural = skills and workflows (how). Semantic = stable facts (what). "Use pnpm" is semantic; "pnpm install → test → deploy" is procedural.
Do coding agents use procedural memory?
Yes — stored tool trajectories, lint-test-commit workflows and codebase interaction patterns. Combine with semantic codebase facts. See coding agents.
Can you store procedural memory in a vector DB?
Yes — embed successful tool-call trajectories and retrieve similar plans by semantic similarity. Tag with memory_type=procedural for filtering.
Fine-tuning vs procedural memory?
Fine-tuning bakes procedures into weights (parametric). Procedural memory keeps workflows external and updatable. See memory vs fine-tuning.
Procedural memory in LangGraph?
LangMem checkpointers store graph state and tool sequences natively. External Engram/Mem0 can store trajectory summaries as procedural memories.
Shared procedural memory in multi-agent systems?
Store team-wide procedures (escalation workflows, deploy runbooks) in a shared namespace; scope user-specific procedures by agent_id. See shared memory.
Procedural memory and memory-as-a-tool?
Agents invoke memory read/write as tools — procedural memories are plans the agent stores and recalls via those tools. See memory as a tool.