Use cases · Coding agents

AI Memory for Coding Agents

A coding agent already has a perfect record of what the code does, because the repository is right there. What it lacks is everything the code cannot say: why an approach was abandoned, which fix failed last week, which test is flaky, which module the team will not let you touch. That gap is what memory is for here, and it makes this the one use case where storing more of the source material actively makes things worse.

Four scopes

1
Repo
Conventions
2
Branch
In progress
3
Task
This change
4
Developer
Preferences

The contents

What should a coding agent remember?

Decisions and outcomes, not code. Which approach was rejected and why, which fix failed against which error, which tests fail intermittently, which conventions the team enforces in review, and which parts of the system nobody wants changed casually.

What to read from the code, what only memory can hold, and what belongs in committed context files.
Figure 1. The left column is derivable from the working tree at any moment. Storing it produces a second, worse copy that immediately begins to diverge.

The test is whether the fact is recoverable from the repository. Function signatures, call sites, imports, directory layout and current behaviour all are, so they belong in an index rebuilt from the tree rather than in a memory store. An index cannot go stale against the code, because it is derived from it.

What is not recoverable is the history of intent. A commit shows that a caching layer was removed; it rarely shows that it was removed because it broke under a specific concurrent workload and that reintroducing it is a known trap. That is exactly the memory that saves the next hour, and it exists nowhere in the tree.

Outcomes are the other high-value category. An agent that knows a fix was already attempted and failed will not propose it again, which is the same episodic value memory delivers in support work, described on episodic memory in AI agents.

All of which raises the objection this use case always attracts, since the source is right there to be read: why reading the codebase is not enough.

The objection

Why is reading the codebase not enough?

Because the code records what the system does now and almost nothing about how it got there. Every abandoned approach, every constraint discovered the hard way and every rule that lives in review comments rather than in a linter is absent from the tree by construction.

There is also a scale problem. A large repository does not fit in any context window, so an agent reads a retrieved subset, and retrieval over code has a specific weakness: the fragments that match a query lexically are frequently not the ones that matter. A question about authentication returns the middleware and misses the configuration two directories away that actually governs it.

Git history looks like it should fill the gap and mostly does not. Commit messages record what changed, occasionally why, and never the three things tried before the one that shipped. The reasoning happened in a conversation, a review thread or an agent session, and unless something captured it, it is gone.

The third gap is the session boundary. A coding agent that spent an hour learning the shape of a module starts the next task knowing none of it, which is the general statelessness described on why AI agents need memory, felt more sharply here because the learning cost per session is so high.

Not all of that belongs in an automatic memory store, though, and the distinction is the one this use case gets wrong most often: context files versus agent memory.

Layers

What is the difference between context files and agent memory?

Context files are written by the team, committed and reviewed; agent memory is captured from sessions automatically and belongs to one developer until it is promoted. Both end up in the prompt, and treating them as one thing is why teams end up with either an unreviewable memory store or a context file nobody maintains.

Three persistence layers for a coding agent: committed context files, learned agent memory, and a derived code index, with promotion between them.
Figure 2. Three layers with different authors, different lifetimes and different review paths. The arrow between the first two is the part most setups are missing.

Context files, the AGENTS.md and CLAUDE.md convention, hold what the team has decided: build and test commands, architectural boundaries, code standards, and what the agent must never do. They are deliberate, version-controlled and reviewed like code, which is exactly right for rules everyone is bound by. Their weakness is that they only contain what someone thought to write down.

A walkthrough of the committed layer: what belongs in a context file and what does not.

Agent memory is captured without anyone deciding to write it: the failure encountered at 2am, the module that turned out to own a behaviour, the command that actually works on this machine. It is comprehensive and unreviewed, which makes it valuable and slightly dangerous.

The connection between them is a promotion step. A learned memory that proves durable and applies to everyone should be moved into a context file, where review governs it and the whole team benefits. Without that step, one developer’s agent knows something the others’ do not, indefinitely.

Both layers need to know which part of the work they apply to, which is a scoping question with more levels here than elsewhere: scoping memory across repositories and branches.

Scoping

How should memory be scoped across repositories and branches?

Four scopes, and the branch one is the trap. Repository, branch, task and developer each have a different lifetime, and collapsing them into one store produces an agent that confidently describes code that was never merged.

Four memory scopes for a coding agent: repository, branch, task and developer.
Figure 3. Only the repository scope should outlive the work that produced it, and only after something has decided it is durable.

Repository scope holds conventions and architecture: stable across branches, worth sharing with everyone, and the destination for anything promoted. Branch scope holds work in progress, the migration half applied or the refactor under way. It has to expire when the branch merges or is abandoned, because a memory describing an intermediate state of an abandoned branch is actively misleading.

Task scope is the current session: what has been tried, what the error was, what the last run produced. Discarded when the task closes, except for outcomes worth keeping. Developer scope is personal preference within the team’s rules, and it should never be promoted to repository scope without being asked, since one person’s habit is not a team standard.

In a monorepo the repository scope needs subdividing further, by package or service, because conventions genuinely differ between them and a rule retrieved from the wrong package is a wrong rule. The general partitioning mechanics are on memory management and partitioning.

Scoping controls where a memory applies. It does not control whether the memory is still true, which is this domain’s hardest problem: keeping code memories from going stale.

Staleness

How do you stop code memories going stale?

Anchor every memory to the commit it was written against and to the symbols it names, then verify those symbols still exist before the memory is used. Code memory decays faster than any other kind, because the subject changes several times a day without notifying the store.

How a code memory goes stale: written against a commit, invalidated by a refactor, retrieved confidently wrong, fixed by anchoring to symbols.
Figure 4. A stale code memory is not a gap in knowledge. It is a confident, specific, wrong assertion, which is considerably worse.

The verification is cheap and rarely implemented. If a memory names requireUser and that symbol no longer exists in the tree, the memory is suspect and should be demoted or re-checked rather than injected. This makes the code index a validator for the memory store as well as a retrieval source.

Preferring durable statements helps more than any invalidation scheme. “Authentication lives in the middleware layer” survives a refactor that renames every function in it; “the token check happens on line 42 of auth.ts” is wrong within a week. Writing memories at the level of the decision rather than the line is the single most effective habit here.

Merges are the natural invalidation trigger: on merge, expire the branch scope and re-check any repository-scope memory naming a symbol the merge touched. That is more work than most setups do and it is the difference between a store that improves over months and one that quietly rots. The general mechanics are on forgetting and eviction and conflicting memories.

Verification assumes the memory should have been written at all, and some of what a coding session produces should never reach a store: what a coding agent should never remember.

Write-time limits

What should a coding agent never remember?

Secrets, file contents, and anything the agent generated itself but never verified. The first is obvious and still happens; the second and third are the ones that quietly degrade a store.

Secrets reach a coding agent constantly: keys pasted into a terminal, connection strings in error output, tokens in environment dumps. An extraction step that is not explicitly told to drop them will store them, and a memory store is rarely covered by the secret-scanning that protects the repository. Filter at write time, not retrieval, for the reasons set out on memory security.

File contents are the domain-specific version of storing everything. Snippets copied into memory are stale the moment the file changes, and they compete at retrieval time with the durable decision memories that are the whole point. Store the claim about the code, and let the index supply the code.

Unverified agent output is the compounding failure. If extraction reads the agent’s own explanations as well as the developer’s messages and the tool results, an invented API becomes a stored memory, is retrieved as evidence and gets reinforced. Restrict extraction to human turns and verified tool output, meaning test runs, build output and file reads, as covered on how agents write and store memories.

What remains after those exclusions still has to be found at the right moment, which raises the relationship with code search: memory alongside retrieval over the codebase.

Two sources

How does memory work alongside retrieval over the codebase?

The index answers what the code is, memory answers what the team knows about it, and a good answer usually needs both. Confusing the two is why “we already have codebase search” is offered as a reason not to add memory.

The index is rebuilt from the tree, is always current and covers symbols, structure and text. Memory is written from interactions, can be wrong, and covers intent and outcome. Their failure modes are opposite: an index is never stale and never explains, while memory explains and can be out of date.

Retrieval over code also benefits from structure that plain chunking destroys. A function split across two chunks retrieves as two fragments that neither compile nor explain, which is why symbol-aware indexing and graph representations of the call structure outperform naive embedding of file contents. The graph approach is covered on knowledge graphs for AI memory.

In the prompt, keep the two labelled separately. A retrieved file is ground truth about the current state; a retrieved memory is a claim that was true at some past commit. An agent that cannot tell them apart will defend a stale memory against the file in front of it. The general pattern is on using RAG together with memory and memory versus RAG.

With the architecture settled, the remaining question is what to build it with: which tools give a coding agent memory.

Tooling

Which tools give a coding agent memory?

Three categories, and most teams end up with one from each: a committed context file, an index over the repository, and a memory store for what neither of those holds.

  • Engram suits teams already running Weaviate. Its hybrid search matters here more than in most domains, because symbol names, error strings and stack frames are lexical and retrieve unreliably by similarity alone. See Engram.
  • Cognee builds a graph from the repository, which fits the structural half of the problem: call relationships and module boundaries are edges rather than passages. See Cognee.
  • Mem0 is the quickest route to per-developer session memory, framework-agnostic and simple to key by repository. See Mem0 and its alternatives.
  • Context files, AGENTS.md and CLAUDE.md, need no tool at all and are where anything durable should end up regardless of what else you run.

Two constraints narrow the field faster than any feature list. Source code is usually the most sensitive asset a company has, so whether the memory layer can run inside your own boundary is frequently decisive, as discussed on open source versus managed memory. And the store must be readable by a person, because a coding memory nobody can inspect cannot be corrected when it goes stale. The wider comparison is on the best AI memory tools.

Whatever you choose, one question decides whether it stays: whether the memory is actually helping.

Measurement

How do you tell whether memory is helping a coding agent?

By counting how often the agent has to re-learn something it already established. That is the behaviour memory exists to remove here, and it is observable without any benchmark.

Three signals are cheap to collect. The first is repeated exploration: how many file reads or searches a session spends re-establishing something a previous session already concluded. If that number does not fall after memory is introduced, either nothing is being written or nothing is being retrieved, and the two are distinguished by reading the store directly rather than by watching the agent. The second is repeated failure, meaning how often the agent proposes a fix that a stored memory records as already tried and failed, which is the clearest possible retrieval miss.

The third is the one that catches the specific danger of this domain: stale assertions. Sample the memories the agent cited over a week and check each against the current tree, which takes minutes and needs no tooling beyond a symbol lookup. Any citation of a symbol that no longer exists is a defect in the anchoring described above, and the rate tells you whether verification is working.

Beyond those, the honest measurement is an ablation: run the same tasks with memory retrieval disabled and compare the number of turns to completion. Developer-satisfaction surveys move for many reasons and turn count moves for fewer. The general method is on how to evaluate agent memory, and the build order for putting any of this in place on the memory build guides.

FAQ

Frequently asked questions

The questions that follow from putting memory behind a coding agent: sharing, privacy and what to do about monorepos.

Is AGENTS.md or CLAUDE.md the same as agent memory?

No. A context file is written by the team, committed and reviewed, so it holds what everyone has agreed. Agent memory is captured automatically from sessions and belongs to one developer until something promotes it. Both reach the prompt, from different places and with different review.

Should coding agent memory be shared across a team?

The durable part should, by promoting it into a committed context file where review governs it. Sharing a raw session memory store across a team spreads one developer's unverified conclusions, including the wrong ones.

How do you handle memory in a monorepo?

Subdivide the repository scope by package or service, because conventions genuinely differ between them and a rule retrieved from the wrong package is a wrong rule. Cross-cutting decisions still belong at the repository level.

Should a coding agent store code snippets in memory?

No. Snippets are stale as soon as the file changes and they compete at retrieval time with the decision memories that matter. Store the claim about the code and let an index rebuilt from the tree supply the code itself.

How do you keep secrets out of a coding agent's memory?

Filter at write time rather than at retrieval, because once a key is in the store it is also in the backups and any derived index. Memory stores are rarely covered by the secret scanning that protects the repository. See memory security.

Does a coding agent need memory if it can search the repository?

Search answers what the code is. Memory answers why it is that way, what was tried before, and which tests cannot be trusted. Neither substitutes for the other, and the useful setups run both.