Use cases · Customer support
AI Memory for Customer Support Agents
A support agent with memory carries the account, the ticket history and the customer’s stated preferences into every new contact, so the person on the other end never has to explain themselves twice. The hard parts are not storage. They are keying memory to the right customer across channels, deciding what must never be written, and handing something auditable to the human who takes the escalation.
Support memory
The contents
What should a customer support agent remember?
Four things: who the customer is and what they are entitled to, what has already been tried, how they want to be dealt with, and which version of a policy applied to them. Each of those is a different memory type with a different write rule, and treating them as one undifferentiated pile is where most support deployments start going wrong.
Identity and entitlement is semantic memory: account ID, plan tier, contracted SLA, products owned. These facts decide what the agent is permitted to offer, and getting them wrong is not an embarrassment but a commercial error. They usually belong to the CRM, with the agent’s memory holding a synchronised copy rather than an independent one. See semantic memory in AI agents.
Ticket history is episodic memory: every prior contact, what was attempted, and how it ended. Its value is not the transcript but the outcome, because the single most irritating support experience is being walked through a fix that has already failed. See episodic memory in AI agents.
Stated preferences are user memory: channel, language, tone, do-not-call windows. Volunteered once, and expected to hold forever after. Policy in force is the temporal case, and it is the one most systems omit: the refund window that applied on the purchase date, not the one published today.
Before any of that is worth building, one objection has to be dealt with, because it is the first thing a support lead will raise: why a bigger context window does not solve this.
The objection
Why does a bigger context window not solve support memory?
Because the context window is per-session and per-conversation, and support is neither. A customer who chats on Tuesday, emails on Thursday and calls the following month is three separate sessions to the model, and no window size changes that. The window is working memory, not a store.
Even inside one long conversation, the window is the wrong instrument. Support transcripts are mostly pleasantries, diagnostics and dead ends, so loading the entire history spends tokens on material that will not affect the answer while pushing the one relevant sentence, stated forty turns ago, toward the middle where attention is weakest. The comparison is worked through on memory versus the context window and long context versus memory.
There is a cost argument as well. Support runs at volume, and a stack that pastes an entire account history into every turn pays for those tokens on every contact, not once. Retrieving four relevant memories instead is cheaper by an order of magnitude and usually more accurate, which is covered on reducing token cost with memory.
Granting that a store is needed, the next question is which store, because support already has one that answers a different question: how memory works alongside RAG.
Two layers
How does memory work alongside RAG in a support stack?
RAG answers what is true of the product, memory answers what is true of this person, and a support reply usually needs both. Most teams already have the first, which is why memory so often gets described as an upgrade to RAG rather than as the separate layer it is.
The retrieval difference matters more than the vocabulary. The help centre index is one corpus shared by everybody, and a mistake in it is a wrong answer. The memory index is partitioned per customer, and a mistake in it is a data breach. That asymmetry should show up in the architecture as separate stores with separate access paths, not as one collection with a customer ID in the metadata and a filter applied at query time.
In practice the two are retrieved in parallel and composed: the documented fix from RAG, filtered by the entitlement and the failed attempts from memory. An agent that has the article but not the history recommends a fix the customer tried last week. An agent that has the history but not the article knows the customer is annoyed and cannot say why the fix failed.
The composition pattern is set out on using RAG together with memory, and the underlying distinction on memory versus RAG.
Both layers assume the agent knows whose memory to open, and in support that assumption is usually false: keeping one memory across chat, email and phone.
Identity
How do you keep one memory across chat, email and phone?
By resolving every channel identifier to a single account ID before a memory is read or written, and refusing to write anything when that resolution fails. This is the step that quietly decides whether a support memory system is an asset or an incident, and it is almost never discussed.
Every channel hands over something different and none of it is reliable on its own. Web chat gives a session cookie and sometimes a logged-in account. Email gives an address, and plenty of customers have two. Phone gives caller ID, which identifies a handset rather than a person, so a household or a shared office line resolves several customers to one number.
Two failure modes follow, and they are not symmetrical. Failing to link identifiers means a returning customer is treated as new, which is the mild version: the agent is merely forgetful. Linking the wrong identifiers means one person’s history is retrieved into another person’s conversation, which is a disclosure incident that no amount of good phrasing recovers from.
The safe default is to treat memory as unavailable until identity is confirmed, exactly as a human agent asks for verification before discussing an account. An unresolved contact still gets a helpful answer from the shared knowledge base, just not a personalised one, and the memory writes from that conversation stay in a holding area until it is linked. Scoping rules are covered on memory management and partitioning.
Once the agent knows whose memory it is writing to, the question becomes what it is allowed to put there: what a support agent should never write to memory.
Write-time limits
What should a support agent never write to memory?
Card numbers, passwords and one-time codes, government identifiers, health details volunteered in passing, and anything a customer shares to prove identity rather than to be remembered. Support conversations carry all of these routinely, and an extraction step that is not told to drop them will store them.
The reason to enforce this at write time rather than at read time is durability. A retrieval filter protects the current model and the current prompt. Once a card number is in the store, it is in the backups, the analytics export and every downstream index built from it, and a deletion request now has to reach all of them. The filter belongs in the extraction pipeline described on how agents write and store memories.
A useful test for the grey cases is to ask whether the customer said it in order to be remembered. A dietary requirement given to a travel service was offered for future use, so storing it is a service. A date of birth given to pass verification was offered as proof, and storing it converts an authentication step into a permanent record. The distinction survives longer than any list of field names.
Two obligations follow from writing anything at all: the customer can ask what is stored, and can ask for it to be removed. Both need to be a supported operation rather than a database query somebody runs by hand, because a support team will receive these requests weekly. The controls, including tenant isolation and deletion paths, are set out on memory security and privacy for AI agents.
Even a clean memory has to survive the moment the conversation stops being handled by the agent: what happens when a ticket escalates to a human.
Escalation
What happens to memory when a ticket escalates to a human?
The human should receive the agent’s memories as a labelled, auditable summary showing what was retrieved, what was attempted and what the agent concluded, with each claim traceable to the contact it came from. Nobody in the published guidance on support memory covers this, and it is the moment the whole system is judged.
The failure is easy to picture. A customer escalates after twenty minutes with a bot, and the human agent opens a ticket containing a fluent paragraph with no indication of which parts came from the CRM, which were inferred by the model, and which the customer actually said. The human either re-asks everything, which wastes the entire interaction, or trusts a summary that may contain a confident invention.
Three properties fix it. Separate what was retrieved from what was generated, so the human can see the difference between a stored fact and a model’s paraphrase of one. Attach a source to every memory, meaning the ticket, date and channel it came from, so a doubtful claim can be checked in seconds. Show the failed attempts explicitly, because the most valuable thing the agent knows at escalation is what has already not worked.
The handover should also run in the other direction. What the human does next, the resolution, the goodwill credit, the promise made, is the highest quality episodic memory the system will ever get, and it is routinely discarded because it happened outside the agent. Writing it back is what stops the next contact starting from nothing. Summarisation at the boundary is covered on memory summarisation.
Escalations frequently turn on a policy the customer was promised, which raises the hardest write problem in support: updating memory when a policy changes.
Change over time
How do you update memory when a policy changes?
By invalidating the old fact with an end date rather than overwriting it, so the agent can still answer what applied to a customer who bought under the previous terms. A refund window that moves from 30 days to 14 does not make the old window untrue, only superseded.
Overwriting breaks the common case. A customer who bought in March under a 30 day window is entitled to it, and an agent that has only the current value will confidently quote 14 and be wrong in a way that costs a complaint. Storing the fact with a validity range keeps both answers available and makes the choice between them a matter of the purchase date rather than of luck.
This is where a bi-temporal graph earns its keep. Zep’s Graphiti models both when a fact was true in the world and when the system learned it, which is exactly the shape of a policy history. In the research introducing it, Zep reports an 18.5% improvement over a full-context baseline on LongMemEval (Rasmussen et al., 2025), the benchmark that has become the standard test for cross-session recall. A vector store can do the same job, but the versioning has to be built explicitly in metadata, and the retrieval filter has to be applied at every query rather than assumed.
Customer-stated facts change too, and less tidily: a preference reversed, an address updated, a contradiction between what was said in March and what is said today. The resolution rules are worked through on handling conflicting memory updates, and expiry on forgetting and eviction.
With the writes under control, the remaining question is whether any of this shows up in the numbers the team is judged on: which support metrics memory actually moves.
Measurement
Which support metrics move when you add memory?
Average handle time, repeat-contact rate, first-contact resolution and CSAT, in that order of how directly memory touches them. Memory vendors report recall accuracy on benchmarks. No support director is measured on recall accuracy, so the benchmark number has to be translated before it means anything internally.
Average handle time falls because the minutes spent reconstructing an account disappear. Repeat-contact rate falls because knowing which fix already failed stops the agent proposing it again. First-contact resolution rises when entitlement is visible at the first turn, since a great many transfers exist only to look something up. CSAT moves last and least predictably, because it responds to not being asked the same question a fourth time rather than to any single retrieval.
Attribution needs care. All four numbers move for reasons unrelated to memory, so the measurement worth running is an ablation: same agent, same period, memory retrieval disabled for a held-out slice of contacts. Anything weaker is a before-and-after comparison that a seasonal shift can explain away.
The retrieval-level numbers still matter as leading indicators, because a memory layer that is slow or that returns the wrong four facts will not move anything downstream. LOCOMO and LongMemEval scores, and the latency targets a support stack should hold to, are covered on memory evaluation metrics.
Which leaves the practical decision, and the one the search queries that reach this page are usually asking: which memory tool fits a support stack.
Selection
Which memory tool fits a support stack?
The choice is decided by two things: whether your facts change over time in ways you must reconstruct, and whether ticket IDs and error codes have to match exactly. Both are more common in support than in the assistant use cases these tools are usually demonstrated on.
- Engram suits a Weaviate stack. Its hybrid search matters more here than elsewhere, because ticket references, order numbers and error codes are lexical strings that pure vector similarity retrieves unreliably. See Engram.
- Zep suits a stack where policy and account state have histories, since the bi-temporal graph makes the validity range native rather than something you maintain. See Zep and its alternatives.
- Mem0 suits the first working version. It is framework-agnostic and quick to stand up, and it scores 66.9 on the LOCOMO J metric. See Mem0 and its alternatives.
- Redis suits the session tier underneath any of them, holding the live conversation while the durable store holds the account. See Redis agent memory.
One decision precedes all four. Support data is customer data, so whether the memory layer can run inside your own boundary is a constraint rather than a preference, and it eliminates candidates faster than any feature comparison. That trade-off is set out on open source versus managed memory, and the full ranking on the best AI memory tools.
Whichever you pick, the build order is the same: identity resolution first, then episodic capture at ticket close, then preferences, then policy versioning. Standing the last three up before the first is the sequence that produces a demo that impresses and a pilot that leaks. The implementation walkthrough is on adding memory to an agent.
FAQ
Frequently asked questions
The questions support teams ask before building: whether the bot remembers at all, which store to pick, and how deletion works.
Does an AI support agent remember past conversations?
Only if it has a memory layer. The model itself starts every session with no history, so recall across a chat on Tuesday and a call in March comes from an external store the agent searches before replying, not from the model. See why AI agents need memory.
What is the best persistent memory for customer support agents?
Engram if you run Weaviate and need exact matching on ticket and order references, Zep if account terms and policies have histories you must reconstruct, Mem0 for the fastest first version. The deciding constraint is usually whether the store can run inside your own boundary.
How do you stop a support bot asking the same question every time?
Write the answer to a durable store keyed to the resolved account, not to the session, and retrieve it at the first turn of the next contact. If the bot still repeats itself after that, the failure is normally identity resolution rather than storage.
Can an AI support agent share memory between customers by mistake?
Yes, if memory is keyed to a channel identifier such as a phone number or a shared email rather than to a resolved account ID. Partition per customer at the store level and treat memory as unavailable until identity is confirmed.
How do you delete what a support agent remembers about a customer?
Deletion has to be a supported operation covering the memory store, its backups and any index derived from it, not a manual query. Keeping payment details and identity proofs out of memory at write time is what makes those requests tractable. See memory security.
Is RAG on the help centre enough for customer support?
No. RAG answers what is true of the product for everybody, and support also needs what is true of this customer: entitlement, ticket history and what has already failed. Both layers are retrieved on the same turn. See RAG with memory.