Use cases · Personal assistants
How Do Personal AI Assistants Use Memory?
A personal assistant is the purest memory product: almost all of its value comes from accumulated knowledge of one person, and almost all of its failures come from remembering the wrong things. This page is about building that memory, covering what to store, how to retrieve it, how to handle facts that change, and where the line sits between helpful and unsettling.
What it accumulates
The requirement
What does a personal assistant need to remember?
Three categories carry almost all the value: stable facts about the person, the history of what has already been discussed or done, and the working preferences that shape how output should look. Everything else that gets stored is usually noise competing with these for retrieval.
Stable facts are the ones users are most annoyed to repeat. Dietary requirements, the names of family members, the city they live in, their role, the tools they use, their working hours. These change rarely, apply broadly, and are cheap to store, which makes them the highest return per memory in the entire system.
History is what makes a follow-up possible. Which task was in progress, what was decided last week, what the assistant already suggested and the person rejected. Without it, an assistant restarts every conversation with the same suggestions, which reads as not listening rather than as forgetting.
Working preferences govern the shape of a response rather than its content. Short answers rather than long ones, bullet points rather than prose, being asked before an action is taken rather than after. These are procedural rather than factual, and they are the category most systems fail to store because they arrive as corrections rather than statements.
A fourth category is worth naming so it can be deliberately excluded. Passing remarks, one-off moods and small talk are the bulk of any conversation and almost never worth storing. An assistant that remembers everything said to it retrieves badly, because the useful memories are competing with hundreds of trivial ones.
Those categories map onto distinct memory types, and the mapping decides the architecture: which memory types a personal assistant needs.
The architecture
Which memory types does a personal assistant need?
Semantic memory does most of the work, episodic memory supplies continuity, and procedural memory is what separates an assistant that improves from one that merely remembers. The three are retrieved for different reasons and should be maintained on different schedules.
Semantic memory holds the stable facts, and it is where most teams start because it delivers the recognisable behaviour quickest. It is also the easiest to get right, since a fact stated plainly can be extracted, deduplicated and stored with little ambiguity. See semantic memory in AI agents.
Episodic memory holds the ordered record of interactions. It is harder, because an episode is longer than a fact, ages faster, and is usually summarised rather than kept whole. The design question is what granularity to store: whole sessions, individual decisions, or only outcomes. See episodic memory in AI agents.
Procedural memory holds the working preferences, and it is the one most personal assistants skip. Skipping it produces the specific complaint that an assistant “never learns”, because a correction given three times has changed nothing. See procedural memory in AI agents.
Working memory, meaning the context window itself, underpins all three, since retrieved memories still have to arrive in the prompt to be used. The distinction is on memory versus the context window.
Before the architecture matters, many people are asking something simpler: whether an assistant that remembers you exists at all.
The state of things
Is there an AI assistant that can remember you?
Yes. Persistent memory is now a standard feature in mainstream consumer assistants, and it works the way this site describes: a filtered set of facts is extracted from conversations, stored outside the model, and retrieved into later sessions. It is not the model remembering; it is a system around the model.
Two things about those products surprise people, and both follow from the mechanism. They remember less than users assume, because extraction discards most of a conversation by design. And what they remember is editable, since each memory is a record rather than a change to the model, which is why these products can show a list of what they know and let you delete an entry.
For anyone evaluating rather than building, the useful questions are not about model quality. Ask whether memories persist across devices and sessions, whether you can see and edit them, whether the assistant tells you when it stores something, and what happens to the memories if you stop using the product. Those four answers separate a genuine memory feature from a longer context window described as one.
For anyone building, the more important observation is that consumer expectations are now set by those products. A new assistant that forgets a stated preference is measured against tools that do not, which moves memory from a differentiator to a baseline requirement.
The comparison against a bigger context window, which is the other thing marketed as memory, is on long context versus memory.
Building it starts at the write path, which is where most of the quality is decided: how an assistant decides what to remember.
The write path
How does an assistant decide what to remember about you?
An extraction step reads each turn and keeps only what looks durable, turning conversation into a small number of clean facts before anything is stored. Getting this filter right matters more than any other decision in a personal assistant, because a store full of noise cannot be rescued by good retrieval.
The filter needs to distinguish three things a conversation contains. Durable facts should be stored. Tasks should not: “book me a table on Friday” is work to do, not something to know. Transient state such as a mood or a one-off constraint should either expire quickly or not be written at all.
Deduplication is the second half of the write path and the more commonly skipped one. Before storing, check whether a similar memory exists and update it rather than adding a near-copy. Without this, a preference restated three different ways occupies three slots at retrieval and the assistant appears to hold three opinions about one thing.
A third element specific to personal assistants is confidence. A fact the user stated plainly should outrank one the assistant inferred from behaviour. Recording which is which is one small field and it prevents an entire class of failure where a guess hardens into a belief. The mechanism is covered on how agents write and store memories.
Storing well is only useful if the right memory comes back at the right moment: how retrieval picks what to surface.
The read path
How does an assistant retrieve the right memory?
It embeds the incoming question, searches the store scoped to that person, ranks candidates on recency, importance and relevance, and injects only the top few into the prompt. The number injected is a real design parameter and smaller is usually better.
The canonical scoring comes from the Generative Agents paper (Park et al., 2023), which sums recency, importance and relevance with each weight set to 1.0 and decays recency exponentially at 0.995 per hour since the memory was last accessed. For a personal assistant the importance signal matters more than in most applications, because a stated dietary requirement should outrank a topically similar restaurant mentioned once.
One retrieval decision is specific to this product shape. A small set of core facts is relevant to almost every conversation: name, working style, key constraints. Many assistants always inject those cheaply, and run the full search only when the query looks historical. That keeps the always-relevant facts present without paying for a full retrieval on every trivial turn.
The failure to watch for is retrieving too much. Every injected memory occupies prompt space and adds a chance of the answer anchoring on the wrong detail, which is the effect Liu et al. documented in “Lost in the Middle” (arXiv:2307.03172). Three well-ranked memories generally beat twelve. The mechanism is on memory retrieval.
Retrieval is also where a personal assistant can go wrong in a way no other product can: the line between helpful and unsettling.
The product line
What makes assistant memory feel unsettling rather than helpful?
Surfacing something the person did not expect the system to have kept, or bringing up a memory in a context where it has no business appearing. Both are retrieval and extraction problems, which means they are engineering decisions rather than matters of taste.
The pattern that works is memory that reduces effort without announcing itself. An assistant that quietly avoids suggesting a restaurant serving nothing the person eats is using memory well. An assistant that says “I remember you mentioned being vegetarian on the fourteenth of March” is using the same memory badly, because it draws attention to surveillance rather than to helpfulness.
Three rules keep it on the right side. Store fewer, more durable things, since the discomfort usually comes from a trivial remark being kept. Use memories to shape answers rather than to narrate them, so recall is visible in behaviour rather than in commentary. Never surface a memory outside its context: something shared while discussing a health question should not appear while planning a holiday, even if similarity search rates it relevant.
That third rule is worth implementing structurally rather than trusting to ranking. Tagging memories with the context they were captured in, and filtering by context at retrieval, prevents a whole category of uncomfortable moments that no amount of ranking tuning will reliably avoid.
The inverse failure is worth stating too. An assistant that stores nothing is not respectful, it is useless, and users experience repeated re-explanation as its own kind of disrespect. The target is a small, accurate, inspectable set of memories rather than either extreme.
Accuracy over time is the next problem, because people change: handling facts that stop being true.
Change over time
How do you handle facts that change about a person?
Mark the old fact superseded rather than deleting it, and rank a direct statement above an inference when the two disagree. A personal assistant accumulates facts about a life, and lives change, so this is not an edge case but the normal condition of the store.
The categories that change most are predictable: where someone lives, what they do, who they work with, what they are currently focused on. Those deserve shorter retention or explicit revalidation, while genuinely stable facts such as a dietary requirement or a family relationship can be held indefinitely.
Keeping the superseded version rather than overwriting it buys two things. The assistant can answer questions about the past, which people do ask. And a correction made in error is reversible, which matters because extraction is imperfect and a misheard statement should not permanently overwrite something true.
Where the ranking cannot separate two versions, asking is better than guessing. “You mentioned Berlin in March and Munich last week, which should I use?” is specific, shows the evidence, and takes a moment to answer. Silently picking one and being wrong is the option users find hardest to forgive. The full procedure is on handling conflicting and stale memories.
Change also means the store grows, which raises the operational question: what assistant memory costs to run.
The economics
What does personal assistant memory cost to run?
A search on every turn, storage that grows with the relationship, and a background job that keeps the store usable, against the alternative of resending the entire history on every call. The comparison favours memory decisively, and the published figures show why.
Mem0’s measurements on the LOCOMO benchmark put selective retrieval at a median search latency of 0.148 seconds and a p95 total latency of 1.44 seconds, against 17.1 seconds for a 26,000-token full-context baseline, with token consumption falling from roughly 26,000 per query to roughly 1,800 (Chhikara et al., 2025). For a product where one person may accumulate years of conversation, that difference is the difference between a viable unit economic and an impossible one.
The accuracy trade is worth stating plainly rather than omitting. The same paper reports an LLM-as-a-judge score of 66.9 for the extract-and-store pipeline against 72.9 for the full-context baseline. Memory gives up a small amount of accuracy for a large amount of cost and latency, and for a personal assistant it becomes the only option once the relationship outgrows any window.
The cost that surprises teams is maintenance rather than serving. A personal assistant’s store grows for as long as the user stays, so consolidation and eviction are not optional refinements: without them the store fills with near-duplicates, retrieval slows, and quality declines gradually enough that nobody attributes it to memory. That is covered on memory consolidation.
Two design questions sit underneath the cost, and both are specific to this product shape: how an assistant starts from nothing.
Cold start
How does an assistant build memory from nothing?
Slowly, and the first two weeks are the period most likely to lose a user, because an assistant with an empty store is strictly worse than one without memory: it has all the cost and none of the benefit. Designing for that window matters more than optimising the steady state.
Three approaches are used, and they combine well. Ask directly, once. A short onboarding that captures the handful of facts that apply to almost every conversation, meaning name, role, working style and any hard constraints, seeds the store with exactly the memories that will be retrieved most often. It feels like setup rather than surveillance because the user is choosing what to share.
Import what already exists. Where a user already has data in a connected system, a calendar, a task list, a profile, the durable facts in it are better seeds than anything inferred from early conversations. This has to be scoped and consented explicitly, but it converts a cold start into a warm one.
Extract aggressively at first, then taper. A store with ten memories can afford a lower bar for what counts as durable than a store with ten thousand, because there is nothing yet for a marginal memory to compete with. Tightening the extraction threshold as the store grows keeps early usefulness without producing long-term noise.
The failure to avoid is announcing memory before it has anything to remember. An assistant that says it will learn your preferences and then demonstrably has not, three sessions later, has spent credibility it did not have. Under-promising during the cold start and letting the behaviour speak once the store fills is the pattern that works.
Where those memories physically live is the other decision specific to personal products: on device or in the cloud.
Placement
Should assistant memories live on the device or in the cloud?
Cloud storage is what makes memory work across devices, and local storage is what makes some users willing to have memory at all. For a personal assistant this is a positioning decision as much as a technical one, and it is worth making deliberately rather than defaulting.
Cloud memory is the straightforward choice. One store, reachable from a phone, a laptop and a browser, with a single consolidation job maintaining it. Everything described on this page assumes it. The cost is that a user’s accumulated personal facts sit in someone else’s infrastructure, which some people will not accept for the exact categories a personal assistant is most useful for.
Local memory keeps the store on the user’s machine, which removes that objection and creates three new problems: memories no longer follow the person between devices, consolidation has to run on consumer hardware, and there is no recovery if the device is lost. A hybrid, where less sensitive memories sync and sensitive categories stay local, is more work than either and is what several privacy-positioned products actually ship.
The technical prerequisite is the same either way and worth building first: memories scoped and categorised at write time. A store where every record already carries a user identifier and a type marker can be split by policy later. A store where it does not cannot be, which turns a positioning change into a migration.
Whichever placement you choose, the retrieval and maintenance mechanics are unchanged, and the storage options are covered on memory storage backends.
Because the store holds personal data, one more requirement is not negotiable: letting people see and control what is remembered.
User control
How do you let users see and control their memories?
Show the list, allow editing and deletion, and make the scope identifier a first-class field from the very first record. This is both a product feature and a compliance requirement, and retrofitting it is painful.
The technical prerequisite is scope. Every memory needs a user identifier attached at write time and applied as a filter during search rather than after it. A store built without that cannot answer “show me everything you hold about this person”, cannot delete it reliably, and cannot guarantee that one person’s memories never reach another’s prompt. Adding the field later means rewriting every row.
The product surface that follows is straightforward and unusually well received. A list of stored memories, each editable and deletable, with an indication of when it was learned. Users who can inspect what an assistant knows are markedly more forgiving of it being occasionally wrong than users who cannot, because a visible error is a correctable one rather than evidence that the system is unreliable.
Two smaller behaviours help disproportionately. Telling the user when something is stored, lightly rather than intrusively, converts memory from something that happens to them into something they participate in. And honouring a deletion completely, including any summaries derived from the deleted memory, avoids the situation where a fact was removed and its shadow keeps influencing answers.
The storage requirements that make all of this possible are covered on the infrastructure cluster, and the scoping detail on the memory layer.
Finally, none of this can be assumed to work without measurement: how to evaluate a personal assistant’s memory.
Measurement
How do you evaluate a personal assistant’s memory?
Build a set of questions whose answers depend on facts stated in an earlier session, then measure three things separately: whether the right memory was retrieved, whether the answer used it, and what it cost. Public benchmarks tell you how a tool performed on somebody else’s conversations.
Separating retrieval from answering is what makes the results actionable. A memory that was never written is an extraction problem. One that exists but was not retrieved is a search or ranking problem. One that was retrieved and ignored is usually a prompt problem, often caused by injecting too many memories. Those three have different fixes and a single quality score conflates them.
Two measurements specific to personal assistants are worth adding. Staleness: how often a retrieved fact is out of date, which is the failure users notice fastest. And store growth: how many memories a typical user accumulates per month, which predicts when consolidation stops being optional.
The public benchmarks remain useful as a filter on tools rather than a verdict on your system. LOCOMO tests very long conversations and LongMemEval tests long-horizon recall, both covered on LOCOMO and LongMemEval, with the metric set on memory metrics.
The implementation of everything above is set out step by step in how to add memory to an AI agent, and the tool choice on the best AI memory tools.
FAQ
Frequently asked questions
The questions that follow: how much an assistant should store, and what happens to memories across devices.
Best memory framework for personal assistants?
Engram for Weaviate-native personal apps; Mem0 for fastest per-user POC (LOCOMO J 66.9); Letta for long multi-week threads with paging. See best AI memory tools.
Does Mem0 work for assistant personalization?
Yes — both Engram and Mem0 extract and retrieve per-user facts automatically. Engram fits Weaviate-native stacks; Mem0 is framework-agnostic.
Does Claude have built-in assistant memory?
Claude's context window covers the current session. Durable cross-session memory requires external tools — Engram, Mem0 or custom vector stores.
Local-only memory for personal assistants?
Self-host Weaviate+Engram, Mem0 OSS, or Redis/pgvector DIY for full data residency. Managed APIs trade control for speed.
Cross-device memory sync for assistants?
Centralize memory in a cloud store keyed by user_id — Engram, Mem0 or Zep. Devices read/write the same collection.
Is Engram good for personal assistants?
Yes — scoped per-user collections, hybrid search and managed write pipeline on Weaviate. GA June 2026. See Engram explained.
GDPR forget for personal assistant memory?
Implement delete-all-memories per user_id. Engram, Mem0 and Zep all support per-user deletion. See forgetting and eviction.