Tools · Managed memory
Engram: Managed Agent Memory on Weaviate
Engram is Weaviate’s managed memory service for agents: you send it conversations, it extracts and reconciles the facts worth keeping, and it serves them back through Weaviate’s hybrid search. It became generally available on 24 June 2026 in Weaviate Cloud. What makes it worth a close look is not the feature list but the write pipeline, because deciding what to store and reconciling it against what is already there is the part most teams underestimate.
The pipeline
Definition
What is Engram?
A memory server for LLM agents, offered as a managed service with a REST API and a Python SDK. Your application sends content and issues searches; the extraction, deduplication and storage happen on the other side of that API.
The distinction worth drawing immediately is between a memory service and a database. A vector database gives you a place to put things and a way to find them, and everything about what goes in remains your problem. Engram takes the write side as well: what counts as a fact, whether it duplicates something stored last week, and what happens when it contradicts something stored last month.
It is built on Weaviate and inherits that database’s retrieval, which matters for memory specifically because memories are full of exact strings that pure vector search handles badly. Weaviate as a database is covered on vector databases; this page is about the memory layer on top of it.
One point of vocabulary, since the name collides with a term from neuroscience. An engram in biology is the physical trace a memory leaves; the product borrows the word, and nothing on this page depends on the analogy. The general warning about borrowed vocabulary is on AI memory versus human memory.
The write side is what distinguishes this category, so it is the right place to start: how Engram decides what to remember.
The write path
How does Engram decide what to remember?
Through a three-stage pipeline that extracts facts from what you send, merges them against what is already stored, and commits the result asynchronously. You can send raw text, a full conversation, or facts you extracted yourself.
The asynchronous part deserves emphasis because it shapes how you integrate. The API returns a run identifier straight away and the processing continues in the background, so the memory is not necessarily searchable the moment your call returns. If you need to know when it landed, you poll the run status. If you do not, you carry on and the hot path is never waiting on memory work.
That design is the right one for a memory system, and it is worth understanding why. Extraction is a model call, deduplication is a search plus another decision, and neither belongs in front of a user waiting for a reply. The general form of this choice, and the trade-off that a memory formed after a turn is not available during it, is on how agents write and store memories.
Accepting pre-extracted facts is worth noticing too. A team that already has a working extraction step can keep it and use Engram for storage, reconciliation and retrieval, which makes adoption incremental rather than a rewrite.
The middle stage of that pipeline is where the genuinely hard problem lives: what happens when a stored fact changes.
Reconciliation
What happens when a stored fact changes?
The transform stage merges the new statement against existing memories rather than appending it, which is how a store avoids filling with five versions of one fact. Weaviate describes the pipeline as handling deduplication, preference changes and time-evolving facts, and reconciling deterministically to avoid drift.
It is worth being clear about why this is the hard part. Appending is trivial and produces a store that gets worse as it grows: the same preference stated across five sessions becomes five near-identical rows, retrieval returns all five, and the context budget is spent restating one thing. Worse, when the preference changes, the old rows do not disappear and the agent now has evidence for both answers.
Doing better means deciding, at write time, whether an incoming statement is new, a duplicate, a refinement, or a replacement, and that decision has no perfect algorithm. The general problem, including why timestamps alone do not solve it, is on conflicting memories and memory consolidation.
For a buyer the practical question is not whether reconciliation exists but whether you can see what it did. A managed pipeline that silently merges two facts into one is doing exactly what you asked and is harder to debug than a rule you wrote, which is a real trade and the reason the fit section below exists.
Once memories are stored and current, the other half of a memory layer is getting them back: how Engram retrieves memories.
Retrieval
How does Engram retrieve memories?
By vector similarity, by BM25 keyword search, or by hybrid retrieval that fuses both, with the choice made per query. That third option is the one that matters for memory, and it comes from the database underneath rather than being bolted on.
Memories are unusually dependent on exact matching. They are short, so a single unusual token carries most of the meaning, and they are full of identifiers: ticket numbers, product names, error codes, account references. Pure vector search handles those worst, and a miss on one is a visible failure rather than a slightly worse ranking. The mechanics and the reasoning are on hybrid search for agent memory.
Choosing the retrieval type per query is more useful than it sounds. A query naming an identifier can lean lexical, a query asking what someone prefers can lean semantic, and routing on that distinction is cheap to implement and noticeably improves recall.
What retrieval cannot do is compensate for what was never written. If the extraction step did not keep the fact, no search will find it, which is why the write pipeline gets the attention on this page. The read side in general is on how agents retrieve memories.
Retrieval also has to return the right person’s memories and nobody else’s, which is a boundary rather than a ranking question: how memory is isolated between users and projects.
Isolation
How is memory isolated between users and projects?
By scopes set at write time: per project, per user, and per custom property such as a conversation identifier. Weaviate’s own framing is that scoping is part of the primitive rather than a flag added later, which is the right architecture for something that is a security boundary rather than a filter.
The distinction between scopes and topics is worth internalising. A scope decides what a caller is allowed to see. A topic categorises memories so that retrieval can be narrowed to a subject. They look similar in an API and they are not interchangeable: filtering by topic where you meant to scope by user is a data leak that passes every functional test.
Custom scope properties are what make the model flexible. A workspace, a tenant, an account, or a single conversation can each become a boundary, which covers the shared and multi-agent cases where several actors read from one pool. That design space is on shared memory.
Scoping is also the mechanism behind most of what a privacy review will ask about, since a question like “show me everything stored about this user” is answerable when memories are scoped to them by construction. The wider requirements are on memory security and privacy.
With the mechanics covered, the questions become commercial: what Engram costs and whether it is generally available.
Availability
What does Engram cost, and is it generally available?
It reached general availability on 24 June 2026 and runs in Weaviate Cloud, with a free tier of 1,000 pipeline runs per month and paid plans starting at $45 per month. Those figures come from the company’s own general availability announcement, and pricing changes, so treat them as the state on that date and check the current page before budgeting.
The unit worth understanding is the pipeline run, since that is what the free tier counts. A run is one pass of the extract, transform and commit sequence over something you sent, so the cost of a memory system on this model tracks how often you write rather than how much you store or how often you search. That is a different shape from a per-record or per-query price and it rewards batching a conversation into one write instead of firing one per turn.
Templates for common use cases ship alongside the primitives. The general availability post describes personalisation as available on day one, with continual learning and multi-agent state templates following in the weeks after, and says that teams who outgrow a template can drop down to direct pipeline control without changing platforms. If a template is central to your plan, verify it has shipped rather than assuming.
One thing this page cannot give you is a benchmark comparison. Engram has no published LOCOMO or LongMemEval result at the time of writing, so any ranking that places it against systems with published scores is comparing a number to an absence. What the benchmarks do and do not settle is on how to evaluate agent memory.
What does exist, unusually for this category, is a published account of the product failing: what Engram’s own team found did not work.
Evidence
What did Engram’s own team find did not work?
Two things, documented by Weaviate in a pre-general-availability write-up of two weeks using Engram in their own daily coding sessions. One was an integration bug and the other is a problem every memory product in this category has.
The agent did not search. In a planning session, relevant memories were stored and the agent’s own instruction file told it to retrieve from memory at the start of a session. It did not, treating the task as forward-looking and working from the prompt alone. The write-up notes that the failure was silent: no error, and no indication that relevant context had been skipped.
That result is worth more than any benchmark, and it generalises. Storing memory is not the same as using it, and an agent given a retrieval tool decides for itself whether to call it. Systems that inject relevant memories before the model runs, rather than offering a tool and hoping, do not have this failure mode. The design choice is on memory as a tool.
The integration blocked on writes. Saving memories added measurable overhead, around ten percent of session time in their test, because the integration waited for processing to complete. The pipeline is eventually consistent and was never meant to be waited on, so the fix was to stop waiting. This is the practical consequence of the asynchronous design described above, and it is an easy mistake to make with any queue-backed API.
A vendor publishing its own failure analysis is rare enough to be worth saying plainly, and it is also the only independent-feeling evidence available about this product, since there is no third-party technical writing about it in the search results at all. Which makes the last question a matter of judgement rather than data: who Engram is a good fit for.
The decision
Who is Engram a good fit for, and when is it the wrong choice?
It fits teams that want the write side solved and are comfortable with a managed pipeline; it is the wrong shape when memory has to live in a database you already run. The deciding question is who owns the rules about what gets stored.
The strongest case is a team already running Weaviate. Memory sits beside the vectors they operate, retrieval is hybrid by default, and there is no second system to deploy, secure and monitor. The second strongest is a team on any stack that has concluded, usually after building it once, that extraction and deduplication are more work than they are worth owning.
Design around the asynchrony. If your application needs a memory written during a turn to be searchable in that same turn, the pipeline’s eventual consistency is a constraint to plan for, with a run poll or a local buffer holding the current turn. This is not unusual for memory systems and it is better known in advance.
Look elsewhere when memories must stay inside a database you already own for compliance or data residency reasons, or when the rules for what is stored must be inspectable and version-controlled per write. Building on a store you run is more work and gives you that control, and the options are on storage backends.
Against the alternatives, the honest framing is that these products differ mainly in who owns the write path and where the data sits. Engram puts the pipeline and the store together on Weaviate. Other memory layers offer the same shape on their own infrastructure, or supply primitives and leave the pipeline to you. The neutral survey is on AI agent memory frameworks and tools, with the head-to-heads on the best AI memory tools.
FAQ
Frequently asked questions
The questions that come up when a team is deciding whether to put its memory layer on someone else’s pipeline.
Is Engram open source?
Engram is a managed service in Weaviate Cloud rather than a package you self-host, though the Python SDK that talks to it is developed in a public repository. Weaviate the database is a separate, open-source product, so the two questions are worth keeping apart.
Do you need to run Weaviate to use Engram?
No. It is offered as a managed service, so the database underneath is operated for you. Teams already running Weaviate get the additional benefit that memory lives beside the vectors they already operate, with one system to secure and monitor.
Can Engram accept facts you extracted yourself?
Yes. The API takes raw text, a full conversation, or pre-extracted facts, which makes adoption incremental: a team with a working extraction step can keep it and use Engram for reconciliation, storage and retrieval.
Is a memory searchable immediately after it is written?
Not necessarily. Processing runs asynchronously and the call returns a run identifier before the memory is committed, so poll the run when you need to know it has landed, or hold the current turn's facts locally if a search in that same turn depends on them.
How does Engram score on LOCOMO or LongMemEval?
There is no published result at the time of writing. Any comparison table that ranks it against systems with published scores is comparing a number to an absence, which is worth knowing before using one to make a decision.
What is the difference between a scope and a topic?
A scope decides who can see a memory, and a topic decides what it is about so retrieval can be narrowed. They look similar in an API, and filtering by topic where you meant to scope by user is a leak that passes every functional test.