Tools · Graph Memory
Cognee: Graph-Vector Memory for AI Agents
Cognee stores agent memory across a graph, a vector index and a relational store at once, and its default retrieval mode uses vector search only to find an entry point into the graph, then traverses relationships to build context, rather than ranking chunks by similarity alone. That one architectural choice, graph traversal as the default rather than an optional extra, is what actually separates it from a conventional RAG-backed memory layer.
Three stores
Architecture
What is Cognee, and how is its memory actually stored?
An open-source memory platform built as a graph-vector hybrid: three storage layers unified behind one engine, rather than a single vector index with metadata attached. That combination is the whole architectural bet the product makes.
The graph store holds entities and relationships and supports structural traversal between them, defaulting to Kuzu with support for Neo4j, FalkorDB, Amazon Neptune and Memgraph. The vector store holds embeddings for semantic similarity, defaulting to LanceDB with support for Qdrant, pgvector, Redis, DuckDB, Pinecone and ChromaDB. The relational store holds documents, chunks and provenance tracking, defaulting to SQLite with support for PostgreSQL. All three defaults are file-based, which means a first installation has no separate service to stand up before anything runs.
The fundamental stored unit is a typed object carrying content and metadata, which a developer can subclass to control exactly which fields get embedded versus which are kept as structural attributes only. That typing is what makes the graph side more than a generic key-value store: relationships between typed objects are what the traversal-based retrieval described further down actually walks.
Three stores working from one write is a design choice with consequences on both sides of the memory lifecycle. On the write side it means an incoming fact can be indexed for similarity and linked into a graph of relationships in the same operation. On the read side it opens retrieval strategies a pure vector store cannot offer at all: how does Cognee separate session memory from permanent memory?
Two tiers
How does Cognee separate session memory from permanent memory?
Session memory is runtime working context, loaded fresh for the current exchange; permanent memory is the durable graph of facts and relationships that persists and keeps growing. Both tiers are backed by the same underlying stores, which is what lets them interact rather than sitting in isolation.
Session memory loads relevant embeddings and graph fragments into runtime context for the duration of an exchange, which is what lets a follow-up question resolve a pronoun against something said earlier in the same session without the user restating it. Permanent memory stores the accumulating long-term artifacts: user data, interaction traces, external documents and the relationships derived from all of them, continuously cross-connected inside the graph while remaining linked to their vector representations.
The practical effect of keeping both tiers on the same underlying graph, rather than as two unconnected systems, is that a fact learned in one session can immediately become part of the structural context available to a completely different session, without a separate synchronisation step. This is the same short-term-versus-long-term distinction covered generally on short-term versus long-term memory, implemented here with both tiers sharing one graph rather than living in separate stores.
Two tiers of storage explain what is kept. What actually happens when a query comes in is a separate and more distinctive design decision: how does Cognee retrieve memories, and how is that different from plain RAG?
Retrieval
How does Cognee retrieve memories, and how is that different from plain RAG?
Fourteen documented search modes are available, and the default uses vector similarity only to locate an entry point into the graph, then traverses relationships from there rather than ranking chunks by similarity alone. That default is the single architectural choice that most separates this product from a conventional vector-backed memory layer.
In a conventional retrieval setup, the top-k passages by cosine similarity are handed to the model as context, and similarity is doing the entire job of deciding what the model sees. In Cognee’s default mode, similarity search finds the handful of graph triplets most relevant to the query, and the system then traverses outward from those triplets through the graph’s explicit relationships, assembling structured context that includes connected facts a similarity search alone would never surface, because they may not be textually similar to the query at all, only structurally connected to something that is.
The other thirteen modes cover cases the default does not fit: chain-of-thought reasoning across multi-hop graph traversals, direct Cypher queries against the graph for developers who want to write the traversal themselves, plain chunk retrieval by similarity for cases where graph structure adds nothing, lexical token-based search for exact-string cases similarity handles poorly, and a mode that lets the system choose automatically. Having this many named, distinct modes is itself informative: it signals that no single retrieval strategy is correct for every query shape, which is the same lesson this site draws generally on hybrid search and on vector versus knowledge graph memory.
A graph gives retrieval more to work with. It also raises a question a pure vector store handles more simply: what happens when many users or agents share the same underlying graph engine? how does Cognee isolate one user’s or agent’s memory from another’s?
Isolation
How does Cognee isolate one user’s or agent’s memory from another’s?
At the graph and trace level, with dataset-scoped permissions, rather than only as a metadata filter applied to a shared vector index. That is a stronger and more specific isolation claim than most competing memory products state, and it follows directly from having a graph store as a first-class citizen rather than an add-on.
Memory graphs can be instantiated per user, per group, or as shared public graphs, and dataset-level permissions cover read, write, delete and share separately, so a caller’s access can be scoped more precisely than a single boolean toggle. This isolation model works across the different graph and vector backends the product supports, including pgvector, Neo4j, Kuzu and LanceDB, rather than being tied to one specific database’s filtering capability.
The distinction from a filter-based approach matters in exactly the failure mode covered on memory security and privacy: a similarity search scoped only by a metadata predicate can still surface a record if the predicate is ever omitted or misconfigured, because the underlying index has no concept of ownership. Isolation enforced at the graph and trace level is a structural property of where the data lives rather than a condition attached to each query, which is a meaningfully different failure mode to defend against.
Architecture and isolation describe the design. Whether the design actually performs better is a separate question, and it is the one the vendor has published its own numbers against: what does Cognee’s own benchmark actually show?
Published results
What does Cognee’s own benchmark actually show?
A self-reported correctness score of 0.93 against a base RAG score of 0.40, on 24 HotPotQA multi-hop questions run 45 times, with the evaluation code published openly. The scale and the source both matter for reading this number correctly.
This is Cognee’s own published comparison against Mem0, Graphiti and LightRAG, run on Modal Cloud. HotPotQA is a multi-hop question-answering dataset, meaning many of its questions require connecting facts across more than one source to answer correctly, which plays to a graph-traversal system’s strength deliberately: the questions are the kind graph relationships are built to answer, not a neutral sample of general queries. The vendor reports the biggest gains coming specifically from chain-of-thought graph traversal on multi-hop questions, which is consistent with the architecture rather than a surprising result.
Read this the way any vendor-published benchmark should be read, following the same checklist this site applies to disputed public benchmarks on the LongMemEval benchmark: it is a self-run comparison, on a specific and modest sample size, on a dataset shape favourable to the architecture being tested, with the code available for independent verification. That the code is open is a genuine point in its favour, since a reader can check the methodology rather than trust the headline number alone; it does not make the 24-question sample a large one.
A benchmark number describes accuracy on one dataset. Whether the underlying difference actually matters day to day is easier to see in a worked example than in a score: what does the difference look like on a real scenario?
A concrete case
What does the difference look like on a real scenario?
A support agent built on Cognee resolves a returning customer’s billing issue using context from a previous, related ticket; the same agent built on stateless retrieval treats every ticket as a fresh conversation with no memory of the prior one. This comparison, built and documented by the vendor as a tutorial, is the clearest illustration of the architecture’s payoff in the capture used for this page.
The setup preprocesses a support-ticket dataset and builds two versions of the same agent side by side: a stateless RAG agent that retrieves from ticket text by similarity alone, and a Cognee-powered agent using the same underlying data. On a comparative billing-issue scenario, the stateless agent answers each new ticket in isolation, unaware that the same customer raised a related complaint previously. The Cognee-powered agent, having connected the prior ticket into its graph, recognises the relationship and responds with the accumulated context already in place. The tutorial also builds a feedback loop on top of this, capturing whether a given interaction resolved the issue and feeding that signal back into what the system treats as reliable going forward.
The scenario is a vendor-authored tutorial rather than an independent case study, so it should be read as a demonstration of the intended behaviour rather than a controlled comparison with a reported statistic. What it demonstrates clearly, regardless of framing, is the concrete shape of the difference a graph connection makes: not a marginally better-ranked chunk, but a qualitatively different answer that required connecting two separate records nobody explicitly linked at write time.
Architecture, isolation, a published benchmark and a worked example all point the same direction. The remaining question is whether the added complexity of three coordinated stores is worth it for a given team: is Cognee’s graph-vector approach worth the added complexity?
The decision
Is Cognee’s graph-vector approach worth the added complexity?
When the memory genuinely needs multi-hop reasoning across connected facts, most of what a graph store buys is real; when it does not, the added moving parts are a cost with no corresponding benefit. The honest answer depends entirely on the shape of the questions a system actually has to answer.
A support or personalisation agent whose queries are mostly “what did this user say” benefits less from graph traversal, since most of the value in that case is retrieval and extraction quality rather than relationship structure, and the case for a simpler vector-only stack, covered on vector databases for agent memory, is stronger there. An agent whose queries genuinely require connecting facts across sources, who reported an issue and how it relates to a policy and who else touched the same account, benefits from exactly the traversal Cognee’s default mode is built around, and the effort of maintaining three coordinated stores is effort spent on a problem that actually exists.
For teams that would rather not build or operate a graph-vector pipeline themselves, a managed pipeline over a hybrid vector store is a real alternative worth weighing directly: Engram runs extraction and reconciliation over a hybrid store without asking a team to run a separate graph database, at the cost of not offering multi-hop graph traversal as a first-class retrieval mode. Which of the two answers you actually need is a question about your queries, not about which architecture is more sophisticated on paper, and the broader landscape of options is on memory tools.
FAQ
Frequently asked questions
The architecture and fit questions that follow once the basic mechanism is understood.
Does Cognee require running a separate graph database?
Not to get started. The default configuration uses file-based Kuzu, LanceDB and SQLite, so a first installation needs no separate infrastructure. Production deployments can swap in Neo4j, pgvector, or another supported backend per store as scale requires.
Is GRAPH_COMPLETION always the best search mode to use?
No. It is the default and suits queries that benefit from relational context, but plain chunk retrieval or lexical search fit better when a query has no relational structure to exploit or needs an exact string match. Choosing among the fourteen modes is a query-shape decision, not a one-size default.
How does Cognee's HotPotQA benchmark compare to LongMemEval or LoCoMo scores?
It is not directly comparable. HotPotQA tests multi-hop question answering over a fixed corpus, while LongMemEval and LoCoMo test conversational long-term memory specifically. A strong score on one says nothing about performance on the others; see the LongMemEval benchmark for what that one measures.
Can multiple agents share the same Cognee graph safely?
Yes, with scoping applied deliberately. Memory graphs can be instantiated per user, per group, or as shared public graphs, with read, write, delete and share permissions set per dataset, rather than relying on every caller correctly applying a filter.
Does adding a graph store fix retrieval quality on its own?
No. A graph only helps when relationships in the data are actually useful to the queries being asked. For memory dominated by personal facts and preferences rather than connected entities, most of the retrieval quality still comes from good extraction, covered on writing memories, not from graph structure.
What happens to isolation if a team uses a vector backend that does not natively support multi-tenancy?
Cognee's documented multi-tenancy support covers pgvector, Neo4j, Kuzu and LanceDB specifically. Using a different backend without verifying its isolation support against your requirements risks falling back to filter-based scoping alone, which is the weaker isolation model this architecture is built to avoid.
Continue exploring
Vector or knowledge graph, generally?
The tradeoff this product’s architecture is built around.
Compare →What does a managed pipeline look like?
Extraction and reconciliation without running a graph database.
Read →How should you read a vendor benchmark?
The same checklist applied to any published score.
Read →