Memory types · Semantic

Semantic Memory in AI Agents

Semantic memory in AI agents holds general facts and learned knowledge: user preferences, product details, domain rules, independent of the specific conversation when they were learned.

Semantic

“User is vegetarian” · “API limit 100/min”

Episodic

“User asked about refund March 3”

Definition

What is semantic memory in AI agents?

Semantic memory is stable, generalized knowledge not tied to one event: preferences, policies, identity facts and domain rules the agent reuses across sessions.

Examples: “User is vegetarian”, “API rate limit is 100 requests/min”, “Refund window is 30 days.” Also called a fact store or knowledge memory. Tulving’s taxonomy: semantic = general knowledge; episodic = particular experiences. The 3-way split most language-agent frameworks use today, semantic, episodic and procedural memory, is borrowed directly from a cognitive-science research paper commonly mapped onto language-agent systems (arXiv:2309.02427), with semantic memory specifically defined as what the agent knows, as distinct from what it has experienced (episodic) or how it should behave (procedural). The boundary with procedural memory is easy to blur in practice but worth keeping distinct: where semantic memory stores “this user works in fintech,” procedural memory stores “when this user asks for a report, format it as a table with totals at the bottom.” One is a fact about the user; the other is learned behavior about how to act, built up from repeated patterns rather than something the agent was ever told once.

→ Episodic memory in AI agents

Comparison

How does semantic memory differ from episodic memory?

Semantic facts are reusable generalizations, retrieved for recall; episodic records are timestamped events, retrieved for precision.

TypeExampleUpdates when
Semantic“User prefers email support”The fact changes (“now prefers chat”)
Episodic“User called Tuesday about error 402”A new event occurs

A support bot stores semantic “user is on Pro plan” and episodic “ticket #8821 opened March 3.” The two types also call for different retrieval strategies, not just different storage: semantic memory favors high-recall, concept-level retrieval, surfacing generally true information that can ground a response, while episodic memory favors precision and strong contextual filters, identity, role, and time, since a specific event only matters if it’s the right event for the right user at the right moment.

→ Episodic memory guide

Storage

What are the storage options for semantic memory?

Vector embeddings for similarity search, knowledge graphs for relationships, or hybrid backends in production.

BackendStrengthFramework examples
Vector storeFast semantic fact lookupEngram, Mem0
Knowledge graphRelationships + temporal validityZep (Graphiti), Cognee
HybridKeyword + vector recallCustom + hybrid search

Each backend makes a different tradeoff, not just a different feature set. A vector store finds what’s related fast and scales well, but similarity isn’t the same as relationship: 2 facts can be semantically distant yet causally connected, and pure vector search has no native way to represent that. A knowledge graph stores facts as entities and relationships instead of isolated chunks, so it can trace how a decision changed over time or connect 2 facts that were never mentioned in the same sentence; adding a temporal dimension lets it also track not just that a relationship exists but when it was true, which matters once a fact from 3 months ago may no longer be current. Most production systems don’t pick one: a common pattern pairs a vector layer for fast semantic retrieval with a graph layer for relationships and temporal reasoning, sometimes backed further by a SQL layer for structured facts that don’t need embedding at all.

→ Embeddings for memory · Knowledge graphs · Vector databases

Vector, graph and hybrid storage for semantic memory, each answering a different kind of question
Vector answers what’s related; graph answers how and when; hybrid combines both.

Pipeline

How is semantic memory written and updated?

Extract fact from conversation, dedup against existing, embed and store, invalidate or supersede on contradiction.

When a user says “I moved to Berlin” after “I live in Munich,” the semantic memory pipeline updates the location fact: Engram and Mem0 handle this via merge/replace on vector records; Zep uses temporal graph invalidation.

Retrieval itself doesn’t have to be a passive, always-on injection. A cleaner pattern exposes semantic memory as an explicit tool the agent calls: the model decides when it needs context and what query to search with, while the application still controls how memory is searched and what’s allowed to come back. This separates reasoning from retrieval as 2 distinct steps rather than blending them, and it also means the same memory-access code path can point at a local development store during testing and a production backend once deployed, without changing how the agent calls it.

Not every extracted candidate fact should actually become a semantic memory update, and treating writes with the same discipline as reads keeps a fact store from silently degrading over time. Most of what a conversation produces should stay history rather than get promoted: a one-off remark, a passing preference the user immediately contradicted, or a detail that has no bearing on future interactions. What earns promotion into semantic memory is the smaller subset that will actually change how the agent behaves the next time it matters, and getting that filter wrong in either direction has a real cost: promote too much and retrieval fills up with stale, low-signal facts; promote too little and the agent keeps re-asking things the user already told it.

→ Writing and storing memories · Conflicting memory updates · Memory as a tool

RAG

How does semantic memory relate to RAG and organizational context?

RAG corpora are static and shared across users; agent semantic memory is dynamic and per-user, and neither is the same as governed organizational knowledge.

RAG answers “What does the manual say?” Semantic memory answers “What does this user prefer?” Overlap is possible; keep org docs in the RAG index and user facts in a separate memory namespace. On LOCOMO, memory-centric systems outperform fixed-chunk RAG baselines in the 50s range (Chhikara et al., 2025).

A third category is worth naming explicitly, because conflating it with either of the first two is where a lot of “just add memory” advice breaks down for enterprise use cases: organizational context. Session memory, the kind this page is mostly about, tracks what an individual user told an agent, their preferences, their past requests; it’s inherently personal and interaction-scoped. Organizational context is different in kind, not just scale: it’s the governed body of knowledge a company maintains about its own processes and policies, access-controlled, versioned, and often subject to compliance requirements unrelated to any single conversation. A support agent remembering that a customer prefers email over chat is session memory. That same agent knowing which refund policy applies to that customer’s specific contract tier is organizational context, and storing that governed fact inside a general-purpose conversation memory layer with no access-control model is a real risk, not a hypothetical one.

→ Memory vs RAG

Session memory versus organizational context: personal per-user facts versus governed, access-controlled company knowledge
A customer’s contact preference is session memory; which policy applies to their contract is organizational context.

Production

How do you scope who can retrieve which semantic facts?

Access control at the retrieval layer, not the application layer, keeps memory isolation enforceable rather than optional.

A real production pattern for this narrows the search space with structured filters, such as memory type, user identity, or time, before semantic retrieval ever runs, which cuts the number of vectors that need to be scored and the amount of context injected into the model; the practical effect is faster retrieval, smaller prompts, and a more focused answer. The stronger version of this applies the same filtering as an actual access-control rule rather than an optimization: a document-level security policy tied to the requester’s identity, so a query from one user automatically excludes another user’s memories at the database level rather than depending on application code to filter correctly every time. Decoupling access control from application logic this way means a bug in the agent’s own code can’t accidentally leak one user’s facts into another user’s context. In practice this means defining roles at the storage layer itself, each with a built-in query filter, so a request tagged with one identity is automatically restricted to memories tagged with that same identity before the search runs, not after the results come back. The agent code never has to remember to apply the filter correctly, because it isn’t given the option to skip it.

→ Memory security

Document-level security scoping semantic memory retrieval by user identity at the database level
Filtering by identity at the retrieval layer means an application bug can’t leak one user’s facts into another’s context.
Semantic memory retrieval strategy compared to episodic memory retrieval strategy: high recall versus high precision
Semantic memory is searched for recall; episodic memory is searched for precision.

FAQ

Frequently asked questions

The storage tradeoffs, the retrieval differences, and where organizational context fits in.

What is an example of semantic memory in AI?

"User is vegetarian", "User prefers email support", "API rate limit is 100/min": stable facts not tied to one conversation event. See definition above.

Semantic vs episodic memory, what's the difference?

Semantic memory is general facts ("user is on Pro plan"), retrieved for high recall. Episodic memory is specific events ("user called Tuesday about error 402"), retrieved for precision. See episodic memory.

Is semantic memory the same as a RAG knowledge base?

No. RAG indexes static org documents shared across users. Semantic agent memory is dynamic, per-user and updated each session. See memory vs RAG.

Is a vector database enough for semantic memory?

Yes for most personalization use cases: embed facts, search by similarity. Graphs add relationship and temporal queries when facts link to entities that change over time. See vector databases.

When do you need a graph for semantic memory?

When facts have relationships, validity windows or conflict resolution: CRM timelines, policy versioning. Zep Graphiti is built for this. See knowledge graphs.

Does Engram store semantic memory?

Yes. Engram extracts and stores semantic facts per user on Weaviate-native stacks, unified vector and memory on one platform. See Engram explained.

Does Mem0 store semantic memory?

Yes. Mem0 specializes in per-user semantic and episodic memory via a vector extraction API. LOCOMO J 66.9 (Chhikara et al., 2025). See Mem0 alternatives.

What role does consolidation play in semantic memory?

Consolidation promotes durable facts from episodic events into stable semantic records, for example summarizing 3 preference mentions into 1 fact. See memory consolidation.

How do you update semantic facts when users correct them?

Dedup on write, then merge or invalidate the old fact. Zep uses temporal graph supersession; vector stores need explicit conflict handling. See conflicting memories.

Semantic memory for customer support bots?

Store stable preferences (language, plan tier) as semantic memory; ticket events as episodic memory. Combine both in retrieval. See customer support use case.

Is organizational context the same as semantic memory?

No. Session memory (a user's own preferences and history) is personal and interaction-scoped. Organizational context (company policies, access-controlled and versioned) is governed and shared. Storing governed facts inside a general-purpose memory layer without access control is a real risk.

Should semantic memory always be retrieved automatically?

Not necessarily. Exposing it as a tool the agent explicitly calls, deciding when and what to search for, separates reasoning from retrieval and keeps the same code path usable for both a local dev store and a production backend.

Should every extracted fact become a semantic memory?

No. Most of what a conversation produces should stay history. Promoting too much fills retrieval with stale, low-signal facts; promoting too little means the agent keeps re-asking what the user already said.

How is semantic memory different from procedural memory?

Semantic memory stores facts, such as "this user works in fintech." Procedural memory stores learned behavior, such as "format reports for this user as a table." One is a fact; the other is a rule built from repeated patterns.

Continue exploring

Three routes onward: episodic memory, long-term persistence, and the vector-vs-graph storage decision.