Compare · Letta alternatives

Top Letta (MemGPT) Alternatives for AI Agent Memory

The best Letta and MemGPT alternatives for AI agent memory are Engram (Weaviate), Zep, LangMem, Cognee and Hindsight: compared on architecture, memory ops, paging vs vector vs graph trade-offs and fit.

MemGPT tiers

1
Core
In-context
2
Page out
To archival
3
Archival
Cold store
4
Recall
On demand
5
Persist
Cross-session

Baseline

What is Letta, and how does it relate to MemGPT?

Letta is the production platform for stateful LLM agents with virtual context: the commercial stack from the MemGPT research project.

MemGPT is the research paper name, “MemGPT: Towards LLMs as Operating Systems” (Packer et al., 2023, arXiv:2310.08560). Letta is what most developers deploy today. Same lineage: the context window acts like RAM; archival memory acts like disk; the agent pages between tiers via memory tools.

The deeper distinction is philosophical, not just architectural: Mem0 and most other tools on this page extract memories passively, the developer’s input decides what goes in, and the extraction pipeline decides what facts get stored, which is predictable and token-efficient but can’t make nuanced in-context judgments. Letta agents self-edit their own memory instead, deciding during their own reasoning loop what’s worth writing to core, recall or archival memory. That’s more adaptive, but memory quality depends entirely on the model’s own judgment, and every memory operation costs inference tokens, since the agent has to reason about what to store before it stores it.

→ Virtual context and MemGPT

Predictability versus intelligence: passive memory extraction versus agent self-edited memory
Letta’s self-editing memory is more adaptive; passive extraction is more predictable and token-efficient.

Mechanism

How does Letta’s virtual context memory actually work?

Five steps, from a hot in-context tier to a persistent cold store the agent controls itself.

1

Tiers

Core (hot) vs archival (cold)

2

Paging

Agent moves segments in and out

3

Recall

Retrieval pulls archived text

4

Persist

Archival survives sessions

5

Agent control

Memory ops exposed as tools

Contrast: Engram is an extraction API to a Weaviate vector store; Zep is a temporal knowledge graph; raw long context is bigger RAM but linear cost and attention dilution. The Agent Development Environment (ADE), a visual interface for inspecting memory state across all 3 tiers and debugging tool calls in real time, is genuinely useful for understanding why an agent made a particular memory decision. → Memory hierarchy

Debate

Do you still need paging with a million-token context window?

Million-token windows exist, but cost scales linearly, attention dilution persists and latency rises; paging is virtual memory for unbounded history at sustainable cost.

Gemini models exceed 1M tokens; Claude 3.7 Sonnet supports 200K. Long context is bigger RAM; Letta paging is virtual memory, still needed when conversation history is unbounded and you cannot afford to resend every token each turn. Memory retrieval typically uses 90% fewer tokens than full-context replay. See the context window problem for the underlying attention-cost math this page does not re-derive.

→ Memory vs context window

When to switch

Why look for Letta alternatives at all?

Letta’s real lock-in cost and pricing, not just architecture fit.

Because agents run inside the Letta runtime itself, switching away means rewriting more than the memory layer: the agent loop, tool execution and state management all live inside Letta too, so unwinding it is a bigger commitment than swapping out an API call, unlike a pluggable memory layer with a clean API boundary. That isn’t a flaw in Letta’s design, tight integration is what makes the self-editing model work, but it’s a real cost worth naming before adopting it. On pricing, Letta’s managed cloud runs $20 to $200 a month depending on usage, including the ADE; self-hosting is free with no feature gating, the same open terms as most alternatives on this page.

  • Weaviate-native vector and memory: Engram (no paging layer to build)
  • Managed personalization API without agent-framework coupling: Engram
  • Temporal graph and fact versioning: Zep
  • LangGraph-only stack: LangMem
  • Document KG ingest: Cognee
  • Multi-strategy retrieval without adopting a full runtime: Hindsight
  • Simple per-user prefs only, not full conversation paging
  • Team lacks appetite for OS-style memory abstractions

Paging is not always required; match architecture to your bottleneck.

Letta's real lock-in cost: switching away means rewriting the agent loop, tool execution and state management, not just the memory layer
Adopting Letta means adopting its runtime; unwinding it is proportionally harder than swapping a memory API.

Side by side

How do the Letta alternatives compare, side by side?

Last updated: September 2026. Architecture, ops and fit at a glance.

ToolArchitectureMemory opsPaging?Best for
Engram (Weaviate)Vector-native layerAsync extract, transform, commitNoWeaviate-native stacks
ZepTemporal KG (Graphiti)Full, temporal invalidationNoChanging facts, CRM timelines
Hindsight4-retriever hybridExtract, reflect, rerankNoRetrieval depth without a runtime
LangMemLangGraph storeExtract, store, retrieveNoLangGraph teams
CogneeEvolving KGGraph write, queryNoDocument-heavy KB agents

Profiles

What does each alternative actually offer?

Real weaknesses against Letta specifically, not a generic feature list.

Engram (Weaviate)

Vector-native dynamic memory layer

Strengths: unified Weaviate stack, per-interaction updates, no paging layer to engineer.
Weaknesses vs Letta: no MemGPT-style agent-controlled tier paging, less suited to unbounded raw chat history.
Best for: Weaviate-native apps needing memory without building virtual context.

Engram explained →

Zep

Temporal knowledge graph (Graphiti)

Strengths: time-changing facts, relationship queries, conflict resolution.
Weaknesses vs Letta: fact graphs vs conversation paging are different problems.
Best for: evolving entity timelines, not raw chat volume.

Zep alternatives →

Hindsight

4-retriever hybrid memory engine

Strengths: a pluggable memory layer rather than a full agent runtime, so it doesn’t carry Letta’s lock-in cost; merges semantic search, BM25 keyword matching, graph traversal and temporal reasoning with a cross-encoder rerank; MIT license, self-hosts as a single Docker container on embedded PostgreSQL.
Weaknesses vs Letta: no agent-controlled self-editing model; a newer project with a smaller community.
Note: its own vendor positions it as addressing retrieval-depth limitations in both Mem0 and Letta, reporting 94.6% on LongMemEval; treat this as that vendor’s own comparison, not an independently reproduced result.
Best for: teams that want Letta’s retrieval depth without adopting its runtime.

LangMem

LangGraph memory store

Strengths: native LangGraph checkpointer, simpler than full Letta for LangChain teams.
Weaknesses vs Letta: no virtual-context OS model, ecosystem lock-in.
Best for: LangGraph persistence without Letta adoption.

LangMem explained →

Cognee

Evolving document KG

Strengths: structured knowledge from corpora.
Weaknesses vs Letta: not a conversation-paging architecture.
Best for: knowledge-heavy agents.

Cognee explained →
Hindsight's reported LongMemEval score of 94.6 percent, positioned by its own vendor as addressing retrieval-depth limitations in both Mem0 and Letta
This is the reporting vendor’s own comparison, not an independently reproduced benchmark.

Scope

Are all “Letta alternatives” really the same kind of thing?

No. Some tools replace your agent; others only extend one you’ve already built.

Every tool profiled on this page, Engram, Zep, Hindsight, LangMem and Cognee, is infrastructure for building an agent: a memory layer or runtime you integrate into your own codebase. That’s a genuinely different category from a plugin that adds memory to an agent tool you already use day to day without writing any code, installing in minutes rather than requiring a server and a database. If what you actually want is the second kind of thing, none of the alternatives profiled here are the right comparison; the question “which memory infrastructure should I build with” and “how do I make my existing coding assistant remember my project” have different answers, and it’s worth knowing which one you’re actually asking before evaluating any of the tools above.

Disambiguation

Letta vs MemGPT: are they the same?

MemGPT is the idea; Letta is the product most developers use to implement it.

The MemGPT paper introduced virtual context paging for LLMs as operating systems. Letta continues that lineage with an agent server, API and tooling for stateful agents in production. When searchers ask for “MemGPT alternatives,” they usually mean Letta alternatives.

→ Virtual context and MemGPT architecture

Head to head

Letta vs Engram: which is actually better for your stack?

Engram wins for simpler API and faster personalization POC on Weaviate; Letta wins when the agent must page unbounded conversation history.

  • Weaviate-native app, per-user facts: Engram
  • 100K+ token multi-session chat: Letta (MemGPT paging)
  • Agent-controlled memory tiers: Letta
  • Framework-agnostic managed memory: Engram

→ Engram explained

Decision

Which one should you actually pick?

Match the scenario, not the brand recognition.

ScenarioPickWhy
Unbounded multi-session chatLettaMemGPT paging built in
Weaviate AI-native appEngramUnified vector and memory
Per-user prefs onlyEngramSimpler than paging
CRM facts that change over timeZepTemporal graph and invalidation
Retrieval depth without runtime lock-inHindsight4-retriever hybrid, pluggable
LangGraph agentLangMemNative checkpointer
Decision matrix mapping the scenario to the right Letta alternative, from Engram to Hindsight to Zep to LangMem
Match the scenario, not the brand recognition.

Stay

When should you stick with Letta?

Five signals its runtime commitment is genuinely earning its lock-in cost.

  • Virtual context and paging is core architecture
  • The agent must control its own memory tiers
  • Very long conversational history is the bottleneck
  • The MemGPT research model fits your product
  • Building stateful agents, not just making memory API calls

→ Best AI memory tools ranking

Migrate

How do you migrate away from Letta?

Plan for rewriting more than the memory layer, per the lock-in cost above.

  1. Export archival and core memory segments
  2. Map tiers to the target schema (vectors for Engram or Mem0, graph for Zep)
  3. Re-summarize or re-embed archival chunks
  4. Run a retrieval POC on long-session recall
  5. Dual-write cutover, noting the loss of agent-controlled paging unless reimplemented

→ Build long-term memory

FAQ

Frequently asked questions

Letta’s real lock-in cost, how the alternatives compare, and how to migrate.

What is the best alternative to Letta?

Depends on need: Engram (Weaviate) for vector-native stacks, Mem0 for a managed personalization API, Zep for temporal facts, LangMem for LangGraph, Hindsight for retrieval depth without runtime lock-in. Letta itself is best when agent-controlled paging for unbounded chat is required. See comparison table above.

Is Letta the same as MemGPT?

Same lineage. MemGPT is the research paper (Packer et al., 2023); Letta is the production platform most developers deploy. See virtual context and MemGPT.

Mem0 vs Letta, which should I choose?

Letta when unbounded multi-session conversation via paging is the bottleneck. Engram or Mem0 when you need simpler per-user fact memory without agent-framework coupling. Mem0 LOCOMO J 66.9 (Chhikara et al., 2025). See Mem0 alternatives.

Letta vs Zep, what's the difference?

Letta pages conversation history across memory tiers. Zep stores temporal knowledge-graph facts with validity windows. Different problems: paging vs evolving entity relationships. See Zep alternatives.

Can Engram replace Letta paging?

Partially. Engram handles dynamic per-user memory on Weaviate but lacks MemGPT-style agent-controlled tier paging. Choose Letta when unbounded raw chat history is the core constraint. See Engram explained and long context vs memory.

Letta vs a long context window, do you still need paging?

Often yes for unbounded history at sustainable cost: long context is bigger RAM, Letta paging is virtual memory. Mem0 uses ~1,800 vs ~26,000 tokens per LOCOMO query (Chhikara et al., 2025). See memory vs context window.

What is virtual context memory?

MemGPT's model: the context window acts as RAM, the archival store acts as disk, and the agent pages segments in and out via tools. Letta implements this in production. See virtual context architecture.

Is Letta open source?

Letta is Apache 2.0 licensed and self-hosts free with no feature gating; managed cloud runs $20 to $200/month depending on usage, including the ADE. See open-source vs managed.

Is Letta good for coding agents?

Letta fits deep multi-session coding conversations via paging. Engram or LangMem may suffice when you need semantic codebase facts without full paging. See coding agents use case.

Is the MemGPT paper still relevant?

Yes. It established virtual context paging for LLMs (Packer et al., 2023, DMR 93.4%). Letta continues the lineage. Zep reports 94.8% on the same DMR task (Rasmussen et al., 2025).

Context memory vs Letta memory?

Context memory often means working memory in the prompt (short-term). Letta spans working (core) plus archival (long-term) memory with agent-controlled paging. See short-term memory.

How do I migrate off Letta?

Export archival and core segments, map tiers to the target schema, re-embed, run a LOCOMO/LongMemEval POC, dual-write cutover. You lose agent-controlled paging unless reimplemented, and the migration is bigger than a typical memory-layer swap since Letta owns the whole agent runtime. See migration section above.

What is the real cost of switching away from Letta?

More than swapping a memory API. Because agents run inside the Letta runtime, migrating means rewriting the agent loop, tool execution and state management too, not just the memory calls.

What is Hindsight, and how does it compare to Letta?

Hindsight is a pluggable memory layer (not a full agent runtime) merging 4 retrieval strategies with an MIT license and single-Docker self-hosting. Its own vendor positions it as addressing retrieval-depth limits in both Mem0 and Letta, reporting 94.6% on LongMemEval; treat that as the vendor's own comparison.

Is every Letta alternative the same kind of tool?

No. Engram, Zep, Hindsight, LangMem and Cognee are all infrastructure for building an agent. Tools that add memory to an existing agent product you already use, without any engineering, are a different category this page does not cover.

Continue exploring

Three routes onward: virtual context and MemGPT, the full benchmarked ranking, and Engram explained.