Fundamentals · Context
What Is Context Engineering?
Context engineering is the practice of curating exactly what an LLM sees each turn: prompts, tools, retrieved docs and memories, so agents stay accurate, on-budget and on-task.
Context assembly
Definition
What is context engineering?
Finding the smallest possible set of high-signal tokens that maximizes the odds of the outcome you want, within a token budget.
Distinct from prompt engineering, which is about wording a single prompt well. Context engineering orchestrates every input an agent sees on a given turn: system instructions, few-shot examples, RAG chunks, memory recalls, tool outputs and conversation history, deciding what to include, in what order and at what priority. Anthropic’s own engineering research frames the underlying reason this discipline exists at all in one line: when an agent fails, it’s more often because the right context wasn’t passed to the model than because the model itself wasn’t capable enough.
This shift in framing reflects a real change in what building with LLMs actually involves day to day. Prompt engineering was the dominant concern when most use cases were one-shot classification or generation tasks with a single input and a single output. As agents began operating over multiple turns of inference and longer time horizons, generating more and more data with every step that could plausibly be relevant to the next decision, the discrete task of writing a good prompt stopped being the bottleneck. What replaced it is iterative: context engineering happens every single time an agent decides what to pass to the model next, not once at build time.
Relationship
How does context engineering relate to AI memory?
Memory decides what to persist; context engineering decides what to inject this turn.
Retrieval selects relevant memories from long-term storage; scoring ranks them (recency times importance times relevance, Park et al., 2023); context engineering formats and orders them under the token budget. Without engineering, retrieved memories overflow the window or bury critical facts. LangChain’s own framework splits this cleanly into 3 data sources an agent reads and writes: Runtime Context (static configuration such as user ID or permissions), State (short-term memory, scoped to the current conversation), and Store (long-term memory, persisted across conversations). Those last 2 map directly onto this site’s own short-term and long-term memory terms; context engineering is the layer that decides how much of each actually reaches the model.
Weaviate’s own engineering research puts the core memory-side challenge bluntly: the worst memory system is the one that faithfully stores everything, since old, low-quality or noisy entries eventually come back through retrieval and contaminate the context with stale assumptions. The fix on the storage side is selectivity, letting the model reflect on an event and assign an importance score before anything gets promoted to long-term storage, then pruning, merging duplicates and retiring outdated facts on an ongoing basis. Context engineering assumes that selectivity already happened upstream; its own job starts one step later, deciding how much of what survived actually belongs in front of the model this turn.
Disambiguation
How is context engineering different from prompt engineering, RAG and memory?
Each discipline controls a different slice of what the model sees; context engineering is the one that orchestrates all of them together.
| Discipline | Scope | Example |
|---|---|---|
| Prompt engineering | Wording of instructions | “You are a helpful assistant…” |
| RAG | Retrieve static documents | Inject product manual chunks |
| Memory | Persistent per-user store | Retrieve “user prefers email” |
| Context engineering | Orchestrate all inputs | Budget 2K tokens: 500 system + 1K memory + 500 history |
Techniques
What techniques actually engineer context for agent memory?
Beyond ranking and summarizing, 2 production engineering teams converged independently on treating the file system itself as external, restorable memory.
- Rank memories by relevance and recency (Park et al., 2023 scoring)
- Right altitude prompting: Anthropic’s own term for the Goldilocks zone between brittle, hardcoded if-else logic and vague guidance that assumes shared context the model doesn’t have
- Just-in-time retrieval: instead of pre-loading everything, maintain lightweight identifiers (file paths, stored queries) and load the actual data only when a tool needs it; Anthropic’s Claude Code and Manus arrived at this independently, both treating the file system as unlimited, persistent context rather than something to compress irreversibly
- KV-cache stability: keep prompt prefixes stable and context append-only, since a single changed token invalidates the cache from that point on; Manus reports Claude Sonnet’s cached input tokens cost $0.30/MTok against $3/MTok uncached, a 10x difference that makes cache hit rate one of the highest-leverage production metrics
- Mask, don’t remove: when tool availability needs to change mid-task, mask which tools the model can select via logit constraints rather than changing the tool definitions themselves, since removing or adding tools invalidates the KV-cache and can confuse the model about actions it already took
- Deterministic serialization: a KV-cache hit requires an identical prompt prefix, so a JSON serializer that doesn’t guarantee stable key ordering can silently break the cache turn after turn without ever throwing an error
- Minimal viable tool set: Anthropic’s own engineering research names bloated tool sets, too much overlapping functionality, as one of the most common failure modes; if a human engineer can’t say definitively which tool applies to a given situation, the agent can’t be expected to either
- Summarize old history to free token budget when compaction is needed
- Separate sections: system, user, memory and tools in distinct blocks
- Token budget allocation: cap each component’s share of the window
Mem0 uses ~1,800 vs ~26,000 tokens per LOCOMO query: context engineering plus selective retrieval (Chhikara et al., 2025).
Long-horizon tasks
What keeps an agent coherent across hours of work?
Three techniques for tasks that outlast the context window entirely, not just a single long prompt.
Anthropic’s own engineering research names 3 techniques for tasks spanning tens of minutes to multiple hours, like large codebase migrations or long research projects, where the token count eventually exceeds the window no matter how large it is. Compaction summarizes a conversation nearing its limit and reinitiates a new context window with that summary; Claude Code implements this by preserving architectural decisions, unresolved bugs and implementation details while discarding redundant tool outputs, then continuing with the compressed summary plus the 5 most recently accessed files. Structured note-taking, sometimes called agentic memory, has the agent regularly write notes to persistent storage outside the context window and read them back after a reset; Claude Playing Pokemon demonstrates this concretely, with the agent tracking things like “for the last 1,234 steps I’ve been training my Pokemon in Route 1, Pikachu has gained 8 levels toward the target of 10” and continuing multi-hour sequences after each context reset by reading its own notes. Manus arrived at a version of this independently: a todo.md file rewritten step by step, reciting the current goal back into the model’s recent attention span to avoid drifting off-topic across roughly 50 tool calls per task on average. Sub-agent architectures take a third approach: specialized sub-agents each explore extensively, sometimes using tens of thousands of tokens, but return only a condensed 1,000 to 2,000-token summary to the lead agent, keeping the detailed search context isolated rather than accumulating in the main thread.
Budget
How does context engineering interact with the context window itself?
Allocate tokens across components; page, summarize or retrieve when over budget.
When total context exceeds the window: summarize the oldest history, page archival memory in (the MemGPT pattern), or retrieve only the top-k memories. Letta’s DMR of 93.4% (Packer et al., 2023) demonstrates virtual context via paging. The underlying reason a window fills up faster than its token count suggests, and why bigger windows don’t fully solve this, is a separate, deeper topic this site covers on its own: see the context window problem for the attention-cost math and the real gap between a model’s stated and effective context size.
Production
What are the best practices for production agents?
Most of these are about measuring and logging what actually reaches the model, not guessing.
- Measure context quality, not just size: LOCOMO recall@k
- Log what was injected each turn (trace IDs)
- A/B memory retrieval strategies in staging
- Monitor context rot symptoms: degraded accuracy on long threads
- Keep the wrong stuff in: Manus’s own engineering research argues against scrubbing failed actions from context, since a model that sees its own error and the resulting observation implicitly updates away from repeating it; erasing the failure erases the evidence it needs to adapt
- Vary few-shot examples deliberately: a context full of near-identical past actions makes an agent mimic that pattern even when it’s no longer the right one. Manus observed this directly when reviewing a batch of 20 resumes: the agent fell into a rhythm, repeating similar actions simply because that’s what it saw in its own recent context, leading to drift and overgeneralization. Their fix was introducing small, structured variation, different serialization templates and phrasing, to break the pattern rather than reinforce it
- Gate deploys on eval regression in CI
FAQ
Frequently asked questions
The techniques above, the long-horizon patterns, and how this discipline maps onto memory terms.
Context engineering vs prompt engineering?
Prompt engineering focuses on wording instructions. Context engineering orchestrates the full input each turn: system prompt, memories, RAG chunks, tools and history, within a token budget.
Is context engineering the same as RAG?
No. RAG is one input source (static docs). Context engineering includes RAG, memory retrieval, tool outputs and conversation history in a unified assembly strategy.
Who coined context engineering?
The term gained traction in 2025-2026 as agents moved beyond single prompts to multi-source context assembly, discussed by Anthropic's own engineering research and referenced in Karpathy's discourse on LLM application design.
Is memory required for context engineering?
Not strictly. You can engineer context with RAG and tools alone. Stateful agents add memory as a persistent input source across sessions.
What is KV-cache hit rate, and why does it matter for context engineering?
It measures how often a model reuses cached computation from an identical prompt prefix instead of recomputing it. Manus reports a 10x price difference between cached and uncached tokens on Claude Sonnet, making cache stability one of the highest-leverage production levers.
What is 'right altitude' prompting?
Anthropic's own term for the balance between brittle, over-specified hardcoded logic and vague guidance that assumes shared context the model doesn't have. The goal is specific enough to guide behavior, flexible enough to generalize.
Should you remove failed actions from an agent's context?
Manus's own engineering research argues against it. A model that sees its own error and the resulting observation implicitly updates away from repeating it; scrubbing the failure removes the evidence it needs to adapt.
What is compaction in context engineering?
Summarizing a conversation nearing its context limit and reinitiating a new window with that summary. Claude Code preserves architectural decisions and unresolved bugs while discarding redundant tool outputs, then continues with the 5 most recently accessed files.
How does context engineering reduce tokens?
Selective memory retrieval, summarization, tool-output pruning, KV-cache-stable prompts and budget caps. Mem0 uses ~1,800 vs ~26,000 tokens per LOCOMO query (Chhikara et al., 2025). See reduce token cost.
Tools for debugging context assembly?
Log injected context per turn with trace IDs, use LangSmith or Langfuse tracing, and run LOCOMO eval in CI. See memory metrics and memory management.
What is 'just-in-time' context retrieval?
Maintaining lightweight identifiers, like file paths or stored queries, and loading the actual data only when a tool needs it, instead of pre-loading everything up front. Anthropic's Claude Code and Manus both converged on this pattern independently, using the file system as unlimited, restorable context.
Why do bloated tool sets hurt agent context?
Anthropic's own engineering research names this a common failure mode: too many overlapping tools make the model more likely to pick the wrong one. If a human engineer can't say definitively which tool applies to a given situation, the agent can't be expected to do better.