Fundamentals · Limits

The Context Window Problem: Why Bigger Isn’t Enough

Longer context windows still hit hard limits: cost scales quadratically with every token, attention degrades over distance (context rot), and nothing persists across sessions unless you add AI memory.

Bigger window

More tokens, higher cost, still session-bound

+ Memory layer

Selective recall, cross-session, lower cost

Definition

What is the context window, and what’s its hard limit?

Maximum tokens per inference: system prompt plus history plus tools plus user message, volatile across sessions.

Before text ever reaches a context window, it’s tokenized, usually via byte-pair encoding, which breaks text into subword units; a rough rule of thumb is that one token represents about 4 characters or three-quarters of a word, though this varies by language and tokenizer. Context windows have expanded roughly 2 orders of magnitude since the original transformer architecture, from a few thousand tokens in early models to the millions available today, according to Redis’s own engineering research.

As of July 2026, the current frontier sits at 1 million tokens across most major vendors (Claude Fable 5, Claude Opus 5, Claude Sonnet 5, OpenAI’s GPT-5.6 family, Google’s Gemini 3.6 and 3.5 Flash, Moonshot’s Kimi K3, DeepSeek V4, and Meta’s Llama 4 Maverick), with Meta’s Llama 4 Scout reaching 10 million tokens on a single GPU and Magic.dev’s LTM-2-Mini publishing 100 million, though Codingscape’s own July 2026 roundup notes plainly that “we still haven’t seen evidence anyone outside of Magic.dev is using this model or its 100 million token context.” Everything in the window exists only for that inference call. A new session starts blank unless you resend the entire history.

→ Memory vs context window

Problems

What are the three problems with relying on context alone?

Cost, accuracy and persistence each break down before a large context window ever fills up.

ProblemWhat happensEvidence
CostAttention cost grows quadratically with every token10K tokens = 100 million attention comparisons; 100K tokens = 10 billion (Redis engineering research)
Context rotAccuracy drops well before the window fillsModels fall short of their own stated maximum by up to 99% (Paulsen, arXiv:2509.21361)
No persistenceNew session = blank slateRequires external memory for LTM

→ Context rot · Long context vs memory

The math

Why does cost scale quadratically with context?

Every token has to be compared against every other token, so doubling the context quadruples the work.

Transformer self-attention is O(n²): a 10,000-token context needs roughly 100 million pairwise comparisons, and a 100,000-token context needs roughly 10 billion, according to Redis’s own engineering research. That quadratic growth, not a vendor pricing decision, is why inference visibly slows down as a conversation lengthens, and why a chatbot that felt fast at the start of a session gets sluggish by the end: the key-value cache that holds every prior token’s representation keeps growing, and eventually GPU memory bandwidth, not raw compute, becomes the bottleneck. Two mitigations are in production use today: FlashAttention, which achieves a 2 to 4x speedup by changing how GPUs access memory rather than reducing the math itself, and sparse attention, which can prune 90 to 99% of attention links by replacing full token-to-token comparison with sliding windows, global tokens and random connections. Neither eliminates the underlying quadratic cost; both push the point where it becomes painful further out.

KV cache quantization is a third lever, and a blunter one: switching the cache from 16-bit to 8-bit or 4-bit precision cuts its memory footprint by roughly 50%, which matters because moving data between a GPU’s fast on-chip memory and its slower main memory, not the attention math itself, typically consumes 70 to 90% of total inference time on long contexts. This is also why large context windows don’t scale for free even on capable hardware: a roughly 14-billion-parameter model running a long context can exhaust the VRAM on a 12GB GPU even with aggressive quantization applied, which is why practical context length in production is often far below whatever the model’s theoretical maximum happens to be.

Quadratic attention cost: 10000 tokens equals 100 million comparisons, 100000 tokens equals 10 billion comparisons
Doubling the context window doesn’t double the cost; it roughly quadruples it.

Context rot

What is context rot, and how bad is it?

Needle-in-haystack failures and attention dilution: a million-token window doesn’t guarantee million-token understanding.

A 2025-2026 study (Paulsen, arXiv:2509.21361) tested this directly rather than trusting vendor specs, defining a “Maximum Effective Context Window” (MECW) distinct from the “Maximum Context Window” (MCW) vendors publish, and measuring the gap across hundreds of thousands of data points on several models. The result: every model tested fell short of its own stated MCW, by as much as 99% in the worst cases; a few top-of-the-line models failed with as little as 100 tokens in context on certain problem types, and most showed severe accuracy degradation by 1,000 tokens. Redis’s own engineering research names a specific, more commonly-cited threshold in the same direction: most long-context models show sharp performance drops past 32,000 tokens, the well-documented “lost-in-the-middle” effect, and even top performers lose significant accuracy at longer contexts than that. Both findings say the same thing from different angles: the number on a model’s spec sheet is not the number that matters for your actual task.

Part of why this gap exists is architectural rather than purely a training issue. Rotary Position Embedding, the mechanism most modern models use to track a token’s position in the sequence, works by applying rotations to a token’s representation based on where it sits, but Redis’s own engineering research notes it “still has severe limitations preventing them from dealing with a context of millions of tokens” without additional extension techniques layered on top. Models also simply fail on positions they never saw during training: a model trained mostly on documents under 50,000 tokens has no guarantee of behaving sensibly at token 800,000, regardless of what its architecture technically permits.

→ Context rot advanced guide

Maximum Context Window versus Maximum Effective Context Window, showing models falling short of their stated maximum by up to 99 percent
The number on the spec sheet is not the number that matters for your actual task.

Trade-offs

How do large and small context windows actually compare?

A larger window trades accuracy, latency and headroom for reach; a smaller one trades reach for consistency.

DimensionLarge windows (128K to 2M+)Small windows (under 32K)
AccuracyDegrades around 32K tokens for most long-context models; the lost-in-the-middle effect is well documentedConsistent attention; reliable on focused tasks
CostPriced from fractions of a dollar to several dollars per million input tokens; prompt caching can cut costs on cached tokens by up to 75%Lower per-token cost; minimal overhead
LatencyCan drop to single-digit tokens per second once a context no longer fits fully in GPU memory50 to 100 tokens per second in-memory; faster inference
MemoryA roughly 14B-parameter model at long context can exhaust a 12GB GPU’s VRAM even with aggressive quantizationMinimal overhead; edge deployment capable

Redis’s own engineering research is blunt about the practical upshot: don’t trust the spec sheet, benchmark your actual use case at your target context length, since even top performers lose meaningful accuracy well before their advertised ceiling.

Solution

Why does memory solve what context cannot?

Selective retrieval, a cross-session store, update and forget, and lower per-turn tokens, benchmarked on LOCOMO and LongMemEval.

  • Retrieve only top-k relevant facts each turn
  • Persist across sessions in an external store
  • Update and forget facts without retraining
  • Mem0 LOCOMO J 66.9; Zep LongMemEval +18.5% vs baseline

Redis’s own engineering research frames the production answer as 3 complementary techniques rather than a single fix: semantic caching to avoid re-running inference on queries that mean the same thing, retrieval-augmented generation to pull only relevant document sections instead of stuffing an entire corpus into context, and agent memory systems for the cross-session, per-user facts a context window structurally cannot hold once the session ends. Most production systems combine all 3 rather than picking one.

→ Why AI agents need memory · How AI memory works

When context wins

When is a large context window still the right tool?

Single-session depth beats retrieval when the whole document, codebase or trace needs to be reasoned over at once.

Codingscape’s own July 2026 research names several cases where the jump to a 1-million-token window has real practical impact rather than diminishing returns: multi-repo codebase analysis, where loading an entire repository plus its documentation and test files avoids losing cross-file dependencies to chunking; long-horizon agentic workflows, where holding every tool call, observation and reasoning step in a single trace eliminates the compaction that used to cause agents to lose the plot mid-task; comprehensive document analysis, where a legal, financial or research corpus stays fully in context instead of being split across a sliding window; multimodal processing, where a single request can reason across text, images, video and audio together rather than through separate pipelines; and enterprise knowledge management, where an entire internal document set, policy library or case archive loads directly rather than depending on a retrieval pipeline to guess which fragments matter. None of these need cross-session memory, since the task starts and ends inside one session; that’s precisely the boundary where a large context window is the right tool and memory would be solving a problem that doesn’t exist yet.

When a large window is not the right fit for the budget, 3 concrete adjustments help before reaching for a bigger model: sizing the context window to the actual input rather than always requesting the maximum available, applying sparse or pruned attention to limit computation to the most relevant tokens, and batching shorter inputs together when the full window isn’t needed for every request. Don’t abandon long context. Use it for the current session’s working set while memory handles persistence and selective recall across sessions.

→ Long context vs memory

Current LLM context window tiers as of July 2026, from 256K token open-weight models to 100 million token research models
Most production traffic runs on the 1-million-token tier; the 10M and 100M tiers remain largely unused outside their own vendors.
Three complementary production techniques for context limits: semantic caching, retrieval-augmented generation, and agent memory systems
Most production systems combine all 3 techniques rather than picking one.

FAQ

Frequently asked questions

The real limits, the current model tiers, and when a large context window is still the right call.

What are max context tokens in 2026?

1 million tokens is the current frontier tier (Claude Fable 5, Claude Opus 5, Claude Sonnet 5, GPT-5.6 family, Gemini 3.6/3.5 Flash, Kimi K3, DeepSeek V4, Llama 4 Maverick), with Llama 4 Scout reaching 10 million on a single GPU. Caps apply per inference call; nothing persists across sessions without external memory.

Is context rot a real problem?

Yes. A 2025-2026 study (Paulsen, arXiv:2509.21361) found models fall short of their own stated maximum context window by as much as 99%, with some failing at just 100 tokens on certain problem types. See context rot.

Is a 1-million-token context window enough?

For single-session depth, often yes. For cross-session personalization and cost control at scale, pair with external memory (Engram, Mem0, Zep).

Is memory cheaper than a large context window?

Usually yes at scale. Mem0 uses ~1,800 vs ~26,000 tokens per query; p95 latency 1.44s vs 17.1s full-context (Chhikara et al., 2025).

Why does a bigger context window cost so much more?

Self-attention is O(n squared): a 10,000-token context needs roughly 100 million comparisons, a 100,000-token context needs roughly 10 billion. Doubling the window roughly quadruples the work.

Does summarization fix context rot?

Helps but can lose facts. Best pattern: summarize plus selective memory retrieval. See memory summarization.

RAG vs a bigger context window?

Different layers. RAG retrieves static docs; memory stores per-user facts. Both beat stuffing everything into context. See memory vs RAG.

Best production pattern for context limits?

Memory retrieves top-k facts, recent messages fill the remaining budget, context engineering allocates tokens. Redis's own engineering research frames this as 3 complementary techniques: semantic caching, RAG, and agent memory.

Is a 100-million-token context window actually usable?

Magic.dev's LTM-2-Mini publishes 100 million tokens, but Codingscape's own July 2026 research states plainly that no evidence exists of anyone outside Magic.dev using it. The 1-million-token tier is where most real production traffic runs.

Why can't models use their full stated context window?

Rotary Position Embedding, the mechanism most models use to track token position, has documented limitations at multi-million-token scale without extension techniques. Models also fail on positions they never saw during training, regardless of what the architecture technically permits.

How many tokens is a word, roughly?

About 0.75 words per token, or roughly 4 characters, for English text using byte-pair encoding. This varies by language and tokenizer, so treat it as a rule of thumb rather than an exact conversion.

Continue exploring

Three routes onward: the full memory-vs-context comparison, adding memory to an agent, or cutting token cost.