Memory Types · Taxonomy

Parametric vs Non-Parametric Memory, Precisely

Parametric memory is knowledge learned into a model’s weights during training; non-parametric memory is knowledge kept outside the model and retrieved at query time, a distinction that traces directly to the original RAG paper rather than being a loose folk taxonomy. Which one actually fails, and when, is measured research, not a matter of opinion.

Parametric

In the weights, fixed at training

Non-parametric

External, retrieved at query time

The actual origin

Where does the parametric vs non-parametric distinction actually come from?

The original Retrieval-Augmented Generation paper, Lewis et al., 2020, which names knowledge stored in a model’s parameters “parametric memory” and knowledge fetched from an external source at query time “non-parametric memory.” The terms aren’t an informal shorthand that emerged organically; they’re a specific piece of academic vocabulary with a traceable origin.

Parametric memory is learned in model weights and is fast but unverifiable, non-parametric memory is retrieved at query time and is updatable and inspectable
Figure 1. The distinction traces to the original RAG paper’s own definitions of knowledge in weights versus retrieved context.

Parametric memory is fast, since no lookup step is needed, the knowledge is already baked into the network’s weights, but it’s also fixed: updating it means retraining or fine-tuning, and it’s fundamentally unverifiable, since there’s no way to point to where a specific fact lives inside billions of parameters or prove it hasn’t drifted. Non-parametric memory inverts every one of those properties: it’s updatable without touching the model at all, inspectable since you can see exactly what got retrieved, and attributable, since an answer can cite the specific source it came from. This is also why “non-parametric” is the right word rather than just “external”: it echoes the same distinction non-parametric statistics makes, where behavior isn’t fixed by a predetermined, fitted set of parameters.

The RAG paper’s own framing was specifically about generation quality, not agent memory: combining a pretrained sequence-to-sequence model with a dense vector index of external text, retrieved and conditioned on for each generation, produced more specific, factual and diverse language than a purely parametric model generating from what it memorized during training alone. Later work, including on the Atlas retrieval-augmented model, dug into a deeper question the original paper didn’t fully resolve: when a model has both kinds of information available at once, which one actually wins, and why. Causal mediation analysis of Atlas’s internal representations answers this directly: when the model can choose between what it already knows and what’s in front of it in context, it relies more on the retrieved context than on its own parametric knowledge, and the analysis further identifies two separate internal mechanisms at work, first the model deciding whether the retrieved context is even relevant to the question, and second, how it computes the output representations that let it copy from that context when it decides the context should win.

Knowing where the term comes from is less useful than knowing precisely where the two approaches actually break down in practice, which turns out to be measured research rather than received wisdom. Where does parametric memory actually fail?

The measured failure

Where does parametric memory actually fail?

Specifically on long-tail, infrequently discussed facts, and scaling the model up barely helps: GPT-3 achieves only 19% accuracy on questions about the least popular entities in a large-scale open-domain knowledge probe, and the accuracy gap between a much smaller model and GPT-3 on those same long-tail questions is only 3 percentage points.

GPT-3 achieves only 19 percent accuracy on the least popular entities in the PopQA benchmark, and scaling to a larger model barely closes the gap
Figure 2. Parametric memory specifically fails on long-tail knowledge, and scaling model size does not fix it.

This comes from PopQA, a 14,000-question open-domain benchmark built specifically to span a wide range of entity popularity, measured by how often an entity is actually discussed online rather than by an arbitrary difficulty label. Across ten models spanning three model families, the pattern held consistently: strong accuracy on well-known entities, a sharp drop on obscure ones, and larger models offering essentially no rescue for that specific gap, since the problem isn’t model capacity, it’s how rarely the fact appeared during training in the first place. This is precise, checkable evidence for something teams often assume rather than measure: a bigger, more capable model is not a substitute for actually having access to the specific fact in question.

The research also found that this popularity effect isn’t uniform across every kind of fact. Some relationship types, like which country an entity belongs to, showed high accuracy regardless of how popular the entity was, while others, like a person’s occupation or which director made a film, showed strong popularity dependence and proved much harder for models to memorize reliably. This matters for a practical reason: it means the decision about whether a given fact needs external memory or can be trusted to parametric recall isn’t just about how obscure the overall subject is, it depends on what kind of relationship is actually being asked about, which is a more precise, and more useful, way to think about the risk than a single blanket rule.

The obvious fix looks like adding retrieval unconditionally to cover that gap. The same research shows that fix has a real cost of its own. Does adding retrieval always fix that failure?

Not a free win

Does adding retrieval always fix that failure?

No. Adding retrieved context caused a base model to get 10% of questions wrong that it had previously answered correctly using parametric memory alone, mostly because of low-quality or irrelevant retrieved passages, often about a different entity sharing the same name. The fix that actually works is retrieving selectively, not universally.

Adding retrieval caused 10 percent of previously correct answers to become wrong, while adaptive retrieval achieved 46.5 percent accuracy at half the inference cost
Figure 3. Adding retrieval indiscriminately can degrade previously correct answers. Adaptive retrieval measurably fixes this.

The same research proposes and measures a direct fix: Adaptive Retrieval, which retrieves external, non-parametric memory only when a question concerns a lower-popularity entity, and relies on parametric memory alone otherwise, using a popularity threshold tuned per relationship type. The measured result on PopQA is a specific, real number: 46.5% accuracy, a 5.3-percentage-point improvement over any single non-adaptive method, while cutting inference cost roughly in half and reducing latency for self-hosted models by up to 9%. Larger, more capable models automatically retrieve less often under this scheme, since their parametric memory is more reliable to begin with, which is exactly the behavior you’d want from a well-designed system rather than something that had to be hand-tuned per model.

The mechanism behind the degradation is worth stating precisely, since it explains why “just always retrieve, it can only help” is the wrong intuition. When a model already knows a fact reliably from training, an added retrieved passage isn’t neutral: if that passage happens to be about a different, similarly-named entity, or is otherwise low-quality, the model has to weigh it against what it already knows, and the research shows this weighing process fails often enough to matter. Retrieval doesn’t just add information; it adds a second source of truth the model has to arbitrate between, and arbitration is itself a place things can go wrong that simply trusting parametric memory alone never had to contend with.

Both approaches so far treat parametric and non-parametric memory as cleanly separate. There’s a genuine third option that blurs the line between them. Is there a middle ground between the two?

A third category

Is there a middle ground between the two?

Yes: model editing techniques like ROME surgically modify only the specific weights responsible for one fact, using causal tracing to locate them, rather than retraining the whole model or leaving the fact external and retrieved.

Model editing techniques like ROME surgically modify only the specific weights encoding one fact, a middle ground between full retraining and external memory
Figure 4. Model editing techniques like ROME surgically target the specific weights encoding one fact, rather than retraining or storing it externally.

ROME (Rank-One Model Editing) treats a specific layer’s feed-forward module as a key-value store, where a “key” represents a subject and a “value” encodes the associated fact, and updates just that key-value pair through a targeted, rank-one change to the weights, identified by systematically corrupting parts of the network and observing which weights actually carry the information in question. This is still parametric, in the strict sense that the edit lives in the model’s weights, but it’s a world away from full retraining: narrowly scoped, fast, and designed specifically to avoid disturbing everything else the model knows.

Full retraining and fine-tuning have a well-documented failure mode this kind of surgical editing tries specifically to avoid: catastrophic forgetting, where teaching the model a new fact degrades or erases something it previously knew, especially when the update touches broad regions of the network rather than the narrow part actually responsible for the fact in question. A rank-one, causally-targeted edit is an attempt to get the update benefit of fine-tuning, a fact permanently embedded in the model with no runtime lookup needed, without paying that broader forgetting cost. It’s worth knowing this category exists mainly so “parametric” doesn’t get treated as synonymous with “expensive and risky to change.” For most production teams, though, this remains a research-stage technique rather than a standard tool reached for alongside fine-tuning or external memory, since the tooling and reliability guarantees around it are far less mature than either of the two more established approaches.

Everything above answers what these terms mean and where each one is measurably strong or weak. The practical question of which one an actual production team should build is a separate, already-covered decision. How does this relate to the fine-tuning-vs-memory decision?

Where to go next

How does this relate to the fine-tuning-vs-memory decision?

This page has covered what the two terms mean and where the research shows each one is actually reliable; the practical question of which to build for a specific agent, and when combining them makes sense, is covered in full elsewhere on this site.

Fine-tuning a model updates its parametric memory in the production sense this page has been discussing academically; adding a memory layer builds non-parametric memory an agent retrieves at runtime. The measured research above supports a specific practical takeaway for that decision: since parametric memory reliably fails on long-tail, specific, or per-user facts, and retrieval isn’t automatically safe to bolt on everywhere either, most production agents need non-parametric memory for anything specific to a user, a project, or a moment in time, reserving parametric updates (through fine-tuning) for stable, general patterns that apply broadly rather than to one user’s specific history.

The popularity-dependence finding above translates directly into a design heuristic worth carrying into that decision: facts about a specific user, a specific account, or a specific project are, by definition, the least “popular” facts a general-purpose model could possibly have seen during training, since a base model was never trained on your specific customer’s specific preferences in the first place. That makes them exactly the category of fact the PopQA research shows parametric memory handles worst, and exactly the category non-parametric memory, built and retrieved at runtime, exists to cover reliably instead.

None of this is a call to avoid fine-tuning generally, only to be precise about which memory problem it’s actually solving. The full build-decision guide, including cost, combining both approaches, and where RAG fits alongside memory, is covered on memory versus fine-tuning; the mechanics of actually building the non-parametric side are covered on vector databases and writing memories.

FAQ

Frequently asked questions

The practical questions that follow once the research above is understood.

Is RAG the same thing as non-parametric memory?

RAG is a pattern that uses non-parametric memory. The original RAG paper defines the retrieved external documents as non-parametric memory specifically, combined with a model's own parametric knowledge during generation.

Does a bigger model need less non-parametric memory?

Not for long-tail facts. Research on PopQA found the accuracy gap between a much smaller and a much larger model on the least popular entity questions was only 3 percentage points. Scale doesn't meaningfully fix that specific gap.

Should every query use retrieval if non-parametric memory is available?

No. Measured research found that adding retrieved context caused 10% of previously correct answers to become wrong, mostly from low quality or irrelevant passages. Retrieving only when parametric memory is likely unreliable performed better and cost less.

Is fine-tuning the same thing as model editing?

No. Fine-tuning updates broad regions of a model's weights and risks catastrophic forgetting. Model editing techniques like ROME target only the specific weights responsible for one fact, a narrower and more surgical kind of parametric update.

Why do per-user facts specifically need non-parametric memory?

A per-user fact is, by definition, something a general-purpose model was never trained on, since it wasn't public knowledge available during training. That makes it exactly the kind of low-popularity fact parametric memory research shows models handle worst.

Does knowing this research change how you should build an agent?

It supports a specific practical default: use non-parametric memory for anything user-specific, project-specific, or time-sensitive, and reserve fine-tuning for stable, general patterns. The full build decision is covered on memory versus fine-tuning.