Infrastructure · RAG

RAG Architecture, Explained Properly

Retrieval-Augmented Generation runs on four converging stages, ingest and chunk documents, embed them, retrieve the relevant pieces, and generate an answer grounded in what was retrieved, and the naive version of this pipeline is only the starting point, not where a production system should stop. Three named variants and three specific production failure modes are worth knowing before treating RAG as solved.

RAG

Static docs · shared corpus · batch index

Memory

Per-user facts · written every turn · dynamic

The mechanics

What does a RAG pipeline actually consist of?

Four stages that converge across nearly every substantive account of the architecture: document ingestion and chunking, embedding generation, a retrieval layer, and prompt construction and generation. This isn’t one vendor’s particular framing; it’s the shape the pipeline actually takes across independent descriptions from AWS, IBM, Databricks and Google Cloud alike.

The four stage RAG pipeline: ingest and chunk documents, embed them, retrieve relevant chunks, then generate an answer
Figure 1. Ingestion and chunking, embedding, retrieval, and generation converge across nearly every substantive source.

Ingestion breaks source documents into chunks sized for retrieval, a decision that matters more than it first appears, covered further below. Embedding converts each chunk into a vector representation capturing its meaning, the same mechanism this site covers in depth on vector databases. The retrieval layer finds the chunks most relevant to a given query, typically via similarity search over those embeddings, sometimes combined with exact keyword matching. Generation injects the retrieved chunks into the model’s prompt and produces a response grounded in that retrieved content rather than relying purely on what the model already knew from training.

What makes RAG worth the added complexity over simply relying on a model’s own knowledge is precisely the failure mode this site covers on parametric versus non-parametric memory: a model’s training-time knowledge measurably fails on specific, less-discussed, or up-to-date facts, and no amount of scaling the model up reliably fixes that gap. Retrieval closes it by fetching the actual, current, or specific information at the moment of answering rather than hoping it was captured accurately in the weights during training. This is also why RAG’s grounding benefit is inseparable from retrieval quality: a pipeline that retrieves the wrong or irrelevant chunks doesn’t just fail to help, it actively grounds the model’s answer in the wrong information, which reads as more confidently wrong than an ungrounded guess would have.

This basic four-stage shape is where most RAG explanations stop. Production systems that actually work well have usually moved past the naive version of it in at least one specific direction. Which RAG architecture variants actually matter for a memory-focused build?

A bounded selection

Which RAG architecture variants actually matter for a memory-focused build?

Of 8 named variants documented in the wild, 3 are worth knowing specifically for a memory-focused system: Simple RAG with Memory, the naive pattern most teams start from; Corrective RAG, which grades retrieved content before trusting it; and Agentic RAG, where the model controls its own retrieval strategy.

Three RAG architecture variants relevant to memory: Simple RAG with Memory as the naive pattern, Corrective RAG for self-grading, Agentic RAG for agent-controlled retrieval
Figure 2. Three of eight named RAG architecture variants, chosen for relevance to a memory-focused build.

Simple RAG with Memory adds a storage component to the naive pipeline so the model can retain something from previous interactions, which is a genuine improvement over pure static retrieval, but it’s worth recognizing as a starting point rather than a finished design: bolting a memory store onto a RAG pipeline without deliberately separating what belongs in each is exactly the conflation this site’s own RAG versus memory comparison exists to untangle. Corrective RAG breaks retrieved documents into smaller “knowledge strips” and grades each one for relevance before generation, falling back to a broader search, sometimes the open web, if nothing retrieved clears a relevance threshold, which matters specifically for domains like legal or financial work where trusting bad retrieval silently is expensive. Agentic RAG goes further still, assigning a dedicated agent to each document and coordinating them through a meta-agent that decides which sources to consult and how, which is the same underlying idea this site covers generally on memory as a tool: giving the system explicit control over its own retrieval and storage decisions rather than running a single fixed pipeline on every query.

The remaining five named variants in the wider taxonomy are worth knowing exist even without covering each in depth here: Branched RAG selects which specific data source to query based on the input rather than searching everything at once, useful when a system spans genuinely separate knowledge domains; HyDE generates a hypothetical answer first and uses that to guide retrieval, which helps with vague or exploratory queries where the literal wording of the question doesn’t closely match how the answer is phrased in the source documents; Adaptive RAG adjusts its retrieval strategy based on how complex a given query actually is rather than running every query through the same fixed path; and Self-RAG lets the model issue its own follow-up retrieval queries mid-generation as it notices gaps in what it already has. Each solves a specific problem worth reaching for when that specific problem actually shows up, not a general upgrade over naive RAG to adopt by default.

Naming the right architecture pattern only helps if the pipeline underneath it is actually well-built. Several specific, well-documented mistakes are common enough to call out directly. What common design mistakes actually degrade RAG in production?

Where it actually breaks

What common design mistakes actually degrade RAG in production?

3 specific ones: treating RAG as a one-time setup rather than something that needs ongoing evaluation, using default chunk sizes that were never tuned against real queries, and over-retrieving context to the point that it overwhelms the model instead of grounding it.

Three ways RAG degrades in production: treating it as a one time setup, using default chunk sizes, and over retrieving context
Figure 3. Three specific, checkable design mistakes rather than a vague call to tune your RAG system.

RAG is not static: as the underlying data and how users actually query it evolve, retrieval quality can degrade silently, since the system keeps running and returning results, they’re just increasingly stale or irrelevant, without continuous evaluation and re-indexing to catch the drift. This is a genuinely easy trap to fall into precisely because nothing visibly breaks: no error gets thrown, no alert fires, the pipeline just gradually starts answering worse questions with confidently wrong information, which is exactly the kind of failure that only shows up if something is actually measuring retrieval quality on an ongoing basis rather than assuming a system validated once stays validated.

Default chunk sizes rarely fit real data: chunks that are too small lose the surrounding context needed to make sense of them, while chunks that are too large add noise that dilutes what’s actually relevant, and the right size has to be tuned against the specific queries a system actually receives rather than assumed from a framework’s default. A legal document with dense, interlocking clauses needs a very different chunking strategy than a support knowledge base of short, self-contained FAQ entries, and applying the same default configuration to both is a common, avoidable source of poor retrieval quality that has nothing to do with the embedding model or the retrieval algorithm being used.

Over-retrieving context compounds both problems: more retrieved documents is not automatically better, and past a certain point additional context overwhelms the model’s ability to focus on what actually answers the query, producing unfocused or diluted answers instead of more accurate ones, a specific instance of the general problem covered on this site’s context rot page. The instinct to retrieve more when an answer seems incomplete is understandable, but it usually treats the wrong symptom: an incomplete answer more often means the wrong chunks were retrieved, not too few of them, and adding more low-relevance context on top rarely fixes a precision problem.

Everything covered so far is about RAG on its own terms. The remaining question is where a separate concern, memory, actually belongs alongside it, rather than blended into the same pipeline by default. Where does memory actually fit alongside RAG?

A separate concern

Where does memory actually fit alongside RAG?

As a genuinely separate system with a different job: RAG serves a shared, mostly static corpus the same way for every user, while memory is dynamic, scoped to one person, and written on nearly every turn rather than indexed in a batch job.

RAG serves a shared mostly static corpus while memory is per user and written every turn, two different jobs often confused
Figure 4. RAG serves a shared, mostly static corpus. Memory is dynamic and specific to one person’s history.

Treating memory as just another document to retrieve, the Simple RAG with Memory pattern covered above at its most naive, works for a prototype but tends to break down as a system matures, since a memory store has genuinely different write patterns, staleness behavior, and per-user scoping requirements than a document corpus does. A document corpus is typically indexed in a batch job and updated on a schedule measured in hours or days; memory is written on nearly every turn, scoped to one specific user or session, and needs to reflect a fact changing within the same conversation it was first stated in, none of which a batch-indexed document pipeline was designed to handle well. Applying the same chunking and re-indexing discipline covered above to a memory store, rather than to the document corpus it was actually designed for, is a common source of confusion once a team notices their “RAG pipeline” is somehow also supposed to remember what a specific user said five minutes ago.

Getting this distinction right early avoids a specific, expensive kind of rework: retrofitting per-user scoping, real-time write support, and staleness handling onto a document-indexing pipeline that was never architected to carry them. The full comparison between the two, including when they genuinely look similar enough to confuse and how each compares against fine-tuning, is covered on RAG versus memory. Most production systems actually need both running together rather than choosing one, and the specific mechanics of combining them, splitting the prompt budget, running both retrievals in a single turn, and handling the case where a memory contradicts a retrieved document, are covered in full on RAG with memory.

FAQ

Frequently asked questions

The practical decisions that follow once the pipeline and its variants above are understood.

Is Simple RAG with Memory a good design to ship in production?

It's a reasonable starting point, not a finished design. Bolting a memory store onto a document-retrieval pipeline tends to break down as a system matures, since memory has different write, staleness, and per-user scoping needs than a document corpus.

Does adding more retrieved documents make RAG answers more accurate?

Not past a certain point. Over-retrieving context overwhelms the model's ability to focus on what actually answers the query, producing diluted or unfocused answers. An incomplete answer usually means the wrong chunks were retrieved, not too few.

How often does a RAG pipeline actually need re-indexing?

On an ongoing basis, not once at launch. Retrieval quality can degrade silently as underlying data and query patterns evolve, with no visible error, just gradually worse answers, unless something is actively measuring retrieval quality over time.

Should the same chunk size be used for every document type?

No. A legal document with dense, interlocking clauses needs different chunking than a support FAQ of short, self-contained entries. Chunk size should be tuned against the actual queries a system receives, not left at a framework's default.

What's the difference between Corrective RAG and standard RAG?

Corrective RAG grades retrieved content for relevance before generation and falls back to a broader search if nothing clears the bar, rather than trusting whatever was retrieved by default. This matters most in domains where acting on bad retrieval is costly.

Can a memory store use the same underlying vector database as a RAG index?

Technically yes, as separate collections or namespaces, but they still need different write pipelines, since memory is written on nearly every turn and scoped per user, while a document index is typically updated in scheduled batches.