RAG or long context: how to actually choose

Context windows got large enough that retrieval was supposed to become unnecessary. It did not. Here is the decision framed by what each approach costs you rather than by what it can do.

Every time context windows get larger, someone declares retrieval dead. The reasoning seems sound: if you can fit the whole corpus into the prompt, why build a pipeline to select fragments of it? The reasoning is wrong, but it is wrong in an interesting way, and understanding why gives you a decision rule rather than a preference.

The two approaches, minus the jargon

Retrieval means keeping your documents outside the model, searching them at question time, and putting only the relevant pieces into the prompt. Long context means putting the documents in the prompt and letting the model find what it needs.

Both are ways of answering the same question: which text should the model see right now? One answers it before the call, the other during.

Long context is not free just because it fits

Fitting inside the window is a necessary condition, not a sufficient one. Three costs come with the choice.

That third point is the one people underestimate. Retrieval is not only about fitting. It is about placement: a short prompt containing exactly the right passage will often beat a huge prompt that contains the same passage somewhere in the middle.

A fact buried in the middle of a very long prompt is in the context window and functionally invisible at the same time.

Retrieval's cost is that it can be wrong before the model runs

The failure mode of retrieval is silent and specific. If your search returns the wrong three passages, the model answers confidently from the wrong three passages. It has no way to know what it was not shown, and neither do you unless you look at what was retrieved.

This is why retrieval systems are harder to operate than they are to build. A first version is an afternoon. Making it reliably surface the right passage across a real range of questions is ongoing work: chunking that does not split ideas in half, handling questions whose answer spans several documents, dealing with terms your embedding model has never seen, and keeping the index in step with the source.

The rule I use

The deciding question is not corpus size, it is how many times you will ask about the same body of text.

For one document and a handful of questions, put it in the prompt. Building a pipeline to search a single file you are asking three things about is pure overhead. This covers a large share of real usage, and it is where the 'retrieval is dead' claim feels true.

For a corpus you will query repeatedly at volume, retrieve. The cost and latency arithmetic stops being close very quickly, and precise placement of the relevant passage is worth more than raw coverage.

For a corpus larger than the window, retrieve, because there is no choice. But note that this is the least interesting case. The genuinely hard decisions are the ones where both would technically work.

The hybrid that usually wins

In practice the best-performing setups are not purely one or the other. Retrieve generously, then let a large window absorb the imprecision. Instead of agonising over returning exactly the right three chunks, return fifteen and let the model discard what it does not need.

Large windows made retrieval easier rather than unnecessary. Precision at the top of the ranking used to be everything, because you could afford so little. With more room, recall matters more than precision, which is a much more forgiving target to build against.

Bigger context windows did not kill retrieval. They lowered the accuracy your retrieval has to hit.

What to measure

Whichever you pick, instrument the retrieval step itself and not only the final answer. Log what was sent to the model for every query. When an answer is wrong, the first question is whether the correct passage was even present, and without that log you cannot tell a retrieval failure from a reasoning failure. Those two have completely different fixes, and confusing them is how teams spend weeks tuning prompts to solve a search problem.

The other thing worth measuring is cost per question, split into input and output tokens. It is the number that decides the architecture, and it is usually the one nobody looks at until the bill arrives.