Modern models accept very large context windows. That raises a fair question: why build a retrieval pipeline when you can paste everything in?

Both approaches have a place. The right choice depends on corpus size, freshness, cost and how precise the answers need to be.

What each approach does

  • Long context – put the relevant material directly in the prompt and let the model find what it needs.
  • RAG (retrieval-augmented generation) – search a corpus first, then give the model only the top matching chunks.

When long context wins

  • The material is small enough to fit comfortably: one contract, one module, one report.
  • The task needs whole-document reasoning: comparing sections, spotting contradictions, summarising structure.
  • You want to ship fast without building indexing infrastructure.

When RAG wins

  • The corpus is much larger than any window: a documentation site, a ticket history, a knowledge base.
  • Content changes often and you need fresh answers without rebuilding prompts.
  • Cost and latency matter. Sending hundreds of thousands of tokens on every request is slow and expensive.
  • You need citations and access control per document.

The quality trap with very long prompts

Models tend to degrade as context grows, and information buried in the middle of long inputs is easier to miss. A larger window is capacity, not a guarantee of attention. Retrieval that surfaces the right five chunks often beats a prompt with the right answer hidden among five hundred.

The hybrid that usually wins

Many production systems combine both:

  1. Retrieve a generous set of candidates (say 50).
  2. Rerank to the best 5–15.
  3. Expand each winner to its full section or neighbouring chunks.
  4. Put that focused, larger context in the prompt.
text
query → hybrid search (keywords + embeddings) → rerank → expand to sections → prompt

Practical tips for RAG

  • Use hybrid search: keyword search catches exact names and error codes that embeddings miss.
  • Chunk by structure (headings, functions), and store the source and heading with each chunk.
  • Evaluate retrieval separately: is the right chunk in the top results at all?

Decision checklist

  • Fits easily and changes rarely → long context.
  • Large, changing or permissioned → RAG.
  • Precise answers from a large corpus with deep reasoning → hybrid.

Key takeaways

  • Long context is simple and great for small, whole-document tasks.
  • RAG scales to large, fresh, permissioned corpora at lower cost.
  • Bigger windows do not guarantee the model uses them well.
  • Retrieve, rerank and expand is a strong default.