RAG patterns: which ones are worth it
In plain terms
There are dozens of named RAG techniques and most teams need four or five. What each pattern does, what the evidence says, and the order I'd reach for them.
Search GitHub for and you'll find HyDE, RAPTOR, CRAG, Self-RAG, GraphRAG, ColPali and a hundred more, each with a paper and a benchmark it wins. What's missing is anyone saying which of them a normal team should actually build.
These are my notes on that, grouped by where each pattern acts. Where the evidence is thin or comes from the vendor, I say so.
Query
Rewrite or split the question, search with a hypothetical answer (HyDE), or decide whether it needs retrieval at all (Adaptive-RAG).
If you do five things
An eval set
Fifty to two hundred real questions, each tagged with the passage that answers it. Everything below is guesswork until you can measure before and after.
Hybrid search
next to vector search, merged with reciprocal rank fusion. No model calls, and it catches IDs, codes and rare words that embeddings blur.
A reranker
Pull 50–150 candidates and let a cross-encoder choose. Open-weights models like bge-reranker-v2-m3 fit on one GPU.
Context on each chunk
Contextual retrieval (a model-written line of context per chunk) or late chunking (no model calls). Details in my chunking comparison.
Routing by difficulty
A cheap classifier picks no retrieval, one step or multi-step. Easy questions stop paying for the expensive path.
Anthropic's contextual retrieval post is one of the few that measures these steps separately. On their tests, contextual embeddings cut top-20 retrieval failures by 35%, adding contextual BM25 took it to 49%, and adding a reranker to 67%.
Query side
HyDE has a model write a hypothetical answer and searches with that. It helps when questions and documents use different vocabulary, and backfires when the model doesn't know the domain. Multi-query / RAG-Fusion rewrites the question several ways and fuses the results; answers get more complete, though the paper notes some drift off-topic. Step-back prompting asks a more general question first and reports strong gains on reasoning benchmarks. Decomposition is less a technique than a habit most agentic systems already have.
Index side
Contextual retrieval and late chunking have the clearest evidence. Parent–child retrieval (match small, read big) is one of my defaults. RAPTOR builds recursive summary trees; it's good for whole-document questions and expensive to update, and I use a structure-based version in my long-document design. Propositions index atomic facts: sharper matches, a much bigger index.
One result worth knowing: a 2024 study found semantic chunking's gains over fixed-size chunks too inconsistent to justify its cost.
Retrieval side
Beyond hybrid search and reranking, two ideas stand out. ColBERT-style late interaction keeps a vector per token instead of per chunk: more accurate, several times the storage. ColPali / ColQwen embed page images directly with a vision-language model, skipping OCR. For slides, forms and chart-heavy reports this now beats text pipelines on the ViDoRe benchmark, and there are small variants for modest hardware.
Google DeepMind's LIMIT paper adds a theoretical reason to care: single-vector embeddings have a ceiling, set by their dimension, on which combinations of documents they can return. Hybrid search, multi-vector retrieval and reranking all work around it.
Graph RAG
GraphRAG (Microsoft) summarises communities of entities and is good at corpus-wide "what are the themes?" questions; its README warns indexing is expensive. LightRAG mixes graph and vector retrieval and updates incrementally. HippoRAG 2 was designed to stop graph RAG losing to plain RAG on simple factual questions.
That last problem is real. An independent benchmark found graph RAG often underperforms plain RAG on real-world tasks. I'd use a graph when questions are about relationships across many documents, and not as a general upgrade.
Adaptive and agentic
Corrective RAG grades what came back and decides to use, refine or discard it; it plugs in without training. Adaptive-RAG routes by question complexity. Self-RAG needs a specially trained model. Speculative RAG drafts with a small model in parallel and verifies with a large one, reporting gains in accuracy and latency. Search-R1 and similar work train models with reinforcement learning to decide when and what to search. That last area moves fastest, and most of its headline numbers are self-reported.
When not to use RAG
| Approach | Quality | Cheap per query | Fresh data | Simple | Best for |
|---|---|---|---|---|---|
| Classic RAG | Large or changing corpora | ||||
| Everything in context, cached | A few hundred pages at most | ||||
| Cache-augmented generation | Small, static, queried constantly | ||||
| Route per query (Self-Route) | Mixed traffic |
Anthropic suggest knowledge bases up to around 200,000 tokens can simply go in the prompt. A 2024 study found long context beats RAG on quality when fully resourced, and that letting the model choose per query keeps most of that quality at much lower cost.
Notes for a team starting out
- Measure retrieval and answers separately. A missing passage can't be fixed by prompting.
- Expect retrieval failures to look like confident answers. Google found large models tend to answer rather than abstain when context is insufficient, so build in "I don't have enough to answer that".
- Build on maintained frameworks. Many well-known research repos are dormant or archived; take the idea and implement it on something still updated.
- For visual PDFs, try page-image retrieval before building an OCR pipeline.
- Run any single-number claim on your own eval set before believing it, especially vendor numbers.
Sources
- Anthropic, Introducing Contextual Retrieval (2024)
- Gao et al., HyDE (2022); Rackauckas, RAG-Fusion (2024); Zheng et al., Take a Step Back (2023)
- Qu et al., Is Semantic Chunking Worth the Computational Cost? (2024)
- Faysse et al., ColPali (2024); Weller et al., On the Theoretical Limitations of Embedding-Based Retrieval (2025)
- Edge et al., GraphRAG (2024); Guo et al., LightRAG (2024); Gutiérrez et al., HippoRAG 2 (2025); Xiang et al., GraphRAG-Bench (2025)
- Yan et al., CRAG (2024); Jeong et al., Adaptive-RAG (2024); Asai et al., Self-RAG (2023); Wang et al., Speculative RAG (2024); Jin et al., Search-R1 (2025)
- Li et al., Self-Route (2024); Chan et al., Don't Do RAG (2024); Joren et al., Sufficient Context (2024)
- Databricks, Reranking in Mosaic AI Vector Search (2025)
- Anthropic, How we built our multi-agent research system (2025)
Repos to explore
Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.
- microsoft/graphragGraph RAG with community summaries.
- HKUDS/LightRAGGraph plus vector retrieval with incremental updates.
- OSU-NLP-Group/HippoRAGGraph retrieval that holds up on simple factual questions.
- illuin-tech/colpaliRetrieval over page images, no OCR.
- stanford-futuredata/ColBERTLate-interaction retrieval.
- lightonai/pylateTraining and running late-interaction models.
- FlagOpen/FlagEmbeddingOpen embedding and reranker models.
- jina-ai/late-chunkingLate chunking reference implementation.
- VectifyAI/PageIndexVector-free, table-of-contents navigation (self-reported benchmarks).
- PeterGriffinJin/Search-R1RL-trained search agents.
- infiniflow/ragflowMaintained RAG engine with deep document parsing.
- deepset-ai/haystackMaintained orchestration framework.
- run-llama/llama_indexMaintained framework with most of these patterns built in.
- explodinggradients/ragasRAG evaluation metrics.
- truera/trulensEvaluation and tracing.
Read next
Chunking strategies, compared
How you cut documents into pieces decides what your AI can find. Seven ways to do it, what each costs, and the order I try them in.
RAG for long documents: give it a map
Ordinary RAG answers 'find me the paragraph' well and 'what does this 400-page file say overall' badly. Keeping the document's structure, with summaries at each level, lets one system answer both.