Skip to content
Ajith Thaduri

RAG for long documents: give it a map

Built & tested4 min readPublished 23 Sept 2026

In plain terms

Ordinary RAG answers 'find me the paragraph' well and 'what does this 400-page file say overall' badly. Keeping the document's structure, with summaries at each level, lets one system answer both.

Standard treats a document as a bag of interchangeable chunks. That's fine for short documents and for questions with one answer in one place. Long documents (an annual report, a contract with schedules, a technical manual) get three kinds of question it handles badly:

  • Global: "What are the main risks this report describes?" No chunk contains the answer.
  • Cross-reference: "Does the exclusion in section 9 apply to schedule C?" Two distant parts need reading together.
  • Structural: "Which section covers termination?" The answer is the document's structure, which chunking discarded.

The design

Long-document pipeline · click a stage
Once per documentPer question

Parse

Extract text with headings, heading levels, page numbers, tables and lists intact. OCR scanned pages first. Everything later depends on this.

Four choices do most of the work.

Keep the structure the document already has. Headings and page numbers tell you what belongs together. If your parser flattens everything to plain text, fix that before anything else.

Summaries answer what chunks can't. A global question needs a view of the whole, and bottom-up summaries give you one at every level. This is close to RAPTOR's recursive summarisation. The difference is that for well-structured documents I build the tree from the headings rather than by clustering.

Route, so simple questions stay simple. Most questions are local. Sending them down the summary path costs more and often answers worse, because summaries drop detail.

Re-sort by position. Retrieval ranks by similarity. Pasted in that order, the model reads references before the thing they refer to. Sorting by document position with a breadcrumb on each piece is one of the cheapest improvements here.

Should you build it?

Often not. Context windows are large, and Anthropic's own guidance is that up to around 200,000 tokens can go straight into the prompt, with caching if you'll ask more than once.

ApproachGlobal questionsLocal precisionCost per querySimplicityBest for
Flat chunk RAGShort documents, lookup questions
Hierarchical (this design)Long, structured, queried often
Whole document in contextFits the window, few questions
Map-reduce summarisationOne-off reports
More dots is better; for cost, cheaper. Prompt caching makes the whole-document option much cheaper for repeated questions.

The hierarchical version earns its complexity when documents are long, structured and queried many times: a contract a team works through over weeks, or a manual thousands of people search.

The query path, sketched

python
def answer(question, doc_id):
    kind = route(question)                  # local | cross_ref | global
    if kind == "global":
        nodes = search(question, doc_id, level="summary", k=6)
    elif kind == "cross_ref":
        nodes = expand_to_sections(search(question, doc_id, level="any", k=12))
    else:
        nodes = search(question, doc_id, level="passage", k=8)

    nodes = sorted(dedupe(nodes), key=lambda n: n.position)
    context = "\n\n".join(f"[{n.section_path} · p.{n.pages}]\n{n.text}" for n in nodes)
    return generate_with_citations(question, context)

Where it goes wrong

Summaries drift: they state things the source doesn't. I keep them as extractive as possible and let a summary hit expand to source text before the final answer. Parsing fails silently: a missed heading merges two sections and nobody notices until answers get strange. And updates are expensive unless you only rebuild summaries on the path from the changed node to the root.

Build your eval set with this in mind. Most RAG eval sets are full of find-the-fact questions because they're easy to write. Make a third of yours global or cross-reference, or you'll never see the failures this design exists for.

Sources

  • Sarthi et al., RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval (2024)
  • Anthropic, Introducing Contextual Retrieval (2024), on skipping RAG for small knowledge bases
  • LlamaIndex docs on hierarchical and auto-merging retrieval
  • Edge et al., From Local to Global: A Graph RAG Approach (2024)
  • Li et al., Retrieval Augmented Generation or Long-Context LLMs? (Self-Route, 2024)
  • Jeong et al., Adaptive-RAG (2024)
  • Kim et al., FABLES: Evaluating faithfulness and content selection in book-length summarization (2024)

Repos to explore

Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day