Chunking strategies, compared
In plain terms
How you cut documents into pieces decides what your AI can find. Seven ways to do it, what each costs, and the order I try them in.
When a system gives a bad answer, I check the chunks before I touch the prompt. More often than not the passage that answers the question exists, but it was split in half, or buried inside a chunk about three other things, or separated from the heading that said what it was about.
The trade-off
Each chunk becomes one . Small chunks match questions precisely but lose context: "the dose was then doubled" means nothing on its own. Large chunks keep context, but their embedding averages several topics together, so they match fewer questions well and cost more tokens every time they're retrieved.
Every strategy below is a different way of getting out of that trade-off.
| Approach | Retrieval quality | Low cost | Simplicity | Best for |
|---|---|---|---|---|
| Fixed-size + overlap | A baseline to beat | |||
| Recursive (separator-aware) | Plain prose | |||
| Structure / layout-aware | Headings, tables, forms, PDFs | |||
| Semantic breakpoints | Long text that changes topic | |||
| Parent–child (small-to-big) | Precise match, but answers need context | |||
| Contextual chunk headers | Chunks that don't stand on their own | |||
| Late chunking | Meaning depends on earlier sections |
The order I try them in
Read twenty real documents first
Headings? Tables? Scans? Long narrative? That hour decides most of what follows.
Structure-aware, with recursive splitting inside long sections
Split along the document's own structure, and put the title and section path at the top of every chunk. That last part is a few lines of code and fixes a surprising number of misses.
Build the eval set now
Fifty to a hundred real questions, each tagged with the passage that answers it. Without it you're comparing strategies by feel.
Parent–child, if answers lack context
If the right chunk comes back but the answer is still weak, retrieve small and hand the model the parent section.
Contextual headers or late chunking, if chunks don't stand alone
If the right chunk doesn't come back at all because it only makes sense in context, give it that context before embedding.
Notes on each
Fixed-size. Cut every N tokens with 10–20% overlap. It ignores meaning entirely. I only keep it as the baseline.
Recursive. Split on the biggest natural boundary that fits (section, paragraph, sentence) and fall back to smaller ones only when needed. Cheap, and it respects the text.
Structure-aware. Parse properly first, so tables stay whole and sections stay with their headings. For PDFs and forms this is the biggest single improvement, and the one teams skip because parsing is dull work.
Semantic breakpoints. Start a new chunk where the similarity between neighbouring sentences drops. It needs an embedding call per sentence and threshold tuning. One 2024 study found the gains over fixed-size chunks too inconsistent to justify the cost, which matches my sense that where you cut matters less than what context the piece carries.
Parent–child. Index small, return big. My usual second step.
Contextual headers. Before embedding, prepend a line or two describing where the chunk sits, optionally written by a model. Anthropic's write-up on contextual retrieval reported large drops in retrieval failures, especially combined with and a . The cost is one model call per chunk at ingestion.
Late chunking. Embed the whole document with a long-context model, then pool token vectors into chunks, so each chunk's vector has seen the rest of the document. No extra model calls, but you need an embedding model that supports it.
Starting sizes
These are first guesses to measure against:
- 200–500 tokens for chunks you match on; larger if questions are summary-style.
- 10–20% overlap only when you split mid-text. None when you split on real boundaries.
- Parents sized so that k of them fit comfortably in the with room for the answer.
A sketch
Helpers like recursive_split and hybrid_search stand in for whatever your stack provides.
def chunk_document(doc, max_tokens=400):
chunks = []
for section in doc.sections: # from a layout-aware parser
header = f"{doc.title} > {' > '.join(section.path)}"
parent_id = store_parent(section.text)
for piece in recursive_split(section.text, max_tokens):
chunks.append({
"text": f"{header}\n\n{piece}", # embedded
"parent_id": parent_id, # what the model reads
"page": section.page, # for citations
})
return chunks
def retrieve(query, k=8):
hits = rerank(query, hybrid_search(query, k=k * 3))[:k]
return [load_parent(p) for p in dedupe(h["parent_id"] for h in hits)]Measuring it
Measure retrieval on its own. If the right passage isn't in the top results, nothing downstream can fix it.
- Recall@k: is the answering passage in the top k? The number I watch most.
- MRR: how high it appears. Matters when you pass only a few chunks on.
- Answer quality, graded against a rubric, once retrieval looks healthy.
Re-run the whole set on every chunking change. A change that helps one kind of question often quietly hurts another.
Sources
- Anthropic, Introducing Contextual Retrieval (2024)
- Günther et al., Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models (Jina AI, 2024)
- Qu et al., Is Semantic Chunking Worth the Computational Cost? (2024)
- Barnett et al., Seven Failure Points When Engineering a Retrieval Augmented Generation System (2024)
- NVIDIA, Finding the Best Chunking Strategy for Accurate AI Responses (2025)
- Chroma, Evaluating Chunking Strategies for Retrieval (2024)
- LlamaIndex, Evaluating the Ideal Chunk Size for a RAG System (2023)
Repos to explore
Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.
- docling-project/doclingStructure-aware parsing of PDFs and Office files into clean sections and tables.
- langchain-ai/langchainRecursive and structure-based text splitters, a practical starting point.
- run-llama/llama_indexHierarchical node parsing and auto-merging (parent–child) retrieval.
- jina-ai/late-chunkingReference code and evaluation for late chunking.
- FlagOpen/FlagEmbeddingOpen embedding and reranker models (BGE) to pair with any chunker.
- explodinggradients/ragasRetrieval and answer metrics for comparing chunking changes.
Read next
RAG for long documents: give it a map
Ordinary RAG answers 'find me the paragraph' well and 'what does this 400-page file say overall' badly. Keeping the document's structure, with summaries at each level, lets one system answer both.
Open-source OCR: route pages, don't pick a model
Open OCR models now rival paid APIs, but the best one for a scanned form is wasted on a clean PDF. This pipeline sends each page to the cheapest tool that reads it well and escalates only when it can't.