Open-source OCR: route pages, don't pick a model
In plain terms
Open OCR models now rival paid APIs, but the best one for a scanned form is wasted on a clean PDF. This pipeline sends each page to the cheapest tool that reads it well and escalates only when it can't.
is where document AI loses accuracy without anyone noticing. The output looks like text, so it gets indexed, and the problem only shows up weeks later as retrieval that keeps missing things.
Open-source OCR has changed a lot in two years. Document models of around a billion parameters now sit at the top of the public benchmarks, ahead of general models many times their size, and they run on one GPU. But the accurate models are far slower than the simple ones, and most pages don't need them.
The pipeline
Triage
Check for an embedded text layer and score its quality: share of real words, garbled characters, how much of the page it covers. Good text layers skip OCR.
Why route
One popular toolkit publishes both numbers for its own pipeline on olmOCR-Bench: 76% at about 3 pages a second with OCR, and 44% at about 24 pages a second when it trusts the text layer instead. That's eight times the speed for a thirty-point drop.
In a typical corpus most pages are born-digital, so OCR only adds errors to them. The scans are the minority, and for them the text layer is empty or wrong. Running the careful path everywhere wastes compute, and running the fast path everywhere ruins the scans.
Picking tools for each tier
| Approach | Accuracy | Speed | Light hardware | Messy input | Best for |
|---|---|---|---|---|---|
| Text-layer extraction | Born-digital PDFs | ||||
| Tesseract, PP-OCR | Clean scans on CPU | ||||
| Docling, Marker, MinerU | Batteries-included conversion | ||||
| Small document VLMs (~1B) | Default model tier | ||||
| Larger VLMs (3–8B) | Handwriting, forms, bad scans |
For the model tiers, the names I'd look at first: PaddleOCR-VL and GLM-OCR (both around 0.9B, layout-first, near the top of OmniDocBench) for the default tier, and olmOCR 2 (7B, Apache-2.0) or dots.ocr for escalation. For toolkits, Docling is MIT-licensed with pluggable engines; Marker reads the text layer first and OCRs only where needed; MinerU offers tiers from model-free to VLM.
Making it fast
Serve, don't loop. Most of these models support vLLM or SGLang. A notebook loop over pages leaves most of the GPU idle; a batched server doesn't.
FP8 is nearly free. AllenAI report olmOCR 2 at 82.4 in FP8 against 82.3 at full precision. Going down to 4-bit is less documented for OCR, so test it on your own pages before relying on it.
Render at the size the model was trained on. olmOCR expects the longest side at 1288 pixels, and Tesseract wants 300 DPI. Bigger costs tokens without helping. DeepSeek-OCR makes the trade explicit: about 97% precision below tenfold compression into vision tokens, around 60% at twentyfold.
Keep the big model rare. AllenAI put olmOCR 2 at under $200 per million pages on their hardware. In a routed pipeline only a fraction of pages ever reach that tier.
The router, sketched
def ocr_document(pdf):
pages = []
for page in pdf.pages:
text = page.text_layer()
if text and text_quality(text, page) > 0.9:
pages.append(from_text_layer(page))
continue
img = page.render(longest_side=1288)
regions = detect_layout(img.downsample())
result = [small_vlm.read(img.crop(r)) for r in regions]
if needs_escalation(result):
result = large_vlm.read_page(img)
pages.append(assemble(result, regions, page.number))
return pages
def text_quality(text, page):
words = tokenize(text)
real = sum(w.lower() in LEXICON for w in words) / max(len(words), 1)
coverage = min(text_area(page) / page.area * 3, 1.0)
return 0.7 * real + 0.3 * coverageThings to know before choosing
- Check the weights licence, not only the code licence. Several strong models have revenue thresholds or extra conditions on the weights even when the repository is Apache-2.0.
- The benchmarks disagree. Pipeline tools trail on OmniDocBench but stay competitive on olmOCR-Bench. Choose the one that resembles your documents, then build a small benchmark from your own pages anyway.
- At the top, rank barely matters. The leading models on olmOCR-Bench are within a few points of each other. Speed, hardware and licence decide it.
- Judge OCR by retrieval. Recent work found that structural errors like merged columns and broken tables hurt retrieval even when character error rates look fine. Run your retrieval on the OCR output; that's the number that counts.
Sources
- AllenAI, olmOCR (2025), olmOCR 2 release notes, olmOCR-Bench
- OpenDataLab, OmniDocBench (CVPR 2025); MinerU2.5 (2025)
- PaddlePaddle, PaddleOCR-VL technical report (2025)
- DeepSeek, DeepSeek-OCR: Contexts Optical Compression (2025)
- Datalab, Marker and Surya benchmark tables (vendor-reported)
- IBM, Docling and the granite-docling model card
- Tesseract docs, Improving the quality of the output
- Hugging Face, FinePDFs dataset write-up (2025)
- Docling technical report (2024)
Repos to explore
Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.
- allenai/olmocrolmOCR toolkit and olmOCR-Bench; a strong escalation-tier model.
- PaddlePaddle/PaddleOCRPP-OCR for fast CPU OCR, plus PaddleOCR-VL for the model tier.
- zai-org/GLM-OCRCompact, layout-first document model.
- docling-project/doclingMIT-licensed conversion with pluggable OCR engines.
- datalab-to/markerText-layer-first PDF conversion; check the weights licence.
- opendatalab/MinerUTiered parsing from model-free to VLM; check the licence terms.
- deepseek-ai/DeepSeek-OCRVision-token compression and resolution modes.
- rednote-hilab/dots.ocrMultilingual layout parsing in one model.
- tesseract-ocr/tesseractThe classic CPU engine, still useful for clean scans.
- opendatalab/OmniDocBenchBenchmark for end-to-end document parsing.
- vllm-project/vllmServing runtime most of these models support.
Read next
Chunking strategies, compared
How you cut documents into pieces decides what your AI can find. Seven ways to do it, what each costs, and the order I try them in.
RAG for long documents: give it a map
Ordinary RAG answers 'find me the paragraph' well and 'what does this 400-page file say overall' badly. Keeping the document's structure, with summaries at each level, lets one system answer both.