# Open-source OCR: route pages, don't pick a model

> Open OCR models now rival paid APIs, but the best one for a scanned form is wasted on a clean PDF. This pipeline sends each page to the cheapest tool that reads it well and escalates only when it can't.

Source: https://www.ajiththaduri.site/blueprints/open-source-ocr-pipeline — CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)

OCR is where document AI loses accuracy without anyone noticing. The output looks like text, so it gets indexed, and the problem only shows up weeks later as retrieval that keeps missing things.

Open-source OCR has changed a lot in two years. Document models of around a billion parameters now sit at the top of the public benchmarks, ahead of general models many times their size, and they run on one GPU. But the accurate models are far slower than the simple ones, and most pages don't need them.

> **From my notes:** Use CPU text extraction for born-digital pages and save GPU models for pages that need them. The gap is large: Marker publishes 23.7 pages a second on CPU without OCR, against 2.9 in its balanced OCR mode, and Docling measured about 1.3 pages a second on an M3 Max CPU. Move off Tesseract-class engines when tables, multi-column layouts or old scans matter. That's where vision-language models now lead, at olmOCR's published cost of $176 per million pages.

## The pipeline

<Flow
  title="Tiered OCR"
  zones={[
    { id: "cheap", label: "Cheap path", tone: "teal" },
    { id: "model", label: "Model path", tone: "accent" },
  ]}
  stages={[
    { id: "triage", label: "Triage", note: "text layer? any good?", zone: "cheap", detail: "Check for an embedded text layer and score its quality: share of real words, garbled characters, how much of the page it covers. Good text layers skip OCR." },
    { id: "extract", label: "Use the text layer", zone: "cheap", detail: "Born-digital pages with healthy text are read directly, with layout. Fastest path, and for these pages usually the most accurate." },
    { id: "layout", label: "Detect layout", zone: "model", detail: "Find text blocks, tables, figures and reading order on a downsampled page. The leading open systems all do this before reading anything." },
    { id: "small", label: "Small VLM", note: "~1B, per region", zone: "model", detail: "Read each region at full resolution with a compact document model. Handles most scanned pages, including tables and columns." },
    { id: "check", label: "Check", zone: "model", detail: "Low confidence, tables whose rows don't line up, text that fails basic language checks, or handwriting: flag the page." },
    { id: "large", label: "Escalate", note: "larger VLM", zone: "model", detail: "Flagged pages go to a larger model that is stronger on handwriting, forms and degraded scans. Expensive per page, cheap overall." },
    { id: "out", label: "Output", note: "Markdown + boxes", detail: "Markdown or JSON with structure preserved, plus page numbers and bounding boxes so any passage can be cited back to its location." },
  ]}
/>

## Why route

One popular toolkit publishes both numbers for its own pipeline on olmOCR-Bench: 76% at about 3 pages a second with OCR, and 44% at about 24 pages a second when it trusts the text layer instead. That's eight times the speed for a thirty-point drop.

In a typical corpus most pages are born-digital, so OCR only adds errors to them. The scans are the minority, and for them the text layer is empty or wrong. Running the careful path everywhere wastes compute, and running the fast path everywhere ruins the scans.

<Callout tone="warn" title="Text layers lie">
Old scanners often embed an invisible, badly OCR'd text layer behind the image. Score the text layer before trusting it.
</Callout>

> **From my notes:** Expect a substantial scanned minority, not a clean split. In Hugging Face's FinePDFs web-scale crawl, about 29% of documents were sent to OCR; enterprise archives will differ, often upwards. Detection doesn't need a large model. FinePDFs used a small XGBoost classifier over eight sampled pages, Marker re-OCRs a page when any of its embedded text looks garbled, and Docling OCRs only the layout regions not backed by real PDF text. The companion code's `text_layer_quality` is the cheap version of this.

## Picking tools for each tier

<Tradeoffs
  columns={["Accuracy", "Speed", "Light hardware", "Messy input"]}
  rows={[
    { name: "Text-layer extraction", cells: [3, 3, 3, 0], best: "Born-digital PDFs" },
    { name: "Tesseract, PP-OCR", cells: [1, 3, 3, 1], best: "Clean scans on CPU" },
    { name: "Docling, Marker, MinerU", cells: [2, 2, 2, 2], best: "Batteries-included conversion" },
    { name: "Small document VLMs (~1B)", cells: [3, 2, 2, 2], best: "Default model tier" },
    { name: "Larger VLMs (3–8B)", cells: [3, 1, 1, 3], best: "Handwriting, forms, bad scans" },
  ]}
  caption="More dots is better. From published benchmarks plus my own testing. Check against your documents."
/>

For the model tiers, the names I'd look at first: **PaddleOCR-VL** and **GLM-OCR** (both around 0.9B, layout-first, near the top of OmniDocBench) for the default tier, and **olmOCR 2** (7B, Apache-2.0) or **dots.ocr** for escalation. For toolkits, **Docling** is MIT-licensed with pluggable engines; **Marker** reads the text layer first and OCRs only where needed; **MinerU** offers tiers from model-free to VLM.

## Making it fast

**Serve, don't loop.** Most of these models support vLLM or SGLang. A notebook loop over pages leaves most of the GPU idle; a batched server doesn't.

**FP8 is nearly free.** AllenAI report olmOCR 2 at 82.4 in FP8 against 82.3 at full precision. Going down to 4-bit is less documented for OCR, so test it on your own pages before relying on it.

**Render at the size the model was trained on.** olmOCR expects the longest side at 1288 pixels, and Tesseract wants 300 DPI. Bigger costs tokens without helping. DeepSeek-OCR makes the trade explicit: about 97% precision below tenfold compression into vision tokens, around 60% at twentyfold.

**Keep the big model rare.** AllenAI put olmOCR 2 at under $200 per million pages on their hardware. In a routed pipeline only a fraction of pages ever reach that tier.

> **From my notes:** Old and low-quality scans remain the weakest category by a wide margin. On olmOCR-Bench's old-scans split, leading systems score roughly 29 to 50, against 94 to 100 on clean pages. Tables come next, ranging from about 61 to 88 across the leaders. OmniDocBench flags handwriting, newspapers, blurry scans, watermarks and rotated pages as its hard cases. If your corpus contains these, benchmark on them specifically; overall scores hide them.

## The router, sketched

```python
def ocr_document(pdf):
    pages = []
    for page in pdf.pages:
        text = page.text_layer()
        if text and text_quality(text, page) > 0.9:
            pages.append(from_text_layer(page))
            continue
        img = page.render(longest_side=1288)
        regions = detect_layout(img.downsample())
        result = [small_vlm.read(img.crop(r)) for r in regions]
        if needs_escalation(result):
            result = large_vlm.read_page(img)
        pages.append(assemble(result, regions, page.number))
    return pages

def text_quality(text, page):
    words = tokenize(text)
    real = sum(w.lower() in LEXICON for w in words) / max(len(words), 1)
    coverage = min(text_area(page) / page.area * 3, 1.0)
    return 0.7 * real + 0.3 * coverage
```

## Things to know before choosing

- **Check the weights licence, not only the code licence.** Several strong models have revenue thresholds or extra conditions on the weights even when the repository is Apache-2.0.
- **The benchmarks disagree.** Pipeline tools trail on OmniDocBench but stay competitive on olmOCR-Bench. Choose the one that resembles your documents, then build a small benchmark from your own pages anyway.
- **At the top, rank barely matters.** The leading models on olmOCR-Bench are within a few points of each other. Speed, hardware and licence decide it.
- **Judge OCR by retrieval.** Recent work found that structural errors like merged columns and broken tables hurt retrieval even when character error rates look fine. Run your retrieval eval set on the OCR output; that's the number that counts.

## Sources

- AllenAI, *olmOCR* (2025), *olmOCR 2* release notes, olmOCR-Bench
- OpenDataLab, *OmniDocBench* (CVPR 2025); *MinerU2.5* (2025)
- PaddlePaddle, *PaddleOCR-VL* technical report (2025)
- DeepSeek, *DeepSeek-OCR: Contexts Optical Compression* (2025)
- Datalab, Marker and Surya benchmark tables (vendor-reported)
- IBM, Docling and the granite-docling model card
- Tesseract docs, *Improving the quality of the output*
- Hugging Face, *FinePDFs* dataset write-up (2025)
- Docling technical report (2024)

## Repos to explore

- [allenai/olmocr](https://github.com/allenai/olmocr) — olmOCR toolkit and olmOCR-Bench; a strong escalation-tier model.
- [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR) — PP-OCR for fast CPU OCR, plus PaddleOCR-VL for the model tier.
- [zai-org/GLM-OCR](https://github.com/zai-org/GLM-OCR) — Compact, layout-first document model.
- [docling-project/docling](https://github.com/docling-project/docling) — MIT-licensed conversion with pluggable OCR engines.
- [datalab-to/marker](https://github.com/datalab-to/marker) — Text-layer-first PDF conversion; check the weights licence.
- [opendatalab/MinerU](https://github.com/opendatalab/MinerU) — Tiered parsing from model-free to VLM; check the licence terms.
- [deepseek-ai/DeepSeek-OCR](https://github.com/deepseek-ai/DeepSeek-OCR) — Vision-token compression and resolution modes.
- [rednote-hilab/dots.ocr](https://github.com/rednote-hilab/dots.ocr) — Multilingual layout parsing in one model.
- [tesseract-ocr/tesseract](https://github.com/tesseract-ocr/tesseract) — The classic CPU engine, still useful for clean scans.
- [opendatalab/OmniDocBench](https://github.com/opendatalab/OmniDocBench) — Benchmark for end-to-end document parsing.
- [vllm-project/vllm](https://github.com/vllm-project/vllm) — Serving runtime most of these models support.
