Skip to content
Ajith Thaduri

Open-source OCR: route pages, don't pick a model

Built & tested4 min readPublished 23 Sept 2026

In plain terms

Open OCR models now rival paid APIs, but the best one for a scanned form is wasted on a clean PDF. This pipeline sends each page to the cheapest tool that reads it well and escalates only when it can't.

is where document AI loses accuracy without anyone noticing. The output looks like text, so it gets indexed, and the problem only shows up weeks later as retrieval that keeps missing things.

Open-source OCR has changed a lot in two years. Document models of around a billion parameters now sit at the top of the public benchmarks, ahead of general models many times their size, and they run on one GPU. But the accurate models are far slower than the simple ones, and most pages don't need them.

The pipeline

Tiered OCR · click a stage
Cheap pathModel path

Triage

Check for an embedded text layer and score its quality: share of real words, garbled characters, how much of the page it covers. Good text layers skip OCR.

Why route

One popular toolkit publishes both numbers for its own pipeline on olmOCR-Bench: 76% at about 3 pages a second with OCR, and 44% at about 24 pages a second when it trusts the text layer instead. That's eight times the speed for a thirty-point drop.

In a typical corpus most pages are born-digital, so OCR only adds errors to them. The scans are the minority, and for them the text layer is empty or wrong. Running the careful path everywhere wastes compute, and running the fast path everywhere ruins the scans.

Picking tools for each tier

ApproachAccuracySpeedLight hardwareMessy inputBest for
Text-layer extractionBorn-digital PDFs
Tesseract, PP-OCRClean scans on CPU
Docling, Marker, MinerUBatteries-included conversion
Small document VLMs (~1B)Default model tier
Larger VLMs (3–8B)Handwriting, forms, bad scans
More dots is better. From published benchmarks plus my own testing. Check against your documents.

For the model tiers, the names I'd look at first: PaddleOCR-VL and GLM-OCR (both around 0.9B, layout-first, near the top of OmniDocBench) for the default tier, and olmOCR 2 (7B, Apache-2.0) or dots.ocr for escalation. For toolkits, Docling is MIT-licensed with pluggable engines; Marker reads the text layer first and OCRs only where needed; MinerU offers tiers from model-free to VLM.

Making it fast

Serve, don't loop. Most of these models support vLLM or SGLang. A notebook loop over pages leaves most of the GPU idle; a batched server doesn't.

FP8 is nearly free. AllenAI report olmOCR 2 at 82.4 in FP8 against 82.3 at full precision. Going down to 4-bit is less documented for OCR, so test it on your own pages before relying on it.

Render at the size the model was trained on. olmOCR expects the longest side at 1288 pixels, and Tesseract wants 300 DPI. Bigger costs tokens without helping. DeepSeek-OCR makes the trade explicit: about 97% precision below tenfold compression into vision tokens, around 60% at twentyfold.

Keep the big model rare. AllenAI put olmOCR 2 at under $200 per million pages on their hardware. In a routed pipeline only a fraction of pages ever reach that tier.

The router, sketched

python
def ocr_document(pdf):
    pages = []
    for page in pdf.pages:
        text = page.text_layer()
        if text and text_quality(text, page) > 0.9:
            pages.append(from_text_layer(page))
            continue
        img = page.render(longest_side=1288)
        regions = detect_layout(img.downsample())
        result = [small_vlm.read(img.crop(r)) for r in regions]
        if needs_escalation(result):
            result = large_vlm.read_page(img)
        pages.append(assemble(result, regions, page.number))
    return pages

def text_quality(text, page):
    words = tokenize(text)
    real = sum(w.lower() in LEXICON for w in words) / max(len(words), 1)
    coverage = min(text_area(page) / page.area * 3, 1.0)
    return 0.7 * real + 0.3 * coverage

Things to know before choosing

  • Check the weights licence, not only the code licence. Several strong models have revenue thresholds or extra conditions on the weights even when the repository is Apache-2.0.
  • The benchmarks disagree. Pipeline tools trail on OmniDocBench but stay competitive on olmOCR-Bench. Choose the one that resembles your documents, then build a small benchmark from your own pages anyway.
  • At the top, rank barely matters. The leading models on olmOCR-Bench are within a few points of each other. Speed, hardware and licence decide it.
  • Judge OCR by retrieval. Recent work found that structural errors like merged columns and broken tables hurt retrieval even when character error rates look fine. Run your retrieval on the OCR output; that's the number that counts.

Sources

  • AllenAI, olmOCR (2025), olmOCR 2 release notes, olmOCR-Bench
  • OpenDataLab, OmniDocBench (CVPR 2025); MinerU2.5 (2025)
  • PaddlePaddle, PaddleOCR-VL technical report (2025)
  • DeepSeek, DeepSeek-OCR: Contexts Optical Compression (2025)
  • Datalab, Marker and Surya benchmark tables (vendor-reported)
  • IBM, Docling and the granite-docling model card
  • Tesseract docs, Improving the quality of the output
  • Hugging Face, FinePDFs dataset write-up (2025)
  • Docling technical report (2024)

Repos to explore

Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day