OCR API for large documents
Default pickAWS Textract (DetectDocumentText)
For large documents, AWS Textract handles the biggest jobs asynchronously — up to 3000 pages and 500MB per PDF or TIFF, re-verified against the quota docs in August 2026 — ahead of Azure (around 2000) and Google (around 200). Mistral caps around 1000 pages and GLM-OCR at 100. These are documented async limits, not throughput or accuracy measures, and structured extraction tiers carry separate costs. A light estimate from provider documentation.
DefaultAWS Textract (DetectDocumentText)max_grootte_asyncAWS Textract (DetectDocumentText)
Provider offerings
| Offering | Price ($/1000 pages) | Score | Max pages | Capabilities |
|---|---|---|---|---|
| Mistral OCR 4Mistral AI | 4 | 85.66 | 1000 | tableshandwritingJSON |
| Google Document AI (Enterprise Document OCR)Google Cloud | 1.5 | — | 200 | tableshandwritingJSON |
| AWS Textract (DetectDocumentText)Amazon Web Services | 1.5 | — | 3000 | tableshandwritingJSON |
| Azure Document Intelligence (Read)Microsoft Azure | 1.5 | — | 2000 | tableshandwritingJSON |
| ReductoReducto | —* | — | — | tableshandwritingJSON |
| LlamaParseLlamaIndex (LlamaCloud) | 3.75* | — | — | tableshandwritingJSON |
| GLM-OCRZ.ai (Zhipu AI) | —* | 95.22 | 100 | tableshandwritingJSON |
* token-/credit-priced — the headline understates real per-unit cost, so it is excluded from the cheapest ranking.
Sources
- Mistral Pricing2026-08-20
- Mistral API Pricing (OCR tiers)2026-08-20
- Mistral OCR 4 — SOTA OCR for Document Intelligence2026-08-20
- Mistral OCR 4.1 model card2026-08-20
- CodeSOTA — OmniDocBench Leaderboard2026-08-20
- OmniDocBench (CVPR 2025) — opendatalab2026-08-20
- Google Cloud Document AI Pricing2026-06-22
- AWS Textract Pricing2026-08-20
- AWS Textract — Set Quotas2026-08-20
- Azure Document Intelligence Pricing2026-06-22
- Reducto Pricing2026-08-20
- LlamaParse / LlamaIndex Pricing2026-08-20
- LlamaParse Pricing — tiers & credits2026-08-20
- GLM-OCR — Z.AI Developer Docs2026-08-20