light estimateLast updated 2026-08-20

OCR API for large documents

Default pickAWS Textract (DetectDocumentText)

For large documents, AWS Textract handles the biggest jobs asynchronously — up to 3000 pages and 500MB per PDF or TIFF, re-verified against the quota docs in August 2026 — ahead of Azure (around 2000) and Google (around 200). Mistral caps around 1000 pages and GLM-OCR at 100. These are documented async limits, not throughput or accuracy measures, and structured extraction tiers carry separate costs. A light estimate from provider documentation.

DefaultAWS Textract (DetectDocumentText)max_grootte_asyncAWS Textract (DetectDocumentText)
Provider offerings compared on Price, Score, Max pages and capabilities
OfferingPrice ($/1000 pages)ScoreMax pagesCapabilities
Mistral OCR 4Mistral AI485.661000tableshandwritingJSON
Google Document AI (Enterprise Document OCR)Google Cloud1.5200tableshandwritingJSON
AWS Textract (DetectDocumentText)Amazon Web Services1.53000tableshandwritingJSON
Azure Document Intelligence (Read)Microsoft Azure1.52000tableshandwritingJSON
ReductoReducto*tableshandwritingJSON
LlamaParseLlamaIndex (LlamaCloud)3.75*tableshandwritingJSON
GLM-OCRZ.ai (Zhipu AI)*95.22100tableshandwritingJSON

* token-/credit-priced — the headline understates real per-unit cost, so it is excluded from the cheapest ranking.