Agent leaderboards / All sectors / Document processing & OCR

Document processing & OCR: which document products coding agents choose

Anthropic Claude took about 24%, the top share in a wide spread.

360 runs12 apps3 agents4 personasupdated 2026-09-14

The interactive board, open on document processing & ocr. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three agents to read the documents that arrive at 12 codebases, 360 runs in all, wording the request several ways and asking as four kinds of user. Azure AI Document Intelligence came second with about 21%, and Amazon Textract third with about 10%.

75 of 120

The agent decided more than the codebase

Claude Code picked Anthropic Claude in 75 of its 120 runs. Both Codex and Cursor led with Azure AI Document Intelligence instead.

34 of 36

Almost every codebase and agent gave mixed answers

Take one codebase with one agent, asked in different words and as different people. In 34 of 36 such cases the runs did not all land on the same product.

15 of 30

One persona carried a product on its own

Enterprise teams picked Ocrolus in 15 of their 30 runs. All came from the mortgage intake, where one upload holds nine documents. No other persona picked it.

0 of 128

Named in hundreds of runs, picked in a handful

Rossum was named 128 times and never chosen. Veryfi was named 141 times and won once. ABBYY was named 119 times and won twice.

  • The simulated user sent the agent back at least once in 274 of 360 runs, and approved every plan in the end.
  • Claude Code picked its own maker's product in about 63% of its runs; Cursor never picked its maker's.
  • Tesseract OCR was named 163 times and never chosen, which is no loss: it is an engine rather than a service.
Explore every run in the interactive board

The ranking360 runs

ProductWinsShare
1 Anthropic Claudeanthropic.com 87 24%
2 Azure AI Document Intelligenceazure.microsoft.com 76 21%
3 Amazon Textractaws.amazon.com 37 10%
4 Google Cloud Document AIcloud.google.com 31 9%
5 Google Geminiai.google.dev 17 5%
6 Ocrolusocrolus.com 15 4%
7 Reductoreducto.ai 13 4%
8 Azure AI Content Understandingazure.microsoft.com 11 3%
9 Anthropic Claude + Azure AI Document Intelligence 10 3%
10 Apache PDFBoxpdfbox.apache.org 10 3%
11 Amazon Textract + Anthropic Claude 6 2%
12 Amazon Bedrock Data Automationaws.amazon.com 6 2%
13 OpenAI modelsopenai.com 6 2%
14 Doclingdocling-project.github.io 5 1%
15 pdf.jsmozilla.github.io 3 1%
16 Mindeemindee.com 3 1%
17 Extendextend.ai 3 1%
18 Google Cloud Document AI + Google Gemini 3 1%
19 Mistral OCRmistral.ai 2 1%
20 Camelotcamelot-py.readthedocs.io 2 1%
21 ABBYYabbyy.com 2 1%
22 Azure AI Document Intelligence + pdf.js 2 1%
23 Anthropic Claude + Google Cloud Document AI 2 1%
24 Anthropic Claude + pdf.js 1 0%
25 LandingAI Agentic Document Extractionlanding.ai 1 0%
26 Amazon Textract + Apache PDFBox 1 0%
27 IntellectAI Magic Submissionintellectai.com 1 0%
28 Camelot + Apache PDFBox 1 0%
29 Apache PDFBox + Tesseract OCR 1 0%
30 Veryfiveryfi.com 1 0%
31 Built in-houseoutcome 1 0%

By agent, by persona, by wording

By agent

Claude Code · Claude Opus 5120 runsAnthropic Claude · 75then Azure AI Document Intelligence · 12
Codex · GPT-5.6 Sol120 runsAzure AI Document Intelligence · 33then Google Cloud Document AI · 27
Cursor · Grok 4.6120 runsAzure AI Document Intelligence · 31then Google Gemini · 15

By persona

Junior developer180 runsAnthropic Claude · 49then Azure AI Document Intelligence · 32
Senior engineer120 runsAzure AI Document Intelligence · 37then Anthropic Claude · 25
Vibe coder30 runsAnthropic Claude · 11then Google Gemini · 6
Enterprise team30 runsOcrolus · 15then Amazon Textract · 7

By what the ask stressed

The plain ask354 runsAnthropic Claude · 87then Azure AI Document Intelligence · 76

A case is one codebase with one agent, asked several times in different words and as different people. 34 of 36 cases did not hold to a single document product.

How this was measured

Every number on this page comes from a controlled experiment. We took 12 small applications, asked 3 coding agents (Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol), Cursor (Grok 4.6)) to read the documents that arrive at each of them, in several wordings and as a junior developer and senior engineer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 274 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown