Agent leaderboards / All sectors / Document processing & OCR
Document processing & OCR: which document products coding agents choose
Anthropic Claude took about 24%, the top share in a wide spread.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three agents to read the documents that arrive at 12 codebases, 360 runs in all, wording the request several ways and asking as four kinds of user. Azure AI Document Intelligence came second with about 21%, and Amazon Textract third with about 10%.
The agent decided more than the codebase
Claude Code picked Anthropic Claude in 75 of its 120 runs. Both Codex and Cursor led with Azure AI Document Intelligence instead.
Almost every codebase and agent gave mixed answers
Take one codebase with one agent, asked in different words and as different people. In 34 of 36 such cases the runs did not all land on the same product.
One persona carried a product on its own
Enterprise teams picked Ocrolus in 15 of their 30 runs. All came from the mortgage intake, where one upload holds nine documents. No other persona picked it.
- The simulated user sent the agent back at least once in 274 of 360 runs, and approved every plan in the end.
- Claude Code picked its own maker's product in about 63% of its runs; Cursor never picked its maker's.
- Tesseract OCR was named 163 times and never chosen, which is no loss: it is an engine rather than a service.
The ranking360 runs
By agent, by persona, by wording
By agent
| Claude Code · Claude Opus 5120 runs | Anthropic Claude · 75then Azure AI Document Intelligence · 12 |
| Codex · GPT-5.6 Sol120 runs | Azure AI Document Intelligence · 33then Google Cloud Document AI · 27 |
| Cursor · Grok 4.6120 runs | Azure AI Document Intelligence · 31then Google Gemini · 15 |
By persona
| Junior developer180 runs | Anthropic Claude · 49then Azure AI Document Intelligence · 32 |
| Senior engineer120 runs | Azure AI Document Intelligence · 37then Anthropic Claude · 25 |
| Vibe coder30 runs | Anthropic Claude · 11then Google Gemini · 6 |
| Enterprise team30 runs | Ocrolus · 15then Amazon Textract · 7 |
By what the ask stressed
| The plain ask354 runs | Anthropic Claude · 87then Azure AI Document Intelligence · 76 |
A case is one codebase with one agent, asked several times in different words and as different people. 34 of 36 cases did not hold to a single document product.
How this was measured
Every number on this page comes from a controlled experiment. We took 12 small applications, asked 3 coding agents (Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol), Cursor (Grok 4.6)) to read the documents that arrive at each of them, in several wordings and as a junior developer and senior engineer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 274 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as MarkdownOther sectors
- Agent sandboxes
- Observability
- Payments
- Deploy
- Auth
- Email providers
- Product analytics
- Databases
- File storage
- LLM evals & observability
- Voice Agents
- Serverless functions
- Cloud
- AI gateway
- Bot protection
- Search
- Agent frameworks
- Performance in CI
- Usage-based billing
- Code review
- Internationalization
- Maps
- AI search
- In-app chat & calls
- Vector search