Agent leaderboards / All sectors / AI search
AI search: which services coding agents choose
Anthropic web search led with about 20%, and no service ran away with it.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to get live web information into 12 codebases, 277 runs in all. We put the same task in different words, and asked as four different people. The rest of the wins spread thin across many services.
Each agent had a different favourite
Claude Code picked Anthropic web search in 55 of its 92 runs. Codex picked Exa most often, 22 times. Cursor picked Tavily most often, 15 times.
New wording, new answer
A case here is one codebase with one agent, asked several times in different words and as different people. In 33 of 36 cases the runs did not all land on the same service.
One agent leaned on its own maker
Claude Code chose Anthropic web search, sold by the company that makes it, in about 60% of its runs. Codex chose OpenAI web search in about 15% of its runs.
- The agents wrote the code themselves instead of picking a service in 13 runs, about 5%.
- The simulated user, who has to approve each plan, sent the agent back at least once in 107 of 277 runs.
- In 18 runs it refused to approve until the agent named a specific service.
- Enterprise teams split almost evenly between Brave Search API and NewsAPI.ai, with eight wins each.
- Mistral web search won seven times, all of them in one Go codebase for supplier screening.
The ranking277 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Anthropic web searchanthropic.com | 56 | 20% | |
| 2 | Exaexa.ai | 38 | 14% | |
| 3 | Brave Search APIbrave.com | 31 | 11% | |
| 4 | Tavilytavily.com | 21 | 8% | |
| 5 | OpenAI web searchopenai.com | 16 | 6% | |
| 6 | Firecrawlfirecrawl.dev | 14 | 5% | |
| 7 | Parallelparallel.ai | 13 | 5% | |
| 8 | Built in-houseoutcome | 13 | 5% | |
| 9 | NewsAPI.ainewsapi.ai | 11 | 4% | |
| 10 | Perplexity Sonarperplexity.ai | 9 | 3% | |
| 11 | Bing groundingazure.microsoft.com | 7 | 3% | |
| 12 | Mistral web searchmistral.ai | 7 | 3% | |
| 13 | Linkuplinkup.so | 6 | 2% | |
| 14 | Perigonperigon.io | 5 | 2% | |
| 15 | Apifyapify.com | 4 | 1% | |
| 16 | NewsAPInewsapi.org | 4 | 1% | |
| 17 | Programmable Searchdevelopers.google.com | 3 | 1% | |
| 18 | GDELTgdeltproject.org | 2 | 1% | |
| 19 | Google Search groundingai.google.dev | 2 | 1% | |
| 20 | Vertex AI Searchcloud.google.com | 2 | 1% | |
| 21 | NewsCatchernewscatcherapi.com | 2 | 1% | |
| 22 | Azure AI Searchazure.microsoft.com | 1 | 0% | |
| 23 | World News APIworldnewsapi.com | 1 | 0% | |
| 24 | PageCrawlpagecrawl.io | 1 | 0% | |
| 25 | Unilogunilogcorp.com | 1 | 0% | |
| 26 | Distributor Data Solutions (DDS)distributordatasolutions.com | 1 | 0% | |
| 27 | GNewsgnews.io | 1 | 0% | |
| 28 | Kagikagi.com | 1 | 0% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol96 runs | Exa · 22then OpenAI web search · 14 |
| Claude Code · Claude Opus 592 runs | Anthropic web search · 55then Brave Search API · 10 |
| Cursor · Grok 4.689 runs | Tavily · 15then Exa · 12 |
By persona
| Junior developer117 runs | Anthropic web search · 25then Exa · 19 |
| Senior engineer68 runs | Anthropic web search · 16then Parallel · 10 |
| Enterprise team48 runs | Brave Search API · 8then NewsAPI.ai · 8 |
| Vibe coder44 runs | OpenAI web search · 10then Anthropic web search · 10 |
A case is one codebase with one agent, asked several times in different words and as different people. 33 of 36 cases did not hold to a single service.
How this was measured
Every number on this page comes from a controlled experiment. We took 12 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to get live web information into each of them, in several wordings and as a junior developer and senior engineer and enterprise team and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 277 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 107 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as MarkdownOther sectors
- Agent sandboxes
- Observability
- Payments
- Deploy
- Auth
- Email providers
- Product analytics
- Databases
- File storage
- LLM evals & observability
- Voice Agents
- Serverless functions
- Cloud
- AI gateway
- Bot protection
- Search
- Agent frameworks
- Performance in CI
- Usage-based billing
- Code review
- Internationalization
- Maps
- In-app chat & calls
- Vector search