Agent leaderboards / All sectors / AI search

AI search: which services coding agents choose

Anthropic web search led with about 20%, and no service ran away with it.

277 runs12 apps3 agents4 personasupdated 2026-09-14

The interactive board, open on ai search. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to get live web information into 12 codebases, 277 runs in all. We put the same task in different words, and asked as four different people. The rest of the wins spread thin across many services.

22 of 96

Each agent had a different favourite

Claude Code picked Anthropic web search in 55 of its 92 runs. Codex picked Exa most often, 22 times. Cursor picked Tavily most often, 15 times.

33 of 36

New wording, new answer

A case here is one codebase with one agent, asked several times in different words and as different people. In 33 of 36 cases the runs did not all land on the same service.

60%

One agent leaned on its own maker

Claude Code chose Anthropic web search, sold by the company that makes it, in about 60% of its runs. Codex chose OpenAI web search in about 15% of its runs.

0 of 62

Named often, picked never

SerpApi came up in 62 runs and won none of them.

  • The agents wrote the code themselves instead of picking a service in 13 runs, about 5%.
  • The simulated user, who has to approve each plan, sent the agent back at least once in 107 of 277 runs.
  • In 18 runs it refused to approve until the agent named a specific service.
  • Enterprise teams split almost evenly between Brave Search API and NewsAPI.ai, with eight wins each.
  • Mistral web search won seven times, all of them in one Go codebase for supplier screening.
Explore every run in the interactive board

The ranking277 runs

ProductWinsShare
1 Anthropic web searchanthropic.com 56 20%
2 Exaexa.ai 38 14%
3 Brave Search APIbrave.com 31 11%
4 Tavilytavily.com 21 8%
5 OpenAI web searchopenai.com 16 6%
6 Firecrawlfirecrawl.dev 14 5%
7 Parallelparallel.ai 13 5%
8 Built in-houseoutcome 13 5%
9 NewsAPI.ainewsapi.ai 11 4%
10 Perplexity Sonarperplexity.ai 9 3%
11 Bing groundingazure.microsoft.com 7 3%
12 Mistral web searchmistral.ai 7 3%
13 Linkuplinkup.so 6 2%
14 Perigonperigon.io 5 2%
15 Apifyapify.com 4 1%
16 NewsAPInewsapi.org 4 1%
17 Programmable Searchdevelopers.google.com 3 1%
18 GDELTgdeltproject.org 2 1%
19 Google Search groundingai.google.dev 2 1%
20 Vertex AI Searchcloud.google.com 2 1%
21 NewsCatchernewscatcherapi.com 2 1%
22 Azure AI Searchazure.microsoft.com 1 0%
23 World News APIworldnewsapi.com 1 0%
24 PageCrawlpagecrawl.io 1 0%
25 Unilogunilogcorp.com 1 0%
26 Distributor Data Solutions (DDS)distributordatasolutions.com 1 0%
27 GNewsgnews.io 1 0%
28 Kagikagi.com 1 0%

By agent, by persona, by wording

By agent

Codex · GPT-5.6 Sol96 runsExa · 22then OpenAI web search · 14
Claude Code · Claude Opus 592 runsAnthropic web search · 55then Brave Search API · 10
Cursor · Grok 4.689 runsTavily · 15then Exa · 12

By persona

Junior developer117 runsAnthropic web search · 25then Exa · 19
Senior engineer68 runsAnthropic web search · 16then Parallel · 10
Enterprise team48 runsBrave Search API · 8then NewsAPI.ai · 8
Vibe coder44 runsOpenAI web search · 10then Anthropic web search · 10

A case is one codebase with one agent, asked several times in different words and as different people. 33 of 36 cases did not hold to a single service.

How this was measured

Every number on this page comes from a controlled experiment. We took 12 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to get live web information into each of them, in several wordings and as a junior developer and senior engineer and enterprise team and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 277 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 107 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown