Agent leaderboards / All sectors
Which tools do coding agents choose?
Which tools coding agents choose, sector by sector: 23 leaderboards from 6455 sandboxed runs with Claude Code, Codex and Cursor, each with its key learnings.
The sectorsone page each
Agent sandboxes
E2B took about 42% of the picks and Modal about 25%.
299 runs · E2B 42%Observability
Sentry took about 37% of picks across 11 small apps.
360 runs · Sentry 37%Payments
Stripe won about 88% of the runs to add payments.
395 runs · Stripe 88%Deploy
Vercel took about 41% of deploy picks, Render about 34%.
270 runs · Vercel 41%Auth
No provider stood out: WorkOS AuthKit led with about 26% of runs.
201 runs · WorkOS AuthKit 26%Email providers
Resend won about 36% of runs, Postmark about 27%.
208 runs · Resend 36%Product analytics
PostHog won about 53% of runs across ten small apps.
359 runs · PostHog 53%Databases
Neon took about 66% of the database picks.
356 runs · Neon 66%File storage
Amazon S3 won about 46% of runs, and led for every agent.
90 runs · Amazon S3 46%LLM evals & observability
Langfuse led the evals boards with about 34% of runs.
288 runs · Langfuse 34%Voice Agents
Vapi led with about 24%, and each agent had its own favorite.
308 runs · Vapi 24%Serverless functions
AWS Lambda took about 24% of runs, Vercel Functions about 23%.
287 runs · AWS Lambda 24%Cloud
AWS took about 62% of the runs to pick a cloud.
215 runs · AWS 62%AI gateway
Portkey and Cloudflare AI Gateway tied on top, each about 21%.
140 runs · Portkey 21%Bot protection
Cloudflare Turnstile took about 57% of the runs.
160 runs · Cloudflare Turnstile 57%Search
Agents wrote search themselves in about 23% of runs.
459 runs · Postgres Full-Text Search 17%Agent frameworks
Agents wrote it themselves in about 25% of runs.
681 runs · Vercel AI SDK 14%Performance in CI
Agents built the gate themselves in about 52% of runs.
216 runs · JMH 10%Code review
Claude Code review took about 26%, and every agent leaned to its own tool.
290 runs · Claude Code review 26%Internationalization
next-intl led with about 25%, Django's framework second.
224 runs · next-intl 25%Maps
Google Maps and Leaflet tie at about 16% each, and no product leads.
316 runs · Google Maps 16%In-app chat & calls
Communication choices depend on the existing stack and call controls.
175 runs · Stream Chat and Video 30%Vector search
Neon and OpenAI benefit from existing integrations.
158 runs · Neon 34%If you sell software in one of these sectors, the library has a page per sector on what the numbers mean for a vendor, and a page per product.