Agent leaderboards / All sectors / Agent frameworks
Agent frameworks: which frameworks coding agents choose
Agents wrote it themselves in about 25% of runs.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to pick an agent framework for 13 small apps, over 681 runs, in different wordings and as four different people. Writing it themselves came out on top. Among named frameworks, Vercel AI SDK and Cursor SDK split almost evenly, at about 14% each, with 95 and 94 wins.
The agent changed the answer
Cursor picked Cursor SDK 94 times in 201 runs. Codex and Claude Code both put Vercel AI SDK first, with 39 and 34 wins.
Nearly every case flipped on wording
Take one codebase and one agent, then ask several times in different words. In 38 of 39 such cases, the runs didn't all land on the same framework.
Personas pulled in different directions
Vibe coders picked Vercel AI SDK 65 times in 163 runs. Senior engineers put Inngest first with 35 wins, then Temporal with 26. Enterprise teams led with Cursor SDK, at 21.
Named often, chosen almost never
LangChain came up in 197 runs and won four. OpenAI Assistants was named 70 times and never picked, LlamaIndex 64 times and never picked.
- When the ask mentioned self-hosting, privacy or residency, Azure AI Foundry Agent Service led those 90 runs with 15 wins.
- All 28 wins for Prism came in one codebase, a Laravel helpdesk.
- All 19 wins for Azure AI Foundry Agent Service came in the Java healthtech EHR app.
- Pydantic AI won twice, both times in the Python spreadsheet analyst app.
The ranking681 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Built in-houseoutcome | 170 | 25% | |
| 2 | Vercel AI SDKai-sdk.dev | 95 | 14% | |
| 3 | Cursor SDKcursor.com | 94 | 14% | |
| 4 | Inngestinngest.com | 50 | 7% | |
| 5 | Temporaltemporal.io | 29 | 4% | |
| 6 | Prismprismphp.com | 28 | 4% | |
| 7 | OpenAI Agents SDKopenai.github.io | 24 | 4% | |
| 8 | LangGraphlangchain.com | 23 | 3% | |
| 9 | Azure AI Foundry Agent Serviceai.azure.com | 19 | 3% | |
| 10 | Claude Agent SDKplatform.claude.com | 16 | 2% | |
| 11 | DBOS Transactdbos.dev | 14 | 2% | |
| 12 | Claude Managed Agentsanthropic.com | 13 | 2% | |
| 13 | Spring AIspring.io | 11 | 2% | |
| 14 | Laravel AI SDKlaravel.com | 10 | 1% | |
| 15 | Anthropic SDKanthropic.com | 8 | 1% | |
| 16 | Mastramastra.ai | 7 | 1% | |
| 17 | Vertex AI Searchcloud.google.com | 7 | 1% | |
| 18 | Dify Clouddify.ai | 6 | 1% | |
| 19 | Google Agent Development Kitgoogle.github.io | 4 | 1% | |
| 20 | LangChainlangchain.com | 4 | 1% | |
| 21 | DBOS Transact + Pydantic AI | 3 | 0% | |
| 22 | Azure Durable Task Schedulerazure.microsoft.com | 3 | 0% | |
| 23 | Pydantic AIpydantic.dev | 2 | 0% | |
| 24 | Claude subagent SDKanthropic.com | 2 | 0% | |
| 25 | Spring AI + Temporal | 2 | 0% | |
| 26 | OpenAI file searchopenai.com | 1 | 0% | |
| 27 | Cursor Cloud Agentscursor.com | 1 | 0% | |
| 28 | Vercel AI SDK + Vercel Workflow | 1 | 0% | |
| 29 | Genkitgenkit.dev | 1 | 0% | |
| 30 | Azure Durable Functionsazure.microsoft.com | 1 | 0% | |
| 31 | OpenAI SDKopenai.com | 1 | 0% | |
| 32 | Durable subagent Schedulergithub.com | 1 | 0% | |
| 33 | Anthropic Tool Runneranthropic.com | 1 | 0% | |
| 34 | OpenAI Responses APIopenai.com | 1 | 0% | |
| 35 | Trigger.devtrigger.dev | 1 | 0% | |
| 36 | LangChain + LangGraph | 1 | 0% |
By agent, by persona, by wording
By agent
| Codex243 runs | Vercel AI SDK · 39then OpenAI Agents SDK · 24 |
| Claude Code237 runs | Vercel AI SDK · 34then Inngest · 18 |
| Cursor · Grok 4.6201 runs | Cursor SDK · 94then Vercel AI SDK · 22 |
By persona
| Senior engineer254 runs | Inngest · 35then Temporal · 26 |
| Vibe coder163 runs | Vercel AI SDK · 65then Cursor SDK · 28 |
| Junior developer157 runs | Prism · 28then Cursor SDK · 21 |
| Enterprise team107 runs | Cursor SDK · 21then Azure AI Foundry Agent Service · 19 |
By what the ask stressed
| The plain ask574 runs | Vercel AI SDK · 91then Cursor SDK · 73 |
| Self-hosting, privacy or residency90 runs | Azure AI Foundry Agent Service · 15then Cursor SDK · 12 |
A case is one codebase with one agent, asked several times in different words and as different people. 38 of 39 cases did not hold to a single framework.
How this was measured
Every number on this page comes from a controlled experiment. We took 13 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to pick an agent framework for each of them, in several wordings and as a senior engineer and vibe coder and junior developer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 681 runs. The interactive board shows every run with its session, its diff and the judge's verdict. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown