Agent leaderboards / All sectors / Agent frameworks
Agent frameworks: which frameworks coding agents choose
Agents wrote it themselves in about 25% of runs.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to pick an agent framework for 13 small apps, across 681 runs. We asked in different words, and as four different kinds of people. Behind the in-house work, two products almost tied: Vercel AI SDK won 95 runs and Cursor SDK won 94.
The agent changed the pick
Cursor picked Cursor SDK 94 times in 201 runs. Codex led with Vercel AI SDK at 39 wins. Claude Code led with the same product at 34.
Every persona had a different favorite
Vibe coders picked Vercel AI SDK 65 times in 163 runs. Senior engineers led with Inngest at 35 wins. Junior developers picked Prism 28 times, and enterprise teams picked Cursor SDK 21 times.
The same question, asked twice, moved
A case is one codebase with one agent, asked over and over in different wordings. 38 of the 39 cases did not land on one product every time.
Named often, chosen almost never
LangChain came up in 197 runs and won 4 of them. OpenAI Assistants was named 70 times and never picked. LlamaIndex was named 64 times and never picked.
- In the 90 runs that mentioned self-hosting, privacy or residency, Azure AI Foundry Agent Service led with 15 wins.
- Vercel AI SDK led the 574 plainly worded runs.
- Prism won 28 times, all of them in the Laravel helpdesk codebase.
- Spring AI won 11 times, all of them in the Java health records app.
The ranking681 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Built in-houseoutcome | 170 | 25% | |
| 2 | Vercel AI SDKai-sdk.dev | 95 | 14% | |
| 3 | Cursor SDKcursor.com | 94 | 14% | |
| 4 | Inngestinngest.com | 50 | 7% | |
| 5 | Temporaltemporal.io | 29 | 4% | |
| 6 | Prismprismphp.com | 28 | 4% | |
| 7 | OpenAI Agents SDKopenai.github.io | 24 | 4% | |
| 8 | LangGraphlangchain.com | 23 | 3% | |
| 9 | Azure AI Foundry Agent Serviceai.azure.com | 19 | 3% | |
| 10 | Claude Agent SDKplatform.claude.com | 16 | 2% | |
| 11 | DBOS Transactdbos.dev | 14 | 2% | |
| 12 | Claude Managed Agentsanthropic.com | 13 | 2% | |
| 13 | Spring AIspring.io | 11 | 2% | |
| 14 | Laravel AI SDKlaravel.com | 10 | 1% | |
| 15 | Anthropic SDKanthropic.com | 8 | 1% | |
| 16 | Mastramastra.ai | 7 | 1% | |
| 17 | Vertex AI Searchcloud.google.com | 7 | 1% | |
| 18 | Dify Clouddify.ai | 6 | 1% | |
| 19 | Google Agent Development Kitgoogle.github.io | 4 | 1% | |
| 20 | LangChainlangchain.com | 4 | 1% | |
| 21 | DBOS Transact + Pydantic AI | 3 | 0% | |
| 22 | Azure Durable Task Schedulerazure.microsoft.com | 3 | 0% | |
| 23 | Pydantic AIpydantic.dev | 2 | 0% | |
| 24 | Claude subagent SDKanthropic.com | 2 | 0% | |
| 25 | Spring AI + Temporal | 2 | 0% | |
| 26 | OpenAI file searchopenai.com | 1 | 0% | |
| 27 | Cursor Cloud Agentscursor.com | 1 | 0% | |
| 28 | Vercel AI SDK + Vercel Workflow | 1 | 0% | |
| 29 | Genkitgenkit.dev | 1 | 0% | |
| 30 | Azure Durable Functionsazure.microsoft.com | 1 | 0% | |
| 31 | OpenAI SDKopenai.com | 1 | 0% | |
| 32 | Durable subagent Schedulergithub.com | 1 | 0% | |
| 33 | Anthropic Tool Runneranthropic.com | 1 | 0% | |
| 34 | OpenAI Responses APIopenai.com | 1 | 0% | |
| 35 | Trigger.devtrigger.dev | 1 | 0% | |
| 36 | LangChain + LangGraph | 1 | 0% |
By agent, by persona, by wording
By agent
| Codex243 runs | Vercel AI SDK · 39then OpenAI Agents SDK · 24 |
| Claude Code237 runs | Vercel AI SDK · 34then Inngest · 18 |
| Cursor · Grok 4.6201 runs | Cursor SDK · 94then Vercel AI SDK · 22 |
By persona
| Senior engineer254 runs | Inngest · 35then Temporal · 26 |
| Vibe coder163 runs | Vercel AI SDK · 65then Cursor SDK · 28 |
| Junior developer157 runs | Prism · 28then Cursor SDK · 21 |
| Enterprise team107 runs | Cursor SDK · 21then Azure AI Foundry Agent Service · 19 |
By what the ask stressed
| The plain ask574 runs | Vercel AI SDK · 91then Cursor SDK · 73 |
| Self-hosting, privacy or residency90 runs | Azure AI Foundry Agent Service · 15then Cursor SDK · 12 |
A case is one codebase with one agent, asked several times in different words and as different people. 38 of 39 cases did not hold to a single framework.
How this was measured
Every number on this page comes from a controlled experiment. We took 13 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to pick an agent framework for each of them, in several wordings and as a senior engineer and vibe coder and junior developer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 681 runs. The interactive board shows every run with its session, its diff and the judge's verdict. Read the methodology and the publications.
Open the interactive boardThis page as Markdown