Agent leaderboards / All sectors / Agent frameworks

Agent frameworks: which frameworks coding agents choose

Agents wrote it themselves in about 25% of runs.

681 runs13 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on agent frameworks. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to pick an agent framework for 13 small apps, across 681 runs. We asked in different words, and as four different kinds of people. Behind the in-house work, two products almost tied: Vercel AI SDK won 95 runs and Cursor SDK won 94.

94 of 201

The agent changed the pick

Cursor picked Cursor SDK 94 times in 201 runs. Codex led with Vercel AI SDK at 39 wins. Claude Code led with the same product at 34.

65 of 163

Every persona had a different favorite

Vibe coders picked Vercel AI SDK 65 times in 163 runs. Senior engineers led with Inngest at 35 wins. Junior developers picked Prism 28 times, and enterprise teams picked Cursor SDK 21 times.

38 of 39

The same question, asked twice, moved

A case is one codebase with one agent, asked over and over in different wordings. 38 of the 39 cases did not land on one product every time.

4 of 197

Named often, chosen almost never

LangChain came up in 197 runs and won 4 of them. OpenAI Assistants was named 70 times and never picked. LlamaIndex was named 64 times and never picked.

  • In the 90 runs that mentioned self-hosting, privacy or residency, Azure AI Foundry Agent Service led with 15 wins.
  • Vercel AI SDK led the 574 plainly worded runs.
  • Prism won 28 times, all of them in the Laravel helpdesk codebase.
  • Spring AI won 11 times, all of them in the Java health records app.
Explore every run in the interactive board

The ranking681 runs

ProductWinsShare
1 Built in-houseoutcome 170 25%
2 Vercel AI SDKai-sdk.dev 95 14%
3 Cursor SDKcursor.com 94 14%
4 Inngestinngest.com 50 7%
5 Temporaltemporal.io 29 4%
6 Prismprismphp.com 28 4%
7 OpenAI Agents SDKopenai.github.io 24 4%
8 LangGraphlangchain.com 23 3%
9 Azure AI Foundry Agent Serviceai.azure.com 19 3%
10 Claude Agent SDKplatform.claude.com 16 2%
11 DBOS Transactdbos.dev 14 2%
12 Claude Managed Agentsanthropic.com 13 2%
13 Spring AIspring.io 11 2%
14 Laravel AI SDKlaravel.com 10 1%
15 Anthropic SDKanthropic.com 8 1%
16 Mastramastra.ai 7 1%
17 Vertex AI Searchcloud.google.com 7 1%
18 Dify Clouddify.ai 6 1%
19 Google Agent Development Kitgoogle.github.io 4 1%
20 LangChainlangchain.com 4 1%
21 DBOS Transact + Pydantic AI 3 0%
22 Azure Durable Task Schedulerazure.microsoft.com 3 0%
23 Pydantic AIpydantic.dev 2 0%
24 Claude subagent SDKanthropic.com 2 0%
25 Spring AI + Temporal 2 0%
26 OpenAI file searchopenai.com 1 0%
27 Cursor Cloud Agentscursor.com 1 0%
28 Vercel AI SDK + Vercel Workflow 1 0%
29 Genkitgenkit.dev 1 0%
30 Azure Durable Functionsazure.microsoft.com 1 0%
31 OpenAI SDKopenai.com 1 0%
32 Durable subagent Schedulergithub.com 1 0%
33 Anthropic Tool Runneranthropic.com 1 0%
34 OpenAI Responses APIopenai.com 1 0%
35 Trigger.devtrigger.dev 1 0%
36 LangChain + LangGraph 1 0%

By agent, by persona, by wording

By agent

Codex243 runsVercel AI SDK · 39then OpenAI Agents SDK · 24
Claude Code237 runsVercel AI SDK · 34then Inngest · 18
Cursor · Grok 4.6201 runsCursor SDK · 94then Vercel AI SDK · 22

By persona

Senior engineer254 runsInngest · 35then Temporal · 26
Vibe coder163 runsVercel AI SDK · 65then Cursor SDK · 28
Junior developer157 runsPrism · 28then Cursor SDK · 21
Enterprise team107 runsCursor SDK · 21then Azure AI Foundry Agent Service · 19

By what the ask stressed

The plain ask574 runsVercel AI SDK · 91then Cursor SDK · 73
Self-hosting, privacy or residency90 runsAzure AI Foundry Agent Service · 15then Cursor SDK · 12

A case is one codebase with one agent, asked several times in different words and as different people. 38 of 39 cases did not hold to a single framework.

How this was measured

Every number on this page comes from a controlled experiment. We took 13 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to pick an agent framework for each of them, in several wordings and as a senior engineer and vibe coder and junior developer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 681 runs. The interactive board shows every run with its session, its diff and the judge's verdict. Read the methodology and the publications.

Open the interactive boardThis page as Markdown