Agent leaderboards / All sectors / Agent frameworks

Agent frameworks: which frameworks coding agents choose

Agents wrote it themselves in about 25% of runs.

681 runs13 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on agent frameworks. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to pick an agent framework for 13 small apps, over 681 runs, in different wordings and as four different people. Writing it themselves came out on top. Among named frameworks, Vercel AI SDK and Cursor SDK split almost evenly, at about 14% each, with 95 and 94 wins.

94 of 201

The agent changed the answer

Cursor picked Cursor SDK 94 times in 201 runs. Codex and Claude Code both put Vercel AI SDK first, with 39 and 34 wins.

38 of 39

Nearly every case flipped on wording

Take one codebase and one agent, then ask several times in different words. In 38 of 39 such cases, the runs didn't all land on the same framework.

65 of 163

Personas pulled in different directions

Vibe coders picked Vercel AI SDK 65 times in 163 runs. Senior engineers put Inngest first with 35 wins, then Temporal with 26. Enterprise teams led with Cursor SDK, at 21.

197 vs 4

Named often, chosen almost never

LangChain came up in 197 runs and won four. OpenAI Assistants was named 70 times and never picked, LlamaIndex 64 times and never picked.

Explore every run in the interactive board

The ranking681 runs

ProductWinsShare
1 Built in-houseoutcome 170 25%
2 Vercel AI SDKai-sdk.dev 95 14%
3 Cursor SDKcursor.com 94 14%
4 Inngestinngest.com 50 7%
5 Temporaltemporal.io 29 4%
6 Prismprismphp.com 28 4%
7 OpenAI Agents SDKopenai.github.io 24 4%
8 LangGraphlangchain.com 23 3%
9 Azure AI Foundry Agent Serviceai.azure.com 19 3%
10 Claude Agent SDKplatform.claude.com 16 2%
11 DBOS Transactdbos.dev 14 2%
12 Claude Managed Agentsanthropic.com 13 2%
13 Spring AIspring.io 11 2%
14 Laravel AI SDKlaravel.com 10 1%
15 Anthropic SDKanthropic.com 8 1%
16 Mastramastra.ai 7 1%
17 Vertex AI Searchcloud.google.com 7 1%
18 Dify Clouddify.ai 6 1%
19 Google Agent Development Kitgoogle.github.io 4 1%
20 LangChainlangchain.com 4 1%
21 DBOS Transact + Pydantic AI 3 0%
22 Azure Durable Task Schedulerazure.microsoft.com 3 0%
23 Pydantic AIpydantic.dev 2 0%
24 Claude subagent SDKanthropic.com 2 0%
25 Spring AI + Temporal 2 0%
26 OpenAI file searchopenai.com 1 0%
27 Cursor Cloud Agentscursor.com 1 0%
28 Vercel AI SDK + Vercel Workflow 1 0%
29 Genkitgenkit.dev 1 0%
30 Azure Durable Functionsazure.microsoft.com 1 0%
31 OpenAI SDKopenai.com 1 0%
32 Durable subagent Schedulergithub.com 1 0%
33 Anthropic Tool Runneranthropic.com 1 0%
34 OpenAI Responses APIopenai.com 1 0%
35 Trigger.devtrigger.dev 1 0%
36 LangChain + LangGraph 1 0%

By agent, by persona, by wording

By agent

Codex243 runsVercel AI SDK · 39then OpenAI Agents SDK · 24
Claude Code237 runsVercel AI SDK · 34then Inngest · 18
Cursor · Grok 4.6201 runsCursor SDK · 94then Vercel AI SDK · 22

By persona

Senior engineer254 runsInngest · 35then Temporal · 26
Vibe coder163 runsVercel AI SDK · 65then Cursor SDK · 28
Junior developer157 runsPrism · 28then Cursor SDK · 21
Enterprise team107 runsCursor SDK · 21then Azure AI Foundry Agent Service · 19

By what the ask stressed

The plain ask574 runsVercel AI SDK · 91then Cursor SDK · 73
Self-hosting, privacy or residency90 runsAzure AI Foundry Agent Service · 15then Cursor SDK · 12

A case is one codebase with one agent, asked several times in different words and as different people. 38 of 39 cases did not hold to a single framework.

How this was measured

Every number on this page comes from a controlled experiment. We took 13 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to pick an agent framework for each of them, in several wordings and as a senior engineer and vibe coder and junior developer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 681 runs. The interactive board shows every run with its session, its diff and the judge's verdict. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown