Agent leaderboards / All sectors / Voice Agents

Voice Agents: which products coding agents choose

Vapi led with about 24%, but each agent had a different favorite.

308 runs10 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on voice agents. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add a voice agent to 10 small apps, 308 runs in all. We asked in different words, and as four different people. Vapi came out on top, with Retell AI second at about 18%.

23 of 94

Three agents, three top picks

Cursor chose Vapi 41 times. Claude Code chose Twilio ConversationRelay most often, 23 times out of 94 runs. Codex chose OpenAI Realtime API 32 times.

2 of 44

Procurement and compliance turned the order around

In 44 runs the ask was about procurement and compliance. OpenAI Realtime API led those with 10 wins. Vapi, the overall leader, won two.

30 of 30

The wording changed the answer every time

We took each app and agent and asked again, phrased another way. That gives 30 cases. In all 30 the runs disagreed with each other about which product to use.

1 of 136

Named in many runs, picked in one

Pipecat came up in 136 runs and was chosen once.

  • The simulated user approved every plan, but sent the agent back at least once in 58 runs.
  • In 15 runs it refused to approve until the agent named a specific product.
  • The agents wrote it themselves in 18 runs, about 6%.
  • Azure Voice Live API won 10 runs, all of them in the .NET insurance platform.
  • Vibe coders split almost evenly between Vapi and Retell AI, with 20 wins each.
Explore every run in the interactive board

The ranking308 runs

ProductWinsShare
1 Vapivapi.ai 74 24%
2 Retell AIretellai.com 55 18%
3 OpenAI Realtime APIopenai.com 37 12%
4 LiveKit Agentslivekit.io 35 11%
5 ElevenLabs Agentselevenlabs.io 30 10%
6 Twilio ConversationRelaytwilio.com 28 9%
7 Built in-houseoutcome 18 6%
8 Azure Voice Live APIazure.microsoft.com 10 3%
9 Cartesia Linecartesia.ai 5 2%
10 Picovoicepicovoice.ai 4 1%
11 Grok Voice Think Fast 2.0x.ai 2 1%
12 Smallest AI Voice Agentssmallest.ai 2 1%
13 Cognigycognigy.com 2 1%
14 Gemini Live APIai.google.dev 1 0%
15 Deepgram Voice Agent APIdeepgram.com 1 0%
16 Fluents.aifluents.ai 1 0%
17 SignalWire AI Agentssignalwire.com 1 0%
18 Pipecatpipecat.ai 1 0%
19 Ultravoxultravox.ai 1 0%

By agent, by persona, by wording

By agent

Codex · GPT-5.6 Sol110 runsOpenAI Realtime API · 32then Retell AI · 28
Cursor · Grok 4.6104 runsVapi · 41then Retell AI · 15
Claude Code · Claude Opus 594 runsTwilio ConversationRelay · 23then LiveKit Agents · 15

By persona

Senior engineer128 runsVapi · 37then Retell AI · 26
Enterprise team76 runsLiveKit Agents · 23then OpenAI Realtime API · 10
Vibe coder62 runsVapi · 20then Retell AI · 20
Junior developer42 runsVapi · 15then OpenAI Realtime API · 7

By what the ask stressed

The plain ask258 runsVapi · 72then Retell AI · 53
Procurement and compliance44 runsOpenAI Realtime API · 10then LiveKit Agents · 8

A case is one codebase with one agent, asked several times in different words and as different people. 30 of 30 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 10 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Cursor (Grok 4.6), Claude Code (Claude Opus 5)) to add a voice agent to each of them, in several wordings and as a senior engineer and enterprise team and vibe coder and junior developer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 308 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 58 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown