Agent leaderboards / All sectors / AI gateway
AI gateway: which products coding agents choose
Portkey and Cloudflare AI Gateway tied on top, each about 21%.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add an AI gateway to eight small apps, 140 runs in all, in different words and as four different people. No product stands out here. The rest of the wins spread over six more products, with LiteLLM third at about 18%.
Each agent led with a different product
Codex picked Cloudflare AI Gateway in 20 of its 48 runs. Claude Code led with Portkey at 14. Cursor split almost evenly between LiteLLM and OpenRouter, 11 wins each.
The same question got different answers
We counted 24 cases, meaning one codebase and one agent asked several times in different wordings. In 21 of them the runs didn't all land on the same product.
Who asked changed the winner
Junior developers picked Vercel AI Gateway in 14 of 33 runs. Senior engineers led with Portkey, 19 of 56. Enterprise teams put LiteLLM first with 10.
One product named often, chosen once
Helicone came up in 103 runs and won one of them. That win was in the SvelteKit indie note-taking app.
- The agents wrote the gateway themselves in 12 runs, about 9%.
- The simulated user, who approves each plan, sent the agent back at least once in 38 of 140 runs.
- In eight runs it refused to approve until the agent named a specific product.
- Envoy AI Gateway won once, in the high-traffic TypeScript checkout platform.
- Amazon Bedrock won three runs.
The ranking140 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Portkeyportkey.ai | 30 | 21% | |
| 2 | Cloudflare AI Gatewaycloudflare.com | 30 | 21% | |
| 3 | LiteLLMlitellm.ai | 25 | 18% | |
| 4 | Vercel AI Gatewayvercel.com | 22 | 16% | |
| 5 | OpenRouteropenrouter.ai | 15 | 11% | |
| 6 | Built in-houseoutcome | 12 | 9% | |
| 7 | Amazon Bedrockaws.amazon.com | 3 | 2% | |
| 8 | Heliconehelicone.ai | 1 | 1% | |
| 9 | Envoy AI Gatewayenvoyproxy.io | 1 | 1% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol48 runs | Cloudflare AI Gateway · 20then LiteLLM · 11 |
| Cursor · Grok 4.647 runs | LiteLLM · 11then OpenRouter · 11 |
| Claude Code · Claude Opus 545 runs | Portkey · 14then Cloudflare AI Gateway · 10 |
By persona
| Senior engineer56 runs | Portkey · 19then LiteLLM · 15 |
| Enterprise team35 runs | LiteLLM · 10then Portkey · 9 |
| Junior developer33 runs | Vercel AI Gateway · 14then Cloudflare AI Gateway · 10 |
| Vibe coder16 runs | Vercel AI Gateway · 5then Cloudflare AI Gateway · 5 |
By what the ask stressed
| The plain ask140 runs | Portkey · 30then Cloudflare AI Gateway · 30 |
A case is one codebase with one agent, asked several times in different words and as different people. 21 of 24 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 8 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Cursor (Grok 4.6), Claude Code (Claude Opus 5)) to add an AI gateway to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 140 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 38 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown