Agent leaderboards / All sectors / AI gateway
AI gateway: which products coding agents choose
Portkey and Cloudflare AI Gateway tie on top at about 21% each.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add an AI gateway to eight small apps, across 140 runs. We varied the wording and asked as four kinds of people. Nothing else pulled ahead: LiteLLM took about 18%, Vercel AI Gateway about 16%, and OpenRouter about 11%.
Each agent had its own favorite
Codex picked Cloudflare AI Gateway in 20 of its 48 runs. Claude Code led with Portkey at 14 wins. Cursor split its top spot evenly between LiteLLM and OpenRouter, 11 wins each.
Rewording the same request changed the answer
We ran each codebase and agent several times, changing how we asked. In 21 of 24 of those groups, the runs did not all land on the same product.
Who asked moved the result too
Junior developers chose Vercel AI Gateway 14 times, its best showing anywhere. Senior engineers chose Portkey 19 times. Enterprise teams leaned to LiteLLM with 10 wins.
One product came up often and won once
Helicone was named in 103 runs and picked in one, on the SvelteKit indie SaaS app.
- The agents wrote the gateway themselves in 12 runs, about 9% of the total.
- The simulated user approved every plan in the end, but pushed back at least once in 38 runs.
- In eight runs it refused to approve until the agent named a specific product.
- Envoy AI Gateway won once, on the high-traffic TS commerce platform.
The ranking140 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Portkeyportkey.ai | 30 | 21% | |
| 2 | Cloudflare AI Gatewaycloudflare.com | 30 | 21% | |
| 3 | LiteLLMlitellm.ai | 25 | 18% | |
| 4 | Vercel AI Gatewayvercel.com | 22 | 16% | |
| 5 | OpenRouteropenrouter.ai | 15 | 11% | |
| 6 | Built in-houseoutcome | 12 | 9% | |
| 7 | Amazon Bedrockaws.amazon.com | 3 | 2% | |
| 8 | Heliconehelicone.ai | 1 | 1% | |
| 9 | Envoy AI Gatewayenvoyproxy.io | 1 | 1% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol48 runs | Cloudflare AI Gateway · 20then LiteLLM · 11 |
| Cursor · Grok 4.647 runs | LiteLLM · 11then OpenRouter · 11 |
| Claude Code · Claude Opus 545 runs | Portkey · 14then Cloudflare AI Gateway · 10 |
By persona
| Senior engineer56 runs | Portkey · 19then LiteLLM · 15 |
| Enterprise team35 runs | LiteLLM · 10then Portkey · 9 |
| Junior developer33 runs | Vercel AI Gateway · 14then Cloudflare AI Gateway · 10 |
| Vibe coder16 runs | Vercel AI Gateway · 5then Cloudflare AI Gateway · 5 |
By what the ask stressed
| The plain ask140 runs | Portkey · 30then Cloudflare AI Gateway · 30 |
A case is one codebase with one agent, asked several times in different words and as different people. 21 of 24 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 8 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Cursor (Grok 4.6), Claude Code (Claude Opus 5)) to add an AI gateway to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 140 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 38 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown