Agent leaderboards / All sectors / Cloud
Cloud: which clouds coding agents choose
AWS took about 62% of the runs to pick a cloud.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to pick a cloud for six small apps, across 215 runs. We asked in different wordings, and as two different people. AWS led for all three agents, and the rest of the picks scattered.
Residency questions turned the order around
When the ask turned to self-hosting, privacy or data residency, Google Cloud led with 12 of 24 runs. AWS led every other theme, including 53 of 83 runs on the plain ask.
Wording changed the pick in almost half the cases
A case here is one codebase with one agent, asked several times in different words. Eight of the 18 cases didn't land on the same product every time.
Enterprise runs split evenly between two clouds
Enterprise teams picked AWS 36 times and Google Cloud 36 times. For senior engineers, Cloudflare came second with 17 wins.
One cloud was named often and never won
Microsoft Azure came up in 76 runs and won none of them.
- Every agent had Google Cloud second, with 12 wins each.
- The simulated user approved every plan, but sent the agent back at least once in 30 runs.
- In 20 runs it refused to approve until the agent named a specific product.
- Inngest and Upstash each won eight runs, all in the Next.js storefront.
- Redis came up in 140 runs and won none, and it's a data store all these clouds offer, not a cloud.
The ranking215 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | AWSaws.amazon.com | 133 | 62% | |
| 2 | Google Cloudcloud.google.com | 36 | 17% | |
| 3 | Cloudflarecloudflare.com | 17 | 8% | |
| 4 | Inngestinngest.com | 8 | 4% | |
| 5 | Upstashupstash.com | 8 | 4% | |
| 6 | Vercel Queuesvercel.com | 3 | 1% | |
| 7 | Cloudflare + Upstash | 2 | 1% | |
| 8 | Renderrender.com | 1 | 0% | |
| 9 | Built in-houseoutcome | 1 | 0% |
By agent, by persona, by wording
By agent
| Claude Code · Claude Opus 572 runs | AWS · 40then Google Cloud · 12 |
| Codex · GPT-5.6 Sol72 runs | AWS · 48then Google Cloud · 12 |
| Cursor · Grok 4.671 runs | AWS · 45then Google Cloud · 12 |
By persona
| Senior engineer143 runs | AWS · 97then Cloudflare · 17 |
| Enterprise team72 runs | AWS · 36then Google Cloud · 36 |
By what the ask stressed
| Volume and cost at scale84 runs | AWS · 44then Google Cloud · 12 |
| The plain ask83 runs | AWS · 53then Google Cloud · 12 |
| Portability, no lock-in24 runs | AWS · 24 |
| Self-hosting, privacy or residency24 runs | Google Cloud · 12then AWS · 12 |
A case is one codebase with one agent, asked several times in different words and as different people. 8 of 18 cases did not hold to a single cloud.
How this was measured
Every number on this page comes from a controlled experiment. We took 6 small applications, asked 3 coding agents (Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol), Cursor (Grok 4.6)) to pick a cloud for each of them, in several wordings and as a senior engineer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 215 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 30 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown