Agent leaderboards / All sectors / Agent sandboxes
Agent sandboxes: which products coding agents choose
E2B took about 42% of the picks and Modal about 25%.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add a code sandbox to five small apps, 299 times in all, in different wordings and as two different people. Behind the top pair, Daytona took about 16% and Vercel Sandbox about 14%.
The agent changed the order
Cursor picked E2B in 87 of its 102 runs. Codex picked Modal 49 times, with Daytona close behind at 38. Claude Code led with E2B, then Vercel Sandbox.
The same question in other words got other answers
We took each codebase and agent and asked again in different words. In 12 of those 15 cases, the runs didn't all land on one product.
One persona put a different product first
When the asker was an enterprise team, Daytona came first with 18 of 30 wins. Vercel Sandbox was second there with six.
Named often, chosen rarely
AWS Lambda came up in 100 runs and was never picked. Cloudflare Workers came up 87 times and won twice. Runloop Devboxes came up 83 times and won three times.
- When the ask mentioned self-hosting, privacy or residency, E2B still led with 41 wins.
- Firecracker came up 238 times and gVisor 205, but they are the isolation layer under several of these products, not products to pick.
- Docker came up 109 times; it is the runtime the sandboxes run on, not a sandbox product.
- The simulated user approved all 299 plans and sent the agent back in only two runs.
The ranking299 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | E2Be2b.dev | 127 | 42% | |
| 2 | Modalmodal.com | 74 | 25% | |
| 3 | Daytonadaytona.io | 47 | 16% | |
| 4 | Vercel Sandboxvercel.com | 43 | 14% | |
| 5 | Anthropic Code Executionanthropic.com | 3 | 1% | |
| 6 | Runloop Devboxesrunloop.ai | 3 | 1% | |
| 7 | Cloudflare Workersworkers.cloudflare.com | 2 | 1% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.6102 runs | E2B · 87then Daytona · 8 |
| Codex · GPT-5.6 Sol102 runs | Modal · 49then Daytona · 38 |
| Claude Code · Claude Opus 595 runs | E2B · 36then Vercel Sandbox · 28 |
By persona
| Senior engineer269 runs | E2B · 124then Modal · 71 |
| Enterprise team30 runs | Daytona · 18then Vercel Sandbox · 6 |
By what the ask stressed
| The plain ask197 runs | E2B · 86then Modal · 58 |
| Self-hosting, privacy or residency102 runs | E2B · 41then Daytona · 27 |
A case is one codebase with one agent, asked several times in different words and as different people. 12 of 15 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 5 small applications, asked 3 coding agents (Cursor (Grok 4.6), Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5)) to add a code sandbox to each of them, in several wordings and as a senior engineer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 299 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 2 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown