Agent leaderboards / All sectors / Agent sandboxes
Agent sandboxes: which products coding agents choose
E2B won about 42% of runs, Modal about 25%.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add a code sandbox to five small apps, across 299 runs. We asked in different words, and as two different people. Behind the top pair, Daytona took about 16% and Vercel Sandbox about 14%.
Each agent had its own favourite
Cursor picked E2B in 87 of its 102 runs. Codex led with Modal at 49 wins, then Daytona at 38. Claude Code picked E2B 36 times and Vercel Sandbox 28 times.
The enterprise runs went elsewhere
Asked as an enterprise team, the agents ran 30 times and picked Daytona in 18 of them. The 269 senior engineer runs led with E2B at 124 wins.
The wording moved the answer
A case is one codebase with one agent, asked the same thing several ways. In 12 of the 15 cases the runs did not all land on the same product.
Named often, chosen almost never
AWS Lambda came up in 100 runs and won none. Cloudflare Workers was named 87 times and won twice. Runloop Devboxes was named 83 times and won three.
- Asking about self-hosting, privacy or residency didn't change the leader; E2B took 41 of those 102 runs.
- Firecracker came up 238 times, but it is the isolation layer under several of these products, not a product.
- The simulated user approved all 299 plans and sent the agent back in only two runs.
The ranking299 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | E2Be2b.dev | 127 | 42% | |
| 2 | Modalmodal.com | 74 | 25% | |
| 3 | Daytonadaytona.io | 47 | 16% | |
| 4 | Vercel Sandboxvercel.com | 43 | 14% | |
| 5 | Anthropic Code Executionanthropic.com | 3 | 1% | |
| 6 | Runloop Devboxesrunloop.ai | 3 | 1% | |
| 7 | Cloudflare Workersworkers.cloudflare.com | 2 | 1% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.6102 runs | E2B · 87then Daytona · 8 |
| Codex · GPT-5.6 Sol102 runs | Modal · 49then Daytona · 38 |
| Claude Code · Claude Opus 595 runs | E2B · 36then Vercel Sandbox · 28 |
By persona
| Senior engineer269 runs | E2B · 124then Modal · 71 |
| Enterprise team30 runs | Daytona · 18then Vercel Sandbox · 6 |
By what the ask stressed
| The plain ask197 runs | E2B · 86then Modal · 58 |
| Self-hosting, privacy or residency102 runs | E2B · 41then Daytona · 27 |
A case is one codebase with one agent, asked several times in different words and as different people. 12 of 15 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 5 small applications, asked 3 coding agents (Cursor (Grok 4.6), Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5)) to add a code sandbox to each of them, in several wordings and as a senior engineer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 299 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 2 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown