Agent leaderboards / All sectors / Agent sandboxes

Agent sandboxes: which products coding agents choose

E2B won about 42% of runs, Modal about 25%.

299 runs5 apps3 agents2 personasupdated 2026-09-01

The interactive board, open on agent sandboxes. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add a code sandbox to five small apps, across 299 runs. We asked in different words, and as two different people. Behind the top pair, Daytona took about 16% and Vercel Sandbox about 14%.

87 of 102

Each agent had its own favourite

Cursor picked E2B in 87 of its 102 runs. Codex led with Modal at 49 wins, then Daytona at 38. Claude Code picked E2B 36 times and Vercel Sandbox 28 times.

18 of 30

The enterprise runs went elsewhere

Asked as an enterprise team, the agents ran 30 times and picked Daytona in 18 of them. The 269 senior engineer runs led with E2B at 124 wins.

12 of 15

The wording moved the answer

A case is one codebase with one agent, asked the same thing several ways. In 12 of the 15 cases the runs did not all land on the same product.

0 of 100

Named often, chosen almost never

AWS Lambda came up in 100 runs and won none. Cloudflare Workers was named 87 times and won twice. Runloop Devboxes was named 83 times and won three.

  • Asking about self-hosting, privacy or residency didn't change the leader; E2B took 41 of those 102 runs.
  • Firecracker came up 238 times, but it is the isolation layer under several of these products, not a product.
  • The simulated user approved all 299 plans and sent the agent back in only two runs.
Explore every run in the interactive board

The ranking299 runs

ProductWinsShare
1 E2Be2b.dev 127 42%
2 Modalmodal.com 74 25%
3 Daytonadaytona.io 47 16%
4 Vercel Sandboxvercel.com 43 14%
5 Anthropic Code Executionanthropic.com 3 1%
6 Runloop Devboxesrunloop.ai 3 1%
7 Cloudflare Workersworkers.cloudflare.com 2 1%

By agent, by persona, by wording

By agent

Cursor · Grok 4.6102 runsE2B · 87then Daytona · 8
Codex · GPT-5.6 Sol102 runsModal · 49then Daytona · 38
Claude Code · Claude Opus 595 runsE2B · 36then Vercel Sandbox · 28

By persona

Senior engineer269 runsE2B · 124then Modal · 71
Enterprise team30 runsDaytona · 18then Vercel Sandbox · 6

By what the ask stressed

The plain ask197 runsE2B · 86then Modal · 58
Self-hosting, privacy or residency102 runsE2B · 41then Daytona · 27

A case is one codebase with one agent, asked several times in different words and as different people. 12 of 15 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 5 small applications, asked 3 coding agents (Cursor (Grok 4.6), Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5)) to add a code sandbox to each of them, in several wordings and as a senior engineer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 299 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 2 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown