# Agent sandboxes: which products coding agents choose

> E2B won about 42% of runs, Modal about 25%.

Source: https://armature.tech/leaderboards/sandboxes (Armature agent leaderboards). 299 runs, 5 apps, 3 agents, 2 personas, updated 2026-09-01. Interactive board with every run: https://armature.tech/leaderboards#app/sandboxes

## Key learnings

We asked three coding agents to add a code sandbox to five small apps, across 299 runs. We asked in different words, and as two different people. Behind the top pair, Daytona took about 16% and Vercel Sandbox about 14%.

### Each agent had its own favourite (87 of 102)

Cursor picked E2B in 87 of its 102 runs. Codex led with Modal at 49 wins, then Daytona at 38. Claude Code picked E2B 36 times and Vercel Sandbox 28 times.

### The enterprise runs went elsewhere (18 of 30)

Asked as an enterprise team, the agents ran 30 times and picked Daytona in 18 of them. The 269 senior engineer runs led with E2B at 124 wins.

### The wording moved the answer (12 of 15)

A case is one codebase with one agent, asked the same thing several ways. In 12 of the 15 cases the runs did not all land on the same product.

### Named often, chosen almost never (0 of 100)

AWS Lambda came up in 100 runs and won none. Cloudflare Workers was named 87 times and won twice. Runloop Devboxes was named 83 times and won three.

Smaller learnings:

- Asking about self-hosting, privacy or residency didn't change the leader; E2B took 41 of those 102 runs.
- Firecracker came up 238 times, but it is the isolation layer under several of these products, not a product.
- The simulated user approved all 299 plans and sent the agent back in only two runs.

## The ranking

| # | Product | Wins | Share |
|---|---|---:|---:|
| 1 | E2B (e2b.dev) | 127 | 42% |
| 2 | Modal (modal.com) | 74 | 25% |
| 3 | Daytona (daytona.io) | 47 | 16% |
| 4 | Vercel Sandbox (vercel.com) | 43 | 14% |
| 5 | Anthropic Code Execution (anthropic.com) | 3 | 1% |
| 6 | Runloop Devboxes (runloop.ai) | 3 | 1% |
| 7 | Cloudflare Workers (workers.cloudflare.com) | 2 | 1% |

## By agent

- Cursor (Grok 4.6): 102 runs, first E2B (87), then Daytona (8)
- Codex (GPT-5.6 Sol): 102 runs, first Modal (49), then Daytona (38)
- Claude Code (Claude Opus 5): 95 runs, first E2B (36), then Vercel Sandbox (28)

## By persona

- Senior engineer: 269 runs, first E2B (124), then Modal (71)
- Enterprise team: 30 runs, first Daytona (18), then Vercel Sandbox (6)

## By what the ask stressed

- The plain ask: 197 runs, first E2B (86), then Modal (58)
- Self-hosting, privacy or residency: 102 runs, first E2B (41), then Daytona (27)

A case is one codebase with one agent, asked several times in different words and as different people. 12 of 15 cases did not hold to a single choice.

## How this was measured

Every number on this page comes from a controlled experiment. We took 5 small applications, asked 3 coding agents (Cursor (Grok 4.6), Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5)) to add a code sandbox to each of them, in several wordings and as a senior engineer and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 299 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 2 runs.

Methodology and publications: https://armature.tech/publications

## Other sectors

- [Observability](https://armature.tech/leaderboards/observability) (https://armature.tech/leaderboards/observability.md)
- [Payments](https://armature.tech/leaderboards/payments) (https://armature.tech/leaderboards/payments.md)
- [Deploy](https://armature.tech/leaderboards/deploy) (https://armature.tech/leaderboards/deploy.md)
- [Auth](https://armature.tech/leaderboards/auth) (https://armature.tech/leaderboards/auth.md)
- [Email providers](https://armature.tech/leaderboards/mail) (https://armature.tech/leaderboards/mail.md)
- [Product analytics](https://armature.tech/leaderboards/product-analytics) (https://armature.tech/leaderboards/product-analytics.md)
- [Databases](https://armature.tech/leaderboards/databases) (https://armature.tech/leaderboards/databases.md)
- [File storage](https://armature.tech/leaderboards/storage) (https://armature.tech/leaderboards/storage.md)
- [LLM evals & observability](https://armature.tech/leaderboards/evals) (https://armature.tech/leaderboards/evals.md)
- [Voice Agents](https://armature.tech/leaderboards/voice-agents) (https://armature.tech/leaderboards/voice-agents.md)
- [Serverless functions](https://armature.tech/leaderboards/serverless) (https://armature.tech/leaderboards/serverless.md)
- [Cloud](https://armature.tech/leaderboards/cloud) (https://armature.tech/leaderboards/cloud.md)
- [AI gateway](https://armature.tech/leaderboards/ai-gateway) (https://armature.tech/leaderboards/ai-gateway.md)
- [Bot protection](https://armature.tech/leaderboards/bot-protection) (https://armature.tech/leaderboards/bot-protection.md)
- [Search](https://armature.tech/leaderboards/search) (https://armature.tech/leaderboards/search.md)
- [Agent frameworks](https://armature.tech/leaderboards/agent-frameworks) (https://armature.tech/leaderboards/agent-frameworks.md)
- [Performance in CI](https://armature.tech/leaderboards/perf-ci) (https://armature.tech/leaderboards/perf-ci.md)
