Playbooks

How to get picked for code sandboxes by coding agents

E2B took 42% of 299 judged code sandboxes sessions. What the numbers say a vendor in this category should do.

Published September 3, 2026 Read as Markdown

If you sell code sandboxes, this page is the part of the market no dashboard shows you: what a coding agent does when a developer asks for code sandboxes and never compares vendors.

The numbers come from 299 judged sessions with Claude Code, Codex and Cursor, spread across 5 realistic codebases, with every session read by a judge.

Read the spread before the totals. 299 sessions is a lot of runs, and 5 codebases are not a lot of variety. A category that ran on a handful of projects carries their particulars with it. The panel had 5 with the code sandboxes seam open, and running on repositories where the request does not make sense would have been worse than running on fewer. Weigh what follows accordingly.

What coding agents choose for code sandboxes

Across 299 judged sessions, E2B was chosen most often, in 42% of runs. Modal was second with 25%.

#ProductRuns wonShare
1E2B12742%
2Modal7425%
3Daytona4716%
4Vercel Sandbox4314%
5Anthropic Code Execution31%
6Runloop Devboxes31%
7Cloudflare Workers21%

Full board, every run replayable: the code sandboxes leaderboard.

What the shape of this category means

The leader takes 42% of runs and there is a real second place. The category has a default but it is not settled.

E2B at 42% with Modal at 25% is a default with a real challenger behind it. The agent is choosing rather than reaching, which means the inputs it uses can move the answer.

The work is to be the easiest correct answer: a quickstart that runs when pasted, documentation that states the current version, and pages that answer the exact configuration questions agents search for.

The agents do not agree with each other

In this category the three agents we ran put different products first.

AgentRunsPicked most often
Claude Code95E2B (36)
Codex102Modal (49)
Cursor102E2B (87)

That split decides where a vendor spends. Codex ran a web search in 53% of decision runs and Claude Code in 1.6%, so the pages you publish are live in half of Codex's code sandboxes sessions and almost none of Claude Code's. Taking Modal's position with Codex is a content problem. Taking E2B's with Claude Code is a repository problem.

Who is asking changes the answer

Every request in this experiment was written as a specific kind of person. In this category the leader changes with the person.

Who is askingRunsPicked most often
Senior engineer269E2B
Enterprise team30Daytona

That is 2 different products winning code sandboxes for 2 kinds of buyer, out of the same 299 sessions. Nobody here is winning code sandboxes. They are each winning one kind of buyer.

If you sell to more than one of them, you need pages for each. See how to win the enterprise persona.

What you are really competing against

In this category agents never chose to build it themselves. Every session ended with a product. That is good news: you are in a straight vendor comparison, and the levers that work are the ones you control.

Considered, and never chosen

Because the judge records every product an agent raised and not only the one it picked, this board also shows who kept reaching the shortlist and losing. In code sandboxes the clearest case is CodeSandbox SDK: on the table in 70 sessions, chosen in none.

ProductRaised inChosen in
CodeSandbox SDK70 sessions0
Fly Machines58 sessions0
Blaxel Sandboxes50 sessions0
AWS CodeBuild42 sessions0
Deno35 sessions0

Being rejected is a better position than being unknown, and a cheaper one to fix. The product is already in the agent's head and on the list. Whatever ended those 255 sessions is recorded in each transcript, one reason at a time.

What to do about it in code sandboxes

  1. Skip the build-versus-buy argument. No code sandboxes session in this experiment ended with the agent writing its own implementation. Every one adopted a product, so the whole contest is against the other names in the table above.
  1. Aim at second place first. E2B holds 42% and Modal holds 25%. The gap between the default and the field is where the reachable sessions are.
  1. Report install share per persona, not as one number. In code sandboxes the leader changes with who is asking, so a blended figure averages markets that behave differently.
  1. Measure per agent. Claude Code put E2B first, Codex put Modal first, Cursor put E2B first. A blended number for code sandboxes describes a market that does not exist.

The work that applies to every category rather than to this one is written up separately: audit your documentation, write a quickstart an agent can follow, and how to measure install share.

Every code sandbox on this board

One page per product, with its install share, the per-agent split, and how often it was raised without being chosen.

Where these numbers come from

299 judged sessions in code sandboxes across 5 codebases, part of a published set of 5,292. Real coding agents at pinned versions, in sandboxes, inside realistic codebases, with a simulated project owner in the loop and a blind judge on every session. The full method is on one page: how we measured this.

Every code sandboxes run can be replayed on the board.

<!-- generated by scripts/write-data-pages.mjs -->

Common questions

How many codebases is this based on?

299 judged sessions across 5 realistic codebases. A category only runs on repositories where its seam is open, so coverage differs: some categories ran on more than ten codebases and some on two.

What code sandbox do coding agents choose?

Across 299 judged sessions, E2B was chosen most often, in 42% of runs. Modal was second with 25%. The result changes by agent and by who is asking.

Do Claude Code and Codex pick the same code sandbox?

No. Claude Code picked E2B, Codex picked Modal, Cursor picked E2B. Measuring one agent tells you about part of the market only.

How often do agents build code sandboxes themselves instead of installing something?

Never, in this category. Every one of the sessions ended with the agent adopting a product rather than writing the code itself.

How can a vendor improve its position here?

Make the quickstart run when pasted, state the current version on the documentation page, use one name across product, package and import, write pages for the symptoms users describe rather than only the category name, and get into the repository through templates and framework integrations.

Which code sandboxes do agents consider but never choose?

CodeSandbox SDK (raised in 70 sessions, chosen in none), Fly Machines (raised in 58 sessions, chosen in none), Blaxel Sandboxes (raised in 50 sessions, chosen in none), AWS CodeBuild (raised in 42 sessions, chosen in none), Deno (raised in 35 sessions, chosen in none). Being considered and not chosen is a different problem from being unknown, and it is usually fixable.

Where this comes from

Armature ran 5,292 judged sessions with Claude Code, Codex and Cursor inside 51 realistic codebases, and published every run. The numbers on this page come from that work.

Read next

All library pages