We built these leaderboards to show which solutions coding agents choose depending on their priors, the expressed needs and the repository they are building in.
We demonstrated through multiple publications that each model has priors but that many parameters could influence their choice:
The dashboard is easier to explore in fullscreen.
We would be happy to set up custom experiments on your category, and to optimize your ranking with you.
Every number comes from a controlled experiment. Same repositories, same prompts, real agents, judged results.
Amplifying's greenfield benchmark measures open-ended recommendations. These boards also require an approved implementation and a production operating path. Community-baseline asks are therefore shown separately from deployment- and repository-constrained work; that seam is the closest like-for-like comparison.
Synthetic repositories that look like real companies. Different languages, stacks and personas. Every run starts from the same code.
Every prompt is written as a persona: vibe coder, junior, senior, enterprise. Prompts are frozen, so results stay comparable between runs.
The real CLIs, at pinned versions: Claude Code, Codex and their peers. The agent is the real product, not a stand-in. It does not stop at a recommendation. It installs the tool and writes the code. If the first choice does not work, it moves to the second one, the way a developer would.
Each agent runs in an isolated sandbox inside the repository, on a real task. Repository and prompt combinations are piloted across agents first, then independently replicated only after their recommendations, code and operating paths pass review.
On this board the person asking is a model, not a human. Gemini 3.7 Flash plays the owner of the project: it sends the ask, reads the recommendation, judges it against written guidance for this sector, and only then asks for the code. Every one of those exchanges is published in the run, in full.
A judge reads every session and records which tool the agent chose, which ones it only talked about, and what it installed.
We publish the whole session behind every number. Open it and replay what the agent did, step by step: the prompt, the searches, the commands and the code it wrote.