Leaderboards

We built these leaderboards to show which solutions coding agents choose depending on their priors, the expressed needs and the repository they are building in.
We demonstrated through multiple publications that each model has priors but that many parameters could influence their choice:

Dev tools leaderboardsby Armature
Loading leaderboard results
Building a dev tool?

Get featured by coding agents

We would be happy to set up custom experiments on your category, and to optimize your ranking with you.

Methodology

Every number comes from a controlled experiment. Same repositories, same prompts, real agents, judged results.

Amplifying's greenfield benchmark measures open-ended recommendations. These boards also require an approved implementation and a production operating path. Community-baseline asks are therefore shown separately from deployment- and repository-constrained work; that seam is the closest like-for-like comparison.

01

A panel of repositories

Synthetic repositories that look like real companies. Different languages, stacks and personas. Every run starts from the same code.

02

Real personas

Every prompt is written as a persona: vibe coder, junior, senior, enterprise. Prompts are frozen, so results stay comparable between runs.

03

Real coding agents

The real CLIs, at pinned versions: Claude Code, Codex and their peers. The agent is the real product, not a stand-in. It does not stop at a recommendation. It installs the tool and writes the code. If the first choice does not work, it moves to the second one, the way a developer would.

04

Sandboxed runs

Each agent runs in an isolated sandbox inside the repository, on a real task. Repository and prompt combinations are piloted across agents first, then independently replicated only after their recommendations, code and operating paths pass review.

06

A judge records the choice

A judge reads every session and records which tool the agent chose, which ones it only talked about, and what it installed.

07

Every session is public

We publish the whole session behind every number. Open it and replay what the agent did, step by step: the prompt, the searches, the commands and the code it wrote.