Agent leaderboards / All sectors / Performance in CI
Performance in CI: which products coding agents choose
Agents built the gate themselves in about 52% of runs.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add a performance gate to eight small apps, 216 runs in all. We varied the wording and asked as different people. When a product did win, JMH led with about 10%.
The language of the codebase set the pick
hyperfine won 13 times, all of them in the Rust CLI tool. PHPBench won 12, all in the PHP government portal. BenchmarkDotNet won seven, all in the .NET field ops app.
Rewording the ask changed the answer
A case here is one codebase with one agent, asked several times in different words and as different people. It counts as flipped when those runs didn't all land on the same product. 18 of 24 cases flipped.
- The simulated user approved all 216 plans, and sent the agent back at least once in 13 runs.
- It never refused to approve until the agent named a specific product.
- Enterprise team runs picked JMH 22 times, out of 60 runs.
- Claude Code split almost evenly at the top, five wins for Autocannon and five for JMH.
- Junior developer had only 12 runs, and no product came out on top there.
The ranking216 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Built in-houseoutcome | 113 | 52% | |
| 2 | JMHopenjdk.org | 22 | 10% | |
| 3 | pytest-benchmarkpytest-benchmark.readthedocs.io | 15 | 7% | |
| 4 | Autocannongithub.com | 14 | 6% | |
| 5 | hyperfinegithub.com | 13 | 6% | |
| 6 | PHPBenchphpbench.readthedocs.io | 12 | 6% | |
| 7 | Grafana k6k6.io | 8 | 4% | |
| 8 | BenchmarkDotNetbenchmarkdotnet.org | 7 | 3% | |
| 9 | CodSpeedcodspeed.io | 3 | 1% | |
| 10 | Tinybenchgithub.com | 3 | 1% | |
| 11 | Vitest Benchvitest.dev | 2 | 1% | |
| 12 | perfgategithub.com | 1 | 0% | |
| 13 | Benchstatgithub.com | 1 | 0% | |
| 14 | Gungraungithub.com | 1 | 0% | |
| 15 | Bencherbencher.dev | 1 | 0% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.672 runs | JMH · 6then pytest-benchmark · 4 |
| Claude Code · Claude Opus 572 runs | Autocannon · 5then JMH · 5 |
| Codex · GPT-5.6 Sol72 runs | JMH · 11then hyperfine · 7 |
By persona
| Senior engineer144 runs | pytest-benchmark · 15then Autocannon · 14 |
| Enterprise team60 runs | JMH · 22then PHPBench · 12 |
| Junior developer12 runs |
By what the ask stressed
| The plain ask216 runs | JMH · 22then pytest-benchmark · 15 |
A case is one codebase with one agent, asked several times in different words and as different people. 18 of 24 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 8 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add a performance gate to each of them, in several wordings and as a senior engineer and enterprise team and junior developer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 216 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 13 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown