Agent leaderboards / All sectors / Performance in CI

Performance in CI: which products coding agents choose

Agents built the gate themselves in about 52% of runs.

216 runs8 apps3 agents3 personasupdated 2026-09-01

The interactive board, open on performance in ci. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add a performance gate to eight small apps, 216 runs in all. We varied the wording and asked as different people. When a product did win, JMH led with about 10%.

13

The language of the codebase set the pick

hyperfine won 13 times, all of them in the Rust CLI tool. PHPBench won 12, all in the PHP government portal. BenchmarkDotNet won seven, all in the .NET field ops app.

18 of 24

Rewording the ask changed the answer

A case here is one codebase with one agent, asked several times in different words and as different people. It counts as flipped when those runs didn't all land on the same product. 18 of 24 cases flipped.

121 vs 3

Named often, chosen almost never

CodSpeed came up in 121 runs and won three. Bencher came up in 83 runs and won one.

  • The simulated user approved all 216 plans, and sent the agent back at least once in 13 runs.
  • It never refused to approve until the agent named a specific product.
  • Enterprise team runs picked JMH 22 times, out of 60 runs.
  • Claude Code split almost evenly at the top, five wins for Autocannon and five for JMH.
  • Junior developer had only 12 runs, and no product came out on top there.
Explore every run in the interactive board

The ranking216 runs

ProductWinsShare
1 Built in-houseoutcome 113 52%
2 JMHopenjdk.org 22 10%
3 pytest-benchmarkpytest-benchmark.readthedocs.io 15 7%
4 Autocannongithub.com 14 6%
5 hyperfinegithub.com 13 6%
6 PHPBenchphpbench.readthedocs.io 12 6%
7 Grafana k6k6.io 8 4%
8 BenchmarkDotNetbenchmarkdotnet.org 7 3%
9 CodSpeedcodspeed.io 3 1%
10 Tinybenchgithub.com 3 1%
11 Vitest Benchvitest.dev 2 1%
12 perfgategithub.com 1 0%
13 Benchstatgithub.com 1 0%
14 Gungraungithub.com 1 0%
15 Bencherbencher.dev 1 0%

By agent, by persona, by wording

By agent

Cursor · Grok 4.672 runsJMH · 6then pytest-benchmark · 4
Claude Code · Claude Opus 572 runsAutocannon · 5then JMH · 5
Codex · GPT-5.6 Sol72 runsJMH · 11then hyperfine · 7

By persona

Senior engineer144 runspytest-benchmark · 15then Autocannon · 14
Enterprise team60 runsJMH · 22then PHPBench · 12
Junior developer12 runs

By what the ask stressed

The plain ask216 runsJMH · 22then pytest-benchmark · 15

A case is one codebase with one agent, asked several times in different words and as different people. 18 of 24 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 8 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add a performance gate to each of them, in several wordings and as a senior engineer and enterprise team and junior developer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 216 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 13 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown