Agent leaderboards / All sectors / Performance in CI

Performance in CI: which products coding agents choose

Agents wrote the gate themselves in about 52% of runs.

216 runs8 apps3 agents3 personasupdated 2026-09-01

The interactive board, open on performance in ci. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add a performance gate to eight small apps, 216 runs in all. We varied the wording and asked as three different people. Among the named products JMH came first, at about 10% of runs.

13 wins

The codebase language picked the tool

hyperfine won 13 times, all in the Rust CLI tool. PHPBench took 12, all in the PHP government portal. The seven wins for BenchmarkDotNet were all in the .NET field ops app.

18 of 24

Most cases did not settle on one product

A case here is one codebase with one agent, asked several times in different words. In 18 of 24 cases the runs did not all land on the same product.

22 of 60

One persona carried the leader

The enterprise team ran 60 times, and all 22 wins for JMH came there. Under senior engineers, pytest-benchmark led with 15.

3 of 121

Named in many plans, chosen in few

CodSpeed came up 121 times and was picked three times. Bencher came up 83 times and was picked once.

  • The simulated user approved all 216 plans, and sent the agent back at least once in 13 of them.
  • Codex picked JMH 11 times, more than any other agent picked any product.
  • Claude Code split its top pick evenly, five wins each for Autocannon and JMH.
  • The junior developer ran only 12 times, and no product came out on top there.
Explore every run in the interactive board

The ranking216 runs

ProductWinsShare
1 Built in-houseoutcome 113 52%
2 JMHopenjdk.org 22 10%
3 pytest-benchmarkpytest-benchmark.readthedocs.io 15 7%
4 Autocannongithub.com 14 6%
5 hyperfinegithub.com 13 6%
6 PHPBenchphpbench.readthedocs.io 12 6%
7 Grafana k6k6.io 8 4%
8 BenchmarkDotNetbenchmarkdotnet.org 7 3%
9 CodSpeedcodspeed.io 3 1%
10 Tinybenchgithub.com 3 1%
11 Vitest Benchvitest.dev 2 1%
12 perfgategithub.com 1 0%
13 Benchstatgithub.com 1 0%
14 Gungraungithub.com 1 0%
15 Bencherbencher.dev 1 0%

By agent, by persona, by wording

By agent

Cursor · Grok 4.672 runsJMH · 6then pytest-benchmark · 4
Claude Code · Claude Opus 572 runsAutocannon · 5then JMH · 5
Codex · GPT-5.6 Sol72 runsJMH · 11then hyperfine · 7

By persona

Senior engineer144 runspytest-benchmark · 15then Autocannon · 14
Enterprise team60 runsJMH · 22then PHPBench · 12
Junior developer12 runs

By what the ask stressed

The plain ask216 runsJMH · 22then pytest-benchmark · 15

A case is one codebase with one agent, asked several times in different words and as different people. 18 of 24 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 8 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add a performance gate to each of them, in several wordings and as a senior engineer and enterprise team and junior developer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 216 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 13 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown