# Do coding agents recommend BenchmarkDotNet?

> BenchmarkDotNet was chosen in 3% of 305 judged performance testing sessions, ranking seventh.

Source: https://armature.tech/library/do-coding-agents-recommend-benchmarkdotnet
Published: 2026-09-23
Publisher: Armature, Inc. (https://armature.tech)

---

> BenchmarkDotNet was chosen in 3% of 305 judged performance testing sessions, ranking seventh. It was also raised as a candidate in 23 further sessions without being chosen.

This page reports what happened when Claude Code, Codex, Cursor, Grok Build CLI and Muse Code had to solve a problem in performance testing inside a realistic codebase. Not what a chat assistant says about BenchmarkDotNet. What an agent actually installed.

One thing to read first: every one of those wins came from a **single codebase**. That is a result about one repository rather than about performance testing in general, and the splits below cannot separate the two. Treat them as a description of that repository.

## The numbers

| | |
| --- | --- |
| Category | Performance testing |
| Sessions in the category | 305 |
| Sessions where BenchmarkDotNet was chosen | 8 |
| Install share | 3% |
| Rank in category | 7 of 15 |
| Codebases it won in | 1 |
| Raised as a candidate, not chosen | 23 |
| Chosen when considered | 26% |
| Site | benchmarkdotnet.org |

## By agent

With 8 wins spread across 5 agents, the rates below are small numbers and a difference between them is not yet a finding. They are here because the direction is worth knowing, not because the gap is established.

| Agent | Sessions | Chose BenchmarkDotNet | Share |
| --- | --- | --- | --- |
| Claude Code | 89 | 3 | 3% |
| Codex | 90 | 2 | 2% |
| Cursor | 90 | 3 | 3% |
| Grok Build CLI | 18 | 0 | 0% |
| Muse Code | 18 | 0 | 0% |

## By who was asking

BenchmarkDotNet performs similarly across the four kinds of buyer, from 0% to 4%. That is unusual: the category leader changed with the persona in 23 of the 29 categories we measured.

| Who is asking | Sessions | Chose BenchmarkDotNet | Share |
| --- | --- | --- | --- |
| Junior developer | 17 | 0 | 0% |
| Senior engineer | 203 | 8 | 4% |
| Enterprise team | 85 | 0 | 0% |

## What BenchmarkDotNet was up against

The full ranking in performance testing, from the same sessions:

| # | Product | Runs won | Share |
| --- | --- | --- | --- |
| 1 | Built in-house (no product adopted) | 166 | 54% |
| 2 | JMH | 28 | 9% |
| 3 | pytest-benchmark | 20 | 7% |
| 4 | Autocannon | 19 | 6% |
| 5 | hyperfine | 17 | 6% |
| 6 | PHPBench | 16 | 5% |
| 7 | Grafana k6 | 14 | 5% |
| 8 | BenchmarkDotNet **(this page)** | 8 | 3% |

## What this means

BenchmarkDotNet was raised in 23 sessions and chosen in 8. That ratio is balanced enough that the ceiling is presence rather than integration: the product converts reasonably when it is on the table, and it is not on the table often enough.

## Where these numbers come from

The 305 sessions in performance testing are part of a published set of 11,878, run with real coding agents inside realistic codebases and judged blind. The full method is on one page: [how we measured this](/library/how-we-measured-this).

Every performance testing run can be replayed on [the board](/leaderboards/perf-ci).

If you work on BenchmarkDotNet: the judge recorded a reason for every session where it was raised and passed over. Those reasons are in the transcripts.

<!-- generated by scripts/write-data-pages.mjs -->

## Common questions

### Do coding agents recommend BenchmarkDotNet?

Yes. BenchmarkDotNet was chosen in 8 of the 305 judged sessions in performance testing, a 3% install share, ranking seventh in its category.

### Does Claude Code recommend BenchmarkDotNet?

In 3 of the 89 sessions in performance testing run with Claude Code, which is 3%.

### Do different coding agents treat BenchmarkDotNet differently?

Not much. Claude Code, Codex and Cursor chose it at similar rates, between 2% and 3% of their runs.

### How was this measured?

Real coding agents at pinned versions were run in sandboxes inside 91 realistic codebases and asked to solve real tasks. A simulated project owner approved or questioned each recommendation before any code was written, and a judge from a model family that builds none of the agents read every session blind.

### How often is BenchmarkDotNet considered but not chosen?

It was raised as a candidate in 23 sessions without being chosen, and chosen in 8. That is a 26% conversion from considered to chosen.

## Read next

- [How to get picked for performance testing by coding agents](https://armature.tech/library/perf-ci-coding-agents-playbook) (Markdown: https://armature.tech/library/perf-ci-coding-agents-playbook.md)
- [Do coding agents recommend JMH?](https://armature.tech/library/do-coding-agents-recommend-jmh) (Markdown: https://armature.tech/library/do-coding-agents-recommend-jmh.md)
- [Do coding agents recommend pytest-benchmark?](https://armature.tech/library/do-coding-agents-recommend-pytest-benchmark) (Markdown: https://armature.tech/library/do-coding-agents-recommend-pytest-benchmark.md)
- [Do Claude Code, Codex and Cursor pick the same tools?](https://armature.tech/library/do-claude-code-and-codex-agree) (Markdown: https://armature.tech/library/do-claude-code-and-codex-agree.md)

---

Armature helps software products get discovered and used by coding agents.
Service: https://armature.tech/discoverability · Results: https://armature.tech/leaderboards/sectors · Contact: contact@armature.tech
