# Do coding agents recommend Promptfoo?

> Promptfoo was chosen in 13% of 288 judged evals sessions, ranking second. Measured with Claude Code, Codex and Cursor.

Source: https://armature.tech/library/do-coding-agents-recommend-promptfoo
Published: 2026-09-03
Publisher: Armature, Inc. (https://armature.tech)

---

> Promptfoo was chosen in 13% of 288 judged evals sessions, ranking second. It was also raised as a candidate in 91 further sessions without being chosen.

This page reports what happened when Claude Code, Codex and Cursor had to solve a problem in evals inside a realistic codebase. Not what a chat assistant says about Promptfoo. What an agent actually installed.

## The numbers

| | |
| --- | --- |
| Category | Evals |
| Sessions in the category | 288 |
| Sessions where Promptfoo was chosen | 38 |
| Install share | 13% |
| Rank in category | 2 of 9 |
| Codebases it won in | 2 |
| Raised as a candidate, not chosen | 91 |
| Chosen when considered | 29% |
| Site | promptfoo.dev |

## By agent

The three agents land within 12 points of each other on Promptfoo, which is closer than most products in this experiment manage.

| Agent | Sessions | Chose Promptfoo | Share |
| --- | --- | --- | --- |
| Claude Code | 96 | 18 | 19% |
| Codex | 96 | 7 | 7% |
| Cursor | 96 | 13 | 14% |

## By who was asking

Promptfoo does much better with one kind of buyer than another. It won 25% of sessions asked as senior engineer and 1% of those asked as junior developer.

| Who is asking | Sessions | Chose Promptfoo | Share |
| --- | --- | --- | --- |
| Junior developer | 145 | 2 | 1% |
| Senior engineer | 143 | 36 | 25% |

## What Promptfoo was up against

The full ranking in evals, from the same sessions:

| # | Product | Runs won | Share |
| --- | --- | --- | --- |
| 1 | Langfuse | 97 | 34% |
| 2 | Built in-house (no product adopted) | 84 | 29% |
| 3 | Promptfoo **(this page)** | 38 | 13% |
| 4 | Arize Phoenix | 23 | 8% |
| 5 | Braintrust | 18 | 6% |
| 6 | Inspect AI | 12 | 4% |
| 7 | LangSmith | 8 | 3% |
| 8 | Helicone | 6 | 2% |

## What this means

Promptfoo was raised in 91 sessions and chosen in 38. That ratio is balanced enough that the ceiling is presence rather than integration: the product converts reasonably when it is on the table, and it is not on the table often enough.

## Where these numbers come from

The 288 sessions in evals are part of a published set of 5,292, run with real coding agents inside realistic codebases and judged blind. The full method is on one page: [how we measured this](/library/how-we-measured-this).

Every evals run can be replayed on [the board](/leaderboards/evals).

If you work on Promptfoo: the judge recorded a reason for every session where it was raised and passed over. Those reasons are in the transcripts.

<!-- generated by scripts/write-data-pages.mjs -->

## Common questions

### Do coding agents recommend Promptfoo?

Yes. Promptfoo was chosen in 38 of the 288 judged sessions in evals, a 13% install share, ranking second in its category.

### Does Claude Code recommend Promptfoo?

In 18 of the 96 sessions in evals run with Claude Code, which is 19%.

### Do different coding agents treat Promptfoo differently?

Yes, and by a wide margin. Claude Code chose it in 19% of its runs and Codex in 7%.

### How was this measured?

Real coding agents at pinned versions were run in sandboxes inside 51 realistic codebases and asked to solve real tasks. A simulated project owner approved or questioned each recommendation before any code was written, and a judge from a model family that builds none of the agents read every session blind.

### How often is Promptfoo considered but not chosen?

It was raised as a candidate in 91 sessions without being chosen, and chosen in 38. That is a 29% conversion from considered to chosen.

## Read next

- [How to get picked for evals by coding agents](https://armature.tech/library/evals-coding-agents-playbook) (Markdown: https://armature.tech/library/evals-coding-agents-playbook.md)
- [Do coding agents recommend Langfuse?](https://armature.tech/library/do-coding-agents-recommend-langfuse) (Markdown: https://armature.tech/library/do-coding-agents-recommend-langfuse.md)
- [Do coding agents recommend Arize Phoenix?](https://armature.tech/library/do-coding-agents-recommend-arize-phoenix) (Markdown: https://armature.tech/library/do-coding-agents-recommend-arize-phoenix.md)
- [Do Claude Code, Codex and Cursor pick the same tools?](https://armature.tech/library/do-claude-code-and-codex-agree) (Markdown: https://armature.tech/library/do-claude-code-and-codex-agree.md)

---

Armature helps software products get discovered and used by coding agents.
Service: https://armature.tech/discoverability · Results: https://armature.tech/leaderboards/sectors · Contact: contact@armature.tech
