# How to get picked for evals by coding agents

> Langfuse took 34% of 288 judged evals sessions. What the numbers say a vendor in this category should do.

Source: https://armature.tech/library/evals-coding-agents-playbook
Published: 2026-09-03
Publisher: Armature, Inc. (https://armature.tech)

---

If you sell eval tools, this page is the part of the market no dashboard shows you: what a coding agent does when a developer asks for evals and never compares vendors.

The numbers come from 288 judged sessions with Claude Code, Codex and Cursor, spread across 2 realistic codebases, with every session read by a judge.

**Read the spread before the totals.** 288 sessions is a lot of runs, and 2 codebases are not a lot of variety. With two codebases, a result is as much about those two projects as about the category. The panel had 2 with the evals seam open, and running on repositories where the request does not make sense would have been worse than running on fewer. Weigh what follows accordingly.

## What coding agents choose for evals

> Across 288 judged sessions, **Langfuse** was chosen most often, in **34%** of runs. Promptfoo was second with 13%. In 29% of runs the agent wrote the code itself and adopted no product at all.

| # | Product | Runs won | Share |
| --- | --- | --- | --- |
| 1 | Langfuse | 97 | 34% |
| 2 | Built in-house (no product adopted) | 84 | 29% |
| 3 | Promptfoo | 38 | 13% |
| 4 | Arize Phoenix | 23 | 8% |
| 5 | Braintrust | 18 | 6% |
| 6 | Inspect AI | 12 | 4% |
| 7 | LangSmith | 8 | 3% |
| 8 | Helicone | 6 | 2% |
| 9 | pydantic-evals | 1 | 0% |
| 10 | DeepEval | 1 | 0% |

Full board, every run replayable: [the evals leaderboard](/leaderboards/evals).

## What the shape of this category means

The leader takes only 34% of runs. This category is genuinely open and the ordering can be moved.

With the top product at 34%, evals is decided in the moment, from what the agent reads and what it finds in the repository. Nothing is locked in, which is the best situation a vendor can be in and the one where the work pays fastest.

The order here is set by the quality of what an agent can read and by whether your product is already present in the codebase. Both are things you can change.

## The agents do not agree with each other

In this category all three agents put Langfuse first, which is less common than it sounds: across the eighteen categories we measured, Claude Code and Codex disagreed on the leader in nine of them.

| Agent | Runs | Picked most often |
| --- | --- | --- |
| Claude Code | 96 | Langfuse (36) |
| Codex | 96 | Langfuse (29) |
| Cursor | 96 | Langfuse (32) |

Even where they agree, they get there differently. Codex ran a web search in 53% of decision runs and Claude Code in 1.6%, so what you publish reaches one of them in half its evals sessions and the other in almost none.

## Who is asking changes the answer

Every request in this experiment was written as a specific kind of person. In this category the leader changes with the person.

| Who is asking | Runs | Picked most often |
| --- | --- | --- |
| Junior developer | 145 | Langfuse |
| Senior engineer | 143 | Promptfoo |

That is 2 different products winning evals for 2 kinds of buyer, out of the same 288 sessions. Nobody here is winning evals. They are each winning one kind of buyer.

If you sell to more than one of them, you need pages for each. See [how to win the enterprise persona](/library/how-to-win-the-enterprise-persona).

## What you are really competing against

In 29% of runs, the agent wrote the code itself. That makes hand-written code the largest single competitor in this category, larger than most vendors in the table above.

This changes the job of your content. Before you argue that you are better than another vendor, you have to argue that the problem is harder than it looks. What breaks at volume. Which edge cases cost a weekend. What the maintenance actually costs after six months.

That argument has to exist as a page an agent can read, with specifics and numbers. "It is harder than you think" is not an argument. "Here are the four failure modes and what each one costs" is.

## Considered, and never chosen

Because the judge records every product an agent raised and not only the one it picked, this board also shows who kept reaching the shortlist and losing. In evals the clearest case is OpenAI Evals: on the table in 72 sessions, chosen in none.

| Product | Raised in | Chosen in |
| --- | --- | --- |
| [OpenAI Evals](/library/do-coding-agents-recommend-openai-evals) | 72 sessions | 0 |
| Ragas | 34 sessions | 0 |
| Traceloop | 23 sessions | 0 |

Being rejected is a better position than being unknown, and a cheaper one to fix. The product is already in the agent's head and on the list. Whatever ended those 129 sessions is recorded in each transcript, one reason at a time.

## What to do about it in evals

1. **Argue that the problem is harder than it looks, before you argue that you are better than a rival.** 29% of evals sessions ended in hand-written code, so in roughly one session in 3 no vendor was in the running at all. Write the failure modes and the year-two maintenance cost, with numbers, in the documentation rather than the blog.

2. **Aim at second place first.** Langfuse holds 34% and Promptfoo holds 13%. The gap between the default and the field is where the reachable sessions are.

3. **Report install share per persona, not as one number.** In evals the leader changes with who is asking, so a blended figure averages markets that behave differently.

The work that applies to every category rather than to this one is written up separately: [audit your documentation](/library/audit-your-docs-for-coding-agents), [write a quickstart an agent can follow](/library/write-a-quickstart-an-agent-can-follow), and [how to measure install share](/library/how-to-measure-install-share).

## Every eval tool on this board

One page per product, with its install share, the per-agent split, and how often it was raised without being chosen.

- [Do coding agents recommend Langfuse?](/library/do-coding-agents-recommend-langfuse) — chosen in 34% of sessions
- [Do coding agents recommend Promptfoo?](/library/do-coding-agents-recommend-promptfoo) — chosen in 13% of sessions
- [Do coding agents recommend Arize Phoenix?](/library/do-coding-agents-recommend-arize-phoenix) — chosen in 8% of sessions
- [Do coding agents recommend Braintrust?](/library/do-coding-agents-recommend-braintrust) — chosen in 6% of sessions
- [Do coding agents recommend Inspect AI?](/library/do-coding-agents-recommend-inspect-ai) — chosen in 4% of sessions
- [Do coding agents recommend LangSmith?](/library/do-coding-agents-recommend-langsmith) — chosen in 3% of sessions
- [Do coding agents recommend OpenAI Evals?](/library/do-coding-agents-recommend-openai-evals) — raised in 72 sessions, chosen in none

## Where these numbers come from

288 judged sessions in evals across 2 codebases, part of a published set of 5,292. Real coding agents at pinned versions, in sandboxes, inside realistic codebases, with a simulated project owner in the loop and a blind judge on every session. The full method is on one page: [how we measured this](/library/how-we-measured-this).

Every evals run can be replayed on [the board](/leaderboards/evals).

<!-- generated by scripts/write-data-pages.mjs -->

## Common questions

### How many codebases is this based on?

288 judged sessions across 2 realistic codebases. A category only runs on repositories where its seam is open, so coverage differs: some categories ran on more than ten codebases and some on two.

### What eval tool do coding agents choose?

Across 288 judged sessions, Langfuse was chosen most often, in 34% of runs. Promptfoo was second with 13%. The result changes by agent and by who is asking.

### Do Claude Code and Codex pick the same eval tool?

Yes. All three agents we tested put Langfuse first in this category, which is unusual: they disagree in half of the categories we measured.

### How often do agents build evals themselves instead of installing something?

In 29% of runs the agent wrote the code itself rather than adopting a product. That makes hand-written code one of the strongest competitors in the category.

### How can a vendor improve its position here?

Make the quickstart run when pasted, state the current version on the documentation page, use one name across product, package and import, write pages for the symptoms users describe rather than only the category name, and get into the repository through templates and framework integrations.

### Which eval tools do agents consider but never choose?

OpenAI Evals (raised in 72 sessions, chosen in none), Ragas (raised in 34 sessions, chosen in none), Traceloop (raised in 23 sessions, chosen in none). Being considered and not chosen is a different problem from being unknown, and it is usually fixable.

## Read next

- [Agent discoverability: the complete guide](https://armature.tech/library/agent-discoverability) (Markdown: https://armature.tech/library/agent-discoverability.md)
- [How coding agents choose tools](https://armature.tech/library/how-coding-agents-choose-tools) (Markdown: https://armature.tech/library/how-coding-agents-choose-tools.md)
- [Do coding agents recommend Langfuse?](https://armature.tech/library/do-coding-agents-recommend-langfuse) (Markdown: https://armature.tech/library/do-coding-agents-recommend-langfuse.md)
- [Do coding agents recommend Promptfoo?](https://armature.tech/library/do-coding-agents-recommend-promptfoo) (Markdown: https://armature.tech/library/do-coding-agents-recommend-promptfoo.md)
- [Do coding agents recommend Arize Phoenix?](https://armature.tech/library/do-coding-agents-recommend-arize-phoenix) (Markdown: https://armature.tech/library/do-coding-agents-recommend-arize-phoenix.md)

---

Armature helps software products get discovered and used by coding agents.
Service: https://armature.tech/discoverability · Results: https://armature.tech/leaderboards/sectors · Contact: contact@armature.tech
