Agent leaderboards / All sectors / LLM evals & observability

LLM evals & observability: which products coding agents choose

Langfuse led the evals boards with about 34% of runs.

288 runs2 apps3 agents2 personasupdated 2026-09-02

The interactive board, open on llm evals & observability. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add LLM evals to two small apps, 288 runs in all. Each run used a different wording, and spoke as one of two people. Langfuse came out on top, and the agents skipped products and wrote the evals themselves in about 29% of runs.

23 of 48

Self-hosting changed which product won

When the ask mentioned self-hosting, privacy or data residency, Arize Phoenix took 23 of those 48 runs. Langfuse won two of them. In the plain ask, with no such constraint, Langfuse won 78 of 192.

36 vs 36

The two personas did not agree

For junior developers, Langfuse won 61 runs. For senior engineers it tied with Promptfoo at 36 wins each.

6 of 6

Every case flipped on wording alone

In each case the codebase and the agent stayed the same and only the words of the ask changed. All six cases came back with more than one product picked.

209 mentions

One product named constantly, picked rarely

LangSmith came up 209 times across the runs and was chosen eight times. OpenAI Evals was named 72 times and never chosen.

  • Codex put Braintrust second with 17 wins, while the other two agents put Promptfoo there.
  • All 12 wins for Inspect AI came in the Python AI analyst app, a FastAPI and pandas codebase.
  • Helicone won six runs, all of them in the Express report builder app.
  • The simulated user approved all 288 plans and sent an agent back only once.
Explore every run in the interactive board

The ranking288 runs

ProductWinsShare
1 Langfuselangfuse.com 97 34%
2 Built in-houseoutcome 84 29%
3 Promptfoopromptfoo.dev 38 13%
4 Arize Phoenixarize.com 23 8%
5 Braintrustbraintrust.dev 18 6%
6 Inspect AIinspect.aisi.org.uk 12 4%
7 LangSmithsmith.langchain.com 8 3%
8 Heliconehelicone.ai 6 2%
9 pydantic-evalspydantic.dev 1 0%
10 DeepEvalconfident-ai.com 1 0%

By agent, by persona, by wording

By agent

Cursor · Grok 4.696 runsLangfuse · 32then Promptfoo · 13
Claude Code · Claude Opus 596 runsLangfuse · 36then Promptfoo · 18
Codex · GPT-5.6 Sol96 runsLangfuse · 29then Braintrust · 17

By persona

Junior developer145 runsLangfuse · 61then LangSmith · 6
Senior engineer143 runsPromptfoo · 36then Langfuse · 36

By what the ask stressed

The plain ask192 runsLangfuse · 78then Promptfoo · 14
Volume and cost at scale48 runsLangfuse · 17then Promptfoo · 11
Self-hosting, privacy or residency48 runsArize Phoenix · 23then Promptfoo · 13

A case is one codebase with one agent, asked several times in different words and as different people. 6 of 6 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 2 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add LLM evals to each of them, in several wordings and as a junior developer and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 288 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 1 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown