Agent leaderboards / All sectors / LLM evals & observability

LLM evals & observability: which products coding agents choose

Langfuse won about 34% of runs, with in-house code second.

288 runs2 apps3 agents2 personasupdated 2026-09-02

The interactive board, open on llm evals & observability. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add LLM evals to two small apps, over 288 runs. We asked in different words and as two different people. Promptfoo came next at about 13%, and the rest of the field was thin.

23 of 48

Privacy asks changed the order

When the ask mentioned self-hosting, privacy or data residency, Arize Phoenix won 23 of those 48 runs. Langfuse won two of them.

36 vs 36

The persona changed the pick

Junior developers chose Langfuse in 61 of their 145 runs. With senior engineers it split almost evenly, 36 runs for Promptfoo and 36 for Langfuse.

6 of 6

Rewording the ask moved every case

A case is one codebase with one agent, asked the same thing several ways. All six cases flipped, meaning the runs in them did not all land on one product.

8 of 209

Named often, chosen rarely

LangSmith came up in 209 runs and was picked in 8 of them. OpenAI Evals came up 72 times and was never picked.

  • The agents wrote evals themselves in about 29% of runs, more than any product but the leader.
  • All three agents led with Langfuse, but Codex put Braintrust second with 17 wins.
  • Inspect AI took all 12 of its wins in the Python AI analyst app, a FastAPI spreadsheet assistant.
  • Helicone won six times, all of them in the TypeScript report builder.
  • The simulated user approved all 288 plans and sent an agent back once.
Explore every run in the interactive board

The ranking288 runs

ProductWinsShare
1 Langfuselangfuse.com 97 34%
2 Built in-houseoutcome 84 29%
3 Promptfoopromptfoo.dev 38 13%
4 Arize Phoenixarize.com 23 8%
5 Braintrustbraintrust.dev 18 6%
6 Inspect AIinspect.aisi.org.uk 12 4%
7 LangSmithsmith.langchain.com 8 3%
8 Heliconehelicone.ai 6 2%
9 pydantic-evalspydantic.dev 1 0%
10 DeepEvalconfident-ai.com 1 0%

By agent, by persona, by wording

By agent

Cursor · Grok 4.696 runsLangfuse · 32then Promptfoo · 13
Claude Code · Claude Opus 596 runsLangfuse · 36then Promptfoo · 18
Codex · GPT-5.6 Sol96 runsLangfuse · 29then Braintrust · 17

By persona

Junior developer145 runsLangfuse · 61then LangSmith · 6
Senior engineer143 runsPromptfoo · 36then Langfuse · 36

By what the ask stressed

The plain ask192 runsLangfuse · 78then Promptfoo · 14
Volume and cost at scale48 runsLangfuse · 17then Promptfoo · 11
Self-hosting, privacy or residency48 runsArize Phoenix · 23then Promptfoo · 13

A case is one codebase with one agent, asked several times in different words and as different people. 6 of 6 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 2 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add LLM evals to each of them, in several wordings and as a junior developer and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 288 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 1 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown