Agent leaderboards / All sectors / LLM evals & observability
LLM evals & observability: which products coding agents choose
Langfuse led the evals boards with about 34% of runs.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add LLM evals to two small apps, 288 runs in all. Each run used a different wording, and spoke as one of two people. Langfuse came out on top, and the agents skipped products and wrote the evals themselves in about 29% of runs.
Self-hosting changed which product won
When the ask mentioned self-hosting, privacy or data residency, Arize Phoenix took 23 of those 48 runs. Langfuse won two of them. In the plain ask, with no such constraint, Langfuse won 78 of 192.
The two personas did not agree
For junior developers, Langfuse won 61 runs. For senior engineers it tied with Promptfoo at 36 wins each.
Every case flipped on wording alone
In each case the codebase and the agent stayed the same and only the words of the ask changed. All six cases came back with more than one product picked.
One product named constantly, picked rarely
LangSmith came up 209 times across the runs and was chosen eight times. OpenAI Evals was named 72 times and never chosen.
- Codex put Braintrust second with 17 wins, while the other two agents put Promptfoo there.
- All 12 wins for Inspect AI came in the Python AI analyst app, a FastAPI and pandas codebase.
- Helicone won six runs, all of them in the Express report builder app.
- The simulated user approved all 288 plans and sent an agent back only once.
The ranking288 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Langfuselangfuse.com | 97 | 34% | |
| 2 | Built in-houseoutcome | 84 | 29% | |
| 3 | Promptfoopromptfoo.dev | 38 | 13% | |
| 4 | Arize Phoenixarize.com | 23 | 8% | |
| 5 | Braintrustbraintrust.dev | 18 | 6% | |
| 6 | Inspect AIinspect.aisi.org.uk | 12 | 4% | |
| 7 | LangSmithsmith.langchain.com | 8 | 3% | |
| 8 | Heliconehelicone.ai | 6 | 2% | |
| 9 | pydantic-evalspydantic.dev | 1 | 0% | |
| 10 | DeepEvalconfident-ai.com | 1 | 0% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.696 runs | Langfuse · 32then Promptfoo · 13 |
| Claude Code · Claude Opus 596 runs | Langfuse · 36then Promptfoo · 18 |
| Codex · GPT-5.6 Sol96 runs | Langfuse · 29then Braintrust · 17 |
By persona
| Junior developer145 runs | Langfuse · 61then LangSmith · 6 |
| Senior engineer143 runs | Promptfoo · 36then Langfuse · 36 |
By what the ask stressed
| The plain ask192 runs | Langfuse · 78then Promptfoo · 14 |
| Volume and cost at scale48 runs | Langfuse · 17then Promptfoo · 11 |
| Self-hosting, privacy or residency48 runs | Arize Phoenix · 23then Promptfoo · 13 |
A case is one codebase with one agent, asked several times in different words and as different people. 6 of 6 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 2 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add LLM evals to each of them, in several wordings and as a junior developer and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 288 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 1 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown