Agent leaderboards / All sectors / LLM evals & observability
LLM evals & observability: which products coding agents choose
Langfuse won about 34% of runs, with in-house code second.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add LLM evals to two small apps, over 288 runs. We asked in different words and as two different people. Promptfoo came next at about 13%, and the rest of the field was thin.
Privacy asks changed the order
When the ask mentioned self-hosting, privacy or data residency, Arize Phoenix won 23 of those 48 runs. Langfuse won two of them.
The persona changed the pick
Junior developers chose Langfuse in 61 of their 145 runs. With senior engineers it split almost evenly, 36 runs for Promptfoo and 36 for Langfuse.
Rewording the ask moved every case
A case is one codebase with one agent, asked the same thing several ways. All six cases flipped, meaning the runs in them did not all land on one product.
Named often, chosen rarely
LangSmith came up in 209 runs and was picked in 8 of them. OpenAI Evals came up 72 times and was never picked.
- The agents wrote evals themselves in about 29% of runs, more than any product but the leader.
- All three agents led with Langfuse, but Codex put Braintrust second with 17 wins.
- Inspect AI took all 12 of its wins in the Python AI analyst app, a FastAPI spreadsheet assistant.
- Helicone won six times, all of them in the TypeScript report builder.
- The simulated user approved all 288 plans and sent an agent back once.
The ranking288 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Langfuselangfuse.com | 97 | 34% | |
| 2 | Built in-houseoutcome | 84 | 29% | |
| 3 | Promptfoopromptfoo.dev | 38 | 13% | |
| 4 | Arize Phoenixarize.com | 23 | 8% | |
| 5 | Braintrustbraintrust.dev | 18 | 6% | |
| 6 | Inspect AIinspect.aisi.org.uk | 12 | 4% | |
| 7 | LangSmithsmith.langchain.com | 8 | 3% | |
| 8 | Heliconehelicone.ai | 6 | 2% | |
| 9 | pydantic-evalspydantic.dev | 1 | 0% | |
| 10 | DeepEvalconfident-ai.com | 1 | 0% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.696 runs | Langfuse · 32then Promptfoo · 13 |
| Claude Code · Claude Opus 596 runs | Langfuse · 36then Promptfoo · 18 |
| Codex · GPT-5.6 Sol96 runs | Langfuse · 29then Braintrust · 17 |
By persona
| Junior developer145 runs | Langfuse · 61then LangSmith · 6 |
| Senior engineer143 runs | Promptfoo · 36then Langfuse · 36 |
By what the ask stressed
| The plain ask192 runs | Langfuse · 78then Promptfoo · 14 |
| Volume and cost at scale48 runs | Langfuse · 17then Promptfoo · 11 |
| Self-hosting, privacy or residency48 runs | Arize Phoenix · 23then Promptfoo · 13 |
A case is one codebase with one agent, asked several times in different words and as different people. 6 of 6 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 2 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add LLM evals to each of them, in several wordings and as a junior developer and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 288 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 1 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown