Agent leaderboards / All sectors / Observability
Observability: which products coding agents choose
Sentry won about 37% of runs and led all three agents.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add observability to 11 small apps, 360 runs in all. We varied the wording, and asked as four different people. Grafana came second at about 14%, and everything after it stayed under 10%.
The enterprise ask changed the leader
Asked as an enterprise team, the agents picked Grafana in 20 of 36 runs. Datadog came next there with seven. With senior engineers, Sentry led again, at 116 of 300 runs.
Wording moved the answer in most cases
A case is one codebase and one agent, asked again in different words. In 28 of 33 cases the runs did not all land on the same product.
- Prometheus was named in 139 runs and never chosen; it comes as part of a Grafana setup, so that isn't a loss.
- The agents wrote it themselves in 22 runs, about 6% of the total.
- The simulated user approved every plan, but sent the agent back at least once in 47 runs.
The ranking360 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Sentrysentry.io | 133 | 37% | |
| 2 | Grafanagrafana.com | 52 | 14% | |
| 3 | Amazon CloudWatchaws.amazon.com | 30 | 8% | |
| 4 | Built in-houseoutcome | 22 | 6% | |
| 5 | Checklychecklyhq.com | 17 | 5% | |
| 6 | Datadogdatadoghq.com | 16 | 4% | |
| 7 | Better Stackbetterstack.com | 15 | 4% | |
| 8 | New Relicnewrelic.com | 15 | 4% | |
| 9 | Honeycombhoneycomb.io | 13 | 4% | |
| 10 | AWS X-Rayaws.amazon.com | 11 | 3% | |
| 11 | GlitchTipglitchtip.com | 9 | 3% | |
| 12 | Azure Monitorazure.microsoft.com | 6 | 2% | |
| 13 | Azure Application Insightsazure.microsoft.com | 6 | 2% | |
| 14 | Axiomaxiom.co | 5 | 1% | |
| 15 | SigNozsignoz.io | 4 | 1% | |
| 16 | Amazon CloudWatch + AWS X-Ray | 3 | 1% | |
| 17 | Bugsinkbugsink.com | 3 | 1% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.6120 runs | Sentry · 60then Grafana · 10 |
| Claude Code · Claude Opus 5120 runs | Sentry · 44then Grafana · 30 |
| Codex · GPT-5.6 Sol120 runs | Sentry · 29then New Relic · 15 |
By persona
| Senior engineer300 runs | Sentry · 116then Grafana · 32 |
| Enterprise team36 runs | Grafana · 20then Datadog · 7 |
| Junior developer12 runs | Sentry · 12 |
| Vibe coder12 runs | Better Stack · 5then Sentry · 1 |
By what the ask stressed
| The plain ask348 runs | Sentry · 133then Grafana · 52 |
A case is one codebase with one agent, asked several times in different words and as different people. 28 of 33 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add observability to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 47 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown