Agent leaderboards / All sectors / Observability
Observability: which products coding agents choose
Sentry took about 37% of picks across 11 small apps.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add observability to 11 small apps, 360 runs in all. We varied the wording and asked as four different people. Grafana came second at about 14%, and the rest of the picks spread thin across a dozen more products.
The runner-up changed with the agent
Claude Code picked Grafana 30 times in 120 runs. Codex put New Relic second with 15. Cursor chose Sentry in 60 of its 120 runs.
One persona put a different product on top
Enterprise teams chose Grafana in 20 of 36 runs. Vibe coders picked Better Stack five times. Junior developers picked Sentry in all 12 of their runs.
Wording changed the answer in most cases
Take one codebase and one agent, then ask in different words. In 28 of 33 such cases the runs did not all land on the same product.
Named often, chosen never
Prometheus came up 139 times and won nothing. It is the open-source metrics engine inside a Grafana setup, named as part of that stack rather than as a product to choose.
- The agents wrote it themselves in 22 runs, about 6%.
- The simulated user approved all 360 plans, and sent the agent back at least once in 47 of them.
- Once it refused to approve until the agent named a specific product.
- AWS X-Ray won 11 times, all of them in the FastAPI SaaS API.
The ranking360 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Sentrysentry.io | 133 | 37% | |
| 2 | Grafanagrafana.com | 52 | 14% | |
| 3 | Amazon CloudWatchaws.amazon.com | 30 | 8% | |
| 4 | Built in-houseoutcome | 22 | 6% | |
| 5 | Checklychecklyhq.com | 17 | 5% | |
| 6 | Datadogdatadoghq.com | 16 | 4% | |
| 7 | Better Stackbetterstack.com | 15 | 4% | |
| 8 | New Relicnewrelic.com | 15 | 4% | |
| 9 | Honeycombhoneycomb.io | 13 | 4% | |
| 10 | AWS X-Rayaws.amazon.com | 11 | 3% | |
| 11 | GlitchTipglitchtip.com | 9 | 3% | |
| 12 | Azure Monitorazure.microsoft.com | 6 | 2% | |
| 13 | Azure Application Insightsazure.microsoft.com | 6 | 2% | |
| 14 | Axiomaxiom.co | 5 | 1% | |
| 15 | SigNozsignoz.io | 4 | 1% | |
| 16 | Amazon CloudWatch + AWS X-Ray | 3 | 1% | |
| 17 | Bugsinkbugsink.com | 3 | 1% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.6120 runs | Sentry · 60then Grafana · 10 |
| Claude Code · Claude Opus 5120 runs | Sentry · 44then Grafana · 30 |
| Codex · GPT-5.6 Sol120 runs | Sentry · 29then New Relic · 15 |
By persona
| Senior engineer300 runs | Sentry · 116then Grafana · 32 |
| Enterprise team36 runs | Grafana · 20then Datadog · 7 |
| Junior developer12 runs | Sentry · 12 |
| Vibe coder12 runs | Better Stack · 5then Sentry · 1 |
By what the ask stressed
| The plain ask348 runs | Sentry · 133then Grafana · 52 |
A case is one codebase with one agent, asked several times in different words and as different people. 28 of 33 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add observability to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 47 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown