Agent leaderboards / All sectors / Observability

Observability: which products coding agents choose

Sentry took about 37% of picks across 11 small apps.

360 runs11 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on observability. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add observability to 11 small apps, 360 runs in all. We varied the wording and asked as four different people. Grafana came second at about 14%, and the rest of the picks spread thin across a dozen more products.

30 of 120

The runner-up changed with the agent

Claude Code picked Grafana 30 times in 120 runs. Codex put New Relic second with 15. Cursor chose Sentry in 60 of its 120 runs.

20 of 36

One persona put a different product on top

Enterprise teams chose Grafana in 20 of 36 runs. Vibe coders picked Better Stack five times. Junior developers picked Sentry in all 12 of their runs.

28 of 33

Wording changed the answer in most cases

Take one codebase and one agent, then ask in different words. In 28 of 33 such cases the runs did not all land on the same product.

0 of 139

Named often, chosen never

Prometheus came up 139 times and won nothing. It is the open-source metrics engine inside a Grafana setup, named as part of that stack rather than as a product to choose.

  • The agents wrote it themselves in 22 runs, about 6%.
  • The simulated user approved all 360 plans, and sent the agent back at least once in 47 of them.
  • Once it refused to approve until the agent named a specific product.
  • AWS X-Ray won 11 times, all of them in the FastAPI SaaS API.
Explore every run in the interactive board

The ranking360 runs

ProductWinsShare
1 Sentrysentry.io 133 37%
2 Grafanagrafana.com 52 14%
3 Amazon CloudWatchaws.amazon.com 30 8%
4 Built in-houseoutcome 22 6%
5 Checklychecklyhq.com 17 5%
6 Datadogdatadoghq.com 16 4%
7 Better Stackbetterstack.com 15 4%
8 New Relicnewrelic.com 15 4%
9 Honeycombhoneycomb.io 13 4%
10 AWS X-Rayaws.amazon.com 11 3%
11 GlitchTipglitchtip.com 9 3%
12 Azure Monitorazure.microsoft.com 6 2%
13 Azure Application Insightsazure.microsoft.com 6 2%
14 Axiomaxiom.co 5 1%
15 SigNozsignoz.io 4 1%
16 Amazon CloudWatch + AWS X-Ray 3 1%
17 Bugsinkbugsink.com 3 1%

By agent, by persona, by wording

By agent

Cursor · Grok 4.6120 runsSentry · 60then Grafana · 10
Claude Code · Claude Opus 5120 runsSentry · 44then Grafana · 30
Codex · GPT-5.6 Sol120 runsSentry · 29then New Relic · 15

By persona

Senior engineer300 runsSentry · 116then Grafana · 32
Enterprise team36 runsGrafana · 20then Datadog · 7
Junior developer12 runsSentry · 12
Vibe coder12 runsBetter Stack · 5then Sentry · 1

By what the ask stressed

The plain ask348 runsSentry · 133then Grafana · 52

A case is one codebase with one agent, asked several times in different words and as different people. 28 of 33 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add observability to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 47 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown