Agent leaderboards / All sectors / Observability

Observability: which products coding agents choose

Sentry won about 37% of runs and led all three agents.

360 runs11 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on observability. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add observability to 11 small apps, 360 runs in all. We varied the wording, and asked as four different people. Grafana came second at about 14%, and everything after it stayed under 10%.

20 of 36

The enterprise ask changed the leader

Asked as an enterprise team, the agents picked Grafana in 20 of 36 runs. Datadog came next there with seven. With senior engineers, Sentry led again, at 116 of 300 runs.

28 of 33

Wording moved the answer in most cases

A case is one codebase and one agent, asked again in different words. In 28 of 33 cases the runs did not all land on the same product.

60 vs 29

One agent leaned twice as hard as another

Cursor chose Sentry in 60 of its 120 runs. Codex chose it 29 times, with New Relic next at 15. Claude Code picked Grafana 30 times.

11 wins

Some products won inside one codebase only

AWS X-Ray took all 11 of its wins in one FastAPI SaaS API. GlitchTip won nine times, all in a Rails claims operations app.

  • Prometheus was named in 139 runs and never chosen; it comes as part of a Grafana setup, so that isn't a loss.
  • The agents wrote it themselves in 22 runs, about 6% of the total.
  • The simulated user approved every plan, but sent the agent back at least once in 47 runs.
Explore every run in the interactive board

The ranking360 runs

ProductWinsShare
1 Sentrysentry.io 133 37%
2 Grafanagrafana.com 52 14%
3 Amazon CloudWatchaws.amazon.com 30 8%
4 Built in-houseoutcome 22 6%
5 Checklychecklyhq.com 17 5%
6 Datadogdatadoghq.com 16 4%
7 Better Stackbetterstack.com 15 4%
8 New Relicnewrelic.com 15 4%
9 Honeycombhoneycomb.io 13 4%
10 AWS X-Rayaws.amazon.com 11 3%
11 GlitchTipglitchtip.com 9 3%
12 Azure Monitorazure.microsoft.com 6 2%
13 Azure Application Insightsazure.microsoft.com 6 2%
14 Axiomaxiom.co 5 1%
15 SigNozsignoz.io 4 1%
16 Amazon CloudWatch + AWS X-Ray 3 1%
17 Bugsinkbugsink.com 3 1%

By agent, by persona, by wording

By agent

Cursor · Grok 4.6120 runsSentry · 60then Grafana · 10
Claude Code · Claude Opus 5120 runsSentry · 44then Grafana · 30
Codex · GPT-5.6 Sol120 runsSentry · 29then New Relic · 15

By persona

Senior engineer300 runsSentry · 116then Grafana · 32
Enterprise team36 runsGrafana · 20then Datadog · 7
Junior developer12 runsSentry · 12
Vibe coder12 runsBetter Stack · 5then Sentry · 1

By what the ask stressed

The plain ask348 runsSentry · 133then Grafana · 52

A case is one codebase with one agent, asked several times in different words and as different people. 28 of 33 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add observability to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 47 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown