# Observability: which products coding agents choose

> Sentry won about 37% of runs and led all three agents.

Source: https://armature.tech/leaderboards/observability (Armature agent leaderboards). 360 runs, 11 apps, 3 agents, 4 personas, updated 2026-09-02. Interactive board with every run: https://armature.tech/leaderboards#app/observability

## Key learnings

We asked three coding agents to add observability to 11 small apps, 360 runs in all. We varied the wording, and asked as four different people. Grafana came second at about 14%, and everything after it stayed under 10%.

### The enterprise ask changed the leader (20 of 36)

Asked as an enterprise team, the agents picked Grafana in 20 of 36 runs. Datadog came next there with seven. With senior engineers, Sentry led again, at 116 of 300 runs.

### Wording moved the answer in most cases (28 of 33)

A case is one codebase and one agent, asked again in different words. In 28 of 33 cases the runs did not all land on the same product.

### One agent leaned twice as hard as another (60 vs 29)

Cursor chose Sentry in 60 of its 120 runs. Codex chose it 29 times, with New Relic next at 15. Claude Code picked Grafana 30 times.

### Some products won inside one codebase only (11 wins)

AWS X-Ray took all 11 of its wins in one FastAPI SaaS API. GlitchTip won nine times, all in a Rails claims operations app.

Smaller learnings:

- Prometheus was named in 139 runs and never chosen; it comes as part of a Grafana setup, so that isn't a loss.
- The agents wrote it themselves in 22 runs, about 6% of the total.
- The simulated user approved every plan, but sent the agent back at least once in 47 runs.

## The ranking

| # | Product | Wins | Share |
|---|---|---:|---:|
| 1 | Sentry (sentry.io) | 133 | 37% |
| 2 | Grafana (grafana.com) | 52 | 14% |
| 3 | Amazon CloudWatch (aws.amazon.com) | 30 | 8% |
| 4 | Built in-house (outcome) | 22 | 6% |
| 5 | Checkly (checklyhq.com) | 17 | 5% |
| 6 | Datadog (datadoghq.com) | 16 | 4% |
| 7 | Better Stack (betterstack.com) | 15 | 4% |
| 8 | New Relic (newrelic.com) | 15 | 4% |
| 9 | Honeycomb (honeycomb.io) | 13 | 4% |
| 10 | AWS X-Ray (aws.amazon.com) | 11 | 3% |
| 11 | GlitchTip (glitchtip.com) | 9 | 3% |
| 12 | Azure Monitor (azure.microsoft.com) | 6 | 2% |
| 13 | Azure Application Insights (azure.microsoft.com) | 6 | 2% |
| 14 | Axiom (axiom.co) | 5 | 1% |
| 15 | SigNoz (signoz.io) | 4 | 1% |
| 16 | Amazon CloudWatch + AWS X-Ray | 3 | 1% |
| 17 | Bugsink (bugsink.com) | 3 | 1% |

## By agent

- Cursor (Grok 4.6): 120 runs, first Sentry (60), then Grafana (10)
- Claude Code (Claude Opus 5): 120 runs, first Sentry (44), then Grafana (30)
- Codex (GPT-5.6 Sol): 120 runs, first Sentry (29), then New Relic (15)

## By persona

- Senior engineer: 300 runs, first Sentry (116), then Grafana (32)
- Enterprise team: 36 runs, first Grafana (20), then Datadog (7)
- Junior developer: 12 runs, first Sentry (12)
- Vibe coder: 12 runs, first Better Stack (5), then Sentry (1)

## By what the ask stressed

- The plain ask: 348 runs, first Sentry (133), then Grafana (52)

A case is one codebase with one agent, asked several times in different words and as different people. 28 of 33 cases did not hold to a single choice.

## How this was measured

Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol)) to add observability to each of them, in several wordings and as a senior engineer and enterprise team and junior developer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 360 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 47 runs.

Methodology and publications: https://armature.tech/publications

## Other sectors

- [Agent sandboxes](https://armature.tech/leaderboards/sandboxes) (https://armature.tech/leaderboards/sandboxes.md)
- [Payments](https://armature.tech/leaderboards/payments) (https://armature.tech/leaderboards/payments.md)
- [Deploy](https://armature.tech/leaderboards/deploy) (https://armature.tech/leaderboards/deploy.md)
- [Auth](https://armature.tech/leaderboards/auth) (https://armature.tech/leaderboards/auth.md)
- [Email providers](https://armature.tech/leaderboards/mail) (https://armature.tech/leaderboards/mail.md)
- [Product analytics](https://armature.tech/leaderboards/product-analytics) (https://armature.tech/leaderboards/product-analytics.md)
- [Databases](https://armature.tech/leaderboards/databases) (https://armature.tech/leaderboards/databases.md)
- [File storage](https://armature.tech/leaderboards/storage) (https://armature.tech/leaderboards/storage.md)
- [LLM evals & observability](https://armature.tech/leaderboards/evals) (https://armature.tech/leaderboards/evals.md)
- [Voice Agents](https://armature.tech/leaderboards/voice-agents) (https://armature.tech/leaderboards/voice-agents.md)
- [Serverless functions](https://armature.tech/leaderboards/serverless) (https://armature.tech/leaderboards/serverless.md)
- [Cloud](https://armature.tech/leaderboards/cloud) (https://armature.tech/leaderboards/cloud.md)
- [AI gateway](https://armature.tech/leaderboards/ai-gateway) (https://armature.tech/leaderboards/ai-gateway.md)
- [Bot protection](https://armature.tech/leaderboards/bot-protection) (https://armature.tech/leaderboards/bot-protection.md)
- [Search](https://armature.tech/leaderboards/search) (https://armature.tech/leaderboards/search.md)
- [Agent frameworks](https://armature.tech/leaderboards/agent-frameworks) (https://armature.tech/leaderboards/agent-frameworks.md)
- [Performance in CI](https://armature.tech/leaderboards/perf-ci) (https://armature.tech/leaderboards/perf-ci.md)
