Agent leaderboards / All sectors / AI SRE

AI SRE: which products coding agents choose

Seer leads. The monitoring stack shapes the choice..

179 runs6 apps3 agents4 personasupdated 2026-09-14

The interactive board, open on ai sre. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked coding agents to add incident investigation and automatic fixes to six applications.

Sentry gives Seer a strong starting point

Sentry Seer is chosen in 83% of storefront runs and 63% of marketplace runs. Both applications already use Sentry.

Datadog keeps investigation inside its platform

Datadog Bits Investigation is chosen to investigate in 93% of commerce runs. The application already has Datadog logs, traces and monitors.

Resolve wins outside the Sentry and Datadog apps

Resolve AI is chosen to investigate in 67% of the Google Cloud fleet app runs and 45% of the CloudWatch analytics app runs.

Explore every run in the interactive board

The ranking179 runs

ProductWinsShare
1 Sentry Seersentry.io 47 26%
2 Resolve AIresolve.ai 35 20%
3 Datadog Bits AI Dev Agent + Datadog Bits Investigation 26 15%
4 Cursor Automationscursor.com 15 8%
5 Claude Code GitHub Actionclaude.com 11 6%
6 Cursor Cloud Agentscursor.com 9 5%
7 incident.io AI SREincident.io 4 2%
8 Cursor Automations + Cursor Cloud Agents 3 2%
9 Grafana Assistant Investigationsgrafana.com 3 2%
10 Clericcleric.ai 3 2%
11 AWS DevOps Agentaws.amazon.com 3 2%
12 Middleware OpsAImiddleware.io 2 1%
13 Claude Code GitHub Action + NeuBird AI 2 1%
14 Datadog Bits Investigationdatadoghq.com 2 1%
15 Claude Code GitHub Action + Cleric 2 1%
16 Claude Code GitHub Action + Resolve AI 2 1%
17 Vercel Agentvercel.com 1 1%
18 OpenAI Codexopenai.com 1 1%
19 Claude Managed Agentsclaude.com 1 1%
20 Kestrelusekestrel.ai 1 1%
21 Gemini Cloud Assist Investigations + Resolve AI 1 1%
22 Claude Code GitHub Action + Datadog Bits Investigation 1 1%
23 Claude Code GitHub Action + Traversal 1 1%
24 Deeptracedeeptrace.com 1 1%
25 Claude Code GitHub Action + HolmesGPT 1 1%
26 AWS DevOps Agent + Kiro 1 1%

By agent, by persona, by wording

By agent

Codex72 runsResolve AI · 25then Sentry Seer · 23
Claude Code71 runsSentry Seer · 21then Claude Code GitHub Action · 11
Cursor · Grok 4.636 runsCursor Automations · 14then Cursor Cloud Agents · 9

By persona

Junior developer60 runsSentry Seer · 22then Claude Code GitHub Action · 11
Enterprise team59 runsDatadog Bits AI Dev Agent + Datadog Bits Investigation · 25then Resolve AI · 11
Vibe coder30 runsSentry Seer · 25then Cursor Automations · 3
Senior engineer30 runsResolve AI · 19then Cursor Automations · 3

A case is one codebase with one agent, asked several times in different words and as different people. 15 of 18 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 6 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to add automatic incident investigation and fixes to each of them, in several wordings and as a junior developer and enterprise team and vibe coder and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 179 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 87 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown