Agent leaderboards / All sectors / AI SRE
AI SRE: which products coding agents choose
Seer leads. The monitoring stack shapes the choice..
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked coding agents to add incident investigation and automatic fixes to six applications.
Sentry gives Seer a strong starting point
Sentry Seer is chosen in 83% of storefront runs and 63% of marketplace runs. Both applications already use Sentry.
Datadog keeps investigation inside its platform
Datadog Bits Investigation is chosen to investigate in 93% of commerce runs. The application already has Datadog logs, traces and monitors.
Resolve wins outside the Sentry and Datadog apps
Resolve AI is chosen to investigate in 67% of the Google Cloud fleet app runs and 45% of the CloudWatch analytics app runs.
The ranking179 runs
By agent, by persona, by wording
By agent
| Codex72 runs | Resolve AI · 25then Sentry Seer · 23 |
| Claude Code71 runs | Sentry Seer · 21then Claude Code GitHub Action · 11 |
| Cursor · Grok 4.636 runs | Cursor Automations · 14then Cursor Cloud Agents · 9 |
By persona
| Junior developer60 runs | Sentry Seer · 22then Claude Code GitHub Action · 11 |
| Enterprise team59 runs | Datadog Bits AI Dev Agent + Datadog Bits Investigation · 25then Resolve AI · 11 |
| Vibe coder30 runs | Sentry Seer · 25then Cursor Automations · 3 |
| Senior engineer30 runs | Resolve AI · 19then Cursor Automations · 3 |
A case is one codebase with one agent, asked several times in different words and as different people. 15 of 18 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 6 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to add automatic incident investigation and fixes to each of them, in several wordings and as a junior developer and enterprise team and vibe coder and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 179 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 87 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as MarkdownOther sectors
- Agent sandboxes
- Observability
- Payments
- Deploy
- Auth
- Email providers
- Product analytics
- Databases
- File storage
- LLM evals & observability
- Voice Agents
- Serverless functions
- Cloud
- AI gateway
- Bot protection
- Search
- Agent frameworks
- Performance in CI
- Document processing & OCR
- Usage-based billing
- Code review
- Internationalization
- Message queues
- Maps
- AI search
- In-app chat & calls
- Vector search