# AI SRE: which products coding agents choose

> Seer leads. The monitoring stack shapes the choice..

Source: https://armature.tech/leaderboards/ai-sre (Armature agent leaderboards). 179 runs, 6 apps, 3 agents, 4 personas, updated 2026-09-14. Interactive board with every run: https://armature.tech/leaderboards#app/ai-sre

## Key learnings

We asked coding agents to add incident investigation and automatic fixes to six applications.

### Sentry gives Seer a strong starting point

Sentry Seer is chosen in 83% of storefront runs and 63% of marketplace runs. Both applications already use Sentry.

### Datadog keeps investigation inside its platform

Datadog Bits Investigation is chosen to investigate in 93% of commerce runs. The application already has Datadog logs, traces and monitors.

### Resolve wins outside the Sentry and Datadog apps

Resolve AI is chosen to investigate in 67% of the Google Cloud fleet app runs and 45% of the CloudWatch analytics app runs.

## The ranking

| # | Product | Wins | Share |
|---|---|---:|---:|
| 1 | Sentry Seer (sentry.io) | 47 | 26% |
| 2 | Resolve AI (resolve.ai) | 35 | 20% |
| 3 | Datadog Bits AI Dev Agent + Datadog Bits Investigation | 26 | 15% |
| 4 | Cursor Automations (cursor.com) | 15 | 8% |
| 5 | Claude Code GitHub Action (claude.com) | 11 | 6% |
| 6 | Cursor Cloud Agents (cursor.com) | 9 | 5% |
| 7 | incident.io AI SRE (incident.io) | 4 | 2% |
| 8 | Cursor Automations + Cursor Cloud Agents | 3 | 2% |
| 9 | Grafana Assistant Investigations (grafana.com) | 3 | 2% |
| 10 | Cleric (cleric.ai) | 3 | 2% |
| 11 | AWS DevOps Agent (aws.amazon.com) | 3 | 2% |
| 12 | Middleware OpsAI (middleware.io) | 2 | 1% |
| 13 | Claude Code GitHub Action + NeuBird AI | 2 | 1% |
| 14 | Datadog Bits Investigation (datadoghq.com) | 2 | 1% |
| 15 | Claude Code GitHub Action + Cleric | 2 | 1% |
| 16 | Claude Code GitHub Action + Resolve AI | 2 | 1% |
| 17 | Vercel Agent (vercel.com) | 1 | 1% |
| 18 | OpenAI Codex (openai.com) | 1 | 1% |
| 19 | Claude Managed Agents (claude.com) | 1 | 1% |
| 20 | Kestrel (usekestrel.ai) | 1 | 1% |
| 21 | Gemini Cloud Assist Investigations + Resolve AI | 1 | 1% |
| 22 | Claude Code GitHub Action + Datadog Bits Investigation | 1 | 1% |
| 23 | Claude Code GitHub Action + Traversal | 1 | 1% |
| 24 | Deeptrace (deeptrace.com) | 1 | 1% |
| 25 | Claude Code GitHub Action + HolmesGPT | 1 | 1% |
| 26 | AWS DevOps Agent + Kiro | 1 | 1% |

## By agent

- Codex: 72 runs, first Resolve AI (25), then Sentry Seer (23)
- Claude Code: 71 runs, first Sentry Seer (21), then Claude Code GitHub Action (11)
- Cursor (Grok 4.6): 36 runs, first Cursor Automations (14), then Cursor Cloud Agents (9)

## By persona

- Junior developer: 60 runs, first Sentry Seer (22), then Claude Code GitHub Action (11)
- Enterprise team: 59 runs, first Datadog Bits AI Dev Agent + Datadog Bits Investigation (25), then Resolve AI (11)
- Vibe coder: 30 runs, first Sentry Seer (25), then Cursor Automations (3)
- Senior engineer: 30 runs, first Resolve AI (19), then Cursor Automations (3)

A case is one codebase with one agent, asked several times in different words and as different people. 15 of 18 cases did not hold to a single choice.

## How this was measured

Every number on this page comes from a controlled experiment. We took 6 small applications, asked 3 coding agents (Codex, Claude Code, Cursor (Grok 4.6)) to add automatic incident investigation and fixes to each of them, in several wordings and as a junior developer and enterprise team and vibe coder and senior engineer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 179 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 87 runs.

Methodology and publications: https://armature.tech/publications

If you sell in this sector, what these numbers mean for a vendor: https://armature.tech/library/ai-sre-coding-agents-playbook (Markdown: https://armature.tech/library/ai-sre-coding-agents-playbook.md)

## Other sectors

- [Agent sandboxes](https://armature.tech/leaderboards/sandboxes) (https://armature.tech/leaderboards/sandboxes.md)
- [Observability](https://armature.tech/leaderboards/observability) (https://armature.tech/leaderboards/observability.md)
- [Payments](https://armature.tech/leaderboards/payments) (https://armature.tech/leaderboards/payments.md)
- [Deploy](https://armature.tech/leaderboards/deploy) (https://armature.tech/leaderboards/deploy.md)
- [Auth](https://armature.tech/leaderboards/auth) (https://armature.tech/leaderboards/auth.md)
- [Email providers](https://armature.tech/leaderboards/mail) (https://armature.tech/leaderboards/mail.md)
- [Product analytics](https://armature.tech/leaderboards/product-analytics) (https://armature.tech/leaderboards/product-analytics.md)
- [Databases](https://armature.tech/leaderboards/databases) (https://armature.tech/leaderboards/databases.md)
- [File storage](https://armature.tech/leaderboards/storage) (https://armature.tech/leaderboards/storage.md)
- [LLM evals & observability](https://armature.tech/leaderboards/evals) (https://armature.tech/leaderboards/evals.md)
- [Voice Agents](https://armature.tech/leaderboards/voice-agents) (https://armature.tech/leaderboards/voice-agents.md)
- [Serverless functions](https://armature.tech/leaderboards/serverless) (https://armature.tech/leaderboards/serverless.md)
- [Cloud](https://armature.tech/leaderboards/cloud) (https://armature.tech/leaderboards/cloud.md)
- [AI gateway](https://armature.tech/leaderboards/ai-gateway) (https://armature.tech/leaderboards/ai-gateway.md)
- [Bot protection](https://armature.tech/leaderboards/bot-protection) (https://armature.tech/leaderboards/bot-protection.md)
- [Search](https://armature.tech/leaderboards/search) (https://armature.tech/leaderboards/search.md)
- [Agent frameworks](https://armature.tech/leaderboards/agent-frameworks) (https://armature.tech/leaderboards/agent-frameworks.md)
- [Performance in CI](https://armature.tech/leaderboards/perf-ci) (https://armature.tech/leaderboards/perf-ci.md)
- [Document processing & OCR](https://armature.tech/leaderboards/document-processing) (https://armature.tech/leaderboards/document-processing.md)
- [Usage-based billing](https://armature.tech/leaderboards/usage-based-billing) (https://armature.tech/leaderboards/usage-based-billing.md)
- [Code review](https://armature.tech/leaderboards/code-review) (https://armature.tech/leaderboards/code-review.md)
- [Internationalization](https://armature.tech/leaderboards/internationalization) (https://armature.tech/leaderboards/internationalization.md)
- [Message queues](https://armature.tech/leaderboards/message-queues) (https://armature.tech/leaderboards/message-queues.md)
- [Maps](https://armature.tech/leaderboards/maps) (https://armature.tech/leaderboards/maps.md)
- [AI search](https://armature.tech/leaderboards/ai-search) (https://armature.tech/leaderboards/ai-search.md)
- [In-app chat & calls](https://armature.tech/leaderboards/in-app-communication) (https://armature.tech/leaderboards/in-app-communication.md)
- [Vector search](https://armature.tech/leaderboards/vector-search) (https://armature.tech/leaderboards/vector-search.md)
