# How to get picked for observability by coding agents

> Sentry took 37% of 360 judged observability sessions. What the numbers say a vendor in this category should do.

Source: https://armature.tech/library/observability-coding-agents-playbook
Published: 2026-09-03
Publisher: Armature, Inc. (https://armature.tech)

---

If you sell observability tools, this page is the part of the market no dashboard shows you: what a coding agent does when a developer asks for observability and never compares vendors.

The numbers come from 360 judged sessions with Claude Code, Codex and Cursor, spread across 11 realistic codebases, with every session read by a judge.


## What coding agents choose for observability

> Across 360 judged sessions, **Sentry** was chosen most often, in **37%** of runs. Grafana was second with 14%.

| # | Product | Runs won | Share |
| --- | --- | --- | --- |
| 1 | Sentry | 133 | 37% |
| 2 | Grafana | 52 | 14% |
| 3 | Amazon CloudWatch | 30 | 8% |
| 4 | Built in-house (no product adopted) | 22 | 6% |
| 5 | Checkly | 17 | 5% |
| 6 | Datadog | 16 | 4% |
| 7 | Better Stack | 15 | 4% |
| 8 | New Relic | 15 | 4% |
| 9 | Honeycomb | 13 | 4% |
| 10 | AWS X-Ray | 11 | 3% |

Full board, every run replayable: [the observability leaderboard](/leaderboards/observability).

## What the shape of this category means

The leader takes 37% of runs and there is a real second place. The category has a default but it is not settled.

Sentry at 37% with Grafana at 14% is a default with a real challenger behind it. The agent is choosing rather than reaching, which means the inputs it uses can move the answer.

The work is to be the easiest correct answer: a quickstart that runs when pasted, documentation that states the current version, and pages that answer the exact configuration questions agents search for.

## The agents do not agree with each other

In this category all three agents put Sentry first, which is less common than it sounds: across the eighteen categories we measured, Claude Code and Codex disagreed on the leader in nine of them.

| Agent | Runs | Picked most often |
| --- | --- | --- |
| Claude Code | 120 | Sentry (44) |
| Codex | 120 | Sentry (29) |
| Cursor | 120 | Sentry (60) |

Even where they agree, they get there differently. Codex ran a web search in 53% of decision runs and Claude Code in 1.6%, so what you publish reaches one of them in half its observability sessions and the other in almost none.

## Who is asking changes the answer

Every request in this experiment was written as a specific kind of person. In this category the leader changes with the person.

| Who is asking | Runs | Picked most often |
| --- | --- | --- |
| Vibe coder | 12 | Better Stack |
| Junior developer | 12 | Sentry |
| Senior engineer | 300 | Sentry |
| Enterprise team | 36 | Grafana |

That is 3 different products winning observability for 4 kinds of buyer, out of the same 360 sessions. Nobody here is winning observability. They are each winning one kind of buyer.

If you sell to more than one of them, you need pages for each. See [how to win the enterprise persona](/library/how-to-win-the-enterprise-persona).

## What you are really competing against

In 6% of runs the agent wrote the code itself rather than adopting a product. That is low enough that your competition is other vendors, but high enough to be worth watching.

## Considered, and never chosen

Because the judge records every product an agent raised and not only the one it picked, this board also shows who kept reaching the shortlist and losing. In observability the clearest case is Prometheus: on the table in 139 sessions, chosen in none.

| Product | Raised in | Chosen in |
| --- | --- | --- |
| [Prometheus](/library/do-coding-agents-recommend-prometheus) | 139 sessions | 0 |
| Pino | 43 sessions | 0 |
| UptimeRobot | 40 sessions | 0 |
| Highlight | 39 sessions | 0 |
| Jaeger | 30 sessions | 0 |

Being rejected is a better position than being unknown, and a cheaper one to fix. The product is already in the agent's head and on the list. Whatever ended those 291 sessions is recorded in each transcript, one reason at a time.

## What to do about it in observability

1. **Aim at second place first.** Sentry holds 37% and Grafana holds 14%. The gap between the default and the field is where the reachable sessions are.

2. **Pick which buyer you are for.** The same observability need written as a vibe coder landed on Better Stack, and written as an enterprise team landed on Grafana. Those are two markets, and the enterprise one needs pages containing the constraint words: audit log, data residency, retention, single sign-on. See [how to win the enterprise persona](/library/how-to-win-the-enterprise-persona).

The work that applies to every category rather than to this one is written up separately: [audit your documentation](/library/audit-your-docs-for-coding-agents), [write a quickstart an agent can follow](/library/write-a-quickstart-an-agent-can-follow), and [how to measure install share](/library/how-to-measure-install-share).

## Every observability tool on this board

One page per product, with its install share, the per-agent split, and how often it was raised without being chosen.

- [Do coding agents recommend Sentry?](/library/do-coding-agents-recommend-sentry) — chosen in 37% of sessions
- [Do coding agents recommend Grafana?](/library/do-coding-agents-recommend-grafana) — chosen in 14% of sessions
- [Do coding agents recommend Amazon CloudWatch?](/library/do-coding-agents-recommend-amazon-cloudwatch) — chosen in 8% of sessions
- [Do coding agents recommend Checkly?](/library/do-coding-agents-recommend-checkly) — chosen in 5% of sessions
- [Do coding agents recommend Datadog?](/library/do-coding-agents-recommend-datadog) — chosen in 4% of sessions
- [Do coding agents recommend Better Stack?](/library/do-coding-agents-recommend-better-stack) — chosen in 4% of sessions
- [Do coding agents recommend New Relic?](/library/do-coding-agents-recommend-new-relic) — chosen in 4% of sessions
- [Do coding agents recommend Honeycomb?](/library/do-coding-agents-recommend-honeycomb) — chosen in 4% of sessions
- [Do coding agents recommend AWS X-Ray?](/library/do-coding-agents-recommend-aws-x-ray) — chosen in 3% of sessions
- [Do coding agents recommend GlitchTip?](/library/do-coding-agents-recommend-glitchtip) — chosen in 3% of sessions
- [Do coding agents recommend OpenTelemetry?](/library/do-coding-agents-recommend-opentelemetry) — raised in 244 sessions, chosen in none
- [Do coding agents recommend Prometheus?](/library/do-coding-agents-recommend-prometheus) — raised in 139 sessions, chosen in none

## Where these numbers come from

360 judged sessions in observability across 11 codebases, part of a published set of 5,292. Real coding agents at pinned versions, in sandboxes, inside realistic codebases, with a simulated project owner in the loop and a blind judge on every session. The full method is on one page: [how we measured this](/library/how-we-measured-this).

Every observability run can be replayed on [the board](/leaderboards/observability).

<!-- generated by scripts/write-data-pages.mjs -->

## Common questions

### How many codebases is this based on?

360 judged sessions across 11 realistic codebases. A category only runs on repositories where its seam is open, so coverage differs: some categories ran on more than ten codebases and some on two.

### What observability tool do coding agents choose?

Across 360 judged sessions, Sentry was chosen most often, in 37% of runs. Grafana was second with 14%. The result changes by agent and by who is asking.

### Do Claude Code and Codex pick the same observability tool?

Yes. All three agents we tested put Sentry first in this category, which is unusual: they disagree in half of the categories we measured.

### How often do agents build observability themselves instead of installing something?

In 6% of runs the agent wrote the code itself rather than adopting a product.

### How can a vendor improve its position here?

Make the quickstart run when pasted, state the current version on the documentation page, use one name across product, package and import, write pages for the symptoms users describe rather than only the category name, and get into the repository through templates and framework integrations.

### Which observability tools do agents consider but never choose?

Prometheus (raised in 139 sessions, chosen in none), Pino (raised in 43 sessions, chosen in none), UptimeRobot (raised in 40 sessions, chosen in none), Highlight (raised in 39 sessions, chosen in none), Jaeger (raised in 30 sessions, chosen in none). Being considered and not chosen is a different problem from being unknown, and it is usually fixable.

## Read next

- [Agent discoverability: the complete guide](https://armature.tech/library/agent-discoverability) (Markdown: https://armature.tech/library/agent-discoverability.md)
- [How coding agents choose tools](https://armature.tech/library/how-coding-agents-choose-tools) (Markdown: https://armature.tech/library/how-coding-agents-choose-tools.md)
- [Do coding agents recommend Sentry?](https://armature.tech/library/do-coding-agents-recommend-sentry) (Markdown: https://armature.tech/library/do-coding-agents-recommend-sentry.md)
- [Do coding agents recommend Grafana?](https://armature.tech/library/do-coding-agents-recommend-grafana) (Markdown: https://armature.tech/library/do-coding-agents-recommend-grafana.md)
- [Do coding agents recommend Amazon CloudWatch?](https://armature.tech/library/do-coding-agents-recommend-amazon-cloudwatch) (Markdown: https://armature.tech/library/do-coding-agents-recommend-amazon-cloudwatch.md)

---

Armature helps software products get discovered and used by coding agents.
Service: https://armature.tech/discoverability · Results: https://armature.tech/leaderboards/sectors · Contact: contact@armature.tech
