Do coding agents recommend reviewdog?
reviewdog was chosen in 3% of 496 judged code review sessions, ranking ninth. Measured with Claude Code, Codex, Cursor, Grok Build CLI and Muse Code.
reviewdog was chosen in 3% of 496 judged code review sessions, ranking ninth. It was also raised as a candidate in 37 further sessions without being chosen.
This page reports what happened when Claude Code, Codex, Cursor, Grok Build CLI and Muse Code had to solve a problem in code review inside a realistic codebase. Not what a chat assistant says about reviewdog. What an agent actually installed.
The numbers
| Category | Code review |
| Sessions in the category | 496 |
| Sessions where reviewdog was chosen | 13 |
| Install share | 3% |
| Rank in category | 9 of 21 |
| Codebases it won in | 4 |
| Raised as a candidate, not chosen | 37 |
| Chosen when considered | 26% |
| Site | github.com |
By agent
With 13 wins spread across 5 agents, the rates below are small numbers and a difference between them is not yet a finding. They are here because the direction is worth knowing, not because the gap is established.
| Agent | Sessions | Chose reviewdog | Share |
|---|---|---|---|
| Claude Code | 130 | 0 | 0% |
| Codex | 130 | 1 | 1% |
| Cursor | 120 | 0 | 0% |
| Grok Build CLI | 30 | 2 | 7% |
| Muse Code | 86 | 10 | 12% |
By who was asking
reviewdog performs similarly across the four kinds of buyer, from 1% to 4%. That is unusual: the category leader changed with the persona in 22 of the 29 categories we measured.
| Who is asking | Sessions | Chose reviewdog | Share |
|---|---|---|---|
| Vibe coder | 100 | 4 | 4% |
| Junior developer | 158 | 3 | 2% |
| Senior engineer | 100 | 1 | 1% |
| Enterprise team | 138 | 5 | 4% |
What reviewdog was up against
The full ranking in code review, from the same sessions:
| # | Product | Runs won | Share |
|---|---|---|---|
| 1 | Claude Code review | 101 | 20% |
| 2 | Cursor Bugbot | 91 | 18% |
| 3 | CodeRabbit | 69 | 14% |
| 4 | Qodo Merge | 46 | 9% |
| 5 | Codex code review | 44 | 9% |
| 6 | Built in-house (no product adopted) | 40 | 8% |
| 7 | GitHub Copilot code review | 31 | 6% |
| 8 | Semgrep | 22 | 4% |
What this means
reviewdog was raised in 37 sessions and chosen in 13. That ratio is balanced enough that the ceiling is presence rather than integration: the product converts reasonably when it is on the table, and it is not on the table often enough.
Where these numbers come from
The 496 sessions in code review are part of a published set of 13,497, run with real coding agents inside realistic codebases and judged blind. The full method is on one page: how we measured this.
Every code review run can be replayed on the board.
If you work on reviewdog: the judge recorded a reason for every session where it was raised and passed over. Those reasons are in the transcripts.
<!-- generated by scripts/write-data-pages.mjs -->
Common questions
Do coding agents recommend reviewdog?
Yes. reviewdog was chosen in 13 of the 496 judged sessions in code review, a 3% install share, ranking ninth in its category.
Does Claude Code recommend reviewdog?
In 0 of the 130 sessions in code review run with Claude Code, which is 0%.
Do different coding agents treat reviewdog differently?
Yes, and by a wide margin. Muse Code chose it in 12% of its runs and Claude Code in 0%.
How was this measured?
Real coding agents at pinned versions were run in sandboxes inside 91 realistic codebases and asked to solve real tasks. A simulated project owner approved or questioned each recommendation before any code was written, and a judge from a model family that builds none of the agents read every session blind.
How often is reviewdog considered but not chosen?
It was raised as a candidate in 37 sessions without being chosen, and chosen in 13. That is a 26% conversion from considered to chosen.
Where this comes from
Armature ran 13,497 judged sessions with Claude Code, Codex, Cursor, Grok Build CLI and Muse Code inside 91 realistic codebases, and published every run. The numbers on this page come from that work.