From the experiment

Do coding agents recommend reviewdog?

reviewdog was chosen in 3% of 496 judged code review sessions, ranking ninth. Measured with Claude Code, Codex, Cursor, Grok Build CLI and Muse Code.

Published September 24, 2026 Read as Markdown

reviewdog was chosen in 3% of 496 judged code review sessions, ranking ninth. It was also raised as a candidate in 37 further sessions without being chosen.

This page reports what happened when Claude Code, Codex, Cursor, Grok Build CLI and Muse Code had to solve a problem in code review inside a realistic codebase. Not what a chat assistant says about reviewdog. What an agent actually installed.

The numbers

CategoryCode review
Sessions in the category496
Sessions where reviewdog was chosen13
Install share3%
Rank in category9 of 21
Codebases it won in4
Raised as a candidate, not chosen37
Chosen when considered26%
Sitegithub.com

By agent

With 13 wins spread across 5 agents, the rates below are small numbers and a difference between them is not yet a finding. They are here because the direction is worth knowing, not because the gap is established.

AgentSessionsChose reviewdogShare
Claude Code13000%
Codex13011%
Cursor12000%
Grok Build CLI3027%
Muse Code861012%

By who was asking

reviewdog performs similarly across the four kinds of buyer, from 1% to 4%. That is unusual: the category leader changed with the persona in 22 of the 29 categories we measured.

Who is askingSessionsChose reviewdogShare
Vibe coder10044%
Junior developer15832%
Senior engineer10011%
Enterprise team13854%

What reviewdog was up against

The full ranking in code review, from the same sessions:

#ProductRuns wonShare
1Claude Code review10120%
2Cursor Bugbot9118%
3CodeRabbit6914%
4Qodo Merge469%
5Codex code review449%
6Built in-house (no product adopted)408%
7GitHub Copilot code review316%
8Semgrep224%

What this means

reviewdog was raised in 37 sessions and chosen in 13. That ratio is balanced enough that the ceiling is presence rather than integration: the product converts reasonably when it is on the table, and it is not on the table often enough.

Where these numbers come from

The 496 sessions in code review are part of a published set of 13,497, run with real coding agents inside realistic codebases and judged blind. The full method is on one page: how we measured this.

Every code review run can be replayed on the board.

If you work on reviewdog: the judge recorded a reason for every session where it was raised and passed over. Those reasons are in the transcripts.

<!-- generated by scripts/write-data-pages.mjs -->

Common questions

Do coding agents recommend reviewdog?

Yes. reviewdog was chosen in 13 of the 496 judged sessions in code review, a 3% install share, ranking ninth in its category.

Does Claude Code recommend reviewdog?

In 0 of the 130 sessions in code review run with Claude Code, which is 0%.

Do different coding agents treat reviewdog differently?

Yes, and by a wide margin. Muse Code chose it in 12% of its runs and Claude Code in 0%.

How was this measured?

Real coding agents at pinned versions were run in sandboxes inside 91 realistic codebases and asked to solve real tasks. A simulated project owner approved or questioned each recommendation before any code was written, and a judge from a model family that builds none of the agents read every session blind.

How often is reviewdog considered but not chosen?

It was raised as a candidate in 37 sessions without being chosen, and chosen in 13. That is a 26% conversion from considered to chosen.

Where this comes from

Armature ran 13,497 judged sessions with Claude Code, Codex, Cursor, Grok Build CLI and Muse Code inside 91 realistic codebases, and published every run. The numbers on this page come from that work.

Read next

All library pages