# Code review: which tools coding agents choose

> Claude Code review took about 26%, and every agent leaned to its own tool.

Source: https://armature.tech/leaderboards/code-review (Armature agent leaderboards). 290 runs, 10 apps, 3 agents, 4 personas, updated 2026-09-08. Interactive board with every run: https://armature.tech/leaderboards#app/code-review

## Key learnings

We asked three agents to add automated code review to 10 small apps, 290 runs in all. We changed the wording each time and asked as four different people. Ten tools got picked, and the ones below the top two split the rest.

### Each agent reached for its maker's tool (76 of 100)

Claude Code picked Claude Code review in 76 of its 100 runs. Cursor picked Cursor Bugbot in about 79% of its 90 runs. Codex picked Codex code review 37 times, the lowest of the three.

### Half the cases did not settle on one tool (15 of 30)

One codebase with one agent, asked in several wordings, often ended on different tools. That happened in 15 of the 30 cases.

### Privacy wording put a different tool first (10 of 38)

When the ask mentioned self-hosting, privacy or data residency, Semgrep came first with 10 wins across 38 runs. In the plain ask, Claude Code review led with 63.

### Other vendors still won a share of runs (87 of 290)

Agents picked a tool from another maker in 87 runs. CodeRabbit was the most common of those, with 34 wins.

Smaller learnings:

- Senior engineers put CodeRabbit on top with 21 wins.
- Agents wrote the review setup themselves in 19 runs, about 7%.
- The simulated user approved every plan, but sent the agent back at least once in 20 runs.
- Both wins for Greptile came in the Kubernetes platform monorepo, and both for GitLab Duo Code Review in the PHP government portal.

## The ranking

| # | Product | Wins | Share |
|---|---|---:|---:|
| 1 | Claude Code review (anthropic.com) | 76 | 26% |
| 2 | Cursor Bugbot (cursor.com) | 71 | 24% |
| 3 | Codex code review (openai.com) | 37 | 13% |
| 4 | CodeRabbit (coderabbit.ai) | 34 | 12% |
| 5 | Qodo Merge (qodo.ai) | 24 | 8% |
| 6 | Built in-house (outcome) | 19 | 7% |
| 7 | Semgrep (semgrep.dev) | 10 | 3% |
| 8 | GitHub Copilot code review (github.com) | 9 | 3% |
| 9 | SonarQube (sonarsource.com) | 6 | 2% |
| 10 | Greptile (greptile.com) | 2 | 1% |
| 11 | GitLab Duo Code Review (gitlab.com) | 2 | 1% |

## By agent

- Claude Code (Claude Opus 5): 100 runs, first Claude Code review (76), then Qodo Merge (8)
- Codex (GPT-5.6 Sol): 100 runs, first Codex code review (37), then CodeRabbit (21)
- Cursor (Grok 4.6): 90 runs, first Cursor Bugbot (71), then CodeRabbit (6)

## By persona

- Junior developer: 116 runs, first Cursor Bugbot (27), then Claude Code review (26)
- Enterprise team: 58 runs, first Claude Code review (17), then Cursor Bugbot (14)
- Senior engineer: 58 runs, first CodeRabbit (21), then Claude Code review (13)
- Vibe coder: 58 runs, first Claude Code review (20), then Cursor Bugbot (18)

## By what the ask stressed

- The plain ask: 234 runs, first Claude Code review (63), then Cursor Bugbot (59)
- Self-hosting, privacy or residency: 38 runs, first Semgrep (10), then Claude Code review (7)

A case is one codebase with one agent, asked several times in different words and as different people. 15 of 30 cases did not hold to a single choice.

## How this was measured

Every number on this page comes from a controlled experiment. We took 10 small applications, asked 3 coding agents (Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol), Cursor (Grok 4.6)) to add automated code review to each of them, in several wordings and as a junior developer and enterprise team and senior engineer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 290 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 20 runs.

Methodology and publications: https://armature.tech/publications

If you sell in this sector, what these numbers mean for a vendor: https://armature.tech/library/code-review-coding-agents-playbook (Markdown: https://armature.tech/library/code-review-coding-agents-playbook.md)

## Other sectors

- [Agent sandboxes](https://armature.tech/leaderboards/sandboxes) (https://armature.tech/leaderboards/sandboxes.md)
- [Observability](https://armature.tech/leaderboards/observability) (https://armature.tech/leaderboards/observability.md)
- [Payments](https://armature.tech/leaderboards/payments) (https://armature.tech/leaderboards/payments.md)
- [Deploy](https://armature.tech/leaderboards/deploy) (https://armature.tech/leaderboards/deploy.md)
- [Auth](https://armature.tech/leaderboards/auth) (https://armature.tech/leaderboards/auth.md)
- [Email providers](https://armature.tech/leaderboards/mail) (https://armature.tech/leaderboards/mail.md)
- [Product analytics](https://armature.tech/leaderboards/product-analytics) (https://armature.tech/leaderboards/product-analytics.md)
- [Databases](https://armature.tech/leaderboards/databases) (https://armature.tech/leaderboards/databases.md)
- [File storage](https://armature.tech/leaderboards/storage) (https://armature.tech/leaderboards/storage.md)
- [LLM evals & observability](https://armature.tech/leaderboards/evals) (https://armature.tech/leaderboards/evals.md)
- [Voice Agents](https://armature.tech/leaderboards/voice-agents) (https://armature.tech/leaderboards/voice-agents.md)
- [Serverless functions](https://armature.tech/leaderboards/serverless) (https://armature.tech/leaderboards/serverless.md)
- [Cloud](https://armature.tech/leaderboards/cloud) (https://armature.tech/leaderboards/cloud.md)
- [AI gateway](https://armature.tech/leaderboards/ai-gateway) (https://armature.tech/leaderboards/ai-gateway.md)
- [Bot protection](https://armature.tech/leaderboards/bot-protection) (https://armature.tech/leaderboards/bot-protection.md)
- [Search](https://armature.tech/leaderboards/search) (https://armature.tech/leaderboards/search.md)
- [Agent frameworks](https://armature.tech/leaderboards/agent-frameworks) (https://armature.tech/leaderboards/agent-frameworks.md)
- [Performance in CI](https://armature.tech/leaderboards/perf-ci) (https://armature.tech/leaderboards/perf-ci.md)
- [In-app chat & calls](https://armature.tech/leaderboards/in-app-communication) (https://armature.tech/leaderboards/in-app-communication.md)
- [Vector search](https://armature.tech/leaderboards/vector-search) (https://armature.tech/leaderboards/vector-search.md)
