Agent leaderboards / All sectors / Code review
Code review: which tools coding agents choose
Claude Code review took about 26%, and every agent leaned to its own tool.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three agents to add automated code review to 10 small apps, 290 runs in all. We changed the wording each time and asked as four different people. Ten tools got picked, and the ones below the top two split the rest.
Each agent reached for its maker's tool
Claude Code picked Claude Code review in 76 of its 100 runs. Cursor picked Cursor Bugbot in about 79% of its 90 runs. Codex picked Codex code review 37 times, the lowest of the three.
Half the cases did not settle on one tool
One codebase with one agent, asked in several wordings, often ended on different tools. That happened in 15 of the 30 cases.
Privacy wording put a different tool first
When the ask mentioned self-hosting, privacy or data residency, Semgrep came first with 10 wins across 38 runs. In the plain ask, Claude Code review led with 63.
Other vendors still won a share of runs
Agents picked a tool from another maker in 87 runs. CodeRabbit was the most common of those, with 34 wins.
- Senior engineers put CodeRabbit on top with 21 wins.
- Agents wrote the review setup themselves in 19 runs, about 7%.
- The simulated user approved every plan, but sent the agent back at least once in 20 runs.
- Both wins for Greptile came in the Kubernetes platform monorepo, and both for GitLab Duo Code Review in the PHP government portal.
The ranking290 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Claude Code reviewanthropic.com | 76 | 26% | |
| 2 | Cursor Bugbotcursor.com | 71 | 24% | |
| 3 | Codex code reviewopenai.com | 37 | 13% | |
| 4 | CodeRabbitcoderabbit.ai | 34 | 12% | |
| 5 | Qodo Mergeqodo.ai | 24 | 8% | |
| 6 | Built in-houseoutcome | 19 | 7% | |
| 7 | Semgrepsemgrep.dev | 10 | 3% | |
| 8 | GitHub Copilot code reviewgithub.com | 9 | 3% | |
| 9 | SonarQubesonarsource.com | 6 | 2% | |
| 10 | Greptilegreptile.com | 2 | 1% | |
| 11 | GitLab Duo Code Reviewgitlab.com | 2 | 1% |
By agent, by persona, by wording
By agent
| Claude Code · Claude Opus 5100 runs | Claude Code review · 76then Qodo Merge · 8 |
| Codex · GPT-5.6 Sol100 runs | Codex code review · 37then CodeRabbit · 21 |
| Cursor · Grok 4.690 runs | Cursor Bugbot · 71then CodeRabbit · 6 |
By persona
| Junior developer116 runs | Cursor Bugbot · 27then Claude Code review · 26 |
| Enterprise team58 runs | Claude Code review · 17then Cursor Bugbot · 14 |
| Senior engineer58 runs | CodeRabbit · 21then Claude Code review · 13 |
| Vibe coder58 runs | Claude Code review · 20then Cursor Bugbot · 18 |
By what the ask stressed
| The plain ask234 runs | Claude Code review · 63then Cursor Bugbot · 59 |
| Self-hosting, privacy or residency38 runs | Semgrep · 10then Claude Code review · 7 |
A case is one codebase with one agent, asked several times in different words and as different people. 15 of 30 cases did not hold to a single tool.
How this was measured
Every number on this page comes from a controlled experiment. We took 10 small applications, asked 3 coding agents (Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol), Cursor (Grok 4.6)) to add automated code review to each of them, in several wordings and as a junior developer and enterprise team and senior engineer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 290 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 20 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown