Agent leaderboards / All sectors / Code review

Code review: which tools coding agents choose

Claude Code review took about 26%, and every agent leaned to its own tool.

290 runs10 apps3 agents4 personasupdated 2026-09-08

The interactive board, open on code review. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three agents to add automated code review to 10 small apps, 290 runs in all. We changed the wording each time and asked as four different people. Ten tools got picked, and the ones below the top two split the rest.

76 of 100

Each agent reached for its maker's tool

Claude Code picked Claude Code review in 76 of its 100 runs. Cursor picked Cursor Bugbot in about 79% of its 90 runs. Codex picked Codex code review 37 times, the lowest of the three.

15 of 30

Half the cases did not settle on one tool

One codebase with one agent, asked in several wordings, often ended on different tools. That happened in 15 of the 30 cases.

10 of 38

Privacy wording put a different tool first

When the ask mentioned self-hosting, privacy or data residency, Semgrep came first with 10 wins across 38 runs. In the plain ask, Claude Code review led with 63.

87 of 290

Other vendors still won a share of runs

Agents picked a tool from another maker in 87 runs. CodeRabbit was the most common of those, with 34 wins.

  • Senior engineers put CodeRabbit on top with 21 wins.
  • Agents wrote the review setup themselves in 19 runs, about 7%.
  • The simulated user approved every plan, but sent the agent back at least once in 20 runs.
  • Both wins for Greptile came in the Kubernetes platform monorepo, and both for GitLab Duo Code Review in the PHP government portal.
Explore every run in the interactive board

The ranking290 runs

ProductWinsShare
1 Claude Code reviewanthropic.com 76 26%
2 Cursor Bugbotcursor.com 71 24%
3 Codex code reviewopenai.com 37 13%
4 CodeRabbitcoderabbit.ai 34 12%
5 Qodo Mergeqodo.ai 24 8%
6 Built in-houseoutcome 19 7%
7 Semgrepsemgrep.dev 10 3%
8 GitHub Copilot code reviewgithub.com 9 3%
9 SonarQubesonarsource.com 6 2%
10 Greptilegreptile.com 2 1%
11 GitLab Duo Code Reviewgitlab.com 2 1%

By agent, by persona, by wording

By agent

Claude Code · Claude Opus 5100 runsClaude Code review · 76then Qodo Merge · 8
Codex · GPT-5.6 Sol100 runsCodex code review · 37then CodeRabbit · 21
Cursor · Grok 4.690 runsCursor Bugbot · 71then CodeRabbit · 6

By persona

Junior developer116 runsCursor Bugbot · 27then Claude Code review · 26
Enterprise team58 runsClaude Code review · 17then Cursor Bugbot · 14
Senior engineer58 runsCodeRabbit · 21then Claude Code review · 13
Vibe coder58 runsClaude Code review · 20then Cursor Bugbot · 18

By what the ask stressed

The plain ask234 runsClaude Code review · 63then Cursor Bugbot · 59
Self-hosting, privacy or residency38 runsSemgrep · 10then Claude Code review · 7

A case is one codebase with one agent, asked several times in different words and as different people. 15 of 30 cases did not hold to a single tool.

How this was measured

Every number on this page comes from a controlled experiment. We took 10 small applications, asked 3 coding agents (Claude Code (Claude Opus 5), Codex (GPT-5.6 Sol), Cursor (Grok 4.6)) to add automated code review to each of them, in several wordings and as a junior developer and enterprise team and senior engineer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 290 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 20 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown