Agent leaderboards / All sectors / Security testing

Security testing: which security tools coding agents choose

ZAP leads. Asking for an AI pen tester changes the choice.

460 runs7 apps4 agents1 personaupdated 2026-09-29

The interactive board, open on security testing. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked coding agents to test the security of seven applications, from a first automated check to a full pen test.

ZAP is the default against a running app

ZAP is chosen in 24% of all runs and in 57% of the access checks between users and tenants.

Source-only reviews go to Semgrep

Semgrep is chosen in 50% of source-only reviews and in 92% of runs on the Python settlement diff. SonarQube follows at 24%.

An AI pen-test ask moves away from ZAP and Burp

A plain pen-test ask goes to ZAP or Burp Suite in 89% of runs. When the ask says AI pen tester, they fall to 14%, and Strix and Escape are each chosen in 18% of runs.

Explore every run in the interactive board

The ranking460 runs

ProductWinsShare
1 ZAPzaproxy.org 111 24%
2 Semgrepsemgrep.dev 78 17%
3 SonarQubesonarsource.com 66 14%
4 Burp Suiteportswigger.net 30 7%
5 GitHub Advanced Securitygithub.com 25 5%
6 Trivytrivy.dev 14 3%
7 StackHawkstackhawk.com 12 3%
8 Strixusestrix.com 11 2%
9 Escapeescape.tech 7 2%
10 OWASP Dependency-Checkowasp.org 6 1%
11 Hurlhurl.dev 5 1%
12 Burp Suite + ZAP 5 1%
13 Shannongithub.com 5 1%
14 govulncheckgo.dev 4 1%
15 OSV-Scannergoogle.github.io/osv-scanner 4 1%
16 Banditbandit.readthedocs.io 4 1%
17 XBOWxbow.com 4 1%
18 Sonatype Lifecyclesonatype.com 3 1%
19 Cobaltcobalt.io 3 1%
20 Snyksnyk.io 3 1%
21 Aikido Securityaikido.dev 3 1%
22 Keygraphkeygraph.io 2 0%
23 CodeQLgithub.com 2 0%
24 Contrast Assesscontrastsecurity.com 2 0%
25 Datadog Code Securitydatadoghq.com 2 0%
26 Codex Securityopenai.com 2 0%
27 Semgrep + ZAP 2 0%
28 Newmanpostman.com 2 0%
29 GitLab Securitygitlab.com 2 0%
30 42Crunch42crunch.com 2 0%
31 Aptori Siftaptori.com 2 0%
32 Bright Securitybrightsec.com 2 0%
33 Promptfoopromptfoo.dev 1 0%
34 Burp Suite + PentestGPT 1 0%
35 Vaadata 1 0%
36 Newman + Postman 1 0%
37 Playwright Testplaywright.dev 1 0%
38 Dependency-Track + SonarQube 1 0%
39 GitHub Agentic Workflowsgithub.com 1 0%
40 Checkmarx Onecheckmarx.com 1 0%
41 govulncheck + Semgrep 1 0%
42 Burp Suite + Nessus + punch-q + ZAP 1 0%
43 SonarQube + Trivy + ZAP 1 0%
44 Trivy + ZAP 1 0%
45 Burp Suite + Invicti 1 0%
46 Built in-house + ZAP 1 0%
47 Psalmpsalm.dev 1 0%
48 Gitleaks + OWASP Dependency-Check + Semgrep + ZAP 1 0%
49 GitLab Security + Kubesec + Semgrep 1 0%
50 GitLab Security + Semgrep 1 0%
51 Invictiinvicti.com 1 0%
52 Find Security Bugs + SpotBugs 1 0%
53 Equixly 1 0%
54 SonarQube + Trivy 1 0%
55 Bishop Foxbishopfox.com 1 0%
56 OWASP Dependency-Check + SonarQube 1 0%
57 Claude Code Security Reviewanthropic.com 1 0%
58 OWASP Coraza + OWASP Core Rule Set 1 0%
59 Psalm + Trivy 1 0%
60 Barrion 1 0%
61 Red Hat Advanced Cluster Security for Kubernetesredhat.com 1 0%
62 Falcofalco.org 1 0%
63 Semgrep + Trivy + ZAP 1 0%
64 Semgrep + Trivy 1 0%
65 Pentest-Tools.com 1 0%

By agent, by persona, by wording

By agent

Claude Code116 runsZAP · 23then Semgrep · 21
Grok Build CLI · Grok 4.7116 runsZAP · 21then Semgrep · 19
Codex · GPT-6 Sol116 runsZAP · 28then Semgrep · 22
Muse Code · Muse Spark 1.3112 runsZAP · 39then SonarQube · 22

By persona

Enterprise team460 runsZAP · 111then Semgrep · 78

By what the ask stressed

The plain ask456 runsZAP · 109then Semgrep · 78

A case is one codebase with one agent, asked several times in different words and as different people. 28 of 28 cases did not hold to a single security tool.

How this was measured

Every number on this page comes from a controlled experiment. We took 7 small applications, asked 4 coding agents (Claude Code, Grok Build CLI (Grok 4.7), Codex (GPT-6 Sol), Muse Code (Muse Spark 1.3)) to test the security of each of them, in several wordings and as an enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 460 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 50 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown