Agent leaderboards / All sectors / Security testing
Security testing: which security tools coding agents choose
ZAP leads. Asking for an AI pen tester changes the choice.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked coding agents to test the security of seven applications, from a first automated check to a full pen test.
ZAP is the default against a running app
ZAP is chosen in 24% of all runs and in 57% of the access checks between users and tenants.
Source-only reviews go to Semgrep
Semgrep is chosen in 50% of source-only reviews and in 92% of runs on the Python settlement diff. SonarQube follows at 24%.
An AI pen-test ask moves away from ZAP and Burp
A plain pen-test ask goes to ZAP or Burp Suite in 89% of runs. When the ask says AI pen tester, they fall to 14%, and Strix and Escape are each chosen in 18% of runs.
The ranking460 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | ZAPzaproxy.org | 111 | 24% | |
| 2 | Semgrepsemgrep.dev | 78 | 17% | |
| 3 | SonarQubesonarsource.com | 66 | 14% | |
| 4 | Burp Suiteportswigger.net | 30 | 7% | |
| 5 | GitHub Advanced Securitygithub.com | 25 | 5% | |
| 6 | Trivytrivy.dev | 14 | 3% | |
| 7 | StackHawkstackhawk.com | 12 | 3% | |
| 8 | Strixusestrix.com | 11 | 2% | |
| 9 | Escapeescape.tech | 7 | 2% | |
| 10 | OWASP Dependency-Checkowasp.org | 6 | 1% | |
| 11 | Hurlhurl.dev | 5 | 1% | |
| 12 | Burp Suite + ZAP | 5 | 1% | |
| 13 | Shannongithub.com | 5 | 1% | |
| 14 | govulncheckgo.dev | 4 | 1% | |
| 15 | OSV-Scannergoogle.github.io/osv-scanner | 4 | 1% | |
| 16 | Banditbandit.readthedocs.io | 4 | 1% | |
| 17 | XBOWxbow.com | 4 | 1% | |
| 18 | Sonatype Lifecyclesonatype.com | 3 | 1% | |
| 19 | Cobaltcobalt.io | 3 | 1% | |
| 20 | Snyksnyk.io | 3 | 1% | |
| 21 | Aikido Securityaikido.dev | 3 | 1% | |
| 22 | Keygraphkeygraph.io | 2 | 0% | |
| 23 | CodeQLgithub.com | 2 | 0% | |
| 24 | Contrast Assesscontrastsecurity.com | 2 | 0% | |
| 25 | Datadog Code Securitydatadoghq.com | 2 | 0% | |
| 26 | Codex Securityopenai.com | 2 | 0% | |
| 27 | Semgrep + ZAP | 2 | 0% | |
| 28 | Newmanpostman.com | 2 | 0% | |
| 29 | GitLab Securitygitlab.com | 2 | 0% | |
| 30 | 42Crunch42crunch.com | 2 | 0% | |
| 31 | Aptori Siftaptori.com | 2 | 0% | |
| 32 | Bright Securitybrightsec.com | 2 | 0% | |
| 33 | Promptfoopromptfoo.dev | 1 | 0% | |
| 34 | Burp Suite + PentestGPT | 1 | 0% | |
| 35 | Vaadata | 1 | 0% | |
| 36 | Newman + Postman | 1 | 0% | |
| 37 | Playwright Testplaywright.dev | 1 | 0% | |
| 38 | Dependency-Track + SonarQube | 1 | 0% | |
| 39 | GitHub Agentic Workflowsgithub.com | 1 | 0% | |
| 40 | Checkmarx Onecheckmarx.com | 1 | 0% | |
| 41 | govulncheck + Semgrep | 1 | 0% | |
| 42 | Burp Suite + Nessus + punch-q + ZAP | 1 | 0% | |
| 43 | SonarQube + Trivy + ZAP | 1 | 0% | |
| 44 | Trivy + ZAP | 1 | 0% | |
| 45 | Burp Suite + Invicti | 1 | 0% | |
| 46 | Built in-house + ZAP | 1 | 0% | |
| 47 | Psalmpsalm.dev | 1 | 0% | |
| 48 | Gitleaks + OWASP Dependency-Check + Semgrep + ZAP | 1 | 0% | |
| 49 | GitLab Security + Kubesec + Semgrep | 1 | 0% | |
| 50 | GitLab Security + Semgrep | 1 | 0% | |
| 51 | Invictiinvicti.com | 1 | 0% | |
| 52 | Find Security Bugs + SpotBugs | 1 | 0% | |
| 53 | Equixly | 1 | 0% | |
| 54 | SonarQube + Trivy | 1 | 0% | |
| 55 | Bishop Foxbishopfox.com | 1 | 0% | |
| 56 | OWASP Dependency-Check + SonarQube | 1 | 0% | |
| 57 | Claude Code Security Reviewanthropic.com | 1 | 0% | |
| 58 | OWASP Coraza + OWASP Core Rule Set | 1 | 0% | |
| 59 | Psalm + Trivy | 1 | 0% | |
| 60 | Barrion | 1 | 0% | |
| 61 | Red Hat Advanced Cluster Security for Kubernetesredhat.com | 1 | 0% | |
| 62 | Falcofalco.org | 1 | 0% | |
| 63 | Semgrep + Trivy + ZAP | 1 | 0% | |
| 64 | Semgrep + Trivy | 1 | 0% | |
| 65 | Pentest-Tools.com | 1 | 0% |
By agent, by persona, by wording
By agent
| Claude Code116 runs | ZAP · 23then Semgrep · 21 |
| Grok Build CLI · Grok 4.7116 runs | ZAP · 21then Semgrep · 19 |
| Codex · GPT-6 Sol116 runs | ZAP · 28then Semgrep · 22 |
| Muse Code · Muse Spark 1.3112 runs | ZAP · 39then SonarQube · 22 |
By persona
| Enterprise team460 runs | ZAP · 111then Semgrep · 78 |
By what the ask stressed
| The plain ask456 runs | ZAP · 109then Semgrep · 78 |
A case is one codebase with one agent, asked several times in different words and as different people. 28 of 28 cases did not hold to a single security tool.
How this was measured
Every number on this page comes from a controlled experiment. We took 7 small applications, asked 4 coding agents (Claude Code, Grok Build CLI (Grok 4.7), Codex (GPT-6 Sol), Muse Code (Muse Spark 1.3)) to test the security of each of them, in several wordings and as an enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 460 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 50 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as MarkdownOther sectors
- Agent sandboxes
- Observability
- AI SRE
- Payments
- Deploy
- Auth
- Email providers
- Product analytics
- Databases
- File storage
- LLM evals & observability
- Voice Agents
- Serverless functions
- Cloud
- AI gateway
- Bot protection
- Search
- Agent frameworks
- Performance in CI
- Document processing & OCR
- Usage-based billing
- Internationalization
- Maps
- Message queues
- AI search
- Code review
- E-signature
- Security platforms
- In-app chat & calls
- Vector search