# Security testing: which security tools coding agents choose

> ZAP leads. Asking for an AI pen tester changes the choice.

Source: https://armature.tech/leaderboards/security-testing (Armature agent leaderboards). 460 runs, 7 apps, 4 agents, 1 persona, updated 2026-09-29. Interactive board with every run: https://armature.tech/leaderboards#app/security-testing

## Key learnings

We asked coding agents to test the security of seven applications, from a first automated check to a full pen test.

### ZAP is the default against a running app

ZAP is chosen in 24% of all runs and in 57% of the access checks between users and tenants.

### Source-only reviews go to Semgrep

Semgrep is chosen in 50% of source-only reviews and in 92% of runs on the Python settlement diff. SonarQube follows at 24%.

### An AI pen-test ask moves away from ZAP and Burp

A plain pen-test ask goes to ZAP or Burp Suite in 89% of runs. When the ask says AI pen tester, they fall to 14%, and Strix and Escape are each chosen in 18% of runs.

## The ranking

| # | Product | Wins | Share |
|---|---|---:|---:|
| 1 | ZAP (zaproxy.org) | 111 | 24% |
| 2 | Semgrep (semgrep.dev) | 78 | 17% |
| 3 | SonarQube (sonarsource.com) | 66 | 14% |
| 4 | Burp Suite (portswigger.net) | 30 | 7% |
| 5 | GitHub Advanced Security (github.com) | 25 | 5% |
| 6 | Trivy (trivy.dev) | 14 | 3% |
| 7 | StackHawk (stackhawk.com) | 12 | 3% |
| 8 | Strix (usestrix.com) | 11 | 2% |
| 9 | Escape (escape.tech) | 7 | 2% |
| 10 | OWASP Dependency-Check (owasp.org) | 6 | 1% |
| 11 | Hurl (hurl.dev) | 5 | 1% |
| 12 | Burp Suite + ZAP | 5 | 1% |
| 13 | Shannon (github.com) | 5 | 1% |
| 14 | govulncheck (go.dev) | 4 | 1% |
| 15 | OSV-Scanner (google.github.io/osv-scanner) | 4 | 1% |
| 16 | Bandit (bandit.readthedocs.io) | 4 | 1% |
| 17 | XBOW (xbow.com) | 4 | 1% |
| 18 | Sonatype Lifecycle (sonatype.com) | 3 | 1% |
| 19 | Cobalt (cobalt.io) | 3 | 1% |
| 20 | Snyk (snyk.io) | 3 | 1% |
| 21 | Aikido Security (aikido.dev) | 3 | 1% |
| 22 | Keygraph (keygraph.io) | 2 | 0% |
| 23 | CodeQL (github.com) | 2 | 0% |
| 24 | Contrast Assess (contrastsecurity.com) | 2 | 0% |
| 25 | Datadog Code Security (datadoghq.com) | 2 | 0% |
| 26 | Codex Security (openai.com) | 2 | 0% |
| 27 | Semgrep + ZAP | 2 | 0% |
| 28 | Newman (postman.com) | 2 | 0% |
| 29 | GitLab Security (gitlab.com) | 2 | 0% |
| 30 | 42Crunch (42crunch.com) | 2 | 0% |
| 31 | Aptori Sift (aptori.com) | 2 | 0% |
| 32 | Bright Security (brightsec.com) | 2 | 0% |
| 33 | Promptfoo (promptfoo.dev) | 1 | 0% |
| 34 | Burp Suite + PentestGPT | 1 | 0% |
| 35 | Vaadata | 1 | 0% |
| 36 | Newman + Postman | 1 | 0% |
| 37 | Playwright Test (playwright.dev) | 1 | 0% |
| 38 | Dependency-Track + SonarQube | 1 | 0% |
| 39 | GitHub Agentic Workflows (github.com) | 1 | 0% |
| 40 | Checkmarx One (checkmarx.com) | 1 | 0% |
| 41 | govulncheck + Semgrep | 1 | 0% |
| 42 | Burp Suite + Nessus + punch-q + ZAP | 1 | 0% |
| 43 | SonarQube + Trivy + ZAP | 1 | 0% |
| 44 | Trivy + ZAP | 1 | 0% |
| 45 | Burp Suite + Invicti | 1 | 0% |
| 46 | Built in-house + ZAP | 1 | 0% |
| 47 | Psalm (psalm.dev) | 1 | 0% |
| 48 | Gitleaks + OWASP Dependency-Check + Semgrep + ZAP | 1 | 0% |
| 49 | GitLab Security + Kubesec + Semgrep | 1 | 0% |
| 50 | GitLab Security + Semgrep | 1 | 0% |
| 51 | Invicti (invicti.com) | 1 | 0% |
| 52 | Find Security Bugs + SpotBugs | 1 | 0% |
| 53 | Equixly | 1 | 0% |
| 54 | SonarQube + Trivy | 1 | 0% |
| 55 | Bishop Fox (bishopfox.com) | 1 | 0% |
| 56 | OWASP Dependency-Check + SonarQube | 1 | 0% |
| 57 | Claude Code Security Review (anthropic.com) | 1 | 0% |
| 58 | OWASP Coraza + OWASP Core Rule Set | 1 | 0% |
| 59 | Psalm + Trivy | 1 | 0% |
| 60 | Barrion | 1 | 0% |
| 61 | Red Hat Advanced Cluster Security for Kubernetes (redhat.com) | 1 | 0% |
| 62 | Falco (falco.org) | 1 | 0% |
| 63 | Semgrep + Trivy + ZAP | 1 | 0% |
| 64 | Semgrep + Trivy | 1 | 0% |
| 65 | Pentest-Tools.com | 1 | 0% |

## By agent

- Claude Code: 116 runs, first ZAP (23), then Semgrep (21)
- Grok Build CLI (Grok 4.7): 116 runs, first ZAP (21), then Semgrep (19)
- Codex (GPT-6 Sol): 116 runs, first ZAP (28), then Semgrep (22)
- Muse Code (Muse Spark 1.3): 112 runs, first ZAP (39), then SonarQube (22)

## By persona

- Enterprise team: 460 runs, first ZAP (111), then Semgrep (78)

## By what the ask stressed

- The plain ask: 456 runs, first ZAP (109), then Semgrep (78)

A case is one codebase with one agent, asked several times in different words and as different people. 28 of 28 cases did not hold to a single choice.

## How this was measured

Every number on this page comes from a controlled experiment. We took 7 small applications, asked 4 coding agents (Claude Code, Grok Build CLI (Grok 4.7), Codex (GPT-6 Sol), Muse Code (Muse Spark 1.3)) to test the security of each of them, in several wordings and as an enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 460 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 50 runs.

Methodology and publications: https://armature.tech/publications

If you sell in this sector, what these numbers mean for a vendor: https://armature.tech/library/security-testing-coding-agents-playbook (Markdown: https://armature.tech/library/security-testing-coding-agents-playbook.md)

## Other sectors

- [Agent sandboxes](https://armature.tech/leaderboards/sandboxes) (https://armature.tech/leaderboards/sandboxes.md)
- [Observability](https://armature.tech/leaderboards/observability) (https://armature.tech/leaderboards/observability.md)
- [AI SRE](https://armature.tech/leaderboards/ai-sre) (https://armature.tech/leaderboards/ai-sre.md)
- [Payments](https://armature.tech/leaderboards/payments) (https://armature.tech/leaderboards/payments.md)
- [Deploy](https://armature.tech/leaderboards/deploy) (https://armature.tech/leaderboards/deploy.md)
- [Auth](https://armature.tech/leaderboards/auth) (https://armature.tech/leaderboards/auth.md)
- [Email providers](https://armature.tech/leaderboards/mail) (https://armature.tech/leaderboards/mail.md)
- [Product analytics](https://armature.tech/leaderboards/product-analytics) (https://armature.tech/leaderboards/product-analytics.md)
- [Databases](https://armature.tech/leaderboards/databases) (https://armature.tech/leaderboards/databases.md)
- [File storage](https://armature.tech/leaderboards/storage) (https://armature.tech/leaderboards/storage.md)
- [LLM evals & observability](https://armature.tech/leaderboards/evals) (https://armature.tech/leaderboards/evals.md)
- [Voice Agents](https://armature.tech/leaderboards/voice-agents) (https://armature.tech/leaderboards/voice-agents.md)
- [Serverless functions](https://armature.tech/leaderboards/serverless) (https://armature.tech/leaderboards/serverless.md)
- [Cloud](https://armature.tech/leaderboards/cloud) (https://armature.tech/leaderboards/cloud.md)
- [AI gateway](https://armature.tech/leaderboards/ai-gateway) (https://armature.tech/leaderboards/ai-gateway.md)
- [Bot protection](https://armature.tech/leaderboards/bot-protection) (https://armature.tech/leaderboards/bot-protection.md)
- [Search](https://armature.tech/leaderboards/search) (https://armature.tech/leaderboards/search.md)
- [Agent frameworks](https://armature.tech/leaderboards/agent-frameworks) (https://armature.tech/leaderboards/agent-frameworks.md)
- [Performance in CI](https://armature.tech/leaderboards/perf-ci) (https://armature.tech/leaderboards/perf-ci.md)
- [Document processing & OCR](https://armature.tech/leaderboards/document-processing) (https://armature.tech/leaderboards/document-processing.md)
- [Usage-based billing](https://armature.tech/leaderboards/usage-based-billing) (https://armature.tech/leaderboards/usage-based-billing.md)
- [Internationalization](https://armature.tech/leaderboards/internationalization) (https://armature.tech/leaderboards/internationalization.md)
- [Maps](https://armature.tech/leaderboards/maps) (https://armature.tech/leaderboards/maps.md)
- [Message queues](https://armature.tech/leaderboards/message-queues) (https://armature.tech/leaderboards/message-queues.md)
- [AI search](https://armature.tech/leaderboards/ai-search) (https://armature.tech/leaderboards/ai-search.md)
- [Code review](https://armature.tech/leaderboards/code-review) (https://armature.tech/leaderboards/code-review.md)
- [E-signature](https://armature.tech/leaderboards/e-signature) (https://armature.tech/leaderboards/e-signature.md)
- [Security platforms](https://armature.tech/leaderboards/security-platforms) (https://armature.tech/leaderboards/security-platforms.md)
- [In-app chat & calls](https://armature.tech/leaderboards/in-app-communication) (https://armature.tech/leaderboards/in-app-communication.md)
- [Vector search](https://armature.tech/leaderboards/vector-search) (https://armature.tech/leaderboards/vector-search.md)
