How to get picked for security testing by coding agents
ZAP took 24% of 460 judged security testing sessions. What the numbers say a vendor in this category should do.
If you sell security testing tools, this page is the part of the market no dashboard shows you: what a coding agent does when a developer asks for security testing and never compares vendors.
The numbers come from 460 judged sessions with Claude Code, Codex, Grok Build CLI and Muse Code, spread across 7 realistic codebases, with every session read by a judge.
What coding agents choose for security testing
Across 460 judged sessions, ZAP was chosen most often, in 24% of runs. Semgrep was second with 17%.
| # | Product | Runs won | Share |
|---|---|---|---|
| 1 | ZAP | 111 | 24% |
| 2 | Semgrep | 78 | 17% |
| 3 | SonarQube | 66 | 14% |
| 4 | Burp Suite | 30 | 7% |
| 5 | GitHub Advanced Security | 25 | 5% |
| 6 | Trivy | 14 | 3% |
| 7 | StackHawk | 12 | 3% |
| 8 | Strix | 11 | 2% |
| 9 | Escape | 7 | 2% |
| 10 | OWASP Dependency-Check | 6 | 1% |
Full board, every run replayable: the security testing leaderboard.
What the shape of this category means
The leader takes only 24% of runs. This category is genuinely open and the ordering can be moved.
With the top product at 24%, security testing is decided in the moment, from what the agent reads and what it finds in the repository. Nothing is locked in, which is the best situation a vendor can be in and the one where the work pays fastest.
The order here is set by the quality of what an agent can read and by whether your product is already present in the codebase. Both are things you can change.
The agents do not agree with each other
In this category Claude Code and Codex put ZAP first, which is less common than it sounds: across the 31 categories we measured, Claude Code and Codex disagreed on the leader in 15 of them.
Grok Build CLI (116 sessions) and Muse Code (112 sessions) join this table with fewer runs, so read their rows as indicative.
| Agent | Runs | Picked most often |
|---|---|---|
| Claude Code | 116 | ZAP (23) |
| Codex | 116 | ZAP (28) |
| Grok Build CLI | 116 | ZAP (21) |
| Muse Code | 112 | ZAP (39) |
Even where they agree, they get there differently. In this category Codex ran a web search in 100% of its sessions and Claude Code in 28%, so what you publish reaches nearly all of Codex's security testing sessions and about a quarter of Claude Code's.
What you are really competing against
In this category agents never chose to build it themselves. All but 2 sessions ended with a product. That is good news: you are in a straight vendor comparison, and the levers that work are the ones you control.
Considered, and never chosen
Because the judge records every product an agent raised and not only the one it picked, this board also shows who kept reaching the shortlist and losing. In security testing the clearest case is SonarQube Cloud: on the table in 37 sessions, chosen in none.
| Product | Raised in | Chosen in |
|---|---|---|
| SonarQube Cloud | 37 sessions | 0 |
Being rejected is a better position than being unknown, and a cheaper one to fix. The product is already in the agent's head and on the list. Whatever ended those 37 sessions is recorded in each transcript, one reason at a time.
What to do about it in security testing
- Skip the build-versus-buy argument. No security testing session in this experiment ended with the agent writing its own implementation. All but 2 adopted a product, so the whole contest is against the other names in the table above.
- Treat the ordering as movable. The leader holds 24%, so agents are deliberating rather than defaulting, and the inputs they use can change the answer. This is the most winnable shape a category comes in.
The work that applies to every category rather than to this one is written up separately: audit your documentation, write a quickstart an agent can follow, and how to measure install share.
Every security testing tool on this board
One page per product, with its install share, the per-agent split, and how often it was raised without being chosen.
- Do coding agents recommend ZAP? — chosen in 24% of sessions
- Do coding agents recommend Semgrep? — chosen in 17% of sessions
- Do coding agents recommend SonarQube? — chosen in 14% of sessions
- Do coding agents recommend Burp Suite? — chosen in 7% of sessions
- Do coding agents recommend GitHub Advanced Security? — chosen in 5% of sessions
- Do coding agents recommend Trivy? — chosen in 3% of sessions
- Do coding agents recommend StackHawk? — chosen in 3% of sessions
- Do coding agents recommend Strix? — chosen in 2% of sessions
- Do coding agents recommend Aikido Security? — chosen in 1% of sessions
- Do coding agents recommend GitLab Security? — chosen in 0% of sessions
- Do coding agents recommend Promptfoo? — chosen in 0% of sessions
- Do coding agents recommend Checkmarx One? — chosen in 0% of sessions
Where these numbers come from
460 judged sessions in security testing across 7 codebases, part of a published set of 15,000. Real coding agents at pinned versions, in sandboxes, inside realistic codebases, with a simulated project owner in the loop and a blind judge on every session. The full method is on one page: how we measured this.
Every security testing run can be replayed on the board.
<!-- generated by scripts/write-data-pages.mjs -->
Common questions
How many codebases is this based on?
460 judged sessions across 7 realistic codebases. A category only runs on repositories where its seam is open, so coverage differs: some categories ran on more than ten codebases and some on two.
What security testing tool do coding agents choose?
Across 460 judged sessions, ZAP was chosen most often, in 24% of runs. Semgrep was second with 17%. Every agent put ZAP first.
Do Claude Code and Codex pick the same security testing tool?
Yes. Claude Code and Codex both put ZAP first in this category, which is less common than it sounds: they disagree on the leader in 15 of the 31 categories we measured.
How often do agents build security testing themselves instead of installing something?
Never, in this category. No session ended with the agent writing the code itself. All but 2 adopted a product; the rest used built-in tools or made no choice.
How can a vendor improve its position here?
Make the quickstart run when pasted, state the current version on the documentation page, use one name across product, package and import, write pages for the symptoms users describe rather than only the category name, and get into the repository through templates and framework integrations.
Which security testing tools do agents consider but never choose?
SonarQube Cloud (raised in 37 sessions, chosen in none). Being considered and not chosen is a different problem from being unknown, and it is usually fixable.
Where this comes from
Armature ran 15,000 judged sessions with Claude Code, Codex, Cursor, Grok Build CLI and Muse Code inside 92 realistic codebases, and published every run. The numbers on this page come from that work.