Agent leaderboards / All sectors / Security platforms

Security platforms: which security platforms coding agents choose

Checkmarx One leads. GitLab apps get GitLab's own suite.

118 runs7 apps4 agents1 personaupdated 2026-09-30

The interactive board, open on security platforms. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked coding agents to choose one security platform for code, dependencies, secrets and live-app tests on seven applications.

Checkmarx One leads on Jenkins and Azure DevOps

Checkmarx One is chosen in 33% of all runs and in 50% on the apps that build in Jenkins or Azure DevOps.

GitLab apps keep GitLab

On the two apps built in GitLab CI, GitLab Security is chosen in 91% of runs. Elsewhere it is chosen in 2%.

GitHub apps do not get GitHub's

On the two apps built in GitHub Actions, GitHub Advanced Security is part of the choice in 4% of runs. Snyk is chosen in 33% there and Aikido Security in 25%.

  • Muse Code chooses Checkmarx One in 4% of its runs, against 37% to 50% for the other agents, and Snyk in 39%.
Explore every run in the interactive board

By agent, by persona, by wording

By agent

Codex · GPT-6 Sol30 runsCheckmarx One · 15then GitLab Security · 5
Grok Build CLI · Grok 4.730 runsCheckmarx One · 12then GitLab Security · 6
Claude Code · Claude Opus 5.530 runsCheckmarx One · 11then GitLab Security · 6
Muse Code · Muse Spark 1.328 runsSnyk · 11then GitLab Security · 5

By persona

Enterprise team118 runsCheckmarx One · 39then GitLab Security · 22

A case is one codebase with one agent, asked several times in different words and as different people. 21 of 27 cases did not hold to a single security platform.

How this was measured

Every number on this page comes from a controlled experiment. We took 7 small applications, asked 4 coding agents (Codex (GPT-6 Sol), Grok Build CLI (Grok 4.7), Claude Code (Claude Opus 5.5), Muse Code (Muse Spark 1.3)) to choose a security platform for each of them, in several wordings and as an enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 118 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 27 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown