Agent leaderboards / All sectors / Bot protection
Bot protection: which products coding agents choose
Cloudflare Turnstile took about 57% of the runs.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three agents to add bot protection to five small apps, over 160 runs. Each time we asked in different words, and as a different kind of person. Below the leader nothing got past 11%, and one of those 11% shares was the agents writing it themselves.
The same ask in other words changed the answer
A case here is one codebase with one agent, asked several times in different wordings and as different people. 13 of 15 cases did not land on the same product every time.
Three products won only inside one codebase
django-axes took all 11 of its wins in the Django learning platform. Google reCAPTCHA won six times, all in that same codebase. Laravel RateLimiter won twice, both in the Laravel helpdesk.
Asked as an enterprise team, the order shifted
In the enterprise team runs, django-axes and Cloudflare Turnstile split almost evenly, 11 wins each. No other persona put anything but Cloudflare Turnstile on top.
Named in most comparisons, picked in none
hCaptcha came up 126 times while the agents compared their options. It won nothing.
- The simulated user approved all 160 plans, and sent the agent back at least once in 16 runs.
- In 18 runs the agents built it in-house instead of picking a product.
- ALTCHA won 12 times, and all of those came from Cursor.
- Codex picked Cloudflare Turnstile in 40 of its 54 runs, the strongest run any agent gave one product.
The ranking160 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Cloudflare Turnstilecloudflare.com | 91 | 57% | |
| 2 | Vercel BotIDvercel.com | 18 | 11% | |
| 3 | Built in-houseoutcome | 18 | 11% | |
| 4 | ALTCHAaltcha.org | 12 | 8% | |
| 5 | django-axesgithub.com | 11 | 7% | |
| 6 | Google reCAPTCHAgoogle.com | 6 | 4% | |
| 7 | Laravel RateLimiterlaravel.com | 2 | 1% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol54 runs | Cloudflare Turnstile · 40then Vercel BotID · 7 |
| Claude Code · Claude Opus 554 runs | Cloudflare Turnstile · 31then django-axes · 3 |
| Cursor · Grok 4.652 runs | Cloudflare Turnstile · 20then ALTCHA · 12 |
By persona
| Junior developer63 runs | Cloudflare Turnstile · 35then Vercel BotID · 18 |
| Vibe coder61 runs | Cloudflare Turnstile · 45then ALTCHA · 8 |
| Enterprise team36 runs | django-axes · 11then Cloudflare Turnstile · 11 |
By what the ask stressed
| The plain ask160 runs | Cloudflare Turnstile · 91then Vercel BotID · 18 |
A case is one codebase with one agent, asked several times in different words and as different people. 13 of 15 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 5 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add bot protection to each of them, in several wordings and as a junior developer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 160 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 16 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown