Agent leaderboards / All sectors / Bot protection
Bot protection: which products coding agents choose
Cloudflare Turnstile won about 57% of 160 runs.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three agents to add bot protection to five small apps, over 160 runs, in different wordings and as three different people. Cloudflare Turnstile led with every agent. Everything else was small, and the agents wrote their own protection in 18 runs.
One persona changed the order
With enterprise teams the leader stopped leading. django-axes and Cloudflare Turnstile split those runs almost evenly, 11 wins each. Vibe coders and junior developers both put Cloudflare Turnstile well on top.
The wording moved the answer
A case here is one codebase with one agent, asked several times in different words. In 13 of 15 cases the runs did not all land on the same product.
Three products belong to one codebase each
All 11 wins for django-axes came in the Django learning platform, and so did the six for Google reCAPTCHA. The two wins for Laravel RateLimiter were all in the Laravel helpdesk.
- The simulated user approved every plan, and sent the agent back at least once in 16 runs.
- Vercel BotID won 18 runs, all of them with junior developers.
- Cursor picked Cloudflare Turnstile in 20 of its 52 runs, the lowest share of the three agents.
- Codex was the most consistent, with 40 of its 54 runs going to Cloudflare Turnstile.
The ranking160 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Cloudflare Turnstilecloudflare.com | 91 | 57% | |
| 2 | Vercel BotIDvercel.com | 18 | 11% | |
| 3 | Built in-houseoutcome | 18 | 11% | |
| 4 | ALTCHAaltcha.org | 12 | 8% | |
| 5 | django-axesgithub.com | 11 | 7% | |
| 6 | Google reCAPTCHAgoogle.com | 6 | 4% | |
| 7 | Laravel RateLimiterlaravel.com | 2 | 1% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol54 runs | Cloudflare Turnstile · 40then Vercel BotID · 7 |
| Claude Code · Claude Opus 554 runs | Cloudflare Turnstile · 31then django-axes · 3 |
| Cursor · Grok 4.652 runs | Cloudflare Turnstile · 20then ALTCHA · 12 |
By persona
| Junior developer63 runs | Cloudflare Turnstile · 35then Vercel BotID · 18 |
| Vibe coder61 runs | Cloudflare Turnstile · 45then ALTCHA · 8 |
| Enterprise team36 runs | django-axes · 11then Cloudflare Turnstile · 11 |
By what the ask stressed
| The plain ask160 runs | Cloudflare Turnstile · 91then Vercel BotID · 18 |
A case is one codebase with one agent, asked several times in different words and as different people. 13 of 15 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 5 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add bot protection to each of them, in several wordings and as a junior developer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 160 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 16 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown