Agent leaderboards / All sectors / Payments
Payments: which providers coding agents choose
Stripe won about 88% of the payment runs across 11 small apps.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add payments to 11 small apps, 395 runs in all. We varied the wording of the ask and the person doing the asking. Second place went to Paddle, with 14 wins.
Named in plans, rarely chosen
Adyen came up in 175 runs and was picked in three. PayPal was named in 139 runs and picked in none. Square was named in 104 and won once.
Half the cases didn't settle on one product
We took each codebase with each agent and asked several times, in different words and as different people. Of 33 such cases, 16 did not come out the same every time.
Who asked moved the second pick
Junior developers chose Stripe in all 108 of their runs. Senior engineers chose it 85 times and gave Paddle 14. Enterprise teams gave GoCardless nine.
One codebase produced names no other did
The .NET utility billing service, a machine-facing billing app in C#, is where Bottomline PTX won three runs. AccessPay won two there. Neither won anywhere else.
- The simulated user, who approves each plan, approved all 395 runs, but sent the agent back at least once in 46.
- In 34 runs it refused to approve until the agent named a specific product.
- Asking with procurement and compliance in mind changed little: Stripe took 35 of those 36 runs.
The ranking395 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Stripestripe.com | 349 | 88% | |
| 2 | Paddlepaddle.com | 14 | 4% | |
| 3 | Molliemollie.com | 13 | 3% | |
| 4 | GoCardlessgocardless.com | 9 | 2% | |
| 5 | Adyenadyen.com | 3 | 1% | |
| 6 | Bottomline PTXbottomline.com | 3 | 1% | |
| 7 | AccessPayaccesspay.com | 2 | 1% | |
| 8 | Helcimhelcim.com | 1 | 0% | |
| 9 | Squaresquareup.com | 1 | 0% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol132 runs | Stripe · 110then Paddle · 13 |
| Claude Code · Claude Opus 5132 runs | Stripe · 126then Bottomline PTX · 2 |
| Cursor · Grok 4.6131 runs | Stripe · 113then Mollie · 9 |
By persona
| Junior developer108 runs | Stripe · 108 |
| Enterprise team108 runs | Stripe · 91then GoCardless · 9 |
| Senior engineer108 runs | Stripe · 85then Paddle · 14 |
| Vibe coder71 runs | Stripe · 65then Mollie · 4 |
By what the ask stressed
| The plain ask359 runs | Stripe · 314then Paddle · 14 |
| Procurement and compliance36 runs | Stripe · 35then GoCardless · 1 |
A case is one codebase with one agent, asked several times in different words and as different people. 16 of 33 cases did not hold to a single provider.
How this was measured
Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add payments to each of them, in several wordings and as a junior developer and enterprise team and senior engineer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 395 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 46 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown