Agent leaderboards / All sectors / Payments

Payments: which providers coding agents choose

Stripe won about 88% of the payment runs across 11 small apps.

395 runs11 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on payments. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add payments to 11 small apps, 395 runs in all. We varied the wording of the ask and the person doing the asking. Second place went to Paddle, with 14 wins.

175 vs 3

Named in plans, rarely chosen

Adyen came up in 175 runs and was picked in three. PayPal was named in 139 runs and picked in none. Square was named in 104 and won once.

16 of 33

Half the cases didn't settle on one product

We took each codebase with each agent and asked several times, in different words and as different people. Of 33 such cases, 16 did not come out the same every time.

108 of 108

Who asked moved the second pick

Junior developers chose Stripe in all 108 of their runs. Senior engineers chose it 85 times and gave Paddle 14. Enterprise teams gave GoCardless nine.

3 wins

One codebase produced names no other did

The .NET utility billing service, a machine-facing billing app in C#, is where Bottomline PTX won three runs. AccessPay won two there. Neither won anywhere else.

  • The simulated user, who approves each plan, approved all 395 runs, but sent the agent back at least once in 46.
  • In 34 runs it refused to approve until the agent named a specific product.
  • Asking with procurement and compliance in mind changed little: Stripe took 35 of those 36 runs.
Explore every run in the interactive board

The ranking395 runs

ProductWinsShare
1 Stripestripe.com 349 88%
2 Paddlepaddle.com 14 4%
3 Molliemollie.com 13 3%
4 GoCardlessgocardless.com 9 2%
5 Adyenadyen.com 3 1%
6 Bottomline PTXbottomline.com 3 1%
7 AccessPayaccesspay.com 2 1%
8 Helcimhelcim.com 1 0%
9 Squaresquareup.com 1 0%

By agent, by persona, by wording

By agent

Codex · GPT-5.6 Sol132 runsStripe · 110then Paddle · 13
Claude Code · Claude Opus 5132 runsStripe · 126then Bottomline PTX · 2
Cursor · Grok 4.6131 runsStripe · 113then Mollie · 9

By persona

Junior developer108 runsStripe · 108
Enterprise team108 runsStripe · 91then GoCardless · 9
Senior engineer108 runsStripe · 85then Paddle · 14
Vibe coder71 runsStripe · 65then Mollie · 4

By what the ask stressed

The plain ask359 runsStripe · 314then Paddle · 14
Procurement and compliance36 runsStripe · 35then GoCardless · 1

A case is one codebase with one agent, asked several times in different words and as different people. 16 of 33 cases did not hold to a single provider.

How this was measured

Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add payments to each of them, in several wordings and as a junior developer and enterprise team and senior engineer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 395 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 46 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown