Agent leaderboards / All sectors / Payments

Payments: which providers coding agents choose

Stripe won about 88% of the runs to add payments.

395 runs11 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on payments. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add payments to 11 small apps, across 395 runs. Each codebase came up several times, worded differently and asked as four different people. Behind Stripe, Paddle came second with 14 wins.

16 of 33

Half the cases didn't settle on one product

Take one codebase with one agent, asked again and again with the words changed. In 16 of 33 such cases, the runs did not all land on the same product.

3 of 175

Named often, chosen rarely

Adyen came up in 175 runs and won three of them. PayPal was named in 139 runs and won none. Square was named in 104 and won one.

3 of 395

A few picks showed up in one codebase only

Bottomline PTX won three runs and AccessPay two, all in the .NET utility billing service. Helcim won once, in the Next.js class booking app.

108 of 108

One persona never varied

Runs as junior developers picked Stripe every time. As senior engineers, Stripe took 85 runs and Paddle 14.

  • The simulated user approved all 395 plans, and sent the agent back at least once in 46 runs.
  • In 34 runs it refused to approve until the agent named a specific provider.
  • The procurement and compliance wording didn't change the order, with Stripe taking 35 of those 36 runs.
  • Second place changed with the agent: Mollie for Cursor, Paddle for Codex, Bottomline PTX for Claude Code.
Explore every run in the interactive board

The ranking395 runs

ProductWinsShare
1 Stripestripe.com 349 88%
2 Paddlepaddle.com 14 4%
3 Molliemollie.com 13 3%
4 GoCardlessgocardless.com 9 2%
5 Adyenadyen.com 3 1%
6 Bottomline PTXbottomline.com 3 1%
7 AccessPayaccesspay.com 2 1%
8 Helcimhelcim.com 1 0%
9 Squaresquareup.com 1 0%

By agent, by persona, by wording

By agent

Codex · GPT-5.6 Sol132 runsStripe · 110then Paddle · 13
Claude Code · Claude Opus 5132 runsStripe · 126then Bottomline PTX · 2
Cursor · Grok 4.6131 runsStripe · 113then Mollie · 9

By persona

Junior developer108 runsStripe · 108
Enterprise team108 runsStripe · 91then GoCardless · 9
Senior engineer108 runsStripe · 85then Paddle · 14
Vibe coder71 runsStripe · 65then Mollie · 4

By what the ask stressed

The plain ask359 runsStripe · 314then Paddle · 14
Procurement and compliance36 runsStripe · 35then GoCardless · 1

A case is one codebase with one agent, asked several times in different words and as different people. 16 of 33 cases did not hold to a single provider.

How this was measured

Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add payments to each of them, in several wordings and as a junior developer and enterprise team and senior engineer and vibe coder, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 395 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 46 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown