Agent leaderboards / All sectors / Usage-based billing

Usage-based billing: which billing platforms coding agents choose

Metronome won about 45% of runs, and the wording changed most answers.

293 runs11 apps3 agents3 personasupdated 2026-09-14

The interactive board, open on usage-based billing. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to bill customers for what they use in 11 codebases, across 293 runs. Each codebase was asked several times, in different words and as different people. Metronome led everywhere, and the rest of the field is spread thin behind it.

31 of 33

Asking the same thing twice changed the pick

A case here is one codebase with one agent, asked again in other words. Of 33 such cases, 31 did not land on the same product every time.

22 of 53

Cost at scale moved the order

When the ask was about volume and cost at scale, Lago came first with 22 wins in 53 runs. The overall leader took 8 of those runs.

242 vs 5

Named everywhere, chosen almost never

Stripe Billing came up in 242 runs and was picked in five. Amberflo was named in 148 runs and picked once. OpenMeter was named in 142 and picked five times.

10 wins

One telecom codebase brought its own shortlist

All 10 wins for CGRateS came in Marnsvik, a Django app rating call records nightly. PortaBilling and Oracle BRM won only there too.

  • The simulated user approved every plan in the end, but sent the agent back at least once in 69 runs.
  • In 22 runs it refused to approve until the agent named a specific product.
  • The agents wrote the billing themselves in 40 runs, about 14%.
  • Cursor put Lago second, while Claude Code and Codex put Orb there.
Explore every run in the interactive board

The ranking293 runs

ProductWinsShare
1 Metronomemetronome.com 131 45%
2 Orbwithorb.com 53 18%
3 Built in-houseoutcome 40 14%
4 Lagogetlago.com 39 13%
5 CGRateSgithub.com 10 3%
6 Stripe Billingstripe.com 5 2%
7 OpenMeteropenmeter.io 5 2%
8 PortaBilling 3 1%
9 Solvimonsolvimon.com 2 1%
10 m3term3ter.com 2 1%
11 Oracle BRMoracle.com 2 1%
12 Amberfloamberflo.io 1 0%

By agent, by persona, by wording

By agent

Cursor · Grok 4.699 runsMetronome · 47then Lago · 14
Codex · GPT-5.6 Sol99 runsMetronome · 54then Orb · 18
Claude Code · Claude Opus 595 runsMetronome · 30then Orb · 24

By persona

Enterprise team132 runsMetronome · 54then Orb · 20
Senior engineer107 runsMetronome · 49then Lago · 22
Junior developer54 runsMetronome · 28then Orb · 15

By what the ask stressed

The plain ask213 runsMetronome · 110then Orb · 36
Volume and cost at scale53 runsLago · 22then Orb · 10

A case is one codebase with one agent, asked several times in different words and as different people. 31 of 33 cases did not hold to a single billing platform.

How this was measured

Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5)) to bill customers for what they use in each of them, in several wordings and as an enterprise team and senior engineer and junior developer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 293 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 69 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown