Agent leaderboards / All sectors / Usage-based billing
Usage-based billing: which billing platforms coding agents choose
Metronome won about 45% of runs, and the wording changed most answers.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to bill customers for what they use in 11 codebases, across 293 runs. Each codebase was asked several times, in different words and as different people. Metronome led everywhere, and the rest of the field is spread thin behind it.
Asking the same thing twice changed the pick
A case here is one codebase with one agent, asked again in other words. Of 33 such cases, 31 did not land on the same product every time.
Cost at scale moved the order
When the ask was about volume and cost at scale, Lago came first with 22 wins in 53 runs. The overall leader took 8 of those runs.
Named everywhere, chosen almost never
Stripe Billing came up in 242 runs and was picked in five. Amberflo was named in 148 runs and picked once. OpenMeter was named in 142 and picked five times.
One telecom codebase brought its own shortlist
All 10 wins for CGRateS came in Marnsvik, a Django app rating call records nightly. PortaBilling and Oracle BRM won only there too.
- The simulated user approved every plan in the end, but sent the agent back at least once in 69 runs.
- In 22 runs it refused to approve until the agent named a specific product.
- The agents wrote the billing themselves in 40 runs, about 14%.
- Cursor put Lago second, while Claude Code and Codex put Orb there.
The ranking293 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | Metronomemetronome.com | 131 | 45% | |
| 2 | Orbwithorb.com | 53 | 18% | |
| 3 | Built in-houseoutcome | 40 | 14% | |
| 4 | Lagogetlago.com | 39 | 13% | |
| 5 | CGRateSgithub.com | 10 | 3% | |
| 6 | Stripe Billingstripe.com | 5 | 2% | |
| 7 | OpenMeteropenmeter.io | 5 | 2% | |
| 8 | PortaBilling | 3 | 1% | |
| 9 | Solvimonsolvimon.com | 2 | 1% | |
| 10 | m3term3ter.com | 2 | 1% | |
| 11 | Oracle BRMoracle.com | 2 | 1% | |
| 12 | Amberfloamberflo.io | 1 | 0% |
By agent, by persona, by wording
By agent
| Cursor · Grok 4.699 runs | Metronome · 47then Lago · 14 |
| Codex · GPT-5.6 Sol99 runs | Metronome · 54then Orb · 18 |
| Claude Code · Claude Opus 595 runs | Metronome · 30then Orb · 24 |
By persona
| Enterprise team132 runs | Metronome · 54then Orb · 20 |
| Senior engineer107 runs | Metronome · 49then Lago · 22 |
| Junior developer54 runs | Metronome · 28then Orb · 15 |
By what the ask stressed
| The plain ask213 runs | Metronome · 110then Orb · 36 |
| Volume and cost at scale53 runs | Lago · 22then Orb · 10 |
A case is one codebase with one agent, asked several times in different words and as different people. 31 of 33 cases did not hold to a single billing platform.
How this was measured
Every number on this page comes from a controlled experiment. We took 11 small applications, asked 3 coding agents (Cursor (Grok 4.6), Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5)) to bill customers for what they use in each of them, in several wordings and as an enterprise team and senior engineer and junior developer, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 293 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 69 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown