Agent leaderboards / All sectors / Product analytics
Product analytics: which products coding agents choose
PostHog won about 53% of runs across three agents.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three agents to add product analytics to 10 small apps, 359 runs in all. We varied the wording and who was asking. In about 23% of runs the agents skipped every product and wrote it themselves.
One theme handed the lead to another product
In 35 runs the ask was to fit an existing data stack. PostHog won none of those. Amplitude led them with six wins.
Most cases did not settle on one product
Take one codebase with one agent, asked the same thing several ways. In 22 of 30 such cases the runs did not all land on the same product.
Who asked changed the answer
Vibe coders picked Vercel Analytics in 27 of their 36 runs. Junior developers picked PostHog 131 times.
Named often, chosen almost never
Mixpanel came up in 222 runs and won five of them. Google Analytics came up in 91 runs and won none.
- The simulated user, who reads each plan before any code is written, approved all 359 runs.
- It sent the agent back at least once in 27 runs, and in three it refused until a product was named.
- Umami won six times, all of them in the Nuxt field service app.
- Datadog won twice, both times in the high-traffic checkout platform.
The ranking359 runs
| Product | Wins | Share | ||
|---|---|---|---|---|
| 1 | PostHogposthog.com | 189 | 53% | |
| 2 | Built in-houseoutcome | 81 | 23% | |
| 3 | Vercel Analyticsvercel.com | 27 | 8% | |
| 4 | Segmentsegment.com | 17 | 5% | |
| 5 | Amplitudeamplitude.com | 14 | 4% | |
| 6 | Umamiumami.is | 6 | 2% | |
| 7 | Mixpanelmixpanel.com | 5 | 1% | |
| 8 | Snowplowsnowplow.io | 2 | 1% | |
| 9 | Metabasemetabase.com | 2 | 1% | |
| 10 | Plausibleplausible.io | 2 | 1% | |
| 11 | Datadogdatadoghq.com | 2 | 1% | |
| 12 | Ahoygithub.com | 1 | 0% | |
| 13 | Mitzumitzu.io | 1 | 0% | |
| 14 | Datadog Product Analyticsdatadoghq.com | 1 | 0% |
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol120 runs | PostHog · 69then Amplitude · 11 |
| Claude Code · Claude Opus 5120 runs | PostHog · 49then Vercel Analytics · 9 |
| Cursor · Grok 4.6119 runs | PostHog · 71then Vercel Analytics · 10 |
By persona
| Junior developer144 runs | PostHog · 131then Amplitude · 1 |
| Senior engineer143 runs | PostHog · 44then Segment · 17 |
| Vibe coder36 runs | Vercel Analytics · 27then PostHog · 4 |
| Enterprise team36 runs | PostHog · 10then Amplitude · 2 |
By what the ask stressed
| The plain ask132 runs | PostHog · 77then Vercel Analytics · 15 |
| Read by someone who is not an engineer84 runs | PostHog · 67then Amplitude · 6 |
| Self-hosting, privacy or residency60 runs | PostHog · 22then Vercel Analytics · 12 |
| Fits an existing data stack35 runs | Amplitude · 6then Segment · 4 |
| Procurement and compliance24 runs | PostHog · 12then Snowplow · 1 |
| Volume and cost at scale24 runs | PostHog · 11then Datadog · 1 |
A case is one codebase with one agent, asked several times in different words and as different people. 22 of 30 cases did not hold to a single product.
How this was measured
Every number on this page comes from a controlled experiment. We took 10 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add product analytics to each of them, in several wordings and as a junior developer and senior engineer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 359 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 27 runs. Read the methodology and the publications.
Open the interactive boardThis page as Markdown