Agent leaderboards / All sectors / Product analytics

Product analytics: which products coding agents choose

PostHog won about 53% of runs across three agents.

359 runs10 apps3 agents4 personasupdated 2026-09-02

The interactive board, open on product analytics. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three agents to add product analytics to 10 small apps, 359 runs in all. We varied the wording and who was asking. In about 23% of runs the agents skipped every product and wrote it themselves.

0 of 35

One theme handed the lead to another product

In 35 runs the ask was to fit an existing data stack. PostHog won none of those. Amplitude led them with six wins.

22 of 30

Most cases did not settle on one product

Take one codebase with one agent, asked the same thing several ways. In 22 of 30 such cases the runs did not all land on the same product.

27 of 36

Who asked changed the answer

Vibe coders picked Vercel Analytics in 27 of their 36 runs. Junior developers picked PostHog 131 times.

222 vs 5

Named often, chosen almost never

Mixpanel came up in 222 runs and won five of them. Google Analytics came up in 91 runs and won none.

  • The simulated user, who reads each plan before any code is written, approved all 359 runs.
  • It sent the agent back at least once in 27 runs, and in three it refused until a product was named.
  • Umami won six times, all of them in the Nuxt field service app.
  • Datadog won twice, both times in the high-traffic checkout platform.
Explore every run in the interactive board

The ranking359 runs

ProductWinsShare
1 PostHogposthog.com 189 53%
2 Built in-houseoutcome 81 23%
3 Vercel Analyticsvercel.com 27 8%
4 Segmentsegment.com 17 5%
5 Amplitudeamplitude.com 14 4%
6 Umamiumami.is 6 2%
7 Mixpanelmixpanel.com 5 1%
8 Snowplowsnowplow.io 2 1%
9 Metabasemetabase.com 2 1%
10 Plausibleplausible.io 2 1%
11 Datadogdatadoghq.com 2 1%
12 Ahoygithub.com 1 0%
13 Mitzumitzu.io 1 0%
14 Datadog Product Analyticsdatadoghq.com 1 0%

By agent, by persona, by wording

By agent

Codex · GPT-5.6 Sol120 runsPostHog · 69then Amplitude · 11
Claude Code · Claude Opus 5120 runsPostHog · 49then Vercel Analytics · 9
Cursor · Grok 4.6119 runsPostHog · 71then Vercel Analytics · 10

By persona

Junior developer144 runsPostHog · 131then Amplitude · 1
Senior engineer143 runsPostHog · 44then Segment · 17
Vibe coder36 runsVercel Analytics · 27then PostHog · 4
Enterprise team36 runsPostHog · 10then Amplitude · 2

By what the ask stressed

The plain ask132 runsPostHog · 77then Vercel Analytics · 15
Read by someone who is not an engineer84 runsPostHog · 67then Amplitude · 6
Self-hosting, privacy or residency60 runsPostHog · 22then Vercel Analytics · 12
Fits an existing data stack35 runsAmplitude · 6then Segment · 4
Procurement and compliance24 runsPostHog · 12then Snowplow · 1
Volume and cost at scale24 runsPostHog · 11then Datadog · 1

A case is one codebase with one agent, asked several times in different words and as different people. 22 of 30 cases did not hold to a single product.

How this was measured

Every number on this page comes from a controlled experiment. We took 10 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add product analytics to each of them, in several wordings and as a junior developer and senior engineer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 359 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 27 runs. Read the methodology and the publications.

Open the interactive boardThis page as Markdown