Agent leaderboards / All sectors / Maps
Maps: which map products coding agents choose
Google Maps and Leaflet tie at about 16% each, with no clear leader.
Read this leaderboard as textrankings, key learnings, method
Key learnings
We asked three coding agents to add a map to 14 small apps, 319 times in all, in different wordings and as four different people. Google Maps and Leaflet each won 52 runs. The rest is spread across more than twenty products and combinations.
Each agent had its own favourite
Cursor picked Leaflet in 30 of its 100 runs. Codex led with Google Maps at 27, and Claude Code led with Apple Maps (MapKit) at 16.
The same job in other words got other answers
We counted 41 cases, each one codebase with one agent asked several times over. In 32 of them the runs did not all land on the same product.
Who asked moved the order
Enterprise teams picked MapLibre 13 times in 36 runs, the only group where it led. Senior engineers picked Apple Maps (MapKit) 24 times.
One product was named often and never chosen
Protomaps came up in 63 runs and won none of them.
- The agents wrote the map themselves in 16 runs, about 5%.
- The simulated user sent the agent back at least once in 16 runs, and in 12 of those it refused until a product was named.
- Nominatim was mentioned in 82 runs and won none, but it turns addresses into coordinates, not maps.
- Azure Maps won 18 runs, all in the .NET app.
The ranking319 runs
By agent, by persona, by wording
By agent
| Codex · GPT-5.6 Sol111 runs | Google Maps · 27then Mapbox · 15 |
| Claude Code · Claude Opus 5108 runs | Apple Maps (MapKit) · 16then Google Maps · 13 |
| Cursor · Grok 4.6100 runs | Leaflet · 30then Apple Maps (MapKit) · 16 |
By persona
| Junior developer147 runs | Google Maps · 32then Leaflet · 29 |
| Senior engineer72 runs | Apple Maps (MapKit) · 24then Mapbox · 15 |
| Vibe coder64 runs | Apple Maps (MapKit) · 22then Leaflet · 13 |
| Enterprise team36 runs | MapLibre · 13then Mapbox · 6 |
By what the ask stressed
| The plain ask307 runs | Leaflet · 51then Google Maps · 48 |
A case is one codebase with one agent, asked several times in different words and as different people. 32 of 41 cases did not hold to a single map product.
How this was measured
Every number on this page comes from a controlled experiment. We took 14 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add a map to each of them, in several wordings and as a junior developer and senior engineer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 319 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 16 runs. Read the methodology and the publications.
If you sell in this sector: what these numbers mean for a vendor.
Open the interactive boardThis page as Markdown