Agent leaderboards / All sectors / Maps

Maps: which map products coding agents choose

Google Maps and Leaflet tie at about 16% each, with no clear leader.

319 runs14 apps3 agents4 personasupdated 2026-09-11

The interactive board, open on maps. Open it full page · this page as Markdown

Read this leaderboard as textrankings, key learnings, method

Key learnings

We asked three coding agents to add a map to 14 small apps, 319 times in all, in different wordings and as four different people. Google Maps and Leaflet each won 52 runs. The rest is spread across more than twenty products and combinations.

30 of 100

Each agent had its own favourite

Cursor picked Leaflet in 30 of its 100 runs. Codex led with Google Maps at 27, and Claude Code led with Apple Maps (MapKit) at 16.

32 of 41

The same job in other words got other answers

We counted 41 cases, each one codebase with one agent asked several times over. In 32 of them the runs did not all land on the same product.

13 of 36

Who asked moved the order

Enterprise teams picked MapLibre 13 times in 36 runs, the only group where it led. Senior engineers picked Apple Maps (MapKit) 24 times.

0 of 63

One product was named often and never chosen

Protomaps came up in 63 runs and won none of them.

  • The agents wrote the map themselves in 16 runs, about 5%.
  • The simulated user sent the agent back at least once in 16 runs, and in 12 of those it refused until a product was named.
  • Nominatim was mentioned in 82 runs and won none, but it turns addresses into coordinates, not maps.
  • Azure Maps won 18 runs, all in the .NET app.
Explore every run in the interactive board

The ranking319 runs

ProductWinsShare
1 Google Mapsmapsplatform.google.com 52 16%
2 Leafletleafletjs.com 52 16%
3 Apple Maps (MapKit)developer.apple.com 46 14%
4 Mapboxmapbox.com 25 8%
5 MapLibremaplibre.org 21 7%
6 react-native-mapsgithub.com 19 6%
7 Azure Mapsazure.microsoft.com 18 6%
8 Built in-houseoutcome 16 5%
9 MapTilermaptiler.com 15 5%
10 OpenStreetMapopenstreetmap.org 10 3%
11 flutter_mapdocs.fleaflet.dev 8 3%
12 Geoapifygeoapify.com 5 2%
13 HEREhere.com 3 1%
14 Amazon Locationaws.amazon.com 3 1%
15 MapLibre + MapTiler 3 1%
16 ArcGIS (Esri)esri.com 2 1%
17 Mapbox + Searoutes 2 1%
18 Leaflet + MapTiler 2 1%
19 Mapbox + react-map-gl (vis.gl) + Searoutes 2 1%
20 Leaflet + Maa-amet (Estonian Land Board) 2 1%
21 GraphHopper + MapLibre 1 0%
22 Maa-amet (Estonian Land Board)maaamet.ee 1 0%
23 Maa-amet (Estonian Land Board) + MapLibre + MapTiler 1 0%
24 Base Adresse Nationale + MapLibre + Protomaps 1 0%
25 Azure Maps + Leaflet 1 0%
26 IGN Géoplateforme + Leaflet 1 0%
27 Geocoder (Ruby gem) + Leaflet + Nominatim + OpenStreetMap 1 0%
28 Base Adresse Nationale + IGN Géoplateforme + Leaflet 1 0%
29 flutter_map + Stadia Maps 1 0%
30 flutter_map + Thunderforest 1 0%
31 Amazon Location + MapLibre 1 0%
32 OpenTopoMapopentopomap.org 1 0%
33 MapLibre + OpenFreeMap 1 0%

By agent, by persona, by wording

By agent

Codex · GPT-5.6 Sol111 runsGoogle Maps · 27then Mapbox · 15
Claude Code · Claude Opus 5108 runsApple Maps (MapKit) · 16then Google Maps · 13
Cursor · Grok 4.6100 runsLeaflet · 30then Apple Maps (MapKit) · 16

By persona

Junior developer147 runsGoogle Maps · 32then Leaflet · 29
Senior engineer72 runsApple Maps (MapKit) · 24then Mapbox · 15
Vibe coder64 runsApple Maps (MapKit) · 22then Leaflet · 13
Enterprise team36 runsMapLibre · 13then Mapbox · 6

By what the ask stressed

The plain ask307 runsLeaflet · 51then Google Maps · 48

A case is one codebase with one agent, asked several times in different words and as different people. 32 of 41 cases did not hold to a single map product.

How this was measured

Every number on this page comes from a controlled experiment. We took 14 small applications, asked 3 coding agents (Codex (GPT-5.6 Sol), Claude Code (Claude Opus 5), Cursor (Grok 4.6)) to add a map to each of them, in several wordings and as a junior developer and senior engineer and vibe coder and enterprise team, and let the agent choose the product. Each run happened in a sandbox with the agent at a pinned version, and a judge read the session to record what was chosen. That is 319 runs. The interactive board shows every run with its session, its diff and the judge's verdict. A simulated user stood in for the owner of the codebase: it read the agent's plan and had to approve it before any code was written; it sent the agent back at least once in 16 runs. Read the methodology and the publications.

If you sell in this sector: what these numbers mean for a vendor.

Open the interactive boardThis page as Markdown