By project
Loading…
Agent usability benchmark · September 2026
We gave Claude Code and Codex three real projects. Each agent had one task: use the official CLI, create a new production project, and return a working public URL.
A verified URL does not prove a working app. We checked the public marker, then audited saved evidence for the full project. Placeholders and apps with broken core features do not count as completed projects.
Results
Ordered by completed-project evidence. The separate URL-workflow score keeps the original rubric: marker proof, setup, configuration, recovery, and handoff. Setup friction counts.
Loading verified results…
Project completion is a post-run evidence audit, not an independent live functional test. Scores use the original URL-workflow rubric out of 100. URL intervals use Wilson 95%. Select a CLI for details.
What changed the result
We used a Vite client with an Express API and Postgres, a Nuxt app with sessions and Postgres, and a Flask service. This tests different build, data and runtime needs. Neither JavaScript project is static-only.
Loading…
Loading…
Method
We ran every combination of five providers, three project types, two agents, and three repeats. Each run started in a clean sandbox. It received one provider credential and no installed provider CLI.
One completed Railway setup check used the credential that was replaced during launch. We excluded it before scoring and reran that cell with the final account token.
Scope limits: Cloudflare used a Pages-scoped token. No existing database credentials were supplied. Some tokens contained trailing whitespace, which agents had to fix. These account and credential choices affect the result. Runs overlapped, so shared rate limits also affected setup friction.
The 90 attempts and 11 setup checks are excluded from discoverability score aggregation. Public test resources were deleted after evidence capture. No claim is made that these URLs remain live.
This study measures agent usability. It is independent from Armature’s product recommendation scores.