← Armature

Agent usability benchmark · September 2026

Which cloud CLI can an agent actually deploy with?

We gave Claude Code and Codex three real projects. Each agent had one task: use the official CLI, create a new production project, and return a working public URL.

90attempts
5cloud CLIs
3project types
2coding agents

A verified URL does not prove a working app. We checked the public marker, then audited saved evidence for the full project. Placeholders and apps with broken core features do not count as completed projects.

Results

Agent deployment usability

Ordered by completed-project evidence. The separate URL-workflow score keeps the original rubric: marker proof, setup, configuration, recovery, and handoff. Setup friction counts.

URL marker proof · 40 Setup and auth · 20 CLI configuration · 15 Error recovery · 15 Handoff and cleanup · 10

Loading verified results…

Project completion is a post-run evidence audit, not an independent live functional test. Scores use the original URL-workflow rubric out of 100. URL intervals use Wilson 95%. Select a CLI for details.

What changed the result

Project and agent effects

We used a Vite client with an Express API and Postgres, a Nuxt app with sessions and Postgres, and a Flask service. This tests different build, data and runtime needs. Neither JavaScript project is static-only.

01

By project

Loading…

02

By agent

Loading…

Method

Real projects. Real accounts. Real URLs.

We ran every combination of five providers, three project types, two agents, and three repeats. Each run started in a clean sandbox. It received one provider credential and no installed provider CLI.

  1. Create. The agent had to make a new project with a unique exact name.
  2. Deploy. The agent had to use the provider’s official CLI and expose the app on a public production URL.
  3. Prove. Armature fetched the URL and looked for a unique marker in the page or same-origin assets.
  4. Clean. The harness tried to delete the exact test resource. Failed automatic cleanup remains in the score. We then audited and removed confirmed test leftovers.
  5. Judge. Gemini 3.7 scored saved, redacted evidence for setup, configuration, recovery, and handoff quality.
  6. Audit. After finding marker-only false positives, we added a separate review of code changes, deployment commands and recorded tests. This checks whether the full app and its required data services were deployed. It does not change the original score.

One completed Railway setup check used the credential that was replaced during launch. We excluded it before scoring and reran that cell with the final account token.

Scope limits: Cloudflare used a Pages-scoped token. No existing database credentials were supplied. Some tokens contained trailing whitespace, which agents had to fix. These account and credential choices affect the result. Runs overlapped, so shared rate limits also affected setup friction.

The 90 attempts and 11 setup checks are excluded from discoverability score aggregation. Public test resources were deleted after evidence capture. No claim is made that these URLs remain live.

Download the result data

This study measures agent usability. It is independent from Armature’s product recommendation scores.