Launching Armature Evals - Agent Evaluations for MCPs and CLIs
Last week we launched Armature Analytics and ended with a promise: Armature would soon verify the fixes to the issues it finds.
Today we keep it. We’re launching Armature Evals, so teams shipping MCPs and CLIs know whether their last change broke agent behavior before their users do.
Your first verdict is three minutes away, and 100 runs a month are free.
Suites, cases, verdicts
Teams that ship agents already live on eval suites: cases describing what the agent must handle, rerun on every change, judged rather than asserted. Your MCP and your CLI are what your end-users’ agents use, and these agent-facing interfaces never got evals of their own.
Armature Evals is agents evals for your MCP and your CLI.
Say your MCP handles payments. You ship a change, you run your suite on the harnesses your users actually use, and one case fails: the refund flow broke on Codex and nowhere else, caught before a single user hit it.
An eval case is one user workflow and the criteria that define success, written in one sentence: a user should be able to get a refund on their last order.

You write the sentence and Armature drafts the case; it can also propose some from your MCP’s tool list or from real sessions Armature Analytics saw.
When the suite is green, you ship.
Evals, not tests
A test asserts: it expects the same output for the same input and treats any difference as a bug. Agents take a different path every time, even on the same prompt, and most of those paths are fine.
So you judge instead of asserting: fix the criteria, run across harnesses and models, read a pass rate. It is how labs measure models, and we point the same method at your product.
The half of the contract you don’t control
We built this because we kept getting burned: we shipped agents at scale, something changed, and we heard about it from a user instead of from a test. Software has 25 years of tooling to catch this before users do. Agent surfaces got none: you ship, you try it by hand once, and you hope.
And your MCP is only one half of the contract. The other half is Claude Code, ChatGPT, Cursor, and whatever ships next. You control none of them: a harness can ship a new version any day and break your product while you changed nothing.
Browsers changed slowly and followed a spec. Models and harnesses change every week and follow nothing.
So you test every change you ship, on every agent your users bring. That is what eval suites are for.
Step two of the Agent Experience
Seeing was step one, verifying is step two.
Armature Analytics shows what broke, Evals prove the fix held, and because Armature exposes an MCP too, your coding agent can close the loop alone: read the failing case, ship the fix, rerun the suite.
In UX you own the client: it only changes when you change it. In AX the client is someone else’s harness, running someone else’s model.
The UX stack had decades to mature. We are rebuilding it for agents, and this is the second piece.
Armature Evals is available today
It is self-serve: connect a remote MCP server or a CLI, describe a use case, read your verdict. Group your cases into a suite and run it on every release. 100 free runs a month!
Please reach out with any question or feedback.
