Armature Evals is agents evals for your MCP and your CLI: describe a user workflow in one sentence, run it on the harnesses and models your users actually use, and know whether your last change broke agent behavior before your users do.
Three lines of code capture intent and turn raw tool-call logs into reconstructed sessions: real user intent, agent reasoning, clustered use cases, and the failures costing you the most.