What we learn from running real coding agents on the repository panel: how the selection works, what a vendor should do about it, and what we ship.
We ran Claude Code and Codex inside 51 realistic repositories, on 481 frozen prompts, across 18 developer tool sectors. A judge read every session. How we built the experiment, what the agents chose, and where the method is weak.
Read the article →
Real agents replay your users' workflows, and a judge scores every run. Catch regressions before you ship.
See what users do on your MCP: sessions rebuilt end to end, grouped into use cases and issues.
The search engine we made to mimic how coding agents search the web, and what it lets us test.
Not published yet
We are still working on . We share drafts and early results with vendors who ask.
Or write to contact@armature.tech.