From the experiment

Do coding agents recommend Inspect AI?

Inspect AI was chosen in 4% of 288 judged evals sessions, ranking fifth. Measured with Claude Code, Codex and Cursor.

Published September 3, 2026 Read as Markdown

Inspect AI was chosen in 4% of 288 judged evals sessions, ranking fifth. It was also raised as a candidate in 30 further sessions without being chosen.

This page reports what happened when Claude Code, Codex and Cursor had to solve a problem in evals inside a realistic codebase. Not what a chat assistant says about Inspect AI. What an agent actually installed.

One thing to read first: every one of those wins came from a single codebase. That is a result about one repository rather than about evals in general, and the splits below cannot separate the two. Treat them as a description of that repository.

The numbers

CategoryEvals
Sessions in the category288
Sessions where Inspect AI was chosen12
Install share4%
Rank in category5 of 9
Codebases it won in1
Raised as a candidate, not chosen30
Chosen when considered29%
Siteinspect.aisi.org.uk

By agent

With 12 wins spread across three agents, the rates below are small numbers and a difference between them is not yet a finding. They are here because the direction is worth knowing, not because the gap is established.

AgentSessionsChose Inspect AIShare
Claude Code9633%
Codex9600%
Cursor9699%

By who was asking

Inspect AI performs similarly across the four kinds of buyer, from 0% to 8%. That is unusual: the category leader changed with the persona in 14 of the 18 categories we measured.

Who is askingSessionsChose Inspect AIShare
Junior developer14500%
Senior engineer143128%

What Inspect AI was up against

The full ranking in evals, from the same sessions:

#ProductRuns wonShare
1Langfuse9734%
2Built in-house (no product adopted)8429%
3Promptfoo3813%
4Arize Phoenix238%
5Braintrust186%
6Inspect AI (this page)124%
7LangSmith83%
8Helicone62%

What this means

Inspect AI was raised in 30 sessions and chosen in 12. That ratio is balanced enough that the ceiling is presence rather than integration: the product converts reasonably when it is on the table, and it is not on the table often enough.

Where these numbers come from

The 288 sessions in evals are part of a published set of 5,292, run with real coding agents inside realistic codebases and judged blind. The full method is on one page: how we measured this.

Every evals run can be replayed on the board.

If you work on Inspect AI: the judge recorded a reason for every session where it was raised and passed over. Those reasons are in the transcripts.

<!-- generated by scripts/write-data-pages.mjs -->

Common questions

Do coding agents recommend Inspect AI?

Yes. Inspect AI was chosen in 12 of the 288 judged sessions in evals, a 4% install share, ranking fifth in its category.

Does Claude Code recommend Inspect AI?

In 3 of the 96 sessions in evals run with Claude Code, which is 3%.

Do different coding agents treat Inspect AI differently?

Not much. The three agents chose it at similar rates, between 0% and 9% of their runs.

How was this measured?

Real coding agents at pinned versions were run in sandboxes inside 51 realistic codebases and asked to solve real tasks. A simulated project owner approved or questioned each recommendation before any code was written, and a judge from a model family that builds none of the agents read every session blind.

How often is Inspect AI considered but not chosen?

It was raised as a candidate in 30 sessions without being chosen, and chosen in 12. That is a 29% conversion from considered to chosen.

Where this comes from

Armature ran 5,292 judged sessions with Claude Code, Codex and Cursor inside 51 realistic codebases, and published every run. The numbers on this page come from that work.

Read next

All library pages