Lab

What we tested, and what came back.

A tool that grades your agent should be measured itself. Every run below is published with its numbers, its limits and the one command that reproduces it.

Measured runs — every figure read from the run's own artifact

Found while building — what we learned and hardened

Not measured yet

  • τ²-bench. Wired up; the first run lost part of its judge matrix to an upstream limit, so it was discarded rather than scored. It appears here when it completes cleanly.
  • Whether the modes it finds are the ones you'd prioritise. Needs teams, not benchmarks.
  • Whether any of it changes what a team ships. The number we most want and least have.

Prefer raw tables? Every metric, threshold and caveat →