Lab
What we tested, and what came back.
A tool that grades your agent should be measured itself. Every run below is published with its numbers, its limits and the one command that reproduces it.
Measured runs — every figure read from the run's own artifact
Found while building — what we learned and hardened
Not measured yet
- τ²-bench. Wired up; the first run lost part of its judge matrix to an upstream limit, so it was discarded rather than scored. It appears here when it completes cleanly.
- Whether the modes it finds are the ones you'd prioritise. Needs teams, not benchmarks.
- Whether any of it changes what a team ships. The number we most want and least have.