← Lab Measured run2026-08-20

Does its taxonomy match one researchers published?

It described these transcripts in 5 concrete, task-level modes. MAST, a published taxonomy for the same corpus, describes process-level properties — so none of the 5 map onto it. Two vocabularies, same transcripts.

A different level of description
5modes discovered, all task-level
19real multi-agent traces
3human annotators each
n/aκ — no matched pair to score

What we set out to measure, and what we measured

MAST coverage (of 11 modes occurring >= 3x)>= 0.50/11under the bar
judge-vs-human kappa (pooled)>= 0.4NaNunder the bar

What we did

19 multi-agent traces from six frameworks, each labelled by three annotators against MAST. We ran the full transcripts blind, then matched our modes to MAST's with a different model family, in both orders.

What came back

Nothing mapped, so agreement with the annotators could not be computed. The matcher is sound — probed directly it pairs "Repeating Identical Plan in a Loop" with MAST's "Step repetition" and rejects unrelated pairs. The difference is real: ours name what went wrong in the task ("Hallucination of Code Artifacts"), MAST names properties of the process ("Lack of result verification").

What it changed

Nothing was tuned toward MAST, deliberately. Bending a discovery pipeline until it reproduces someone else's list defeats the point of reading your traces.

What it doesn't show

19 traces is small and every interval is wide. These transcripts are plain text, so tool calls aren't linked to results — tool-correctness modes are out of reach here by construction.

MODEL google/gemini-3.7-flash · SAMPLE 19 traces, seed 42 · COST $0.92
REPRODUCE npm run validate:mad

Full metric tables and every caveat

Run it on your own traces.

Free, no account, parsed in your browser.

Upload a trace file