A 24-minute investigative documentary was analyzed independently by 23 AI analyst seats — 7 model families across four effort tiers — each extracting every sourced claim, its on-screen evidence, and the video's source ledger, blind to one another. A human then adjudicated the disagreements moment-by-moment in a watch-along interview, and every on-screen citation was audited against the film's own 151-entry bibliography. This page scores the fleet against that ground truth.
Each seat received the full transcript plus a 68-frame visual kit and produced structured claims: kind, origin, on-screen evidence role, and a self-enumerated source ledger — with per-row confidence. Outputs were clustered across seats into 108 claim clusters; the contested ones became a 66-question watch-along interview answered by a human against the actual video.
Ground truth here = the 50 adjudicated moments (collapsed to the one visual distinction that changes a production script: real evidence on screen vs illustrative footage) plus the adjudicated 16-source ledger. Separately, all 80 question-moments' corner citations were traced into the film's own numbered bibliography and read at the source — surfacing citation drift, one stuffed citation reused across four beats, and two claims their own sources don't support.
Seats are labelled by the model that actually answered, read from each call's own usage record — not by the alias requested on the command line. That distinction matters here: the four seats originally labelled haiku were silently served by claude-sonnet-5, because --permission-mode plan overrides a weak model without erroring. They ran clean and reported success for two trials. They appear here as sonnet 5 ·B — a second, independent run of sonnet 5 at the same four effort tiers, kept rather than deleted because it is the study's only measure of run-to-run variance. haiku 4.5 is the corrected re-run and is the first real measurement of that family. Compare sonnet 5 against sonnet 5 ·B before reading any small gap between two different models as a difference in capability.