19 blind analysts, one 51-minute investigation, 43 adjudicated moments.

Nexpo's 51-minute ALLATRA investigation was analyzed independently by 19 AI analyst seats — 8 model families across effort tiers — each extracting every sourced claim, its on-screen evidence, and the video's source ledger. Trial 3 of the five-video study, and the first SOURCED trial: seats received the creator's own published 46-entry bibliography and grounded their ledgers against it.

Leaderboard

Click a column to sort. Accuracy = correct ÷ judged on the adjudicated visual axis (evidence-on-screen vs illustrative B-roll); abstentions ("my frames didn't cover this moment") are honest and excluded. Shown-hits = the genuine evidence-on-screen moments caught. Phantoms = entities adjudication rejected. Δ = declared source count vs the adjudicated truth.

Does effort buy accuracy?

Visual accuracy per family across its effort ladder, low → med → high → xhigh. Mostly: no — mid-effort seats match or beat their expensive siblings.

Two skills, rarely together

Seeing the screen correctly (y) vs finding the real sources (x). Bubble area = moments judged. The empty top-right corner is the finding.

How many sources does the film have?

Each seat enumerated the source ledger blind; adjudication settled the count (amber line).

Every seat, every moment

One row per seat, one cell per adjudicated moment in video order. ▲ columns are evidence-on-screen moments. Hover any cell; click a time label to open that moment in the watch-along player.
correct wrong abstained no vote

Do they know when they're right?

Mean self-reported confidence when correct (green) vs when wrong (red). Rows sorted by calibration gap — bottom rows were more confident while wrong.

Method

Each seat received the full transcript, a 150-frame visual kit (107 OCR-escalated — this video is wall-to-wall screen recordings), and the published bibliography with stable ids. Outputs were clustered across seats into 81 claim clusters; contested ones became the sourcing half of a watch-along interview answered by a human against the actual video.

Ground truth = the adjudicated moments plus the adjudicated source ledger, scored on the script axis (real evidence on screen vs illustrative), with unknown as an abstention. New this trial: bibliography grounding diagnostics (fan-out, collisions, per-seat coverage) and owner bib-audit rulings.

A blind shot census closes the interview: 43 detector intervals sampled from 296 with recorded inclusion weights, plus 12 random sentinels — no seat votes shown. It scores the four visual-specialist seats, which tiled the full 51 minutes with the abstention role live for the first time.