← FERN · Hansa watch-along · calibration study

23 blind analysts, one documentary, 50 adjudicated moments.

A 24-minute investigative documentary was analyzed independently by 23 AI analyst seats — 7 model families across four effort tiers — each extracting every sourced claim, its on-screen evidence, and the video's source ledger, blind to one another. A human then adjudicated the disagreements moment-by-moment in a watch-along interview, and every on-screen citation was audited against the film's own 151-entry bibliography. This page scores the fleet against that ground truth.

Leaderboard

Click a column to sort. Accuracy = correct ÷ judged on the adjudicated visual axis (evidence-on-screen vs illustrative B-roll); abstentions ("my frames didn't cover this moment") are honest and excluded. Shown-hits = the genuine evidence-on-screen moments caught. Phantoms = entities adjudication rejected. Δ = declared source count vs the adjudicated truth.

Does effort buy accuracy?

Visual accuracy per family across its effort ladder, low → med → high → xhigh. Mostly: no — mid-effort seats match or beat their expensive siblings.

Two skills, rarely together

Seeing the screen correctly (y) vs finding the real sources (x). Bubble area = moments judged. The empty top-right corner is the finding.

How many sources does the film have?

Each seat enumerated the source ledger blind; adjudication settled the count (amber line).

Every seat, every moment

One row per seat, one cell per adjudicated moment in video order. ▲ columns are evidence-on-screen moments. Hover any cell; click a time label to open that moment in the watch-along player.
correct wrong abstained no vote

Do they know when they're right?

Mean self-reported confidence when correct (green) vs when wrong (red). Rows sorted by calibration gap — bottom rows were more confident while wrong.

Method

Each seat received the full transcript plus a 68-frame visual kit and produced structured claims: kind, origin, on-screen evidence role, and a self-enumerated source ledger — with per-row confidence. Outputs were clustered across seats into 108 claim clusters; the contested ones became a 66-question watch-along interview answered by a human against the actual video.

Ground truth here = the 50 adjudicated moments (collapsed to the one visual distinction that changes a production script: real evidence on screen vs illustrative footage) plus the adjudicated 16-source ledger. Separately, all 80 question-moments' corner citations were traced into the film's own numbered bibliography and read at the source — surfacing citation drift, one stuffed citation reused across four beats, and two claims their own sources don't support.

Seats are labelled by the model that actually answered, read from each call's own usage record — not by the alias requested on the command line. That distinction matters here: the four seats originally labelled haiku were silently served by claude-sonnet-5, because --permission-mode plan overrides a weak model without erroring. They ran clean and reported success for two trials. They appear here as sonnet 5 ·B — a second, independent run of sonnet 5 at the same four effort tiers, kept rather than deleted because it is the study's only measure of run-to-run variance. haiku 4.5 is the corrected re-run and is the first real measurement of that family. Compare sonnet 5 against sonnet 5 ·B before reading any small gap between two different models as a difference in capability.