Plain-language explainer for the 2026-08 cross-vendor calibration study (nexpo-h2h + the FERN trials: room1046, cryptoqueen, cocacola). Saved as a candidate for a public website page. Names models because it is empirical model-study data (calibration carve-out).
What's the cheapest way to get a good video analyst? — the "analyst" being the component that watches a video and extracts every sourced claim, then judges whether the real evidence is on screen (vs. illustrative B-roll).
It answers that by varying two knobs and scoring every combination on accuracy per dollar:
on-demand frame-grabber → actual vision.
The lanes are just those combinations. Each lane pair isolates one effect.
A lane is one setting of the "how much help the model gets" knob — from nothing but the transcript up to actually watching the video. The lane fixes what the model can work with; the models themselves (codex, GLM, DeepSeek…) race inside the lanes.
| Lane | What the model works from | Who runs in it |
|---|---|---|
A1 | The transcript only — no source list | codex55 · GLM-4.7 · GLM-5.3 · DeepSeek-pro |
A2 | The transcript + the creator's source list | codex55 · GLM-4.7 · GLM-5.3 · DeepSeek-pro |
B1 | The transcript + an on-demand "grab me a frame" tool | codex |
C1 | The real video frames (+ sources) | DeepSeek-vision |
C2 | The real video frames, visual-only specialist prompt | DeepSeek-vision |
A → B → C climbs from reads the words → can pull a frame when it wants → watches the video. The number is the variant within a lane — e.g. A1 vs A2 is the same models, minus vs. plus the creator's source list.
Each comparison reads one pair of the lanes above and answers one product question.
| Comparison | Lanes that isolate it | The product question it answers |
|---|---|---|
| Vendor head-to-head | A1 / A2: codex55 vs GLM-4.7 vs GLM-5.3 vs DeepSeek-pro, same task | Could a cheaper vendor do the analyst job nearly as well as the expensive one? |
| Does the source list help? | A1 (no sources) vs A2 (+ the creator's bibliography), same seats | Is it worth building the feature that hands the model the creator's source list? |
| Does the frame-blaster help? | B1: codex with an on-demand "pull me a frame" tool | Is it worth building the tool that lets the analyst grab video frames when it wants to look? |
| Does seeing buy anything? | C1 / C2 (DeepSeek-vision) vs the text-only lanes | Do we need a vision model at all, or is reading the transcript enough to judge on-screen evidence? |
It's the cheap "can actually watch the video" option. The text models (codex / GLM / DeepSeek-pro) only read the transcript — they guess what's on screen. DeepSeek-vision looks at the frames. So including it answers: for the hardest part of the job — "is the real evidence shown, or is it just B-roll?" — does a model that sees the frames beat one that only reads the words, and is it worth the cost? (C2 is a stripped-down "just describe what's on screen" prompt vs. C1's full prompt — testing whether a specialist visual pass does better.)
It's a targeted design, not a full factorial: the vendor race runs in the text lanes; the blaster is codex-only (it needs a strong tool-user); vision is DeepSeek-only (the cheap vision option). So you can't read every cell — but each of the four questions above has the pair it needs.
Every one of those is scored in dollars (see cost.py / the cross-vendor cost reports), so "nearly as good for half the price" wins.
The admin-roster trials (fern / brofessor / nexpo-20260807, 19-seat) answer a different question — "which of our own candidate models to ship as the analyst." The cross-vendor h2h asks "which vendor + which features are worth paying for." They share the same videos (and therefore the same owner-adjudicated ground truth), but they contest different moments, so they are mostly disjoint in questions — the h2h is new information, not a rehash.