‹ back to Calibrating ScriptGen

What the cross-vendor head-to-head measures

Plain-language explainer for the 2026-08 cross-vendor calibration study (nexpo-h2h + the FERN trials: room1046, cryptoqueen, cocacola). Saved as a candidate for a public website page. Names models because it is empirical model-study data (calibration carve-out).

The one question it's really asking

What's the cheapest way to get a good video analyst? — the "analyst" being the component that watches a video and extracts every sourced claim, then judges whether the real evidence is on screen (vs. illustrative B-roll).

It answers that by varying two knobs and scoring every combination on accuracy per dollar:

on-demand frame-grabber → actual vision.

The lanes are just those combinations. Each lane pair isolates one effect.

The lanes

A lane is one setting of the "how much help the model gets" knob — from nothing but the transcript up to actually watching the video. The lane fixes what the model can work with; the models themselves (codex, GLM, DeepSeek…) race inside the lanes.

LaneWhat the model works fromWho runs in it
A1The transcript only — no source listcodex55 · GLM-4.7 · GLM-5.3 · DeepSeek-pro
A2The transcript + the creator's source listcodex55 · GLM-4.7 · GLM-5.3 · DeepSeek-pro
B1The transcript + an on-demand "grab me a frame" toolcodex
C1The real video frames (+ sources)DeepSeek-vision
C2The real video frames, visual-only specialist promptDeepSeek-vision

A → B → C climbs from reads the words → can pull a frame when it wants → watches the video. The number is the variant within a lane — e.g. A1 vs A2 is the same models, minus vs. plus the creator's source list.

The four comparisons

Each comparison reads one pair of the lanes above and answers one product question.

ComparisonLanes that isolate itThe product question it answers
Vendor head-to-headA1 / A2: codex55 vs GLM-4.7 vs GLM-5.3 vs DeepSeek-pro, same taskCould a cheaper vendor do the analyst job nearly as well as the expensive one?
Does the source list help?A1 (no sources) vs A2 (+ the creator's bibliography), same seatsIs it worth building the feature that hands the model the creator's source list?
Does the frame-blaster help?B1: codex with an on-demand "pull me a frame" toolIs it worth building the tool that lets the analyst grab video frames when it wants to look?
Does seeing buy anything?C1 / C2 (DeepSeek-vision) vs the text-only lanesDo we need a vision model at all, or is reading the transcript enough to judge on-screen evidence?

Where DeepSeek-vision fits

It's the cheap "can actually watch the video" option. The text models (codex / GLM / DeepSeek-pro) only read the transcript — they guess what's on screen. DeepSeek-vision looks at the frames. So including it answers: for the hardest part of the job — "is the real evidence shown, or is it just B-roll?" — does a model that sees the frames beat one that only reads the words, and is it worth the cost? (C2 is a stripped-down "just describe what's on screen" prompt vs. C1's full prompt — testing whether a specialist visual pass does better.)

Why it isn't a clean grid

It's a targeted design, not a full factorial: the vendor race runs in the text lanes; the blaster is codex-only (it needs a strong tool-user); vision is DeepSeek-only (the cheap vision option). So you can't read every cell — but each of the four questions above has the pair it needs.

What the data can show

Every one of those is scored in dollars (see cost.py / the cross-vendor cost reports), so "nearly as good for half the price" wins.

How this differs from the admin-roster study

The admin-roster trials (fern / brofessor / nexpo-20260807, 19-seat) answer a different question — "which of our own candidate models to ship as the analyst." The cross-vendor h2h asks "which vendor + which features are worth paying for." They share the same videos (and therefore the same owner-adjudicated ground truth), but they contest different moments, so they are mostly disjoint in questions — the h2h is new information, not a rehash.