Dataset · v1
Labeled video shot descriptions for evaluating automatic judges
The data behind the decision-model benchmark: 75 shot descriptions from my own video pipeline, labeled strong or generic, with what each judge answered. Free to use if you cite it.
- Version
- 1 · 7 October 2026
- License
- CC-BY-4.0: any use, including commercial, with attribution
- Content
- 75 shot descriptions (44 strong, 31 generic) and 13 five-shot lists
- Text language
- English (that is how the prompts reach the image generator)
- Judges
- Strands Decider 2B v19 and v21 (noul and choice), qwen2.5 14B, Claude Opus, Claude Haiku, MiniMax M3, plus 2 model-free baselines
- Author
- Efrain Garay
What is inside
Three files plus a readme. Every value comes from the raw results; nothing was typed by hand.
| File | Columns |
|---|---|
planos.csvone row per shot (75) | id, stratum, source, text, subject, action, framing (0/1 each), label, label_loose, borderline |
juicios.csvone row per shot × judge (750) | id, judge, model (exact identifier), p_strong, latency_ms, latency_kind |
listas.jsonone entry per five-shot list (13) | kind, variant, shots, p_repeated per judge, embedding similarity, distinct-text ratio |
How it was built
22 production scenes from real runs of my shot-list step, 41 outputs of a weak model (qwen2.5 14B and gemma3 4B running the same step with the unimproved prompt, where the real failures show up: one scene pasted ten times, the style prefix with no scene, placeholder subjects) and 12 hand-written adversarial ones with trap words or camera jargon and no content.
Each shot is annotated with three features: a concrete subject, a visible action or state, and an explicit camera decision. A shot is strong if it has all three. The judges answered the same question, and their probabilities are stored as they came, without thresholds.
How to cite it
Garay, E. (2026). Labeled video shot descriptions for evaluating automatic judges (v1) [Dataset]. efraingaray.com. https://efraingaray.com/en/datasets/juez-decision-75/. License CC-BY-4.0.
Limits
- 75 items is a small bench: compare judges with paired tests, as the article does.
- One annotator, the author. Another annotator could move some borderline cases.
- The texts are in English and come from one pipeline and its visual styles.
- The repetition lists are easy: exact copies and light paraphrases.