12°

Dataset · v1

Labeled video shot descriptions for evaluating automatic judges

The data behind the decision-model benchmark: 75 shot descriptions from my own video pipeline, labeled strong or generic, with what each judge answered. Free to use if you cite it.

Version
1 · 7 October 2026
License
CC-BY-4.0: any use, including commercial, with attribution
Content
75 shot descriptions (44 strong, 31 generic) and 13 five-shot lists
Text language
English (that is how the prompts reach the image generator)
Judges
Strands Decider 2B v19 and v21 (noul and choice), qwen2.5 14B, Claude Opus, Claude Haiku, MiniMax M3, plus 2 model-free baselines
Author
Efrain Garay

What is inside

Three files plus a readme. Every value comes from the raw results; nothing was typed by hand.

FileColumns
planos.csvone row per shot (75)id, stratum, source, text, subject, action, framing (0/1 each), label, label_loose, borderline
juicios.csvone row per shot × judge (750)id, judge, model (exact identifier), p_strong, latency_ms, latency_kind
listas.jsonone entry per five-shot list (13)kind, variant, shots, p_repeated per judge, embedding similarity, distinct-text ratio

How it was built

22 production scenes from real runs of my shot-list step, 41 outputs of a weak model (qwen2.5 14B and gemma3 4B running the same step with the unimproved prompt, where the real failures show up: one scene pasted ten times, the style prefix with no scene, placeholder subjects) and 12 hand-written adversarial ones with trap words or camera jargon and no content.

Each shot is annotated with three features: a concrete subject, a visible action or state, and an explicit camera decision. A shot is strong if it has all three. The judges answered the same question, and their probabilities are stored as they came, without thresholds.

How to cite it

Garay, E. (2026). Labeled video shot descriptions for evaluating automatic judges (v1) [Dataset]. efraingaray.com. https://efraingaray.com/en/datasets/juez-decision-75/. License CC-BY-4.0.

Limits

  • 75 items is a small bench: compare judges with paired tests, as the article does.
  • One annotator, the author. Another annotator could move some borderline cases.
  • The texts are in English and come from one pipeline and its visual styles.
  • The repetition lists are easy: exact copies and light paraphrases.