# Labeled video shot descriptions for evaluating automatic judges (v1)

75 shot descriptions for an image generator (in English), labeled strong or generic, and
the probabilities 8 judges and 2 baselines gave them. Plus 13 five-shot lists to measure
detection of repeated scenes.

- Page: https://efraingaray.com/en/datasets/juez-decision-75/
- Article with the analysis: https://efraingaray.com/en/blog/juez-decision-2b/
- Scripts: https://gist.github.com/EfrainGaray/108dda29e4796288af8a63197fd42e36
- License: CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/)

## Files
- `planos.csv`: id, stratum (production, weak-model, adversarial), source, text, subject,
  action, framing (rubric features, 0/1), label (literal: strong if all three),
  label_loose (subject and action or framing), borderline (1 = borderline case).
- `juicios.csv`: id, judge, model (exact identifier), p_strong (probability of "strong"),
  latency_ms, latency_kind (local_call = timer around the call on the same GPU;
  api_reported = latency reported by the API).
- `listas.json`: the 13 lists (kind: distinct or repeated; variant), with every judge's
  probability of "repeated", the mean embedding similarity and the ratio of distinct scenes
  by text.

## How it was built
22 scenes from real production runs, 41 outputs of a weak model (qwen2.5 14B and gemma3 4B
running the real node with the unimproved prompt) and 12 hand-written adversarial ones. The
features of every shot were annotated before running the judges; the rule that aggregates
them into a label was corrected afterwards to match the question (both variants are
published). Annotations by the author; 8 shots are marked borderline.

## Cite
Garay, E. (2026). Labeled video shot descriptions for evaluating automatic judges (v1) [Dataset]. efraingaray.com. https://efraingaray.com/en/datasets/juez-decision-75/. License CC-BY-4.0.
