
Strands Decider 2B as the judge of my shots: 37 ms, and a keyword rule ties it
I measured Strands Decider 2B as a judge on 75 shots: with choice it hits AUC 0.91 in 34 ms per judgment, vs 0.99 for Claude Opus; a simple rule got 0.90. With its default primitive, noul, it caught none of the 31 generic shots; the v21 checkpoint does not improve it, and repetition across shots, my real failure, is caught by every judge. Data, labels and scripts are published.
My short-video pipeline has a step that writes the shot list before rendering. When a small local model does that step, it fails in a specific way: it pastes the same scene into all ten shots, or returns only the style prefix with no scene. The result is a video made of a single image. I wanted a gate that caught it before spending the GPU on rendering.
The usual answer is to put an LLM in as a judge. It works, but every judgment costs money and time. So I tried another category: decision models, models that do not write text and return a probability directly. The candidate was Strands Decider 2B, 2 billion parameters, Apache-2.0 licensed.
My first probe was useless, and a regex proved it
The first version of this article used eight hand-written shot descriptions, four strong and four generic. The decision model got all eight in 33 ms; Opus seven, Haiku six and MiniMax four. It looked like a clean result.
An external audit took it apart with three lines of code. The eight negatives literally said “generic”, “interchangeable” or “no subject”. A regular expression that looks for those words also gets eight out of eight. The bench measured whether the text declared itself generic, not whether the shot was.
What a decision model is and how I put it to judge
A decision model receives a state and a question, and instead of writing an answer it returns numbers: the probability of a yes (noul), a choice among options with its probability (choice) or a score. There is no text to parse, which is why it answers in milliseconds. I had already measured another one from the same category, Intern-Decision predicting the rain without training it.
I set it against two kinds of rivals: LLMs used as judges and rules that use no model at all.
The bench
75 shot descriptions, in English because that is how they reach the image generator.
- 22 from production. Scenes my pipeline generated with the large model in real runs, pulled from the Postgres checkpoint where each run’s state lives.
- 41 from a weak model. I ran the real node that writes the shots with qwen2.5 14B and with gemma3 4B, using the original unimproved prompt, over eight profiles. The three real failures showed up there: the same scene pasted ten times, the style prefix with no scene, and filler subjects like “a character” or “a mysterious figure”. Good shots showed up too: the weak model is not always wrong.
- 12 hand-written adversarial ones. Six strong shots that contain trap words like same, typical or generic, and six generic ones dressed up in camera jargon with no subject or action.
The rubric. I annotated each shot with three features: whether it names a concrete subject, whether it describes a visible action or state, and whether it sets a camera decision (shot size, angle, position in the frame or movement). A shot is strong if it has all three. That left 44 strong and 31 generic.
The judges. Strands Decider 2B with two primitives (noul and choice), qwen2.5 14B running locally, Claude Opus, Claude Haiku and MiniMax M3, plus two baselines: the length of the text and a rule that only looks for framing words like close-up, wide shot or third. Every judge gets the same shot and the same question.
Three stumbles before the numbers
The CLI contaminated the judge. I called the LLMs from the command line, and by default the tool loads my own instructions file: about 22,000 tokens of context unrelated to the task. With that in place, Haiku did not even return the JSON; it replied that it needed more context. I ran it in a clean mode, with no user settings, a minimal system prompt and a single tool. Opus’s cost per judgment dropped from about 15 cents to 1.1.
The 404 seconds of load were the download. In the first probe the model took 404 s to load and I reported it as a cold start. It was the first download of the weights. From disk it loads in 10.7 s.
The model ran on the slow kernels. On load, the library warns that causal_conv1d and flash-linear-attention are missing and that it falls back to the reference implementation. I did not install them, so the 37 ms is a conservative bound.
The numbers
| Judge | AUC | Generic caught | Good rejected | Median latency |
|---|---|---|---|---|
| Claude Opus | 0.988 | 30/31 | 2/44 | 1.52 s (API) |
| Claude Haiku | 0.941 | 16/31 | 0/44 | 5.50 s (API) |
| MiniMax M3 | 0.938 | 26/31 | 4/44 | 2.56 s (API) |
Decider 2B v19 · choice | 0.905 | 16/31 | 4/44 | 34 ms (local) |
| Framing rule | 0.903 | 25/31 | 0/44 | ~0 |
Decider 2B v21 · choice | 0.881 | 7/31 | 0/44 | 41 ms (local) |
| qwen2.5 14B local | 0.849 | 13/31 | 2/44 | 291 ms (local) |
Decider 2B v19 · noul | 0.819 | 0/31 | 0/44 | 37 ms (local) |
Decider 2B v21 · noul | 0.750 | 0/31 | 0/44 | 40 ms (local) |
| Text length | 0.727 | n/a | n/a | 0 |
AUC measures how well each judge ranks the strong shots above the generic ones, independently of the threshold; 1 is perfect and 0.5 is chance. “Generic caught” and “good rejected” use the 0.5 threshold, which is what an uncalibrated gate would do. The 95% confidence intervals, from 2,000 resamples, are in the bench below.
- Claude Opus0.98895% CI: 0.962 to 1.000 · catches 30/31 · rejects 2/44
- Claude Haiku0.94195% CI: 0.863 to 0.994 · catches 16/31 · rejects 0/44
- MiniMax M30.93895% CI: 0.877 to 0.980 · catches 26/31 · rejects 4/44
- Decider 2B v19 · choice0.90595% CI: 0.831 to 0.963 · catches 16/31 · rejects 4/44
- Framing rule0.90395% CI: 0.833 to 0.969 · catches 25/31 · rejects 0/44
- Decider 2B v21 · choice0.88195% CI: 0.800 to 0.946 · catches 7/31 · rejects 0/44
- qwen2.5 14B local0.84995% CI: 0.757 to 0.929 · catches 13/31 · rejects 2/44
- Decider 2B v19 · noul0.81995% CI: 0.707 to 0.912 · catches 0/31 · rejects 0/44
- Decider 2B v21 · noul0.75095% CI: 0.633 to 0.854 · catches 0/31 · rejects 0/44
- Text length0.72795% CI: 0.597 to 0.852 · a bias of the bench, not a judge
75 shots (63 captured from the pipeline, 12 adversarial), features annotated before running the judges. Local judges on an RTX 4070 Ti SUPER; LLMs through their API in clean mode.
Opus ranks almost perfectly and, at the 0.5 threshold, catches 30 of the 31 generic shots while rejecting only 2 of the 44 good ones. Since every judge scored the same 75 shots, the fair comparison is paired, resampling the shots together: Opus beats the 2B with choice by 0.083 AUC (95% interval between 0.022 and 0.154), the rule by 0.084 (0.023 to 0.158) and qwen by 0.139 (0.054 to 0.233). It is the best judge in the bench and the gap is not noise.
The threshold that does not travel
The result that surprised me most is the noul row. It ranks acceptably, AUC 0.82, but at the 0.5 threshold it catches no generic shot. Every probability landed between 0.54 and 0.84: the generic ones between 0.58 and 0.82, the strong ones with a median of 0.80. Everything falls to the right of the threshold, and the best-scored generic shot sits above half of the strong ones.
generic strong
In my first probe, the same probabilities separated cleanly, 0.71 to 0.74 against 0.21 to 0.36. That separation was a property of eight easy sentences in Spanish, not of the model. Its own card warns about it: calibration was fitted on short classification, “noul and score transfer poorly to rubrics”, and you have to measure on your own traffic before trusting a threshold.
With choice, the primitive its calibration is fitted on, the model improves: AUC 0.91 and it catches 16 of 31 at 0.5. If the threshold is also tuned on labeled examples, leaving each shot out while the cut is chosen with the other 74, it reaches 0.84 balanced accuracy. It works, but it requires labeling before you use it.
A keyword rule ties it
The framing rule uses no model: it marks as strong any shot that says close-up, wide shot, third, centered or a camera move. It reaches an AUC of 0.90, catches 25 of 31 generic shots and rejects no good shot.
The reason is the rubric: a shot with no camera decision is never strong, and that part a word search can detect. What the rule cannot see is the subject. The 6 generic shots that slip past it are style-only prompts that mention “vertical framing” or “rule of thirds” without saying what is in the frame. Text length also predicts something, AUC 0.73, because good shots are longer: a bias of the bench worth keeping in view.
Repetition, which was my real failure
The gate that motivated all this was not about single shots but about lists: the weak model that pastes the same scene ten times. A judge that gets one shot at a time cannot see that, because each shot on its own can be fine.
I built 13 lists of five shots: 6 with distinct scenes from one production run, 6 with one repeated scene (exact copies or light paraphrases) and the real list qwen emitted, “a hero in a dramatic pose” five times. I asked each judge whether the list repeats the same scene.
The 2B decision model got all 13 right, as did Opus and Haiku; MiniMax 12 and qwen 10. A 300-million-parameter embedding (embeddinggemma), comparing the shots against each other, separated them perfectly: every repeated list came out more self-similar than every distinct one, with no judge at all. The lexical check my pipeline already uses, counting distinct scenes by their text, missed one of the paraphrases.
This test was easy: the paraphrases changed few words. What it does make clear is that repetition does not need an LLM.
Speed, load time and cost
The 2B is 8 times faster than the 14B on the same card.
Measured from my network, round trip included. Haiku came out slower than Opus in both runs.
Different animation scale per scenario; the printed figures are the measured ones.
The 2B answers in 37 ms and takes 3.7 GB of VRAM. qwen2.5 14B, on the same card and with the same timer, takes 291 ms. Opus takes 1.5 s median through its API. Since the 2B loads in 10.7 s, it pays off against Opus after about 7 judgments: it makes sense as a service that stays loaded, not as a process that starts to judge one shot and exits.
On cost, one Opus judgment in clean mode comes to 1.1 cents, MiniMax 0.7 and Haiku 0.6. A thousand judgments with Opus is about 11 dollars. For the 2B, the card drew about 200 W while the local models were inferring: 37 ms is about 7 joules per judgment. A million judgments uses about 2 kWh, in the order of 40 cents of electricity at 0.20 dollars per kWh.
What this measurement does not say
- I made the annotations myself, with a rubric and before running the judges; the rule that turns them into a label I fixed afterwards, as told above. Eight shots were marked as borderline cases. Another annotator could move some of them; the annotations are published item by item.
- 75 shots is a small bench. That is why every comparison in this article is paired. The 2B with
choiceand the rule end up +0.001 AUC apart (95% interval between −0.097 and 0.099): with this data they cannot be told apart, and neither can the 2B withchoiceand qwen (+0.056, between −0.029 and 0.147). - A single question. A different wording could move every judge, especially the 2B, whose card says it reads questions less carefully than documents.
- LLM latencies are not comparable with local ones. Some are API latencies with the network included and the others a local call. That is why the clean comparison is the 2B against the 14B on the same GPU.
- The repetition test was easy. It lacks lists with distinct but similar shots, which is where a judge really earns its keep.
When I would use each one
For repetition across shots, a metric, not a judge. The embedding, or even a text comparison, catches it for free and in milliseconds. The 2B also gets it right if I pass it the whole list.
For judging the quality of each shot at low volume, an LLM. Opus separated almost perfectly, and at a cent per judgment a twenty-shot video costs twenty cents.
For high volume or critical latency, a rule first and the 2B after it, calibrated. The framing rule filters the obvious cases for free. The decision model only adds value if I use choice and tune its threshold on labeled shots from my own pipeline; with its factory threshold it lets everything through.
The decision model is fast, light and installs with one pip install. What it does not bring from the factory is the judgment of my rubric, and speed does not make up for that.
Measured on 6 and 7 October 2026. Local judges on an RTX 4070 Ti SUPER with 16 GB, torch 2.14.1+cu130, strands-decider 0.1.0 (checkpoint StrandsAgents/strands-decider-2B-hobson-v19) and qwen2.5 14B through ollama. LLMs through their API in clean mode: claude-opus-5-5, claude-haiku-4-5-20251001 and MiniMax-M3, the identifiers each API returned. The first eight-sentence probe ran earlier on a rented RTX 2080 Ti with torch 2.7.1. I measured the hobson-v21 checkpoint afterwards, on the same 7 October, with the same harness and GPU.
Frequently asked questions
What is Strands Decider 2B?
A decision model: a LoRA adapter on Qwen3.5-2B-Base, Apache-2.0 licensed, that does not generate text but typed probabilities. You give it a state and a question, and it answers with a number (noul), a choice among options (choice) or a score. The card of hobson-v19, the checkpoint I measured, reports 0.723 accuracy on public JevBench (167 of 231 tasks); the card of hobson-v21, released later, 0.762 (176 of 231).
Can it replace an LLM as a quality judge?
In my 75-shot bench, not as a drop-in replacement. It ranked the shots with an AUC of 0.82 using noul and 0.91 using choice; Claude Opus reached 0.99, Claude Haiku 0.94 and MiniMax M3 0.94. A rule that only looks for framing words reached 0.90. It is much faster and almost free, but it does not judge better.
Why did it catch no generic shot?
Because with noul every probability landed between 0.54 and 0.84, so the 0.5 threshold accepts everything. Its own card warns that noul and score transfer poorly to rubrics and that you should measure on your own traffic before trusting a threshold. With choice and a threshold tuned on labeled examples it reached 0.84 balanced accuracy.
Does the v21 checkpoint change the result?
No. hobson-v21 goes from 167 to 176 of 231 on public JevBench, but on my 75 shots it ranks worse with noul (AUC 0.750 vs 0.819; paired difference −0.069, with a 95% interval between −0.137 and −0.010) and stays the same with choice (0.881 vs 0.905, no significant difference). With its default threshold and noul it still catches no generic shot.
How fast is it?
37 ms per judgment, median, on an RTX 4070 Ti SUPER with the model already loaded. qwen2.5 14B on the same GPU and with the same timer took 291 ms. Claude Opus, through its API, 1.5 s median. The 2B loads in 10.7 s from disk and takes 3.7 GB of VRAM.
Do I need a GPU to run it?
Not necessarily. According to its card it runs on cuda, mps (Apple Silicon) and cpu. I only measured it on CUDA, so the published latencies are GPU ones; on a CPU they will be higher.
Can I reproduce the numbers?
Yes. The 75 shots, the rubric, the per-item annotations, both label variants, the scripts and the raw outputs of the judges are published; the data also as a downloadable CC-BY-4.0 dataset at efraingaray.com/en/datasets/juez-decision-75/. I annotated the three features of every shot before running any judge; later I found that my first way of turning those features into strong or generic did not match the question, and I fixed it. Both variants are published and the order of the judges does not change.
Sources
- Strands Decider 2B hobson-v19 model card on Hugging Face: the checkpoint measured; primitives, public JevBench result (167 of 231), calibration and limitations, including the warning about rubrics.
- hobson-v21 model card: the later checkpoint, 176 of 231 on public JevBench, trained with question paraphrases and distillation from Qwen3.5-4B.
- strands-labs/strands-decider: the code of the inference engine I used (
pip install strands-decider). - Qwen3.5-2B-Base: the base model of the adapter.
- The data, labels, scripts and raw outputs: the 75 shots, the rubric, the annotations, both label variants and the results of the eight judges.
python3 analyze.pyreproduces every figure. - The downloadable dataset: the 75 shots with their labels, every judge’s judgments and the 13 lists in CSV and JSON, under CC-BY-4.0, with how to cite it.
- Intern-Decision predicting the rain: another decision model measured on this blog, with its own compressed-probability problem.
- AWS’s agent harness, measured: the same Strands ecosystem, on the agent side.
Comments
No comments yet. The first one is yours.