
Eight decision models against my data: none beat what I already use, and the index's second place ranked fraud backwards
I ran eight decision models on the Decision Index, which I reproduced, and on three banks of my own with real labels: video shot descriptions, card fraud and rain. None clearly beat each bank's classic reference. On fraud, the only bank where they separate clearly, d1-3B, second on the index, ranked backwards and JEV-9B was the clear best; on shots and rain the top ones tied.
In three weeks almost ten decision models came out: models that don’t write, that read a state and answer typed questions with a probability per option, usually in a single pass. Each arrived with its place on the same leaderboard, the Decision Index. The latest, Liquid AI’s d1-3B, came out on October 7, presented by its lab as the best under 10 billion parameters on edition 0.2.1 of the index.
I had already measured one, Strands Decider 2B, as a judge of my video pipeline’s shot descriptions, and a keyword rule had tied it. This time I wanted to answer a different question: if I take the eight I can run and put them on data of my own with real labels, does the leaderboard tell me which one to use?
How I ran them
Eight models: JEV-9B from autotrust, a student distilled from TypeSafe’s Jev 1.13 (the original Jev is only available through an API), Clef-flash from Cloudflare, d1-3B and d1-omni-600M from Liquid AI, Decider 2B and 4B from Mapika, Strands Decider 2B from AWS and Laya from Convai Innovations. Left out are the ones I can’t run on home hardware or cheaply rented hardware, such as Perplexity’s 27-billion-parameter pplx-decider.
The rule I set myself was to write no prompt. Each model answers through its own call: the engine its lab published for the Decision Index when there is one (d1, JEV-9B, Clef) and, if not, the system_one function its card documents (Decider, Laya, Strands). And the same rules the published engines apply: nothing is shortened, a request that doesn’t fit the model’s window or the GPU counts as unanswered, and unanswered counts as wrong.
Six fit in the 16 GB RTX 4070 Ti SUPER for my banks. The two 9B ones don’t: JEV-9B takes about 15 GB just to load, so JEV-9B and Clef-flash ran on a rented 48 GB A6000. For the Decision Index I also moved d1-omni, both Deciders and Strands to the A6000, because their long requests didn’t fit in 16 GB; d1-3B and Laya ran it on the 4070 Ti SUPER. Everything rented cost US$3.86, and latencies from one machine aren’t compared with the other’s.
Three setup snags, because each model ships its own package: Laya, Decider and Strands silently shorten the text when it doesn’t fit their window, and I had to detect it so as not to reward the one that shortens; Strands rejects options that aren’t text; and Decider’s engine compiles every new request shape with torch.compile, which on the public suite left it at one request every several seconds. Its own constructor accepts turning that off, and that’s what I did.
The public index, reproduced
Decision Index 0.3 scores a decision model on 37 public benchmarks in five areas, chance-corrected: 0 is guessing and 100 is perfect. That public part weighs 20% of the score that orders the site’s leaderboard; the rest comes from private tests only its maintainers run, so what I reproduced is the public index, not the full leaderboard. Its suite isn’t redistributed, so I rebuilt it with the public kit, which downloads each pinned source and checks that the files come out identical to the official ones; all four matched. That’s 140,178 scoreable requests, too many for eight models on my budget, so I ran a stratified sample of 6,000 with the kit’s own scorer.
Before trusting the sample, I checked it against the labs. Liquid and autotrust publish the answers of their full runs, so I scored them on my same 6,000 rows: d1-3B gets 49.04 on the sample and 48.99 on the whole suite. And my runs reproduce theirs: d1-3B 49.11 against 49.04, d1-omni 14.02 against 14.17, and JEV-9B differs by 0.25 points per benchmark on average (leaving out GSM8K, which their run, made on the previous edition, doesn’t include).
- Clef-flash 9B55.4
- d1-3B49.1their run: 49.0
- JEV-9B42.9
- Decider-4b40.9
- Decider-2b27.0
- Strands Decider 2B20.5269 requests don’t fit
- d1-omni-600M14.01,758 exceed its instruction and option budget
- Laya3.81,451 don’t fit
Stratified sample of the public suite, scored with the kit at 62d2f51 on the sampled rows (lab/decision-judge/di/score_sample.py). Unsupported requests count as wrong.
Clef-flash comes first and d1-3B second, ahead of JEV-9B, which is three times its size. Liquid presented d1-3B as the best under 10B on edition 0.2.1; on 0.3 and on this sample, the 9.4-billion-parameter Clef-flash comes out above it. Reading limits punish the small ones: Laya reads 512 tokens and leaves 1,451 requests unanswered; d1-omni cuts each question’s instructions and options to a fixed budget, and its engine refuses the 1,758 it would read in part. All of those count as wrong.
Bank 1: my pipeline’s shot descriptions
The same 75 shots as the judge post: 44 strong and 31 generic, labelled before any judge ran, with the same question. The reference is the rule that looks for framing keywords, which reached 0.903.
- Claude Opus (reference)0.988
- JEV-9B0.938catches 8 of 31 generic at 0.5
- Decider-4b0.904
- Keyword rule (reference)0.903catches 25 of 31
- d1-3B0.886catches 19 of 31
- Clef-flash0.855
- Strands Decider 2B0.8190.909 with choice
- Decider-2b0.739
- d1-omni-600M0.580
- Laya0.331
Measured on October 9, 2026; JEV-9B and Clef-flash on an RTX A6000, the rest on an RTX 4070 Ti SUPER. Bootstrap 95% intervals in lab/decision-judge/results/summary_engines.json.
None has its interval wholly above the rule’s. JEV-9B ranks best, but at 0.5 it lets 23 of the 31 generic shots through; d1-3B catches 19 at the cost of rejecting 4 good shots. On the bank’s other question, whether five shots repeat the same scene, all get the 13 lists right except d1-omni (11) and Laya (7).
Bank 2: card fraud, against the Flink forest
Last month’s Flink experiment left 101,668 labelled synthetic payments, with three kinds of injected anomaly: high amounts, country jumps and bursts of back-to-back payments. I took the 2,506 anomalies and 7,494 normal payments at random. Each model sees what the rules and the forest see: the amount, the country, the seconds since the card’s previous payment, whether the country changed and how many payments the card made in the last minute. I removed the identifiers, because the generator names burst payments with a mark that gives the label away. The question: “Is this card payment suspicious enough to flag for fraud review?”
On that sample, the Flink post’s random cut forest reaches an AUC of 0.982, and the rules catch 81.4%, the same recall as then; there are fewer false alarms (701) because the sample holds fewer normal payments than the full stream.
- Random cut forest (reference)0.982trained on the stream
- JEV-9B0.966without seeing an example
- Decider-4b0.924
- Strands Decider 2B0.877
- Clef-flash0.820
- Laya0.730
- d1-omni-600M0.682
- Decider-2b0.662
- d1-3B0.314ranks backwards
Measured on October 9, 2026 on 2,506 anomalies and 7,494 normal payments from the Flink experiment. lab/decision-judge/results/summary_banks.json.
JEV-9B lands 0.016 below the forest without having seen a single training payment, and that’s the most interesting result in this post. But AUC isn’t the decision: at the 0.5 threshold, JEV-9B flags 7,177 of the 7,494 normal payments as suspicious. To compare it with the forest on the forest’s terms, I gave each model the same false-alarm budget the forest has, 336, and found the threshold on its own probabilities.
With that budget, JEV-9B catches 89.5% of the anomalies and the forest 94.8%. Decider-4b, 73%. Those 5.3 points between JEV-9B and the forest are not small in fraud, and the forest runs in microseconds on a CPU.
d1-3B deserves its own paragraph, because it’s second on the public index and here it ranks backwards. It flags high amounts with a median probability of 0.84 and country jumps with 0.68, fine. But a burst of payments gets 0.41, less than a normal payment (0.56), and bursts are two thirds of the bank’s anomalies. The few normal payments that come less than a second after the previous one also get 0.41. I didn’t isolate which part of the text drives it, but the pattern is consistent: a payment right after the previous one looks less suspicious to it, not more.
Bank 3: rain tomorrow
The Intern-Decision post’s bank: the 27,256 test days from 2016 to 2026 in seven cities, with the same state and the same question, “Will it rain at least 1 mm tomorrow at this station?”. The series’ logistic regression reaches 0.847 on those days, and Intern-Decision 2B 0.823. I added two references I didn’t have then: using today’s rain in mm alone as the score gives 0.813, and the yes/no rule “at least 1 mm today” gives 0.709.
- Logistic regression (reference)0.847
- Intern-Decision 2B0.823
- JEV-9B0.820
- Decider-2b0.816
- Clef-flash0.814
- Today’s rain in mm alone (reference)0.813
- d1-3B0.810
- Decider-4b0.805
- Strands Decider 2B0.793
- d1-omni-600M0.593
- Laya0.577
Measured on October 9, 2026 on the series' test days (NASA POWER, 7 cities). lab/decision-judge/results/summary_banks.json.
None reaches the logistic regression, and four of them fall inside that single variable’s interval. I didn’t test which part of the state they use, so I can’t say they mostly read today’s rain, although the closeness suggests it.
The order changes from bank to bank
With the four banks side by side, the public index gets the broad strokes right and the details wrong.
It gets the two weakest right: d1-omni and Laya end up at the bottom on shots and rain, which is why its order looks like my shots’ (Spearman rank correlation 0.76) and rain’s (0.74). On fraud it says nothing (0.10). And on fraud, the only bank where the differences are clear, it doesn’t anticipate the order: d1-3B, second on the index, is last, and JEV-9B, third, is the clear best. On shots and rain the top ones tie within the noise, so there was no best one for the index to pick.
What I can’t claim
It’s one sample per bank and one machine per model size. The payments are synthetic, generated for the Flink post, and the rain comes as text, worded the way the series worded it; a different wording of the state or the question can move a decision model’s results, and I didn’t try variants. I gave none of them examples or tuned anything, which is how they come out of the box and not how they’d be used in production. Of the Decision Index I ran a stratified sample of 6,000 requests from its public part, not the full suite of 140,000 nor the private tests that complete the leaderboard. And the confidence intervals resample single rows, while the payments in a burst and the rain days of one city aren’t independent: they’re narrower than they should be, so I don’t use them to claim small differences.
My take
Decision models are a good idea and JEV-9B is the proof: without a single example, it ranks fraud almost like a forest trained on the stream, and it does it by reading a sentence. For prototyping a new decision, with no data yet, that’s worth a lot.
But the public index doesn’t replace measuring on your own data. On shots and rain it ruled out the two weakest; it didn’t help choose among the good ones: the same model was second on the index and ranked one of my banks backwards, and the factory threshold is almost never the one that works. If you already have a rule, a classic model or a model trained on your data, compare them on the same set before switching.
When I’d use one and when not
- Yes, for a new decision without labelled data, with JEV-9B or Decider-4b, and as a cheap second opinion next to a rule.
- Carefully, calibrating the threshold on a few hundred examples of your own before using the probability as a cut-off.
- No, instead of a model trained on your data when you already have one, nor taking first place on the index as a guarantee for your case.
Frequently asked questions
What is a decision model?
A model that doesn't generate text: it takes a state (text or JSON) and typed questions, yes or no (noul) or a pick among options with their criteria (choice), and returns a probability per option in a single pass (JEV-9B's engine uses several when a question has more than 16 options). Jev, from TypeSafe, opened the category in September 2026 as an API service; in the following weeks came Clef from Cloudflare, Strands Decider from AWS, Decider from Mapika, Laya, Liquid AI's d1 and others. The JEV-9B I measured is an open student of Jev, published by autotrust and distilled from its answers, not TypeSafe's Jev.
Which one was best on my data?
It depends on the bank. JEV-9B was the clear best on fraud (AUC 0.966, against 0.924 for the next one); on shots (0.938) and rain (0.820) it tied with others within the measurement's noise. But on none of the three did it clearly beat the reference I already use: a keyword rule on shots (0.903), a random cut forest trained on the stream for fraud (0.982) and a logistic regression for rain (0.847).
Why does d1-3B rank fraud backwards?
Because a burst of payments looks less suspicious to it than a normal payment: median probability 0.41 against 0.56. It flags high amounts (0.84) and country jumps (0.68) well, but bursts are two thirds of the bank's anomalies, so the overall order comes out inverted (AUC 0.31).
Does the 0.5 threshold work?
Not without calibration. On fraud, JEV-9B at 0.5 flags 7,177 of the 7,494 normal payments. Given the same false-alarm budget as the forest (336), it catches 89.5% of the anomalies, against the forest's 94.8%. The threshold has to be tuned on your own examples.
Do I need a big GPU?
For the 9B ones in bf16, more than 16 GB: JEV-9B took about 15 GB just to load and ran out of memory on an RTX 4070 Ti SUPER with long requests, so I ran JEV-9B and Clef-flash on a rented 48 GB A6000. The rest fit in 16 GB: d1-3B peaked at 5.8 GB, Decider-4b 8.0 GB, Strands 3.6 GB, Laya 2.3 GB and d1-omni 1.1 GB.
Can I reproduce it?
Yes. The three banks, each model's engine, the per-request results and the scripts that compute every number are published in a gist linked in the sources. The Decision Index is rebuilt with the public kit, which checks that the suite comes out identical to the official one.
Sources
- Decision Index and its reproduction kit, version 0.3 at
62d2f51. - d1-3B and d1-omni-600M, Liquid AI’s announcement, and their Decision Index runs.
- JEV-9B and its Decision Index runs.
- Clef-flash and the engine published with Clef’s run.
- Decider 2B and Decider 4B.
- Strands Decider 2B and Laya.
- The Habr article comparing their architectures (in Russian).
- My own banks: the judge post, the Flink one and the Intern-Decision one.
- The banks, the engines, the per-request results and the scripts.
Measured on October 9 and 10, 2026. On my banks, six models on a 16 GB RTX 4070 Ti SUPER (torch 2.14.1, transformers 5.19.0) and JEV-9B and Clef-flash on a rented 48 GB RTX A6000 (torch 2.8.0, transformers 5.19.0). On Decision Index 0.3 (kit at 62d2f51), d1-3B and Laya on the RTX 4070 Ti SUPER and the other six on the A6000.
Comments
No comments yet. The first one is yours.