17°
Portada del artículo: Eight decision models against my data: none beat what I already use, and the index's second place ranked fraud backwards
IABenchmarksInferencia local

Eight decision models against my data: none beat what I already use, and the index's second place ranked fraud backwards

I ran eight decision models on the Decision Index, which I reproduced, and on three banks of my own with real labels: video shot descriptions, card fraud and rain. None clearly beat each bank's classic reference. On fraud, the only bank where they separate clearly, d1-3B, second on the index, ranked backwards and JEV-9B was the clear best; on shots and rain the top ones tied.

Efrain Garay 10 October 2026

Playing summary

In three weeks almost ten decision models came out: models that don’t write, that read a state and answer typed questions with a probability per option, usually in a single pass. Each arrived with its place on the same leaderboard, the Decision Index. The latest, Liquid AI’s d1-3B, came out on October 7, presented by its lab as the best under 10 billion parameters on edition 0.2.1 of the index.

I had already measured one, Strands Decider 2B, as a judge of my video pipeline’s shot descriptions, and a keyword rule had tied it. This time I wanted to answer a different question: if I take the eight I can run and put them on data of my own with real labels, does the leaderboard tell me which one to use?

Narrated: eight decision models against the public index and against my own data. d1-3B, second on the index, ranks fraud backwards; JEV-9B reaches 0.966 without seeing an example, but at the factory threshold it flags almost every normal payment.Watch it in the reel viewer →

How I ran them

Eight models: JEV-9B from autotrust, a student distilled from TypeSafe’s Jev 1.13 (the original Jev is only available through an API), Clef-flash from Cloudflare, d1-3B and d1-omni-600M from Liquid AI, Decider 2B and 4B from Mapika, Strands Decider 2B from AWS and Laya from Convai Innovations. Left out are the ones I can’t run on home hardware or cheaply rented hardware, such as Perplexity’s 27-billion-parameter pplx-decider.

The rule I set myself was to write no prompt. Each model answers through its own call: the engine its lab published for the Decision Index when there is one (d1, JEV-9B, Clef) and, if not, the system_one function its card documents (Decider, Laya, Strands). And the same rules the published engines apply: nothing is shortened, a request that doesn’t fit the model’s window or the GPU counts as unanswered, and unanswered counts as wrong.

How the eight decision models were measuredEight decision models with their measured footprint, the same request path and rules for all, and four banks with the reference each model has to beat.JEV-9Bautotrust9.0B parameters · base Qwen3.5-9BSystem 1 LoRA · 24-slot headtemperature per kindjev_engine.py (autotrust)RTX A6000 48 GB · window 262,144GPU peak 15.3 GB · 94.9 ms per questionClef-flash 9BCloudflare9.4B parameters · base Qwen3.5-9Bone backbone prefill · joint schema headsoftmax per questionclef.py (lucataco)RTX A6000 48 GB · window 65,536GPU peak 17.8 GB · 101.6 ms per questiond1-3BLiquid AI3.1B parameters · base LFM2.5-VL-3Bdecoder, one pass · text and imageno tokens generatedd1_engine.py (Liquid AI)RTX 4070 Ti SUPER 16 GB · window 32,768GPU peak 5.8 GB · 14.7 ms per questiond1-omni-600MLiquid AI587M parameters · base LFM2.5-Encoder-350Mbidirectional encoder · text with image or audioexperimental released1_engine.py (Liquid AI)RTX 4070 Ti SUPER 16 GB · window 16,384GPU peak 1.1 GB · 7.0 ms per questionDecider-4bMapika4.2B parameters · base Qwen3.5-4B-Basesupervised stages and RL · reads the option letterstemperature per typeDecider.system_oneRTX 4070 Ti SUPER 16 GB · window 32,768GPU peak 8.0 GB · 19.2 ms per questionDecider-2bMapika1.9B parameters · base Qwen3.5-2B-Basesupervised stages and RL · reads the option letterstemperature per typeDecider.system_oneRTX 4070 Ti SUPER 16 GB · window 32,768GPU peak 3.5 GB · 7.7 ms per questionStrands Decider 2BStrands (AWS)1.9B parameters · base Qwen3.5-2B-Basepointer over the options · reads at the last tokenquestion before stateSystemOneEngine.askRTX 4070 Ti SUPER 16 GB · window 4,096GPU peak 3.6 GB · 30.9 ms per questionLayaConvai Innovations421M parameters · base ModernBERT-large0.4B encoder · agent trained with RL512-token windowAgent.system_one (laya 0.4.1)RTX 4070 Ti SUPER 16 GB · window 512GPU peak 2.3 GB · 11.5 ms per questionOne requeststate: text or JSONtyped questions: noul(yes / no) or choiceOne forward passno tokens, the model’s own callJEV-9B: several passeswith more than 16 optionsA probability per optionnoul: P(yes)choice: a distributionthat sums to 1Same rules for every modelnever shortened: longer than the window → unsupporteddoes not fit in GPU memory → unsupportedunsupported counts as wrongno examples, no tuning: factory 0.5 and AUCDecision Index 0.3, public part6,000 requests, 43 benchmarksstratified sample of the public suiteits scorer: chance-corrected indexreference: each lab’s own run75 shot descriptions44 strong, 31 generic“strong, distinctive shot?”AUC · framing-keyword rule 0.903from the judge postCard risk (Flink)10,000 events, 2,506 anomalies“suspicious enough to flag?”AUC · forest 0.982rules: 81.4 % caughtRain tomorrow27,256 days, 7 cities“at least 1 mm tomorrow?”AUC · logistic 0.847today’s rain in mm alone: 0.813

Six fit in the 16 GB RTX 4070 Ti SUPER for my banks. The two 9B ones don’t: JEV-9B takes about 15 GB just to load, so JEV-9B and Clef-flash ran on a rented 48 GB A6000. For the Decision Index I also moved d1-omni, both Deciders and Strands to the A6000, because their long requests didn’t fit in 16 GB; d1-3B and Laya ran it on the 4070 Ti SUPER. Everything rented cost US$3.86, and latencies from one machine aren’t compared with the other’s.

Three setup snags, because each model ships its own package: Laya, Decider and Strands silently shorten the text when it doesn’t fit their window, and I had to detect it so as not to reward the one that shortens; Strands rejects options that aren’t text; and Decider’s engine compiles every new request shape with torch.compile, which on the public suite left it at one request every several seconds. Its own constructor accepts turning that off, and that’s what I did.

The public index, reproduced

Decision Index 0.3 scores a decision model on 37 public benchmarks in five areas, chance-corrected: 0 is guessing and 100 is perfect. That public part weighs 20% of the score that orders the site’s leaderboard; the rest comes from private tests only its maintainers run, so what I reproduced is the public index, not the full leaderboard. Its suite isn’t redistributed, so I rebuilt it with the public kit, which downloads each pinned source and checks that the files come out identical to the official ones; all four matched. That’s 140,178 scoreable requests, too many for eight models on my budget, so I ran a stratified sample of 6,000 with the kit’s own scorer.

Before trusting the sample, I checked it against the labs. Liquid and autotrust publish the answers of their full runs, so I scored them on my same 6,000 rows: d1-3B gets 49.04 on the sample and 48.99 on the whole suite. And my runs reproduce theirs: d1-3B 49.11 against 49.04, d1-omni 14.02 against 14.17, and JEV-9B differs by 0.25 points per benchmark on average (leaving out GSM8K, which their run, made on the previous edition, doesn’t include).

Decision Index 0.3, public index, sample of 6,000 requestsindex · higher is better
  1. Clef-flash 9B55.4
  2. d1-3B49.1their run: 49.0
  3. JEV-9B42.9
  4. Decider-4b40.9
  5. Decider-2b27.0
  6. Strands Decider 2B20.5269 requests don’t fit
  7. d1-omni-600M14.01,758 exceed its instruction and option budget
  8. Laya3.81,451 don’t fit

Stratified sample of the public suite, scored with the kit at 62d2f51 on the sampled rows (lab/decision-judge/di/score_sample.py). Unsupported requests count as wrong.

Clef-flash comes first and d1-3B second, ahead of JEV-9B, which is three times its size. Liquid presented d1-3B as the best under 10B on edition 0.2.1; on 0.3 and on this sample, the 9.4-billion-parameter Clef-flash comes out above it. Reading limits punish the small ones: Laya reads 512 tokens and leaves 1,451 requests unanswered; d1-omni cuts each question’s instructions and options to a fixed budget, and its engine refuses the 1,758 it would read in part. All of those count as wrong.

Bank 1: my pipeline’s shot descriptions

The same 75 shots as the judge post: 44 strong and 31 generic, labelled before any judge ran, with the same question. The reference is the rule that looks for framing keywords, which reached 0.903.

75 shots: AUC of P(strong)AUC · higher is better
  1. Claude Opus (reference)0.988
  2. JEV-9B0.938catches 8 of 31 generic at 0.5
  3. Decider-4b0.904
  4. Keyword rule (reference)0.903catches 25 of 31
  5. d1-3B0.886catches 19 of 31
  6. Clef-flash0.855
  7. Strands Decider 2B0.8190.909 with choice
  8. Decider-2b0.739
  9. d1-omni-600M0.580
  10. Laya0.331

Measured on October 9, 2026; JEV-9B and Clef-flash on an RTX A6000, the rest on an RTX 4070 Ti SUPER. Bootstrap 95% intervals in lab/decision-judge/results/summary_engines.json.

None has its interval wholly above the rule’s. JEV-9B ranks best, but at 0.5 it lets 23 of the 31 generic shots through; d1-3B catches 19 at the cost of rejecting 4 good shots. On the bank’s other question, whether five shots repeat the same scene, all get the 13 lists right except d1-omni (11) and Laya (7).

Last month’s Flink experiment left 101,668 labelled synthetic payments, with three kinds of injected anomaly: high amounts, country jumps and bursts of back-to-back payments. I took the 2,506 anomalies and 7,494 normal payments at random. Each model sees what the rules and the forest see: the amount, the country, the seconds since the card’s previous payment, whether the country changed and how many payments the card made in the last minute. I removed the identifiers, because the generator names burst payments with a mark that gives the label away. The question: “Is this card payment suspicious enough to flag for fraud review?”

On that sample, the Flink post’s random cut forest reaches an AUC of 0.982, and the rules catch 81.4%, the same recall as then; there are fewer false alarms (701) because the sample holds fewer normal payments than the full stream.

Fraud, 10,000 payments: AUCAUC · higher is better
  1. Random cut forest (reference)0.982trained on the stream
  2. JEV-9B0.966without seeing an example
  3. Decider-4b0.924
  4. Strands Decider 2B0.877
  5. Clef-flash0.820
  6. Laya0.730
  7. d1-omni-600M0.682
  8. Decider-2b0.662
  9. d1-3B0.314ranks backwards

Measured on October 9, 2026 on 2,506 anomalies and 7,494 normal payments from the Flink experiment. lab/decision-judge/results/summary_banks.json.

JEV-9B lands 0.016 below the forest without having seen a single training payment, and that’s the most interesting result in this post. But AUC isn’t the decision: at the 0.5 threshold, JEV-9B flags 7,177 of the 7,494 normal payments as suspicious. To compare it with the forest on the forest’s terms, I gave each model the same false-alarm budget the forest has, 336, and found the threshold on its own probabilities.

Where the threshold leaves each model on the risk bankJev-9B at the factory threshold of 0.5 flags 7,177 of 7,494 normal payments. Allowed the forest's 336 false alarms it catches 89 % of anomalies against the forest's 95 %; Decider-4b 73 %, d1-3B 11 %.0 %50 %100 %01003361,0003,0007,494false alarms (normal payments flagged, of 7,494)Jev-9BDecider-4bd1-3Beach model at the factory 0.5forest: 336 false alarms, 95 % caught
Jev-9B at the factory threshold of 0.5 flags 7,177 of 7,494 normal payments. Allowed the forest's 336 false alarms it catches 89 % of anomalies against the forest's 95 %; Decider-4b 73 %, d1-3B 11 %.

With that budget, JEV-9B catches 89.5% of the anomalies and the forest 94.8%. Decider-4b, 73%. Those 5.3 points between JEV-9B and the forest are not small in fraud, and the forest runs in microseconds on a CPU.

d1-3B deserves its own paragraph, because it’s second on the public index and here it ranks backwards. It flags high amounts with a median probability of 0.84 and country jumps with 0.68, fine. But a burst of payments gets 0.41, less than a normal payment (0.56), and bursts are two thirds of the bank’s anomalies. The few normal payments that come less than a second after the previous one also get 0.41. I didn’t isolate which part of the text drives it, but the pattern is consistent: a payment right after the previous one looks less suspicious to it, not more.

Bank 3: rain tomorrow

The Intern-Decision post’s bank: the 27,256 test days from 2016 to 2026 in seven cities, with the same state and the same question, “Will it rain at least 1 mm tomorrow at this station?”. The series’ logistic regression reaches 0.847 on those days, and Intern-Decision 2B 0.823. I added two references I didn’t have then: using today’s rain in mm alone as the score gives 0.813, and the yes/no rule “at least 1 mm today” gives 0.709.

Rain, 27,256 days: AUCAUC · higher is better
  1. Logistic regression (reference)0.847
  2. Intern-Decision 2B0.823
  3. JEV-9B0.820
  4. Decider-2b0.816
  5. Clef-flash0.814
  6. Today’s rain in mm alone (reference)0.813
  7. d1-3B0.810
  8. Decider-4b0.805
  9. Strands Decider 2B0.793
  10. d1-omni-600M0.593
  11. Laya0.577

Measured on October 9, 2026 on the series' test days (NASA POWER, 7 cities). lab/decision-judge/results/summary_banks.json.

None reaches the logistic regression, and four of them fall inside that single variable’s interval. I didn’t test which part of the state they use, so I can’t say they mostly read today’s rain, although the closeness suggests it.

The order changes from bank to bank

With the four banks side by side, the public index gets the broad strokes right and the details wrong.

Each model’s place on four banksThe eight decision models ranked on the public Decision Index and on three banks of my own with real labels. A model that leads one bank can sit near the bottom of another.Decision Indexpublic, sampleShotsAUCCard riskAUCRainAUC42.90.9380.9660.820JEV-9B55.40.8550.8200.814Clef-flash 9B49.10.8860.3140.810d1-3B14.00.5800.6820.593d1-omni-600M40.90.9040.9240.805Decider-4b27.00.7390.6620.816Decider-2b20.50.8190.8770.793Strands Decider 2B3.80.3310.7300.577Laya

It gets the two weakest right: d1-omni and Laya end up at the bottom on shots and rain, which is why its order looks like my shots’ (Spearman rank correlation 0.76) and rain’s (0.74). On fraud it says nothing (0.10). And on fraud, the only bank where the differences are clear, it doesn’t anticipate the order: d1-3B, second on the index, is last, and JEV-9B, third, is the clear best. On shots and rain the top ones tie within the noise, so there was no best one for the index to pick.

What I can’t claim

It’s one sample per bank and one machine per model size. The payments are synthetic, generated for the Flink post, and the rain comes as text, worded the way the series worded it; a different wording of the state or the question can move a decision model’s results, and I didn’t try variants. I gave none of them examples or tuned anything, which is how they come out of the box and not how they’d be used in production. Of the Decision Index I ran a stratified sample of 6,000 requests from its public part, not the full suite of 140,000 nor the private tests that complete the leaderboard. And the confidence intervals resample single rows, while the payments in a burst and the rain days of one city aren’t independent: they’re narrower than they should be, so I don’t use them to claim small differences.

My take

Decision models are a good idea and JEV-9B is the proof: without a single example, it ranks fraud almost like a forest trained on the stream, and it does it by reading a sentence. For prototyping a new decision, with no data yet, that’s worth a lot.

But the public index doesn’t replace measuring on your own data. On shots and rain it ruled out the two weakest; it didn’t help choose among the good ones: the same model was second on the index and ranked one of my banks backwards, and the factory threshold is almost never the one that works. If you already have a rule, a classic model or a model trained on your data, compare them on the same set before switching.

When I’d use one and when not

  • Yes, for a new decision without labelled data, with JEV-9B or Decider-4b, and as a cheap second opinion next to a rule.
  • Carefully, calibrating the threshold on a few hundred examples of your own before using the probability as a cut-off.
  • No, instead of a model trained on your data when you already have one, nor taking first place on the index as a guarantee for your case.

Frequently asked questions

What is a decision model?

A model that doesn't generate text: it takes a state (text or JSON) and typed questions, yes or no (noul) or a pick among options with their criteria (choice), and returns a probability per option in a single pass (JEV-9B's engine uses several when a question has more than 16 options). Jev, from TypeSafe, opened the category in September 2026 as an API service; in the following weeks came Clef from Cloudflare, Strands Decider from AWS, Decider from Mapika, Laya, Liquid AI's d1 and others. The JEV-9B I measured is an open student of Jev, published by autotrust and distilled from its answers, not TypeSafe's Jev.

Which one was best on my data?

It depends on the bank. JEV-9B was the clear best on fraud (AUC 0.966, against 0.924 for the next one); on shots (0.938) and rain (0.820) it tied with others within the measurement's noise. But on none of the three did it clearly beat the reference I already use: a keyword rule on shots (0.903), a random cut forest trained on the stream for fraud (0.982) and a logistic regression for rain (0.847).

Why does d1-3B rank fraud backwards?

Because a burst of payments looks less suspicious to it than a normal payment: median probability 0.41 against 0.56. It flags high amounts (0.84) and country jumps (0.68) well, but bursts are two thirds of the bank's anomalies, so the overall order comes out inverted (AUC 0.31).

Does the 0.5 threshold work?

Not without calibration. On fraud, JEV-9B at 0.5 flags 7,177 of the 7,494 normal payments. Given the same false-alarm budget as the forest (336), it catches 89.5% of the anomalies, against the forest's 94.8%. The threshold has to be tuned on your own examples.

Do I need a big GPU?

For the 9B ones in bf16, more than 16 GB: JEV-9B took about 15 GB just to load and ran out of memory on an RTX 4070 Ti SUPER with long requests, so I ran JEV-9B and Clef-flash on a rented 48 GB A6000. The rest fit in 16 GB: d1-3B peaked at 5.8 GB, Decider-4b 8.0 GB, Strands 3.6 GB, Laya 2.3 GB and d1-omni 1.1 GB.

Can I reproduce it?

Yes. The three banks, each model's engine, the per-request results and the scripts that compute every number are published in a gist linked in the sources. The Decision Index is rebuilt with the public kit, which checks that the suite comes out identical to the official one.

Sources

Measured on October 9 and 10, 2026. On my banks, six models on a 16 GB RTX 4070 Ti SUPER (torch 2.14.1, transformers 5.19.0) and JEV-9B and Clef-flash on a rented 48 GB RTX A6000 (torch 2.8.0, transformers 5.19.0). On Decision Index 0.3 (kit at 62d2f51), d1-3B and Laya on the RTX 4070 Ti SUPER and the other six on the A6000.

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.