17°
Portada del artículo: A model that does not write, it only decides: I had Intern-Decision predict the rain without training it
ModelsBenchmarksGPUData

A model that does not write, it only decides: I had Intern-Decision predict the rain without training it

Intern-Decision came out on 26 September 2026: 0.8B to 4B models that do not generate text and return one probability per question in a single forward pass. I installed it the same day, ran into three failures along the way, one of which made it between 2 and 3.2 times slower, and had it predict tomorrow's rain over 27,256 days in seven Chilean cities without training it. The 2B reaches an AUC of 0.823, almost the same as a Naive Bayes trained on 74,145 rows, and stays below logistic regression.

Efrain Garay 26 September 2026

Playing summary

Last week a new category of language model showed up: one that does not write. Jev, from TypeSafe AI, takes text and returns numbers: probabilities for categories, yes-or-no questions and ratings. But Jev, according to that introduction, works only as an API.

Today, 26 September, Shanghai AI Lab uploaded Intern-Decision to Hugging Face: the same idea with open weights, the Apache 2.0 license and three sizes. The largest takes about 9 GiB of video memory and the other two fit on modest cards. When I downloaded it, it had zero downloads.

I installed it that same morning and asked it the question I had already asked nine models in the algorithms series: will it rain tomorrow? The difference is that I did not train this one.

bench · resultsdone · exit 0

$ python banco.py 2B lluvia

0.823intern-2B · no training, prompt only
0.847logistic · trained, 74,145 rows
0.883boosting · trained, 74,145 rows

verdict > Without training it on this data, the 2B landed 0.024 of AUC below the trained logistic regression and tied the Naive Bayes, at 21.8 ms per query.▋

What it is

A regular language model answers by generating text, token by token. Intern-Decision does something else: it takes a state (text, and optionally images) and a schema of closed questions, builds a template with one <decision> marker per question, and runs a single forward pass. At the position before each marker it reads the logits, keeps only those of the allowed options (A, B, C…) and applies softmax after dividing them by a temperature fixed at the factory.

There is no generate() and no sampling. That is why it answers in milliseconds, and why it can answer several questions for the price of one.

How Intern-Decision answers: one forward pass, no text generatedThe same request the rain test sends, drawn on a real test day, saved with the measurement data. Model shown: 2B.
How Intern-Decision answers: one forward pass, no text generatedstate of the dayTemuco · 2021-08-07max 10.5 °C · humidity 86 %templatequestion: rain tomorrow?answer: <decision>one forward passfine-tuned Qwen3.5no generate(), no samplingmarker logitsposition before <decision>only A = no, B = yessoftmax(logits / T)T = 2.10factory calibrationtyped answerrain: P(yes) = 0.32plus the decision labelHow Intern-Decision answers: one forward pass, no text generatedstate of the dayTemuco · 2021-08-07max 10.5 °C · humidity 86 %templatequestion: rain tomorrow?answer: <decision>one forward passfine-tuned Qwen3.5no generate(), no samplingmarker logitsposition before <decision>only A = no, B = yessoftmax(logits / T)T = 2.10factory calibrationtyped answerrain: P(yes) = 0.32plus the decision label

No token is ever sampled. The whole answer is two logits read at one position, which is why a 0.8B model answers in about 20 ms on a GPU.

All three sizes are fine-tuned from Qwen3.5, which mixes regular attention with linear attention. That architectural detail is the one that cost me the most during installation.

In 59 seconds, with narration: an oracle that does not speak, it only raises a card with a number. What broke when installing it, the 19.9 ms per query, the AUC of 0.823 predicting rain without training, and the 0.62 ceiling on its probabilities. Every figure is a measured one. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

Installation, and what broke

The README asks for Python 3.12, pip install -r requirements.txt and importing DecisionEngine from the model folder. I followed it to the letter, inside an isolated environment on my machine with the RTX 4070 Ti SUPER, and I tripped three times.

1. The model cannot find its own tokenizer. The first load fails with Couldn't instantiate the backend tokenizer, even though tokenizer.json is right there, also with the transformers version the model pins (5.14.1). The cause is in its inference.py: it uses Path(__file__).resolve().parent as the checkpoint folder. In the Hugging Face cache, inference.py is a symbolic link to blobs/<hash>, so resolve() points to the blobs folder, where no file keeps its name. The fix is to pass the snapshot folder explicitly: DecisionEngine(checkpoint=snap). I saved the reproduction with both loads: the one that fails comes from the blobs folder, and the one that passes the snapshot loads without problems.

2. It runs, but between 2 and 3.2 times slower than it should. On load, transformers warns in a line that is easy to miss: The fast path is not available because one of the required library is not installed. Qwen3.5 uses linear attention, and its fast path needs flash-linear-attention and causal-conv1d. Neither is in the model’s requirements.txt. Without them, transformers falls back to a pure PyTorch implementation.

3. causal-conv1d does not compile with the CUDA I had. For this combination of torch and CUDA it is built from source. With nvcc 13.0 it fails in the headers of Fedora 43’s glibc (mathcalls.h: error: exception specification is incompatible with that of previous function "rsqrt"). With nvcc 13.1 it compiles; I checked it afterwards with a minimal file, and it makes no difference whether the host compiler is gcc 14 or gcc 15. What I used, limited to my card’s architecture:

export CUDA_HOME=/usr/local/cuda-13.1 PATH=/usr/local/cuda-13.1/bin:$PATH
export TORCH_CUDA_ARCH_LIST=8.9 NVCC_CCBIN=/usr/bin/g++-14 CC=gcc-14 CXX=g++-14 MAX_JOBS=8
uv pip install flash-linear-attention
uv pip install --no-build-isolation causal-conv1d

The two paths use different compute kernels and can give slightly different probabilities; I did not quantify that difference over the full test. All the rain predictions that follow were made with the fast path.

How long each query takes

I measured inside a container with pod limits: 8 CPUs, 24 GB of RAM and the GPU through CDI. Two hundred queries per model after warming up, timing from when the dictionary goes in until the answer comes out. I ran two kinds: the README example (three questions at once) and the rain query (one question).

Milliseconds per query, from request to answerThe rain query inside the pod with 8 CPUs and 24 GB: 200 times per model on GPU, 30 on CPU.
0.8B with the fast path19.81 ms
2B with the fast path21.87 ms
4B with the fast path43.57 ms
0.8B without the two libraries63.22 ms
2B without the two libraries64.22 ms
4B without the two libraries85 ms

Without flash-linear-attention or causal-conv1d, transformers falls back to pure PyTorch for linear attention. One second of animation is 20 measured ms.

0.8B in bf16298.12 ms
2B in bf16647.65 ms
4B in bf161635.06 ms
0.8B in float32533.41 ms
2B in float321175.88 ms
4B in float323213.97 ms

Median of 30 queries on a Ryzen 7 7800X3D with an 8-CPU quota and no visible GPU. One second of animation is one measured second.

On GPU, mean of 200 queries: with the fast path the p95 stays within 0.3 ms of the mean, and without it within 1 ms. On CPU, median of 30, because the spread is larger: the 4B in float32 goes from 3.2 s at the median to 5.6 s at the p95. In the README's table, measured on an RTX 4090, the author reports 34, 33 and 44 ms.

ModelWith the fast pathWithout the two librariesPeak memory (MiB allocated by PyTorch)Reported (RTX 4090)
Intern-Decision-0.8B19.9 ms63.3 ms1,94834.0 ms
Intern-Decision-2B21.8 ms64.1 ms4,25533.3 ms
Intern-Decision-4B42.5 ms85.4 ms9,05644.2 ms

Mean latency of the three-question example, which is the one the README reports. The rain query, with a single question, stays within 1.1 ms of that figure; the chart above uses the rain query.

On CPU, in the same pod (8-CPU quota) with no visible GPU, the 0.8B’s median is 0.30 seconds per query in bf16 and 0.53 in float32. The 4B reaches 1.6 seconds in bf16 and 3.2 in float32, with a p95 of 5.6. One query at a time, the 0.8B in bf16 gives about 200 decisions per minute; the same model on the GPU, about 3,000. The engine loads in bf16 by default and accepts dtype="float32"; on this Ryzen, which does have bf16 instructions, bf16 was the fastest.

Two things caught my attention. The README query, with three questions, and the rain query, with one, take almost the same time; it is not a clean proof that three cost the same as one, because the states have different lengths, but it points in the direction the idea promises. And on my card, more modest than the author’s, the 0.8B and the 2B come in below what he reports. I do not know what installation he measured with.

My own test: will it rain tomorrow?

The algorithms series always uses the same table: daily data from NASA POWER for seven Chilean cities, from La Serena to Punta Arenas. It trains on 1984-2012 (74,145 rows) and tests on 2016-2026 (27,256 days). The question is whether at least 1 mm will fall tomorrow.

I did not give Intern-Decision the training rows. For each test day I sent it a text state with the same weather information the trained models use, but written out: rounded values, the wind in eight compass directions, the full date instead of its sines and cosines, and the city name. Plus a yes-or-no question:

Weather station in Temuco, Chile (latitude -38.7). Date: 2021-08-07.
Today: max temperature 10.5 °C, min temperature 0.2 °C, precipitation 0.7 mm
(yesterday 9.6 mm), relative humidity 86 %, dew point 3.3 °C, surface pressure
99.16 kPa (+0.39 kPa since yesterday), wind 4.9 m/s from the S, solar radiation
12.3 MJ/m².

rain (noul): Will it rain at least 1 mm tomorrow at this station?

I wrote it in English because its benchmark is in English. The probability of “yes” is the prediction. From that I compute the AUC, the same metric used across the whole series, over the same 27,256 days.

AUC over the same 27,256 test daysAUC · higher is better
  1. Boosting, 594 rounds (trained)0.883
  2. Forest of 200 trees (trained)0.882
  3. Logistic regression (trained)0.847
  4. Intern-Decision-2B (untrained)0.823
  5. Gaussian Naive Bayes (trained)0.822
  6. Intern-Decision-4B (untrained)0.821
  7. Intern-Decision-0.8B (untrained)0.803
  8. Persistence: tomorrow same as today0.709

Trained models: the series' unified table (transformer/vs_models.json), trained on 1984-2012. Intern-Decision: no training, one query per day, inside the pod with the RTX 4070 Ti SUPER.

Read carefully: a model I did not train on these stations ranks the days better than the naive “tomorrow same as today” rule, and lands 0.0004 away from a Naive Bayes that was trained on 74,145 rows. But the simplest logistic regression, a 1.8 KB file that answers in under a millisecond on CPU (0.4 ms, measured in the series with one thread and outside this container), beats it by 0.024. And the tree models, the forest and boosting, by 0.06.

The 4B does not improve on the 2B. It gets 0.821 against 0.823, with twice the memory and twice the latency. On the average of the author’s benchmarks the 4B does win, so this looks specific to the task.

The problem is not the ranking, it is the scale

AUC only checks whether the rainy days end up above the dry ones. It does not check whether the probability tells the truth. And that is where the thing that surprised me most shows up.

In the 10% of days each model considers most likely to rain, it rained 63.2% of the time with the 2B, 63.6% with the 4B and 63.3% with logistic regression. At that end they hit in the same proportion. What changes is the number they put next to it.

Compressed probabilities: the 2B never goes above 0.62, the 4B above 0.68Each point is a band of predicted probability; its height is how often it actually rained the next day. Point size is the number of days in the band.
0.00.00.20.20.40.40.60.60.80.81.01.0ceiling 2B 0.62ceiling 4B 0.68predicted probability of rainobserved rain
  • trained logistic regression
  • Intern-Decision-2B, no training
  • Intern-Decision-4B, no training

Same 27,256 test days for the three curves. Above the diagonal the model is under-confident: in the 2B 0.5–0.6 band, the highest with at least 30 days, it predicted 0.52 on average and it rained 79% of the time.

The highest value the 2B emitted across the 27,256 days was 0.62. Only 2% of the days reached 0.5 or more, and barely 1.3% went above it; in the 0.5 to 0.6 band, with 547 days, it predicted 0.52 on average and it rained 79% of the time. Its decision label (the decision field, which picks the most likely option and says “no” on an exact tie) marks rain on 346 of 27,256 days: it catches 5% of the days it did rain. The 0.8B is worse: its maximum was exactly 0.50, on 6 days, and since the tie resolves to “no”, its label did not mark rain on a single day.

Where it wins without training

Split by city, the result is uneven. The 2B beats logistic regression in three of the seven cities, and they do not follow a simple frequency pattern: La Serena, where it almost never rains, Santiago, and Punta Arenas, where it rains the most.

CityDays with rain the next dayLogistic regressionIntern-Decision-2B
La Serena2.2%0.8360.886
Santiago9.7%0.7350.771
Punta Arenas39.2%0.5600.617
Valparaíso7.4%0.8500.809
Concepción20.5%0.8600.815
Temuco26.1%0.8390.779
Puerto Montt36.5%0.8280.780

I have a hypothesis I did not test. Logistic regression learns a single set of weights for all seven cities; it gets the latitude (its third strongest coefficient in absolute value), but not the city name. The language model reads “La Serena” as context, and in La Serena, where it almost never rains, that context can be worth more than a shared weight. Punta Arenas, where it rains on 39% of the days, does not fit that explanation. It could also be something else: per-city AUC measures how it ranks within each city, and there any local advantage counts.

What this measurement does not say

  • It is a single task. Tabular, numeric, and in a domain where a model trained on 74,000 rows has every advantage. In free-text classification, which is what the author shows it for, the result may be different.
  • I did not tune the wording of the state. I wrote a reasonable prompt and did not change it after seeing results, on purpose: tuning the text while looking at the test would be training by hand on the test.
  • The model may have seen weather data during pretraining. There is no way to know how much of the 2016-2026 period Qwen3.5 knows.
  • Latency is per query, one at a time. According to its README the engine takes one request per call; I did not measure batches or parallel queries.
  • I measured the CPU in the same pod with an 8-CPU quota (Ryzen 7 7800X3D), not on an old CPU. My other machine, an i5-9400F, was busy with other work and I discarded what I measured there.

When I would use it and when I would not

I would use it for decisions over text where I do not have labeled data yet: classifying tickets, routing requests, deciding whether a message needs a human reply. Twenty milliseconds per query on a consumer card, several questions in a single pass, and typed JSON instead of text that has to be parsed. I did not measure it on those tasks; that is where I would try it first, as an initial classifier while labels are collected.

I would not use it where I already have data to train something simple. In this test, with 74,145 training rows, a 1.8 KB logistic regression beat it on quality and latency, and it runs anywhere. Nor would I use its probability as is to decide with a threshold: in my test it was compressed and the decision label almost never said yes.

And if you install it, also install flash-linear-attention and causal-conv1d. Without them you pay between 2 and 3.2 times the latency with no warning other than a line in the log.


Measured on 26 September 2026 on a 16 GB RTX 4070 Ti SUPER, inside a container with 8 CPUs and 24 GB (Fedora 43), with torch 2.14.0+cu130 (the README pins 2.9.1), transformers 5.14.1, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 and NVIDIA driver 595.91. Data from NASA POWER, the same table used across the series.

Sources

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.