
Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwards
What a Transformer is, what attention does and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. The chosen network has 9,633 parameters and 49 KB, and landed below boosting, the forest and a one-layer network. The ranking is not the lesson: among the six architectures that competed, the validation order inverted on the test period.
A network of 9,633 parameters and 49 KB reached AUC 0.876 on the 2016-2026 test period. A 2.1 MB boosting model reached 0.883, a 124 MB forest reached 0.882, and a network with one hidden layer, 81 KB, reached 0.883. The paired difference against all three does not cross zero with any of the seeds I tried. The Transformer, the architecture everything else is built on today, came out underneath on fifteen columns of weather.
This is the eleventh and final post in the series, with the same daily NASA POWER data from seven Chilean cities I used for Random Forest, linear regression, K-Means, logistic regression, decision trees, gradient boosting, SVM, KNN, Naive Bayes and neural networks. The question is still whether at least 1 mm of rain will fall tomorrow. I trained on 1984-2012 and evaluated on 2016-2026, in a container with 8 CPUs, PyTorch 2.9.1 on CPU and scikit-learn 1.7.2.
The contract, which matters more here than in the earlier posts: the architecture was chosen without looking at the test period. The duration of each fit came from a frame inside the training period, 2008-2009; the architecture was chosen by scoring 2010-2012, one reading per candidate; and the test period was read once with the chosen one. Everything else that appears about 2016-2026 is descriptive and chose nothing. There were also two earlier readings of the test period, with two other architectures, and they are published further down rather than hidden.
What is a Transformer on a table?
The weights w and b are learned per column, so each one has its own way of turning a number into a vector. There is no time window and no order between columns: the model sees exactly the same information as the other ten in the series. The [CLS] token starts as a learned vector, collects from the rest through attention, and the probability of rain comes out of it.
Each column becomes a vector of its own: token = value × w + b, where w and b are learned per column. Fifteen columns give fifteen tokens. A [CLS] token that does not come from the data is added, attention lets every token look at the others and decide how much each one weighs, and the prediction comes out of the [CLS]. It is the architecture Vaswani and coauthors published in 2017 for translation, adapted to tables by Gorishniy and coauthors in 2021.
There is no time window and no order between columns, on purpose: that way the model sees exactly the same information as the other ten in the series. Giving it the last few days would have been giving it data none of the others had.
One difference from the original paper worth declaring: they drop the normalization of the first block and I kept it in every block. That decision, which looks minor, is what explains the units section.
The grid that did not survive the test period
rank correlation 0.63 · linear 0.84 · n 11rank correlation -0.49 · linear -0.39 · n 6
With all 11, validation looks right: rank correlation 0.63 and linear 0.84, held up by the width-256 cells sitting at the bottom of both periods. Take them out and the ranking among the 6 that competed inverts: -0.49. Their moves run from -0.0010 to 0.0034 with no relation to what validation had ranked — d128_b4 (534,273 parameters) sat 7th on validation and ends up first, while the best on validation (d64_b4) ends at 0.8801 behind it. The chosen cell (d32_b1, 9,633 parameters) moves -0.0003. The search separated what did not compete and did not order what did.
I tried eleven architectures, from 2,769 to 2,117,121 parameters, moving one axis at a time: the width of the token and the number of blocks, at sixteen dimensions per head. Each one with three seeds, the duration fixed on the inner frame and a single reading of 2010-2012.
On validation the ranking was clear: the width-64 architectures on top, the best at 0.8802, and the width-256 ones at the bottom, collapsing to 0.872, two of them stopping after four epochs. It looks like the conclusion of the previous post, that growing does not help.
On the test period that ranking falls apart, and it is worth being precise about where. Looking at all eleven together, validation got it right: the rank correlation between the two periods is 0.63 and the linear one 0.84, because the three width-256 architectures sit at the bottom of both and hold that relation up. What validation separated well is what was not competing.
Among the six that were competing, the width-64 and width-128 ones, the order does not hold: it inverts. The rank correlation among them is −0.49. The three width-256 architectures rise between 0.0025 and 0.0038, and among the six at the top the moves run from −0.0010 to +0.0034 with no relation to what validation had ranked: the one with 534,273 parameters rises 0.0034 and goes from seventh to first, while the best on validation drops and ends up third.
This is not a failure of the bench, it is the result. On this table and under this budget, the search told apart what was not competing and failed to order what was.
The margin validation could not resolve
Both intervals cross zero and both are wider than the margin they were meant to respect: 0.0037 against 0.0055, two to three times 0.002. The rule did not pick a worse architecture, it picked the smallest among architectures this frame cannot tell apart, which is what it exists for. What has to change next time is where the margin comes from: I set it at the scale of the seed noise, about 0.001, when it should come from the sampling error of the validation frame. One warning: the interval of the absolute AUC is 0.071 wide on average, twenty times more, and using that one to compare architectures would be wrong — comparing two of them on the same rows cancels the common noise.
The rule was written before running: among the architectures within 0.002 AUC of the best, keep the smallest. It chose one with 9,633 parameters where the best had 136,065, giving up 0.0018 of validation AUC for fourteen times fewer parameters.
The problem is that 0.002 was finer than that validation frame can measure. Resampling the city-year blocks, the interval of the difference between the chosen one and the best runs from −0.0045 to +0.0010: a width of 0.0055, almost three times the margin, and it crosses zero. Against the runner-up the interval is 0.0037 wide and also crosses.
That does not invalidate the rule, it explains it: it picked the smallest among architectures validation could not tell apart, which is exactly what it is for. What has to change next time is where the margin comes from. I set it at the scale of the noise the seed moves, about 0.001, when it should have come from the sampling error of the validation frame, which is five times larger.
A detail that nearly cost me: the interval of the absolute validation AUC is 0.071 wide, twenty times more. It is tempting to use that one to say validation cannot tell anything apart, and it would be false: comparing two architectures on the same rows cancels the sample noise, and the only honest number is the paired difference.
Attention is not importance
They agree on the first one: rain today takes 0.241 of the attention against the 0.067 an even split would give, and shuffling it costs 0.103. Below that they part ways. pressure change gets 0.097 of attention and costs 0.029, while wind (cos) gets 0.067 and costs 0.034 — more damage with less attention. dew point receives 0.062 and shuffling it costs 0.002, which is nothing. The top five by attention and the top five by damage are not the same set.
The attention the [CLS] pays to the fifteen columns can be read, and it is the flashiest part of the model. In the chosen network, today’s rain takes 0.241 where an even split would be 0.067, and it is also the column whose shuffling does the most damage: AUC falls 0.103. There they agree.
Below that they stop agreeing. Pressure change gets 0.097 of attention and shuffling it costs 0.029; latitude gets 0.095 and costs 0.036; the cosine of the wind gets 0.067 and costs 0.034, more damage with less attention. Dew point gets 0.062 and shuffling it costs 0.002, which is nothing. The top five by attention and the top five by damage are not the same set.
This is what Jain and Wallace showed in 2019: attention weights explain how information gets mixed, not how much the model needs it. An attention map describes the mechanism; it does not measure importance.
Taking the attention out did not make it worse
Eighteen per-seed paired differences that run from -0.0063 to 0.0044, 8 of them positive, and the sign flips between architectures and even between seeds of the same one. In the chosen cell (d32_b1) both replacements win; in d64_b4 both lose; in d128_b2 they disagree with each other. The even-weight arm keeps V and the output projection, so the parameter count that actually runs is the same; the plain mean drops them, which is why it is the weaker control. On fifteen weather columns, attention does not move the AUC more than changing the seed does.
If attention is the piece that defines a Transformer, the obvious question is what happens without it. I tried two replacements, on three architectures and with three seeds each: one that keeps the whole block but spreads the weights evenly across the sixteen tokens, the fifteen columns and the [CLS], and one that simply averages the tokens and skips the projections.
In the chosen architecture, removing it wins: 0.0026 with even weights and 0.0040 with the mean. In the 136,065-parameter one it loses, 0.0048 and 0.0011. In the 269,313-parameter one the two replacements point in different directions. And within a single architecture there are seeds that flip the sign.
Six median differences running from −0.005 to +0.004, with no stable sign, say one thing, and it is not that attention gets in the way: on fifteen columns of weather, attention does not move the AUC more than changing the seed does.
The layer that saves it from the trap and switches a column off
Unscaled, the raw token travels 7.62 and what reaches the attention travels 0.011: the logit moves 0.00017. In pascals the raw token travels 7,652 and the logit does not move at all. Standardized, the normalized token travels 9.84 and the logit moves 0.59. The cosine between normalized tokens of different days is 1.000000 unscaled against 0.997476 standardized: the column enters and says the same thing every day. Remove that first normalization and the unit trap bites again — AUC falls to 0.573 in pascals against 0.874 in kPa.
The units trap of this series means writing the same pressure in pascals instead of kilopascals, unscaled. It sank the SVM and sent the neural network’s log loss to 6.81. Here nothing happens: 0.8688 in kilopascals and 0.8688 in pascals, identical to four decimals.
My first explanation was that each column has its own weight and absorbs the scale. The measurement said otherwise: the pressure weight measures 0.5108 in kilopascals and 0.5128 in pascals, and stays at 0.5128 multiplying by a million and by a billion. It does not shrink. Nobody compensates anything.
What happens is something else. With the weights frozen, moving pressure from its lowest to its highest value, the raw token travels 7.6 units and the already normalized token moves 0.011. The output logit changes by 0.0002. In pascals the raw token travels 7,651 and the normalized one 0.00001: the logit does not change at all. Standardize the column instead and the normalized token moves 9.8 and the logit changes by 0.59.
The normalization ahead of attention, the one Ba, Kiros and Hinton proposed in 2016, receives a token whose typical value is enormous next to its variation between days, and returns practically the same vector for every row: the cosine between normalized tokens of different days is 1.000000. The column goes in, and says nothing. And it can be checked by removing that layer: without it, the same trap drops AUC from 0.874 to 0.573.
So the piece that protects the model from the unit change is the same one that switches the column off. Scaling is still mandatory, even though the trap no longer does damage.
Probabilities that drift upward
The network ended with a log loss of 0.356, the worst among the models that compete at the top: boosting lands at 0.332. It predicted 27.3 % rain on average when 20.3 % of the days rained. Band by band: among the days it gave between 0.6 and 0.7 it predicted 0.648 and 0.491 rained; between 0.2 and 0.3 it predicted 0.249 and 0.135 rained. Only in the highest band, with 568 days, does it come close: 0.938 against 0.907. For ranking days it works; for reading the number as a probability, not without calibrating it.
With all 74,145 rows, the network ended with a log loss of 0.356, the worst among the models that compete at the top: boosting lands at 0.332 and the neural network at 0.338. It predicted 27.3 % rain on average when 20.3 % of the days rained.
Band by band you can see where: among the days it gave between 0.6 and 0.7 it predicted 0.648 and 0.491 rained; between 0.2 and 0.3 it predicted 0.249 and 0.135 rained. It overestimates across almost the whole scale, and only in the highest band does it come close. For ranking days it works; for reading the number as a probability, not without calibrating.
The cost
At a quarter of real time.
The 27,256 rows, in real time.
All from the same one-thread run. Serialized, the Transformer takes 49 KB, the neural network 81 KB, boosting 2.1 MB and the forest 124 MB, without counting that the Transformer needs PyTorch alongside. Scoring one row took the Transformer 0.62 ms, the network 0.42 and boosting 3.11. In the cost run, fitting with eight threads dropped from 41.3 s to 23.9 s: here threads do help, unlike the neural network and the SVM.
It is the most expensive of the eleven to train and the third slowest to score, to land 0.007 below the best. In exchange it is the lightest of the four at the top, and fourth of the eleven, though that number misleads: the 49 KB are the weights, and PyTorch has to be installed next to them.
For numbers, and the ceiling it does not break
Predicting tomorrow’s maximum temperature, the Transformer missed by 1.56 °C on average, between boosting (1.44 °C) and linear regression (1.70 °C).
The series’ extrapolation test leaves it worse off than anyone. Trained only on Santiago’s April to September rows and asked about the December to February summers, it predicted at most 23.09 °C when the highest next-day maximum it saw while training was 32.97 °C, and missed by 6.71 °C. Boosting, which also stays inside, reached 29.56 °C and missed by 5.48 °C. Linear regression went up to 36.97 °C and missed by 1.72 °C. The neural network of the previous post reached 38.1 °C.
It is consistent with the normalization result: bounded normalized tokens produce bounded outputs. I do not measure that directly, so I leave it as consistent and not as demonstrated.
The two earlier readings, which I am not hiding
Before this grid there was another one, with a protocol error: it used the AUC on 2010-2012 for two things at once, stopping each fit and choosing among candidates. That is a maximum over twenty per-epoch evaluations and then a maximum over eight candidates, on the same noise. It is the selection bias Cawley and Talbot documented in 2010.
That grid chose an architecture of 269,313 parameters with 0.8851 on validation. Repeating the fit ten times with only the seed changed moved the AUC between 0.8785 and 0.8851, and the seed the whole grid had run on turned out to be the best of the ten. Under the corrected protocol that same architecture gives 0.8791, and its range across seeds is 0.0015. I will not put those two figures side by side, because one comes from ten seeds and the other from three and a range grows with the number of draws; what can be said is that much of that 0.0066 was not the training moving, it was the procedure keeping the best of twenty moments and of eight candidates.
That architecture was refit and read on the test period: it gave 0.8805, with 200.7 s of refitting. I then fixed the protocol but still without measuring the lower edge of the grid, and under that second rule the choice was a different one, of 35,649 parameters, which gave 0.8810 on the test period; there the differences against the neural network, boosting and the forest crossed zero with one of the three seeds, so they could not be told apart consistently.
With the complete grid, lower edge included, the choice is the 9,633-parameter one and the test period gives 0.876, with no difference crossing zero on any seed. The more favourable numbers came out of the worse procedures, and that is why the procedure is fixed beforehand and not changed after seeing the result.
Where a Transformer lives in a real system
Measured on one thread with all 74,145 rows: 9,633 parameters, 41.3 s to fit and 49 KB serialized, scoring one row in 0.62 ms and the whole test period in 0.362 s, at AUC 0.8760 and log loss 0.3555. Boosting reaches AUC 0.8830 in 2.6 s. With eight threads the fit drops to 23.9 s, so here threads do help. And the artifact is not the file: those 49 KB are the weights, with PyTorch to install next to them.
- Outside the table. Where this architecture pulls ahead is text, images, audio or long series, where order and context are the signal. Fifteen numeric columns are not that problem.
- The preprocessing inside the artifact. Not because the units trap does damage, but because without scaling a column can go in switched off and nobody finds out.
- PyTorch alongside. The weights take 49 KB and the runtime that executes them, hundreds of megabytes. The artifact is not the file.
- A pinned seed and several runs. With three seeds per architecture, the gap between two neighbouring architectures fits inside the noise.
When I would choose it: when the problem stops being a table. When not: for fifteen numeric columns where boosting gets further in a sixteenth of the time.
Eleven models, one table
This is where the series ends. Eleven algorithms on the same 74,145 rows and the same question, and the ceiling sat near 0.883: boosting, the neural network and the forest reach it, and the Transformer stayed 0.007 below.
What separated the models was not the AUC. It was the cost of training and of scoring against the accuracy gained, the size of the artifact, whether you have to scale, whether you have to choose a size, how much the seed moves them, and how much you have to measure before you can choose honestly. The training/scoring cost against accuracy is exactly what Grinsztajn and coauthors measured in 2022 across many tabular datasets; the rest, now, with my own numbers.
And what I take from the eleven is not which one won. It is the two findings that showed up once the measuring was done carefully: that a badly built procedure hands you the most favourable number, and that a grid can tell apart what is not competing while failing to order what is.
Limits on all of the above: a single table, CPU, one optimizer and one learning rate, three seeds per architecture, and a budget of sixty epochs.
Sources
- Vaswani, A. et al. (2017). “Attention Is All You Need”. NeurIPS 2017.
- Gorishniy, Y., Rubachev, I., Khrulkov, V. and Babenko, A. (2021). “Revisiting Deep Learning Models for Tabular Data”. NeurIPS 2021. The FT-Transformer this variant comes from.
- Jain, S. and Wallace, B. C. (2019). “Attention is not Explanation”. NAACL 2019.
- Ba, J. L., Kiros, J. R. and Hinton, G. E. (2016). “Layer Normalization”.
- Cawley, G. C. and Talbot, N. L. C. (2010). “On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation”. JMLR, 11, 2079–2107.
- Grinsztajn, L., Oyallon, E. and Varoquaux, G. (2022). “Why do tree-based models still outperform deep learning on typical tabular data?”. NeurIPS Datasets and Benchmarks.
- Loshchilov, I. and Hutter, F. (2019). “Decoupled Weight Decay Regularization”. ICLR 2019. The AdamW the bench uses.
- NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.