
Neural network explained: it matched boosting with 2,177 weights, and changing the seed moved its AUC by 0.004
What a neural network is, what each layer does and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. One hidden layer of 128 units, 2,177 weights and 81 KB, reached AUC 0.883 on the 2016-2026 test period: the same figure as a 2.1 MB boosting model, at 0.40 ms per row against 3.09. Training it cost five times more, and repeating the training with only the seed changed moved the validation AUC between 0.881 and 0.884.
One hidden layer of 128 units, 2,177 weights and 81 KB serialized, reached AUC 0.8829 on the 2016-2026 test period. A boosting model of 594 rounds and 2.1 MB reached 0.8830. The paired difference between them is −0.00002, with a 95 % interval that crosses zero: at ranking days, the bootstrap found no conclusive difference. Everything else about them can be told apart.
This is the tenth post in the series, with the same daily NASA POWER data from seven Chilean cities I used for Random Forest, linear regression, K-Means, logistic regression, decision trees, gradient boosting, SVM, KNN and Naive Bayes. The question is still whether at least 1 mm of rain will fall tomorrow. I trained on 1984-2012 and evaluated on the 2016-2026 test period, in a container with 8 CPUs and scikit-learn 1.7.2. I chose the size of the network by fitting on 1984-2009 (up to 30 December, because the label of the 31st is rain on 1 January 2010) and scoring on 2010-2012; the regularization, the step, the activation and the batch size I swept in that same period and settled on alpha=1e-4, learning_rate_init=1e-3, ReLU and the batch of 256 I pinned for the whole bench (scikit-learn’s default is auto, which here would be 200 rows). The comparison models use the settings from their own posts, and the 2016-2026 curves and examples are descriptive.
What is a neural network?
Each unit in a layer linearly combines what it receives, adds a bias and applies an activation: here, ReLU in the hidden layers and the sigmoid on the binary output. With no hidden layer, what is left computes the same thing as a logistic regression. What changes is that there are many in parallel, that the output of one layer feeds the next, and that fitting corrects all the weights at once by propagating the error backwards, the method Rumelhart, Hinton and Williams popularized in 1986. The single unit is older: it is Rosenblatt’s perceptron (1958).
A network with no hidden layer is exactly a logistic regression: 3 weights, a straight boundary, AUC 0.838 on those days and 0.803 on 2010-2012. With 32 units the boundary bends and 2010-2012 improves to 0.806. With two layers of 32, 1,185 weights, the fit on the 200 days climbs to 0.886 while 2010-2012 falls to 0.800: a wider gap, and a sign of more overfitting in this run. The 200 days are a balanced sample, 100 rainy and 100 dry, not the real frequency of rain.
With two variables and the same 200 days as the earlier posts, a network with no hidden layer has 3 weights and draws a straight line: it is exactly a logistic regression. With 32 units the boundary bends and AUC on 2010-2012 goes from 0.803 to 0.806. With two layers of 32, 1,185 weights for 200 days, the fit on those days climbs to 0.886 and 2010-2012 falls to 0.800: the gap widens, a sign of overfitting. Those 200 days are balanced on purpose, 100 rainy and 100 dry, so they do not reproduce the real frequency.
That the boundary can bend this much is what Cybenko proved in 1989 and Hornik generalized in 1991: with one hidden layer and enough units, a network can approximate any continuous function on a compact domain. The theorem says the shape exists, not that it can be estimated from the data at hand.
Bigger is not better
A logistic regression, which is the same network with no hidden layer, has 16 weights and reached 0.850. One hidden layer of 128 units, 2,177 weights, reached 0.882 and was the choice. Past that, the larger architectures evaluated under this optimizer and budget scored lower: 3 layers of 64, 9,409 weights, fell to 0.850, level with the logistic regression in this validation period. Refit on 1984-2012, the chosen network scored 0.883 on 2016-2026.
I tried eight sizes on the 66,466 rows of 1984-2009, three initializations each, scoring on 2010-2012. Logistic regression, which is this same network with no hidden layer, has 16 weights and scored 0.850. One layer of 4 units already reaches 0.873, and 32 units 0.882. The choice was one layer of 128, with 2,177 parameters counting weights and biases and 0.882 on validation, which refit on 1984-2012 scored 0.883 on the test period.
Beyond that, the larger architectures I evaluated under this optimizer and this budget landed lower: 512 units dropped to 0.879, two layers of 128 (18,689 weights) to 0.855 and three layers of 64 to 0.850, tying with logistic regression. Several of them used up all 200 epochs available. With these 66,466 selection rows and 15 variables, the network worth having was small, and that is the one I later refit on all 74,145. It is what Grinsztajn and coauthors measured in 2022 across many tabular datasets: tree-based models remain hard to beat on this terrain.
Scale before training
Unscaled and with pressure in kPa, the network reached AUC 0.874 with a log loss of 0.364; standardized, 0.884 and 0.349. Writing the same pressure in pascals, without scaling, AUC only drops to 0.845 but the log loss explodes to 6.81: it stopped after 30 epochs predicting an average of 1.9 % rain against the 22.9 % observed over 2010-2012, so the ranking survives and the probability does not. Scaled, the two unit choices give exactly the same numbers.
Unscaled with pressure in kPa, the network reached 0.874 with log loss 0.364; standardized, 0.883 and 0.349. The units trap of the series, writing the same pressure in pascals, behaves differently here than in the earlier posts: AUC barely drops to 0.845, but log loss explodes to 6.81. The fit stopped after 30 epochs and the model ended up predicting 1.9 % rain on average against the 22.9 % that actually rained over 2010-2012. It ranks days much the same and lies about how much, which is a harder failure to notice than falling to chance. Scaled, the two unit choices give exactly the same numbers.
Epoch by epoch
- 4
- 32
- 128 · 128
- fitting rows
- early-stopping epoch
The final loss came out well below the first one for all three, and that says nothing about whether the model got better outside: the two layers of 128 (18,689 parameters, weights and biases) reach the lowest loss, 0.313, and at epoch 40 they score 0.914 on the fitting rows against 0.883 on 2010-2012, a gap of 0.031. With one layer of 32 the gap is 0.018. The mark gives the epoch and nothing else, and it comes from a different protocol, not from a stop on this curve: a separate fit with early stopping, which holds out 10 % of the fitting rows, watches accuracy on them with a patience of 5 and restores the best weights. That one halted at epoch 16 for the network of 32 and landed at 0.874 on 2010-2012, against 0.882 here.
The final loss came out well below the first one for all three sizes, even though it does not fall in every single epoch, and it does not say whether the model got better outside. The two layers of 128 reach the lowest loss, 0.313, and are the ones that separate the most: 0.914 on the fitting rows against 0.883 on 2010-2012, a gap of 0.031. With one layer of 32 the gap is 0.018.
Scikit-learn’s early_stopping holds out 10 % of the fitting rows at random —after the scaler has already seen all 66,466— and stops when accuracy on that piece stops improving; it does not watch AUC and it does not respect the time cut. With a patience of 5 it halted at epoch 16 for the network of 32 and landed at 0.874 on 2010-2012, against the 0.882 of the long fit. These are two different protocols, not the same run stopped sooner: the rows that update the weights, the patience and the stopping rule all change, and at the end it restores the best weights, which are not necessarily those of epoch 16. The curve on the right is not a continuous training either: I measured it refitting one epoch at a time, which restarts the optimizer on every pass.
The same network, ten times
Ten runs of the same hidden layer of 128 units gave AUC between 0.8806 and 0.8842, a spread of 0.0037, with a median of 0.8825. They also stopped at different times: between 159 and 200 epochs. Nothing else changed. Reporting a single number from a single run hides that range, which here is as wide as the gap between two different models.
Ten runs with the same architecture, the same rows and a single difference, the seed that sets the initial weights and the order of the batches, gave AUC between 0.8806 and 0.8842, and stopped between 159 and 200 epochs. That spread of 0.0037 is the same order as the gap between two different architectures. It is why the seed belongs in the artifact and why several runs beat one: a single run is fine for comparing point results, but it does not say how much the training itself moves. And pinning the seed is not enough to reproduce a fit if the row order changes, as shown further down.
The other knobs
L2 regularization, which scikit-learn calls alpha, moved validation from 0.881 at 1e-6 to 0.884 at 0.01, and sank it to 0.851 at 10, where the fit on the training rows themselves also falls to 0.860. The initial step matters less than it seems within a wide range: between 1e-4 and 1e-2 validation moves between 0.881 and 0.883, and only at 0.1 does it break, down to 0.868 in 32 epochs. Among the three activations, the logistic sigmoid gave the best validation (0.8846) and the slowest fit (17.2 s), against 0.8823 and 14.7 s for relu. In all four sweeps the largest improvement over the base setting fell inside the range the seed alone moves, 0.0037: 0.0037 for regularization, 0.0021 for the step, 0.0024 for the activation and 0.0031 for the batch. The drops do go well past it, like the 0.029 of alpha at 10, so that range bounds the gains and not the losses. Comparing against the range of ten seeds is a heuristic call, not an equivalence test, and the activations use medians of three seeds while the other sweeps use a single one. On that basis I kept the base setting.
Batches of 32 rows took 0.215 s per epoch, 4.4 times more than batches of 2,048 (0.049 s), and scored lower: 0.879 against 0.884. The full batch, one update per epoch over all 66,466 rows, is the worst of the four at 0.870. The smaller number of updates is a plausible explanation, not one isolated here: the epoch budget was the same for all four, the number of updates was not.
Batch size is the knob that changes the time the most. Batches of 32 rows took 0.215 s per epoch, 4.4 times more than batches of 2,048, and they also landed lower on validation. The full batch, with a single update per epoch, was the worst of the four at 0.8704. The smaller number of updates is a plausible explanation, not one isolated here: I matched the epoch budget, not the update budget.
Slightly overconfident probabilities
With all 74,145 rows, the network ended with log loss 0.338, the third best of the series, after boosting (0.332) and the forest (0.335), and well below logistic regression (0.378). Looking at it by band, the confidence shows at the top: among days it gave more than 0.9, it predicted 0.947 on average and 0.872 rained; among those between 0.7 and 0.8, it predicted 0.749 and 0.594 rained. It needs no calibration to rank, and it is worth checking before the number is used as a probability.
For numbers, and the ceiling it does break
- neural network
- linear regression
- boosting
Trained only on the April to September rows, the highest next-day maximum the models saw while training was 32.97 °C, and the real December to February summers averaged 29.63 °C. On the real summer days the network went up to 38.10 °C, past everything it had seen, and missed by 4.03 °C (the line above, with the other variables fixed at their median, stays well below that); boosting never went above 29.56 °C and missed by 5.48 °C. Linear regression, the simplest of the three, was the most accurate here: 1.72 °C. Extrapolating is not the same as extrapolating well: the network leaves its range, and that alone does not make it right. On the normal task, predicting tomorrow's maximum with all the rows, it missed by 1.46 °C against 1.44 °C for boosting.
Predicting tomorrow’s maximum temperature with all the rows, the network missed by 1.46 °C on average, between boosting (1.44 °C) and linear regression (1.70 °C).
The series’ extrapolation test is more interesting. Trained only on Santiago’s April to September rows, and asked about the December to February ones of 2016-2026, the highest next-day maximum they saw while training was 32.97 °C. The network predicted up to 38.1 °C, past everything it had seen, and missed by 4.03 °C; boosting never went above 29.56 °C and missed by 5.48 °C. Leaving the range does not make it right: linear regression, the simplest of the three, was the most accurate here at 1.72 °C. It is a seasonal cut, not a clean out-of-range test: summer changes several variables at once.
The network against the whole series
At a quarter of real time.
The 27,256 rows, in real time.
Serialized, the network takes 81 KB, boosting 2.1 MB, the forest 124 MB and KNN 9.5 MB. Scoring one row took the network 0.40 ms, boosting 3.09 ms and KNN 1.92 ms.
With the same rows, I resampled the city-year blocks of 2016-2026 2,000 times to compare AUC in pairs. Against boosting the network landed at −0.00002 with an interval from −0.0025 to 0.0022, and against the forest at 0.0006, from −0.0021 to 0.0030: both intervals cross zero. Against KNN it was 0.013 above, against the tree 0.021, against logistic regression 0.036 and against Naive Bayes 0.061, and those four intervals do not cross zero.
Fitting the network took 13.4 s on one thread in the comparison run. With eight threads it took the same, 13.4 s: the threads did not help. In the cost run it took 8.5 s, because with a different row order it stopped at epoch 104.
Where a neural network lives in a real system
Measured on one thread with all 74,145 rows, all of it from the comparison run: one hidden layer of 128 units takes 13.4 s, weighs 81 KB serialized, scores one row in 0.40 ms and the whole test period in 0.007 s, at AUC 0.883 and log loss 0.338. Boosting reaches AUC 0.883 in 2.6 s, but takes 2.1 MB and 3.09 ms per row. Logistic regression: 0.049 s, AUC 0.847. With eight threads the fit took 13.4 s, the same as that one thread. A separate cost run, with the rows in another order, stopped at epoch 104 in 8.5 s and 80 KB; it does not measure quality.
- Online prediction with many queries. Scoring a row takes 0.40 ms, the lowest latency among the models that landed at the top on AUC, and the whole artifact weighs 81 KB.
- The scaler inside the artifact. The artifact has to keep the preprocessing it was trained with. In this experiment, training unscaled and with pressure in pascals hurt the log loss far more than the AUC.
- A pinned seed and several runs. Without them the same code returns different numbers and no change can be judged.
- Data that does not fit in a table. Where a network really pulls ahead is text, images or audio, not fifteen numeric columns.
When I would choose it: when latency per row and artifact size matter, when there is non-linear signal a linear model leaves out, or when the problem stops being tabular. When not: for a medium table where boosting reaches the same AUC in a fifth of the time and with no scaling, using the boosting setting I pinned for this comparison.
Sources
- Rosenblatt, F. (1958). “The perceptron: A probabilistic model for information storage and organization in the brain”. Psychological Review, 65(6), 386–408. DOI 10.1037/h0042519.
- Rumelhart, D. E., Hinton, G. E. and Williams, R. J. (1986). “Learning representations by back-propagating errors”. Nature, 323, 533–536. DOI 10.1038/323533a0.
- Cybenko, G. (1989). “Approximation by superpositions of a sigmoidal function”. Mathematics of Control, Signals and Systems, 2(4), 303–314. DOI 10.1007/BF02551274.
- Hornik, K. (1991). “Approximation capabilities of multilayer feedforward networks”. Neural Networks, 4(2), 251–257. DOI 10.1016/0893-6080(91)90009-T.
- Kingma, D. P. and Ba, J. (2015). “Adam: A Method for Stochastic Optimization”. ICLR 2015. It is the default optimizer of
MLPClassifier. - Grinsztajn, L., Oyallon, E. and Varoquaux, G. (2022). “Why do tree-based models still outperform deep learning on typical tabular data?”. NeurIPS Datasets and Benchmarks.
- scikit-learn 1.7, supervised neural networks.
- NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.