17°
Portada del artículo: Gradient boosting explained: 594 chained trees that train 9 times faster than a forest
Machine learningAlgorithmsPythonData

Gradient boosting explained: 594 chained trees that train 9 times faster than a forest

What gradient boosting is, how each tree corrects the previous ones and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. With rate 1.0 it destabilized and after 3,000 rounds ended at AUC 0.681; with rate 0.03 and 594 rounds chosen on validation it reached 0.883, with no conclusive difference from a 200-tree forest.

Efrain Garay 16 September 2026

Playing summary

With learning rate 1.0, gradient boosting found its best point at round 7 and then destabilized. After 3,000 rounds, on 2016-2026 days it ended at AUC 0.681, below a decision tree with no depth limit. The same algorithm with rate 0.03 and 594 rounds, chosen on a validation period, reached 0.883.

This is the sixth part of the series, with the same NASA POWER daily data from seven Chilean cities I used for Random Forest, linear regression, K-Means, logistic regression and the decision tree. The question is still whether it will rain at least 1 mm tomorrow. I trained on 1984-2012 and evaluated on 2016-2026, inside a container with 8 CPUs and scikit-learn 1.7.2. I chose the rate and the rounds by fitting on 1984-2009 and scoring on 2010-2012; everything else stayed at its default or was fixed for each comparison.

In 51 seconds and narrated: with learning rate 1.0, gradient boosting destabilized and after 3,000 rounds ended at AUC 0.681, below a decision tree with no depth limit; it adds small trees in a chain and each one fits what the previous ones missed; with rate 0.03, validation chose 594 rounds and it reached 0.883; against a 200-tree forest the difference was not conclusive, but on one thread it trained 9 times faster and weighs 2.1 MB against 124 MB; and with one-question trees it needed almost 20,000 rounds. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

What is gradient boosting?

Gradient boosting is a sum of small decision trees. The first one makes a rough prediction. The second uses the same variables, but it is not fitted to the original label; it is fitted to what the first one missed. The third, to what the first two missed, and so on for hundreds of rounds. Each tree adds its contribution multiplied by a learning rate.

A random forest does the opposite: it trains large, independent trees, each on a different sample of rows, and averages them. A forest’s trees can be trained at the same time; boosting’s trees go in a line, because each one depends on the previous ones.

A very influential general formulation is Jerome Friedman’s, in “Greedy Function Approximation: A Gradient Boosting Machine” (2001): each tree is fitted to the negative gradient of the loss function. If the loss is half the squared error, that gradient is exactly the residual, what is left to reach the true value. With the log loss of a rain or no-rain problem, it is the difference between what happened (1 or 0) and the probability the model gave. A predecessor is AdaBoost, by Freund and Schapire (1997), which instead of fitting residuals gave more weight to misclassified examples.

Each tree fits what is left

With a single variable the mechanism is visible. Today’s max temperature predicts tomorrow’s, with one-question trees, called stumps, and squared error.

Each stump fits what the previous rounds left overToday's max temperature predicts tomorrow's with one-question trees and squared error. Top: the prediction after each round, over 80 sample days. Bottom: those days' residuals, which the next stump is fitted to.
learning rate
-1001020304050035-200+20tomorrow's max (°C)residual (°C)today's max (°C)round 0 · train MSE 42.27round 1 · train MSE 17.48round 2 · train MSE 13.84round 3 · train MSE 11.51round 5 · train MSE 9.09round 10 · train MSE 7.32round 20 · train MSE 6.31round 50 · train MSE 5.67round 100 · train MSE 5.52round 100 · train MSE 5.52round 0 · train MSE 42.27round 1 · train MSE 29.63round 2 · train MSE 22.56round 3 · train MSE 17.45round 5 · train MSE 12.01round 10 · train MSE 7.33round 20 · train MSE 5.87round 50 · train MSE 5.70round 100 · train MSE 5.64round 100 · train MSE 5.64

The first stump splits at 17.3 °C: it takes 4.69 °C off the days below and adds 5.29 °C to the days above. With rate 0.3 the same stump adds only 0.3 of that (-1.41 and +1.59). Training MSE goes from 42.27 with no trees to 5.52 after 100 rounds at rate 1.0 and 5.64 at 0.3; on 2016-2026 they miss by 1.86 and 1.88 °C on average. With half the squared error as the loss, the residual is exactly the negative gradient; with the log loss of rain it is the label minus the probability, as the text explains.

With no trees, the prediction is the mean: 17.0 °C for every day. The first stump splits at 17.3 °C, takes 4.69 °C off the colder days and adds 5.29 °C to the warmer ones. The second still uses today’s temperature as input, but its target is now the residuals the first one left, and it looks for the cut that reduces them most. After 100 rounds at rate 1.0, training mean squared error went from 42.27 to 5.52.

At rate 0.3, each stump adds only 0.3 times its value, so the same first cut moves the prediction 1.41 °C down and 1.59 °C up. After 100 rounds it ended at 5.64. The path is slower, and on the rain task the low rates coincided with the best results.

The learning rate buys rounds

For rain I used HistGradientBoostingClassifier with trees of up to 31 leaves, its default, and 3,000 rounds with no automatic stopping. I recorded validation log loss round by round for four rates.

The learning rate buys roundsValidation loss (2010-2012) against rounds, log scale. Dashed: the same model on its training data. Dot: best round for each rate.
1101001,0003,0000.20.30.40.50.60.7rounds (log scale)1101001,0003,0000.650.700.750.800.850.90rounds (log scale)
  • rate 1.0 · best 7
  • rate 0.3 · best 30
  • rate 0.1 · best 188
  • rate 0.03 · best 594

rate 1.0 destabilizes and by round 1,000 is stuck at 3.25

With rate 1.0 the best round on validation is round 7; left running to 3000 rounds and scored on 2016-2026 it falls to AUC 0.681, below a decision tree with no depth limit (0.689 in the previous post). Rate 0.3 bottoms out at round 30, 0.1 at 188 and 0.03 at 594, with the lowest loss. That choice, refitted on 1984-2012, scores AUC 0.883 and log loss 0.332 on 2016-2026; the same rate with 3000 rounds, 0.8805 and 0.338.

With rate 1.0 validation loss bottomed out at round 7 and then the fit destabilized: it rose in jumps and by round 1,000 it was stuck at 3.25. Training loss did too: it fell to 0.33 at round 15 and ended at 3.16. That is not ordinary overfitting, where training keeps improving; the step is so large that the model stops converging. With the other three rates training loss kept falling up to round 3,000. Rate 0.3 bottomed out on validation at round 30, 0.1 at 188 and 0.03 at 594, with the lowest loss of the four (0.347). I chose that combination and refitted on all of 1984-2012: on 2016-2026 it scored AUC 0.883 and log loss 0.332.

Overshooting the rounds with a low rate cost little. With rate 0.03 and 3,000 rounds, five times the chosen number, AUC on 2016-2026 dropped to 0.8805 and loss rose to 0.338. With rate 1.0 and the same 3,000 rounds, to 0.681 and 3.15.

scikit-learn has automatic early stopping that switches itself on above 10,000 rows: it sets aside a random 10% of the training data and stops when 10 rounds in a row bring no improvement. With daily data, setting days aside at random leaves neighbours of training days in the validation set, so I expected it to run too long. It did not: with five seeds it stopped between 335 and 466 rounds, and the same rule applied to 2010-2012 stopped at 344. On 2016-2026, AUC was between 0.8817 and 0.8826 with automatic stopping and 0.8824 with the temporal one. On this data automatic stopping worked; with more autocorrelated series it is worth measuring before trusting it.

How big each tree should be

The rate is not the only thing that decides how many rounds are needed. The size of each tree matters too.

How big each tree in the chain should beRate 0.03. Rounds chosen on 2010-2012; scores on 2016-2026. Dots: maximum leaves per tree. Highlighted: 31 leaves, the default used in the rest of the post.
2 leaves
19,968 rounds
0.869AUC
4 leaves
5,571 rounds
0.882AUC
8 leaves
2,154 rounds
0.882AUC
15 leaves
1,312 rounds
0.883AUC
31 leaves
594 rounds
0.883AUC
63 leaves
427 rounds
0.883AUC

Stumps, trees with a single question, cannot combine two variables in one tree: their best validation round was 19,968, close to the 20,000 cap, and they stayed at AUC 0.869. Four leaves were enough to reach 0.882 with 5,571 rounds, and 63 leaves got 0.883 with 427. Bigger trees need fewer rounds; past four leaves the score barely moves.

A stump can only ask about one variable, so a boosting of stumps is a sum of one-variable effects, with no combinations. Its best validation round was 19,968, close to the 20,000 cap I set, so with more rounds it could have improved somewhat; it ended at AUC 0.869. With four-leaf trees, which can already combine two questions, it reached 0.882 with 5,571 rounds. From there, larger trees cut rounds while barely moving the result: 0.883 with 31 leaves and 594 rounds, 0.883 with 63 leaves and 427. I kept the default 31 leaves; on validation, 63 leaves came out just 0.0007 lower in loss.

Why the histogram version is fast

scikit-learn has two implementations. GradientBoostingClassifier looks for cuts among all the values of each variable and works on one thread. HistGradientBoostingClassifier, which its documentation describes as inspired by LightGBM, first splits each variable into bins and spreads the work across threads.

Boxes instead of valuesBefore the first round, each variable is split into boxes holding similar numbers of days; to look for a cut, the tree scans boxes, not values. Here: 32 real days of today's max falling into 8 boxes.
11733today's max (°C)

AUC 2016-2026 with 200 rounds of depth-3 trees

4
0.874
16
0.880
64
0.880
255
0.880

boxes per variable

With 255 boxes, the scikit-learn default, AUC was 0.8802; with 64, 0.8799; with 16, 0.8796; with only 4 boxes per variable, 0.8736. The loss from grouping values is small here, and it is what makes each round cheap.

Classic boosting against histogram boostingMedian of 3 fits (classic) and 5 (histograms) on 74,145 rows. Both with 200 rounds, depth 3, rate 0.1 and at least 20 days per leaf.
Histograms, 8 threads · AUC 0.8800.174 s
Histograms, 1 thread · AUC 0.8800.639 s
Classic, 50% of rows per round · AUC 0.87916.09 s
Classic · AUC 0.87929.72 s

At a quarter of real time.

The classic version is single-threaded by design; the histogram one was measured with its OpenMP threads limited. With this experiment's main configuration (up to 31 leaves and 200 rounds), the histogram version took 0.94 s on one thread and 0.34 s on eight.

On one thread, the histogram version was 46.5 times faster than the classic one, with a barely higher AUC (0.880 against 0.879). Using half the rows in each round, the stochastic variant Friedman proposed in 2002, took the classic version from 29.7 to 16.1 s without changing AUC at three decimals. On eight threads, the histogram version went from 0.64 to 0.17 s: the trees go in a line, but inside each tree the search for cuts is split by variable.

Wrong labels

If each tree is fitted to what the previous ones got wrong, a mislabeled day is an error the model will try to correct. I measured it by flipping 5, 10 and 20% of the training labels at random, with five seeds per level (and a single run with no noise). I chose the boosting rounds on an untouched 2010-2012 validation set and scored on 2016-2026.

Wrong labelsShare of training labels flipped at random (rain to dry and back), five seeds per noisy level and one run with no noise; AUC on 2016-2026. Dot: mean; bar: worst to best seed.
0 %5 %10 %20 %0.840.870.89labels flipped
  • boosting, rounds on validation · 20 %: 0.868
  • forest, 200 trees · 20 %: 0.867
  • boosting, 3,000 rounds · 20 %: 0.851

With 10% flipped, the forest averaged 0.876 and boosting 0.871, which ranged from 0.868 to 0.877 across seeds. With 20%, they were very close on average (0.868 and 0.867), but boosting went from 0.854 to 0.874: validation picked between 632 and 2,982 rounds. Left at 3,000 rounds, boosting was the worst at every noise level.

With 5% flipped, boosting and forest were very close on average: AUC 0.880 and 0.879. With 10%, the forest came out ahead, 0.876 against 0.871, and it also barely moved across seeds (0.875 to 0.876), while boosting went from 0.868 to 0.877. With 20% they were again very close on average (0.868 and 0.867), but boosting depended on the seed: from 0.854 to 0.874.

What changes across seeds is how many rounds validation picks. With 20% noise it picked between 632 and 2,982, and the 2,982 run was the worst. Left at 3,000 rounds with no choice at all, boosting was the worst of the three at every noise level: 0.851 at 20%. On log loss, though, boosting with chosen rounds stayed below the forest at all three levels (0.434 against 0.443 at 20%). With dirty labels, the forest’s AUC was more predictable across seeds; boosting’s depended on how many rounds validation picked.

Boosting, forest, tree and logistic regression

Tree, logistic regression, forest and boosting: training on one threadMedian of 10 fits (3 for the forest) on one thread, after a warm-up fit, on 74,145 rows.
Logistic regression · AUC 0.8470.0488 s
Tree, depth 7 · AUC 0.8620.3719 s
Boosting, 594 rounds · AUC 0.8832.6161 s
Random Forest, 200 trees · AUC 0.88224.0411 s

At a quarter of real time.

Serialized with pickle, the tree takes 21.6 KB, logistic regression 1.8 KB, boosting 2.1 MB and the forest 124 MB. That is size on disk, not memory in use.

On AUC, boosting and forest ended 0.0006 apart. To see whether that difference means anything, I resampled the 77 city-year blocks of 2016-2026 1,000 times and recomputed it: the 95% interval went from −0.0007 to +0.0021 and crosses zero, so the difference is not conclusive. On log loss boosting did come out better, by a little: −0.0025, with an interval from −0.0042 to −0.0008.

Where they parted was cost. On one thread, boosting trained in 2.6 s and the forest in 24.0 s. At prediction time the advantage partly flips: scoring one day took 3.0 ms for boosting and 3.1 ms for the forest, but the 27,256 test days in one batch took 0.58 s for boosting and 0.32 s for the forest. I did not measure where that difference comes from.

Both predicted more rain than fell: on average, 23.4% for boosting and 23.8% for the forest, against 20.3% observed. That overestimate coincided with a drop in rain frequency between periods: it rained on 26.6% of training days and 20.3% of test days, as I explained in the Random Forest post. Among the days boosting gave between 50 and 60%, it rained on 46%; in the same band for the forest, on 48%. If the probability is going to be used as is, it is worth recalibrating it on recent data, as the logistic regression post showed.

For numbers, the same ceiling as the forest

With tomorrow’s max temperature as the target and all variables, boosting with rate 0.1 and 933 rounds chosen on validation missed by 1.44 °C on average on 2016-2026. The 200-tree forest, by 1.50 °C, and linear regression, by 1.70 °C.

I also repeated the Random Forest post’s extrapolation test, here with all fifteen features, day of the year and latitude included: I trained on Santiago’s April to September months between 1984 and 2012 and predicted the 2016-2026 summers. The warmest training day reached 32.97 °C and the summers averaged 29.63 °C. Boosting predicted no more than 29.56 °C, averaged 24.2 °C and missed by 5.48 °C; the forest, by 4.30 °C; linear regression, by 1.72 °C. There the forest and the line read 3.76 and 1.86 °C because that controlled case drops the day of the year and latitude; the order between the three models does not change. Past the cuts it learned, each tree returns a constant value, so the sum does not follow the trend toward temperatures it never saw. It is the same ceiling I measured on the forest.

Built in: monotonic constraints and missing values

Built in: monotonic constraints and missing values

Mean probability of rain tomorrow when today's rain is set to each value on 500 real training days.

0510152010 %38 %65 %today's rain (mm)freeconstrained

Unconstrained, the curve went down in 4 of 40 steps; constrained, it can only rise or stay flat, and AUC on 2016-2026 was 0.883.

Humidity and pressure blanked at random in 20% of days, each on its own; then blanked on every test day.

clean0.883 boosting
20% missing0.882 boosting0.881 forest
all missing at test0.880 boosting0.879 forest

Boosting learns at each split which side days with a missing value go to, with no imputer: 0.883 clean, 0.882 with 20% missing, 0.880 with both variables missing on every test day. The forest, which also accepts missing values in this version of scikit-learn: 0.881 and 0.879. The blanks are synthetic and at random; real sensor gaps may not be.

Sometimes you know in advance that a relationship goes in only one direction. Unconstrained, the mean probability of rain tomorrow went down in 4 of the 40 steps as today’s rain rose from 0 to 20 mm; with monotonic_cst on today’s rain and humidity it can only rise or stay flat, and AUC on 2016-2026 was 0.883, the same as unconstrained.

I blanked humidity and pressure at random on 20% of days, each on its own. HistGradientBoostingClassifier learns at each split which branch to send days with a missing value to: AUC went from 0.883 to 0.882, and to 0.880 with both variables missing on every test day. scikit-learn’s forest, which also accepts missing values since version 1.4, scored 0.881 and 0.879. The gaps in this experiment are random; those of a sensor that fails in bad weather may not be, and there the result may change.

Where gradient boosting lives in a real system

Where gradient boosting livesThe number of rounds is decided on a separate stretch of time, so the split is part of the system.
Where gradient boosting livestable74,145 rowsfit | stop1984-2009 | 2010-12trainrate 0.03 · 594 roundsartifact2.1 MBbatch scoring27,256 days · 0.58 sonline API1 day · 3.03 msretrain + watchcalibration, rounds

Measured on one thread: boosting with 594 rounds trains in 2.62 s, its pickle takes 2.1 MB, it scores one day in 3.03 ms and the 27,256 test days in 0.58 s. The 200-tree forest: 24.0 s, 124.3 MB, 3.08 ms and 0.32 s. For reference, the depth-7 tree scores one day in 0.25 ms.

  • Scoring tables. Risk, demand, fraud, any table with dozens of columns. It is the first model I would try there: it trains in seconds, accepts missing values without imputation and can handle categorical columns when they are marked as such.
  • Batch scoring. A nightly job that scores millions of rows. It is worth timing the full batch: here 27,256 days took 0.58 s on one thread, almost twice the forest.
  • Online API. One row took 3.0 ms with scikit-learn, similar to the forest. If that is not enough, try fewer rounds with larger trees or exporting the model to a compiled format, and measure again.
  • Monitored retraining. Rounds are chosen on a separate stretch of time, and that stretch has to move as new data arrives. Calibration drifts too: here the overestimate of rain coincided with a drop in rain frequency between periods.

When I would choose it: for tabular data where accuracy and training cost matter, and when one variable must push in only one direction. When not: when you need to extrapolate beyond the range seen, when labels have many errors and there is no clean data to choose rounds, or when a readable rule is worth more than two to four hundredths of AUC, where a small decision tree or a logistic regression is enough. With a GPU and tables of up to 32,000 rows, a pretrained tabular model beat tuned XGBoost on all fourteen tables I tried in TabPFN against XGBoost.

Sources

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.