
Decision tree explained: 15 nodes you can read, 8,524 leaves that memorize
What a decision tree is, how it picks its questions and when it memorizes, measured on daily weather from seven Chilean cities between 1984 and 2026. With no depth limit it got everything it saw right and dropped to AUC 0.689 on new days; at depth 7, chosen on validation, it reached 0.862.
A decision tree with no depth limit got every day it learned from right: AUC 1.000 over 74,145 rows. On 2016-2026 days, which it had not seen, it dropped to 0.689, worse than a logistic regression. The same algorithm, stopped at the seventh question, reached 0.862.
This is the fifth part of the series, with the same NASA POWER daily data from seven Chilean cities I used for Random Forest, linear regression, K-Means and logistic regression. The question is the usual one: will it rain at least 1 mm tomorrow? I trained on 1984-2012, chose every parameter on a validation period and evaluated on 2016-2026, inside a container with 8 CPUs.
What is a decision tree?
A decision tree is a chain of questions. The first question splits all days into two groups; each group gets its own question, and so on until a leaf. The leaf returns the share of rainy days that ended there during training, which is used as the probability.
Each question compares a single variable with a threshold: “did it rain 0.44 mm or less today?”. To choose it, the algorithm tries each variable with the candidate cut points taken from its observed values and keeps the split that leaves the two purest groups, measured with the Gini index. That is the idea behind CART, the method Breiman, Friedman, Olshen and Stone published in 1984, and what scikit-learn implements.
swipe the tree sideways →
- rain today 3.15 > 0.44 → no
- rain today 3.15 ≤ 3.54 → yes
- wind N-S (cos) 0.19 > -0.15 → no
probability 53 % · rained · right
- rain today 0.46 > 0.44 → no
- rain today 0.46 ≤ 3.54 → yes
- wind N-S (cos) 0.35 > -0.15 → no
probability 53 % · rained · right
- rain today 0.05 ≤ 0.44 → yes
- latitude -36.83 > -37.79 → no
- rain today 0.05 > 0.04 → no
probability 12 % · dry · right
- rain today 0.92 > 0.44 → no
- rain today 0.92 ≤ 3.54 → yes
- wind N-S (cos) 0.30 > -0.15 → no
probability 53 % · dry · wrong
Each day goes through three questions: with this tree, 38% of 2016-2026 days end in the leaf of dry days from Concepción northward (rain in 3% of its training days). AUC in 2016-2026: 0.831. The fourth day shows its limit: 0.92 mm of rain and a north-south wind component above -0.15 send it to a leaf with 53% rain, and it did not rain.
At depth 3 the tree has 15 nodes and fits on a page. Its first question is today’s rain. If it rained little, the second one uses latitude to separate the cities from Temuco southward from those from Concepción northward; if it rained, it asks how much. The last questions look at rain again or at the north-south component of wind direction. With those three questions it reached AUC 0.831 in 2016-2026.
Reading it also shows its limit. The fourth day in the example, in Punta Arenas, had 0.92 mm and a wind that sent it to a leaf with 53% rain. It did not rain. Eight leaves are eight possible probabilities, and every day in a leaf gets the same one.
The deeper it asks, the more it memorizes
Nothing forces the tree to stop at three questions. If I let it go on, it splits until every leaf holds days of a single class.
leaves (log scale)
- days it learned from
- validation
- 2016-2026
depth 1: 2 leaves · AUC training 0.776 · validation 0.731 · 2016-2026 0.741 · log loss 0.423
depth 2: 4 leaves · AUC training 0.833 · validation 0.790 · 2016-2026 0.803 · log loss 0.399
depth 3: 8 leaves · AUC training 0.855 · validation 0.826 · 2016-2026 0.831 · log loss 0.380
depth 4: 16 leaves · AUC training 0.866 · validation 0.842 · 2016-2026 0.843 · log loss 0.371
depth 5: 32 leaves · AUC training 0.875 · validation 0.853 · 2016-2026 0.855 · log loss 0.361
depth 6: 64 leaves · AUC training 0.883 · validation 0.859 · 2016-2026 0.860 · log loss 0.359
depth 7: 127 leaves · AUC training 0.890 · validation 0.863 · 2016-2026 0.862 · log loss 0.362
depth 8: 245 leaves · AUC training 0.898 · validation 0.862 · 2016-2026 0.862 · log loss 0.370
depth 9: 461 leaves · AUC training 0.907 · validation 0.857 · 2016-2026 0.853 · log loss 0.412
depth 10: 804 leaves · AUC training 0.918 · validation 0.846 · 2016-2026 0.839 · log loss 0.518
depth 11: 1,283 leaves · AUC training 0.930 · validation 0.827 · 2016-2026 0.822 · log loss 0.659
depth 12: 1,879 leaves · AUC training 0.943 · validation 0.800 · 2016-2026 0.802 · log loss 0.841
depth 13: 2,568 leaves · AUC training 0.956 · validation 0.780 · 2016-2026 0.777 · log loss 1.067
depth 14: 3,367 leaves · AUC training 0.967 · validation 0.766 · 2016-2026 0.748 · log loss 1.349
depth 15: 4,131 leaves · AUC training 0.977 · validation 0.736 · 2016-2026 0.727 · log loss 1.577
depth 16: 4,903 leaves · AUC training 0.985 · validation 0.724 · 2016-2026 0.708 · log loss 1.834
depth 17: 5,590 leaves · AUC training 0.990 · validation 0.711 · 2016-2026 0.695 · log loss 2.066
depth 18: 6,233 leaves · AUC training 0.994 · validation 0.704 · 2016-2026 0.702 · log loss 2.297
depth 19: 6,808 leaves · AUC training 0.997 · validation 0.695 · 2016-2026 0.694 · log loss 2.517
depth 20: 7,280 leaves · AUC training 0.998 · validation 0.687 · 2016-2026 0.695 · log loss 2.654
depth no limit: 8,524 leaves · AUC training 1.000 · validation 0.700 · 2016-2026 0.689 · log loss 3.107
Validation peaked at depth 7 (127 leaves, AUC 0.863); in 2016-2026 that tree scored 0.862. Without a depth limit the tree grew 8,524 leaves, reached AUC 1.000 on the days it learned from and fell to 0.689 in 2016-2026, with a log loss of 3.107: pure leaves give probabilities of 0 or 1, and a wrong 0 costs a lot.
With each extra level, AUC on the training days rises. On validation it rises up to depth 7 and then falls. I chose the depth with that curve, without looking at 2016-2026: the depth-7 tree has 127 leaves and gave AUC 0.862 in 2016-2026.
With no limit, the tree reached depth 31 and 8,524 leaves, under nine training days per leaf on average. Many leaves are pure, so they return probabilities of 0 or 1. When they are wrong, they are wrong with full certainty, and it shows in the log loss: 3.11 against 0.36 for the depth-7 tree.
Capping depth is not the only brake. I tried two more, also chosen on validation.
- 10.700 · 8,524 leaves
- 50.767 · 5,047 leaves
- 200.839 · 1,889 leaves
- 500.857 · 892 leaves
- 1000.864 · 478 leaves
- 2000.863 · 256 leaves
- 5000.860 · 110 leaves
Chosen setting in 2016-2026: 478 leaves, AUC 0.863, log loss 0.384.
- 0.0e+00.767 · 4,519 leaves
- 4.7e-60.743 · 4,101 leaves
- 1.2e-50.764 · 3,554 leaves
- 1.7e-50.775 · 3,060 leaves
- 2.5e-50.796 · 2,513 leaves
- 3.3e-50.813 · 1,880 leaves
- 4.2e-50.834 · 1,201 leaves
- 5.9e-50.856 · 458 leaves
- 7.7e-50.863 · 205 leaves
Chosen setting in 2016-2026: 167 leaves, AUC 0.861, log loss 0.358.
- 20.790 · 4 leaves
- 40.842 · 16 leaves
- 60.859 · 64 leaves
- 70.863 · 127 leaves
- 80.862 · 245 leaves
- 100.846 · 804 leaves
- 140.766 · 3,367 leaves
- ∞0.700 · 8,524 leaves
Chosen setting in 2016-2026: 127 leaves, AUC 0.862, log loss 0.362.
The three brakes land on AUC between 0.861 and 0.863 in 2016-2026, with 127 to 478 leaves. What matters is stopping; the exact brake mattered little here.
Requiring at least 100 days per leaf left 478 leaves and AUC 0.863 in 2016-2026. Cost-complexity pruning, which cuts branches that add little (starting from a tree with at least 5 days per leaf), left 167 leaves and AUC 0.861. All three ended in the same place.
The purity criterion mattered little. At depth 8, Gini gave AUC 0.8615 and entropy 0.8655, with the same variable at the root.
The root holds; the predictions move
Trees are known as unstable: change the data a little and the whole tree changes. I measured it in two different ways, because a changing structure is not the same as changing answers.
first question in 50 of 50 resamples: rain today ≤ threshold (mm)
range across block resamples: 0.40 to 0.71 mm
days without agreement on rain / no rain
The structure at the top barely moves: rain today was always the first question. The predictions do move: with the depth chosen on validation, the 50 resampled trees disagreed on 34.8 % of 2016-2026 days, against 17.6 % at depth 3. Averaging that noise is what a random forest does.
I resampled the data 50 times by city-year blocks, so neighboring days would not be spread out as if they were independent. The first question was today’s rain in all 50 cases; the threshold moved between 0.40 and 0.71 mm. Training one tree per consecutive block of up to three years (the last one, 2011-2012, has two) gave the same: always today’s rain, with thresholds from 0.39 to 0.97 mm.
The predictions did move. The 50 depth-7 trees did not agree on the rain or no-rain label on 34.8% of 2016-2026 days; the depth-3 ones, on 17.6%. With a different measure, not comparable with those, two trees with no limit trained on random halves of the rows disagreed on 24.1% of days. A random forest exists precisely to average those differences out, as the Random Forest post showed.
Which variables matter? I permuted families of variables together on the validation period, so one correlated variable would not hide another. Shuffling today’s and yesterday’s rain lowered the depth-7 tree’s AUC by 0.133; season and latitude, by 0.059; wind, by 0.056; pressure, by 0.029. Humidity and dew point together barely moved it (0.002): this tree’s predictions hardly depended on them. That describes this tree, not what causes rain.
A regression tree predicts in steps
A tree can also predict numbers. Each leaf returns the mean of its days, so the prediction is a staircase.
Depth 2 has 4 steps and misses 2016-2026 by 2.50 °C on average; depth 4, 16 steps and 1.89 °C; depth 8, 252 steps and 1.85 °C. Past 39.5 °C, the warmest day it learned from, the prediction no longer rises; the Random Forest post measured the same ceiling on a forest.
To predict tomorrow’s max temperature from today’s, a depth-2 tree has 4 steps and misses by 2.50 °C on average in 2016-2026; at depth 4, 16 steps and 1.89 °C; at depth 8, 252 steps and 1.85 °C. Beyond the 39.5 °C of the warmest training day, the prediction is flat: a tree does not extrapolate. In the Random Forest post I measured the same ceiling on a forest.
How well it calibrates and what it costs
Each leaf returns a single probability, so the tree gives only 118 distinct values. On average it predicted 23.4 % rain and it rained on 20.3 % of days. Only the lowest band matches (13,554 days under 10%, rain on 3.5 %); from there up the predictions sit above the observed rate: on the days it gave 50 to 60%, it rained on 44 %.
The depth-7 tree gives 118 distinct probabilities. It predicted 23.4% rain on average and it rained on 20.3% of 2016-2026 days. Except in the lowest band, it predicted more rain than fell: on the days it gave 50 to 60%, it rained on 44%.
At half of real time.
Serialized with pickle, the tree takes 21.6 KB, logistic regression 1.8 KB and the forest 124 MB. That is size on disk, not memory in use.
The tree landed between logistic regression and the forest on almost everything: it ranked better than logistic regression and worse than the forest, trained 7 times slower than logistic regression and 67 times faster than the forest. Its advantage is different: it can be read, and it can be rewritten.
Where a decision tree lives in a real system
the depth-3 tree exported as rules
|--- precip <= 0.44
| |--- lat <= -37.79
| | |--- wind_cos <= -0.09
| | | |--- class: 0.0
| | |--- wind_cos > -0.09
| | | |--- class: 0.0
| |--- lat > -37.79
| | |--- precip <= 0.03
| | | |--- class: 0.0
| | |--- precip > 0.03
| | | |--- class: 0.0
|--- precip > 0.44
| |--- precip <= 3.53
| | |--- wind_cos <= -0.15
| | | |--- class: 0.0
| | |--- wind_cos > -0.15
| | | |--- class: 1.0
| |--- precip > 3.53
| | |--- wind_cos <= 0.76
| | | |--- class: 1.0
| | |--- wind_cos > 0.76
| | | |--- class: 1.0
Measured on one thread: the depth-7 tree trains in 0.37 s and its pickle takes 21.6 KB; logistic regression, 0.05 s and 1.8 KB; the 200-tree forest, 24.7 s and 124 MB. Rewritten as plain Python if/else, the depth-3 tree scores one day in 0.0042 ms, against 0.253 ms through scikit-learn, with the same result on the first 2,000 days.
- Exported rules. A small tree translates to if/else or a SQL
CASEand runs where there are no machine learning libraries. Rewritten in plain Python, the depth-3 tree scored one day in 0.0044 ms, against 0.25 ms through scikit-learn. - Explainable triage. When someone has to justify a decision, the path through the tree is the justification: three questions, three answers.
- Building block of other models. Random forests and boosting combine many trees; understanding one means understanding the piece they repeat.
- Exploration. A depth-2 or depth-3 tree quickly shows which variable separates best, before training something bigger.
When I would choose it: when the rule must be readable or exportable, or as a first look at the data. When I would not: when fine-grained probabilities or the stability of each prediction matter, where a 200-tree forest did better here, or when extrapolation is needed.
Sources
- Breiman, L., Friedman, J. H., Olshen, R. A. and Stone, C. J. (1984). Classification and Regression Trees. Wadsworth. Reissued by Routledge, 2017. DOI 10.1201/9781315139470.
- Quinlan, J. R. (1986). “Induction of decision trees”. Machine Learning, 1(1), 81–106. DOI 10.1007/BF00116251.
- scikit-learn, decision trees,
DecisionTreeClassifierand cost-complexity pruning. - NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.