17°
Portada del artículo: Decision tree explained: 15 nodes you can read, 8,524 leaves that memorize
Machine learningAlgorithmsPythonData

Decision tree explained: 15 nodes you can read, 8,524 leaves that memorize

What a decision tree is, how it picks its questions and when it memorizes, measured on daily weather from seven Chilean cities between 1984 and 2026. With no depth limit it got everything it saw right and dropped to AUC 0.689 on new days; at depth 7, chosen on validation, it reached 0.862.

Efrain Garay 16 September 2026

Playing summary

A decision tree with no depth limit got every day it learned from right: AUC 1.000 over 74,145 rows. On 2016-2026 days, which it had not seen, it dropped to 0.689, worse than a logistic regression. The same algorithm, stopped at the seventh question, reached 0.862.

This is the fifth part of the series, with the same NASA POWER daily data from seven Chilean cities I used for Random Forest, linear regression, K-Means and logistic regression. The question is the usual one: will it rain at least 1 mm tomorrow? I trained on 1984-2012, chose every parameter on a validation period and evaluated on 2016-2026, inside a container with 8 CPUs.

In 42 seconds and narrated: a decision tree with no limit reached AUC 1.000 on the days it learned from and dropped to 0.689 on new days; it asks one variable at a time; unchecked it grew more than 8,500 leaves, and at depth 7, chosen on validation, it did better; the first question barely changes, but the predictions disagree on a third of the days; and for numbers it predicts in steps, flat beyond what it saw. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

What is a decision tree?

A decision tree is a chain of questions. The first question splits all days into two groups; each group gets its own question, and so on until a leaf. The leaf returns the share of rainy days that ended there during training, which is used as the probability.

Each question compares a single variable with a threshold: “did it rain 0.44 mm or less today?”. To choose it, the algorithm tries each variable with the candidate cut points taken from its observed values and keeps the split that leaves the two purest groups, measured with the Gini index. That is the idea behind CART, the method Breiman, Friedman, Olshen and Stone published in 1984, and what scikit-learn implements.

A day falling down the treeDepth-3 tree trained on 1984-2012: 15 nodes, three questions per day. Pick one of four 2016-2026 days drawn at random; left branch means yes.
A day falling down the treeyesnoyesnoyesnoyesnoyesnoyesnoyesnorain today ≤ 0.4474,145 dayslatitude ≤ -37.7948,487 dayswind N-S (cos) ≤ -0.0914,042 days15 %9,778 days37 %4,264 daysrain today ≤ 0.0434,445 days3 %25,975 days12 %8,470 daysrain today ≤ 3.5425,658 dayswind N-S (cos) ≤ -0.1515,259 days35 %5,466 days53 %9,793 dayswind N-S (cos) ≤ 0.7610,399 days62 %4,228 days82 %6,171 days

swipe the tree sideways →

  1. rain today 3.15 > 0.44 → no
  2. rain today 3.15 ≤ 3.54 → yes
  3. wind N-S (cos) 0.19 > -0.15 → no

probability 53 % · rained · right

  1. rain today 0.46 > 0.44 → no
  2. rain today 0.46 ≤ 3.54 → yes
  3. wind N-S (cos) 0.35 > -0.15 → no

probability 53 % · rained · right

  1. rain today 0.05 ≤ 0.44 → yes
  2. latitude -36.83 > -37.79 → no
  3. rain today 0.05 > 0.04 → no

probability 12 % · dry · right

  1. rain today 0.92 > 0.44 → no
  2. rain today 0.92 ≤ 3.54 → yes
  3. wind N-S (cos) 0.30 > -0.15 → no

probability 53 % · dry · wrong

Each day goes through three questions: with this tree, 38% of 2016-2026 days end in the leaf of dry days from Concepción northward (rain in 3% of its training days). AUC in 2016-2026: 0.831. The fourth day shows its limit: 0.92 mm of rain and a north-south wind component above -0.15 send it to a leaf with 53% rain, and it did not rain.

At depth 3 the tree has 15 nodes and fits on a page. Its first question is today’s rain. If it rained little, the second one uses latitude to separate the cities from Temuco southward from those from Concepción northward; if it rained, it asks how much. The last questions look at rain again or at the north-south component of wind direction. With those three questions it reached AUC 0.831 in 2016-2026.

Reading it also shows its limit. The fourth day in the example, in Punta Arenas, had 0.92 mm and a wind that sent it to a leaf with 53% rain. It did not rain. Eight leaves are eight possible probabilities, and every day in a leaf gets the same one.

The deeper it asks, the more it memorizes

Nothing forces the tree to stop at three questions. If I let it go on, it splits until every leaf holds days of a single class.

The deeper it asks, the more it memorizesAUC by maximum depth. Training: 1984-2012. Validation: tree trained up to 2009, scored on 2010-2012. Test: 2016-2026, read only after choosing the depth on validation.
0.70.80.91.0chosen on validation14710131619∞

leaves (log scale)

  • days it learned from
  • validation
  • 2016-2026

depth 1: 2 leaves · AUC training 0.776 · validation 0.731 · 2016-2026 0.741 · log loss 0.423

depth 2: 4 leaves · AUC training 0.833 · validation 0.790 · 2016-2026 0.803 · log loss 0.399

depth 3: 8 leaves · AUC training 0.855 · validation 0.826 · 2016-2026 0.831 · log loss 0.380

depth 4: 16 leaves · AUC training 0.866 · validation 0.842 · 2016-2026 0.843 · log loss 0.371

depth 5: 32 leaves · AUC training 0.875 · validation 0.853 · 2016-2026 0.855 · log loss 0.361

depth 6: 64 leaves · AUC training 0.883 · validation 0.859 · 2016-2026 0.860 · log loss 0.359

depth 7: 127 leaves · AUC training 0.890 · validation 0.863 · 2016-2026 0.862 · log loss 0.362

depth 8: 245 leaves · AUC training 0.898 · validation 0.862 · 2016-2026 0.862 · log loss 0.370

depth 9: 461 leaves · AUC training 0.907 · validation 0.857 · 2016-2026 0.853 · log loss 0.412

depth 10: 804 leaves · AUC training 0.918 · validation 0.846 · 2016-2026 0.839 · log loss 0.518

depth 11: 1,283 leaves · AUC training 0.930 · validation 0.827 · 2016-2026 0.822 · log loss 0.659

depth 12: 1,879 leaves · AUC training 0.943 · validation 0.800 · 2016-2026 0.802 · log loss 0.841

depth 13: 2,568 leaves · AUC training 0.956 · validation 0.780 · 2016-2026 0.777 · log loss 1.067

depth 14: 3,367 leaves · AUC training 0.967 · validation 0.766 · 2016-2026 0.748 · log loss 1.349

depth 15: 4,131 leaves · AUC training 0.977 · validation 0.736 · 2016-2026 0.727 · log loss 1.577

depth 16: 4,903 leaves · AUC training 0.985 · validation 0.724 · 2016-2026 0.708 · log loss 1.834

depth 17: 5,590 leaves · AUC training 0.990 · validation 0.711 · 2016-2026 0.695 · log loss 2.066

depth 18: 6,233 leaves · AUC training 0.994 · validation 0.704 · 2016-2026 0.702 · log loss 2.297

depth 19: 6,808 leaves · AUC training 0.997 · validation 0.695 · 2016-2026 0.694 · log loss 2.517

depth 20: 7,280 leaves · AUC training 0.998 · validation 0.687 · 2016-2026 0.695 · log loss 2.654

depth no limit: 8,524 leaves · AUC training 1.000 · validation 0.700 · 2016-2026 0.689 · log loss 3.107

Validation peaked at depth 7 (127 leaves, AUC 0.863); in 2016-2026 that tree scored 0.862. Without a depth limit the tree grew 8,524 leaves, reached AUC 1.000 on the days it learned from and fell to 0.689 in 2016-2026, with a log loss of 3.107: pure leaves give probabilities of 0 or 1, and a wrong 0 costs a lot.

With each extra level, AUC on the training days rises. On validation it rises up to depth 7 and then falls. I chose the depth with that curve, without looking at 2016-2026: the depth-7 tree has 127 leaves and gave AUC 0.862 in 2016-2026.

With no limit, the tree reached depth 31 and 8,524 leaves, under nine training days per leaf on average. Many leaves are pure, so they return probabilities of 0 or 1. When they are wrong, they are wrong with full certainty, and it shows in the log loss: 3.11 against 0.36 for the depth-7 tree.

Capping depth is not the only brake. I tried two more, also chosen on validation.

Three ways to stop a tree from memorizingEach setting chosen on validation (trained up to 2009, scored on 2010-2012). Bar: validation AUC; the marked row is the one chosen.
  • 10.700 · 8,524 leaves
  • 50.767 · 5,047 leaves
  • 200.839 · 1,889 leaves
  • 500.857 · 892 leaves
  • 1000.864 · 478 leaves
  • 2000.863 · 256 leaves
  • 5000.860 · 110 leaves

Chosen setting in 2016-2026: 478 leaves, AUC 0.863, log loss 0.384.

  • 0.0e+00.767 · 4,519 leaves
  • 4.7e-60.743 · 4,101 leaves
  • 1.2e-50.764 · 3,554 leaves
  • 1.7e-50.775 · 3,060 leaves
  • 2.5e-50.796 · 2,513 leaves
  • 3.3e-50.813 · 1,880 leaves
  • 4.2e-50.834 · 1,201 leaves
  • 5.9e-50.856 · 458 leaves
  • 7.7e-50.863 · 205 leaves

Chosen setting in 2016-2026: 167 leaves, AUC 0.861, log loss 0.358.

  • 20.790 · 4 leaves
  • 40.842 · 16 leaves
  • 60.859 · 64 leaves
  • 70.863 · 127 leaves
  • 80.862 · 245 leaves
  • 100.846 · 804 leaves
  • 140.766 · 3,367 leaves
  • ∞0.700 · 8,524 leaves

Chosen setting in 2016-2026: 127 leaves, AUC 0.862, log loss 0.362.

The three brakes land on AUC between 0.861 and 0.863 in 2016-2026, with 127 to 478 leaves. What matters is stopping; the exact brake mattered little here.

Requiring at least 100 days per leaf left 478 leaves and AUC 0.863 in 2016-2026. Cost-complexity pruning, which cuts branches that add little (starting from a tree with at least 5 days per leaf), left 167 leaves and AUC 0.861. All three ended in the same place.

The purity criterion mattered little. At depth 8, Gini gave AUC 0.8615 and entropy 0.8655, with the same variable at the root.

The root holds; the predictions move

Trees are known as unstable: change the data a little and the whole tree changes. I measured it in two different ways, because a changing structure is not the same as changing answers.

The root holds; the predictions moveTop: first question of the depth-3 tree, 50 city-year block resamples and each consecutive block of up to three years. Bottom: share of 2016-2026 days where the trees did not agree on the label.

first question in 50 of 50 resamples: rain today ≤ threshold (mm)

0.00.40.81.2

range across block resamples: 0.40 to 0.71 mm

days without agreement on rain / no rain

  • depth 3, block resampling17.6 %
  • depth 7, block resampling34.8 %
  • depth 3, two random halves5.8 %
  • no limit, two random halves24.1 %

The structure at the top barely moves: rain today was always the first question. The predictions do move: with the depth chosen on validation, the 50 resampled trees disagreed on 34.8 % of 2016-2026 days, against 17.6 % at depth 3. Averaging that noise is what a random forest does.

I resampled the data 50 times by city-year blocks, so neighboring days would not be spread out as if they were independent. The first question was today’s rain in all 50 cases; the threshold moved between 0.40 and 0.71 mm. Training one tree per consecutive block of up to three years (the last one, 2011-2012, has two) gave the same: always today’s rain, with thresholds from 0.39 to 0.97 mm.

The predictions did move. The 50 depth-7 trees did not agree on the rain or no-rain label on 34.8% of 2016-2026 days; the depth-3 ones, on 17.6%. With a different measure, not comparable with those, two trees with no limit trained on random halves of the rows disagreed on 24.1% of days. A random forest exists precisely to average those differences out, as the Random Forest post showed.

Which variables matter? I permuted families of variables together on the validation period, so one correlated variable would not hide another. Shuffling today’s and yesterday’s rain lowered the depth-7 tree’s AUC by 0.133; season and latitude, by 0.059; wind, by 0.056; pressure, by 0.029. Humidity and dew point together barely moved it (0.002): this tree’s predictions hardly depended on them. That describes this tree, not what causes rain.

A regression tree predicts in steps

A tree can also predict numbers. Each leaf returns the mean of its days, so the prediction is a staircase.

A regression tree predicts in stepsToday's max temperature predicts tomorrow's with one variable. Each leaf is a flat step; the loop goes through depths 2, 4 and 8. Shaded: temperatures the tree never saw in 1984-2012.
-15-5515253545-51535depth 2 · 4depth 4 · 16depth 8 · 252today's max (°C)outside training range

Depth 2 has 4 steps and misses 2016-2026 by 2.50 °C on average; depth 4, 16 steps and 1.89 °C; depth 8, 252 steps and 1.85 °C. Past 39.5 °C, the warmest day it learned from, the prediction no longer rises; the Random Forest post measured the same ceiling on a forest.

To predict tomorrow’s max temperature from today’s, a depth-2 tree has 4 steps and misses by 2.50 °C on average in 2016-2026; at depth 4, 16 steps and 1.89 °C; at depth 8, 252 steps and 1.85 °C. Beyond the 39.5 °C of the warmest training day, the prediction is flat: a tree does not extrapolate. In the Random Forest post I measured the same ceiling on a forest.

How well it calibrates and what it costs

Leaves that round probabilitiesDepth-7 tree in 2016-2026: ten fixed bands of predicted probability, observed rain share with its 95% interval. Dot size grows with the number of days.
0.000.000.250.250.500.500.750.751.001.00predicted probability

Each leaf returns a single probability, so the tree gives only 118 distinct values. On average it predicted 23.4 % rain and it rained on 20.3 % of days. Only the lowest band matches (13,554 days under 10%, rain on 3.5 %); from there up the predictions sit above the observed rate: on the days it gave 50 to 60%, it rained on 44 %.

The depth-7 tree gives 118 distinct probabilities. It predicted 23.4% rain on average and it rained on 20.3% of 2016-2026 days. Except in the lowest band, it predicted more rain than fell: on the days it gave 50 to 60%, it rained on 44%.

Tree, logistic regression and forest: training on one threadMedian of 10 fits (3 for the forest) on one thread, after a warm-up fit, on 74,145 rows.
Logistic regression · AUC 0.8470.0519 s
Tree, depth 7 · AUC 0.8620.3712 s
Random Forest, 200 trees · AUC 0.88224.7329 s

At half of real time.

Serialized with pickle, the tree takes 21.6 KB, logistic regression 1.8 KB and the forest 124 MB. That is size on disk, not memory in use.

The tree landed between logistic regression and the forest on almost everything: it ranked better than logistic regression and worse than the forest, trained 7 times slower than logistic regression and 67 times faster than the forest. Its advantage is different: it can be read, and it can be rewritten.

Where a decision tree lives in a real system

Where a decision tree livesTrained offline, stored as a small artifact, and used in three places.
Where a decision tree livestable74,145 rowstraindepth 7artifact21.6 KBrules / SQL CASEif … then …explainable triagepath = reasonforests, boostingmany trees

the depth-3 tree exported as rules

|--- precip <= 0.44
|   |--- lat <= -37.79
|   |   |--- wind_cos <= -0.09
|   |   |   |--- class: 0.0
|   |   |--- wind_cos >  -0.09
|   |   |   |--- class: 0.0
|   |--- lat >  -37.79
|   |   |--- precip <= 0.03
|   |   |   |--- class: 0.0
|   |   |--- precip >  0.03
|   |   |   |--- class: 0.0
|--- precip >  0.44
|   |--- precip <= 3.53
|   |   |--- wind_cos <= -0.15
|   |   |   |--- class: 0.0
|   |   |--- wind_cos >  -0.15
|   |   |   |--- class: 1.0
|   |--- precip >  3.53
|   |   |--- wind_cos <= 0.76
|   |   |   |--- class: 1.0
|   |   |--- wind_cos >  0.76
|   |   |   |--- class: 1.0

Measured on one thread: the depth-7 tree trains in 0.37 s and its pickle takes 21.6 KB; logistic regression, 0.05 s and 1.8 KB; the 200-tree forest, 24.7 s and 124 MB. Rewritten as plain Python if/else, the depth-3 tree scores one day in 0.0042 ms, against 0.253 ms through scikit-learn, with the same result on the first 2,000 days.

  • Exported rules. A small tree translates to if/else or a SQL CASE and runs where there are no machine learning libraries. Rewritten in plain Python, the depth-3 tree scored one day in 0.0044 ms, against 0.25 ms through scikit-learn.
  • Explainable triage. When someone has to justify a decision, the path through the tree is the justification: three questions, three answers.
  • Building block of other models. Random forests and boosting combine many trees; understanding one means understanding the piece they repeat.
  • Exploration. A depth-2 or depth-3 tree quickly shows which variable separates best, before training something bigger.

When I would choose it: when the rule must be readable or exportable, or as a first look at the data. When I would not: when fine-grained probabilities or the stability of each prediction matter, where a 200-tree forest did better here, or when extrapolation is needed.

Sources

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.