17°
Portada del artículo: Random Forest explained: how a forest of trees votes, measured on 42 years of Chilean weather
Machine learningAlgorithmsPythonData

Random Forest explained: how a forest of trees votes, measured on 42 years of Chilean weather

What Random Forest is, how bootstrap sampling and feature subsampling work, and why its trees train in parallel instead of in a chain. I measured it predicting rain in seven Chilean cities: 7.5 times faster on 8 CPUs, a plateau past 100 trees, a noise column that the default importance ranks above real features, and a ceiling when extrapolating temperatures.

Efrain Garay 16 September 2026

Playing summary

A widely shared carousel of machine learning cheat sheets showed up on my Instagram feed recently. Its Random Forest card draws the trees joined by arrows: tree 1, tree 2, tree T, one after another. That drawing describes a different algorithm. A random forest trains its trees without any of them seeing the others, and that independence is exactly what defines it.

Instead of explaining it with another drawing, I measured it. I trained forests to predict whether it rains tomorrow in seven Chilean cities, from La Serena to Punta Arenas, with daily data since 1984, inside a container with CPU and memory limits. Everything below comes from those runs.

What is Random Forest?

Random Forest is a supervised learning algorithm that trains many decision trees, each on a different sample of the data and a different random draw of features, and combines their answers. It works for classification (will it rain or not?) and for predicting numbers (how hot?). Leo Breiman formalized it in 2001; before that, Tin Kam Ho had already tried diversifying trees by training them on random subsets of features.

A single tree memorizes: it follows every quirk of the data until its leaves hold a handful of examples. A hundred trees that make different mistakes, averaged, make fewer mistakes. To make them err differently, you force them to look at different data.

In 64 seconds and narrated: the viral Random Forest diagram draws the trees in a chain, and that is boosting. On eight cores, 200 trees drop from 26.1 to 3.5 seconds; AUC climbs from 0.69 with one tree to 0.88 with a hundred; a column of pure noise looks important with the default measure; and trained on cold months from 1984 to 2012, the forest never predicted above 28.9 °C in the summers of 2016 to 2026. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

How does a random forest work?

There are three steps, and two of them are random draws:

  1. Bootstrap. Each tree gets a sample the same size as the original data, drawn with replacement. Some rows repeat and 36.8% never show up (I checked by drawing 36 thousand rows twenty times); those left-out rows give the out-of-bag error, a validation you get for free.
  2. Feature subsampling. At each split, the tree only looks at a random subset of features. In scikit-learn’s classifier the default is the square root of the total; in the regressor, all of them. That way no tree can lean on the same dominant feature every time.
  3. Combination. For classification the trees’ answers are pooled; for regression they are averaged.

Since no step depends on another tree’s result, the trees can be trained at the same time. Boosting is the opposite: each new tree is trained on the previous one’s errors, so the chain is mandatory.

Random Forest in parallel, boosting in a chainTop: every tree gets its own bootstrap sample and never sees the others. Bottom: every tree learns from the previous one’s mistakes.
Random Forest in parallel, boosting in a chainRandom Forest · baggingGradient Boosting · sequentialtrainingdatasample 1tree 1sample 2tree 2sample 3tree 3sample 4tree 4average of probabilitiestree 1tree 2tree 3tree 4each tree receives the previous one’s errorsum of correctionsRandom Forest in parallel, boosting in a chaintraining dataRandom Forest · baggingtree 1tree 2tree 3tree 4average of probabilitiesGradient Boosting · sequentialtree 1tree 2tree 3tree 4sum of corrections
  • forest, 200 trees: 26.1 s on 1 CPU → 3.50 s on 8 (7.5×)
  • histogram boosting, 200 rounds: 0.94 s → 0.34 s (2.8×), rounds run in a row and only some work inside each one is spread across cores

The difference is measurable. With 200 trees and 74 thousand rows, the forest took 26.1 seconds on one CPU and 3.50 seconds on eight: 7.5 times faster, 93% of what perfect scaling would give. scikit-learn’s histogram boosting, with 200 rounds, went from 0.939 to 0.341 seconds: 2.8 times. Its rounds run in a row, and only some operations inside each round are spread across cores.

Training the same thing on 1, 2, 4 and 8 CPUsEach lane fills in the median of three training runs, measured in the container, on 74,145 rows.
1 CPU26.13 s
2 CPUs13.19 s
4 CPUs6.76 s
8 CPUs3.5 s

At a quarter of real time. Every time the cores double, the time almost halves.

1 CPU0.939 s
2 CPUs0.602 s
4 CPUs0.419 s
8 CPUs0.341 s

Five times slower than real time. Doubling the cores pays off less each time.

On 8 CPUs the forest reaches 93% of perfect scaling; boosting, 34%. Even so, boosting on one CPU finishes before the forest on eight.

Don’t read that as “the forest is faster”. On a single CPU the histogram boosting was 28 times faster than the forest, and slightly more accurate (AUC 0.882 versus 0.880). What scales with cores is the forest; what is already fast to begin with is modern boosting. The classic GradientBoostingClassifier, without histograms and without parallelism, took 29.2 seconds.

The data: 42 years of daily weather in seven cities

I downloaded the daily data for La Serena, Valparaíso, Santiago, Concepción, Temuco, Puerto Montt and Punta Arenas from NASA POWER: maximum and minimum temperature, precipitation, wind (speed and direction), radiation, humidity, pressure and dew point. Each row is one day in one city, and the question is whether at least 1 mm will fall the next day.

I trained on 1984 to 2012 (74,145 days) and tested on 2016 to 2026 (27,256 days). The three years in between are left out on purpose: one day’s weather looks a lot like the next, and without that gap the model would be tested on days it almost saw. I use those same years for every comparison in the post, so the figures describe that period; they are not a final test kept untouched.

Training, a three-year gap, then the test yearsShare of days followed by at least 1 mm of rain.
Training, a three-year gap, then the test years · by yeartraining 1984–2012gaptest 2016–20260 %10 %20 %30 %198519952005201520252026 up to August
  • seven cities
  • Santiago
Training, a three-year gap, then the test years · by city0 %20 %40 %60 %La Serena2.8 → 2.2 %Valparaíso9.9 → 7.4 %Santiago13.6 → 9.7 %Concepción28.0 → 20.5 %Temuco35.6 → 26.1 %Puerto Montt52.9 → 36.5 %Punta Arenas43.3 → 39.2 %
  • training 1984–2012
  • test 2016–2026

Training 26.6%, test 20.3%. The drop shows in all seven cities; it may come from the climate, from the period mix or from the data source changes, and this chart doesn't separate them.

In training it rains on 26.6% of days; in testing, 20.3%. The drop shows up in all seven cities, so it isn’t an error in preparing one of them (Santiago goes from 13.6% to 9.7% of days, Puerto Montt from 52.9% to 36.5%). In Santiago the decline shows mostly from 2009, consistent with the megadrought affecting central Chile since 2010, although the source changes I explain below may also play a part; I didn’t separate one from the other.

The stumble before measuring

I started with Open-Meteo and the first table already looked wrong: Arica, the driest city in Chile, showed frequent rain, and the rain rate collapsed between the training and test years. Requesting single years revealed why. Open-Meteo’s default option, best_match, blends several models depending on the date: Arica 2005 came out with 169 mm and Arica 2022 with 19 mm. The model would have learned the change of source, not the weather. Pinning ERA5 fixed that, but ERA5 gave Santiago 75 rainy days in 2005 and 50 in 2022, and the hourly quota limit finished pushing me out. NASA POWER delivers each city’s series in seconds and, for Santiago, 58 rainy days in 2005 and 25 in 2022. It isn’t a single source either: it uses MERRA-2 for the historical record and GEOS-IT for the most recent months, and radiation goes through several satellite products. What I needed was for rain not to change scale halfway through the series, as it did for Arica on Open-Meteo, and I didn’t see that jump in Santiago’s yearly rate. Arica was left out: 50 rainy days in 2005 aren’t credible for the coastal desert.

A caveat that applies to the whole post: this predicts the precipitation of a corrected reanalysis (MERRA-2, with GEOS-IT in the most recent stretch), not what a rain gauge records, the same kind of data I use to compare forecasts in my weather scoreboard. For learning how a forest behaves it works just the same; for deciding whether to take an umbrella, it doesn’t.

How many trees do you need?

Before the forest, two dumb baselines. Always saying “no rain” is right on 79.7% of test days, so accuracy alone is misleading. Saying “it rains tomorrow if it rained today” is right 81.2% of the time with a rain F1 of 0.535. Any model has to beat both.

Adding trees: fast gains, then a plateauAverage of 5 seeds up to 50 trees and 2 seeds from 100; the band spans the worst and best seed. Log scale on the X axis.
Adding trees: fast gains, then a plateau · AUC0.630.700.760.830.90125102550100200400treesone decision tree

AUC: one tree 0.688 · forest of 1 0.685 · forest of 400 0.881.

Adding trees: fast gains, then a plateau · F1 on rain0.380.440.500.560.61125102550100200400treesone decision tree

F1 on rain: one tree 0.494 · forest of 1 0.490 · forest of 400 0.596.

A single decision tree reaches an AUC of 0.688. A forest of just one tree does the same or slightly worse, 0.685 (the seed bands overlap): that tree sees a bootstrap sample and fewer features at each split, so there’s no reason to expect it to win. With 10 trees AUC rises to 0.849; with 100, to 0.879; with 400, to 0.881. Almost all the gain happens before 50 trees, and between 200 and 400 the gap is under a thousandth. Above 25 trees the AUC band nearly disappears (I ran five seeds up to 50 trees and two from 100); in F1 it still moves by a point.

There’s an odd detail in the F1 tab: with 2 trees F1 drops to 0.429, below the single tree (0.494), even though accuracy rises from 0.774 to 0.815. With full-depth trees each one answers 0 or 1, and a 1–1 tie resolves as dry, so a two-tree forest only predicts rain when both agree: it gets more dry days right and misses more rainy ones.

Against the baselines, the 400-tree forest is right 84.3% of the time with a rain F1 of 0.596. It beats “it rains tomorrow if it rained today”, but by 3 points of accuracy, not 30. Predicting rain from single-point variables is hard, and an algorithm doesn’t invent information the data doesn’t carry.

The out-of-bag validation needs care on a time series. With 200 trees it gave an AUC of 0.894, versus 0.880 on the test years: more optimistic. Accuracy seems to say the opposite (0.834 versus 0.840), but it can’t be compared, because it rains more in the training period and always saying “dry” is right 73.4% of the time there versus 79.7% in testing. One explanation is that neighboring days look alike, so each out-of-sample row usually has a twin inside the sample; another is that the test years are simply different. This experiment doesn’t separate the two causes. I wouldn’t use it as the only validation on data that has dates.

How does the forest vote?

The classic explanation says each tree votes and the majority wins. Breiman described it that way, but scikit-learn doesn’t do it: its user guide explains that it averages the probability each tree gives, and each tree returns the share of rainy days in the leaf the example lands in. I measured how much that matters, and the answer depends on depth. With trees grown to the end, the default, every leaf is pure and every tree answers 0 or 1: counting votes and averaging give exactly the same result on all 27,256 test days. With trees limited to depth 12, only 16% of each tree’s answers are 0 or 1, and the two rules disagree on 348 days, 1.3%.

One hundred trees vote on tomorrowEvery tree votes at the same time: none of them waits for another.
Santiago · April 15, 2016next day: it rained
94 of 100 trees vote rainaverage of the probabilities: 0.94
Santiago · July 17, 2016next day: it did not rain
50 of 100 trees vote rainaverage of the probabilities: 0.50
Santiago · January 19, 2016next day: it did not rain
0 of 100 trees vote rainaverage of the probabilities: 0.00

scikit-learn does not count votes: it averages each tree’s probability. With full-depth trees every tree answers 0 or 1, so over the 27,256 test days both rules always agree (on 203 days the forest split exactly 50/50). With trees limited to depth 12 they disagree on 348.

The share of trees voting rain behaves like a reasonable probability. On days where fewer than 10% of trees voted rain, it rained 2% of the time; where 90% or more voted rain, it rained 90% of the time. In between it rises in order, although in every bin it rained somewhat less than the share suggests (with 60 to 70% of votes it rained 59% of the time), as expected if the test years are drier than the training ones. That shows the share ranks days well from most to least likely; it doesn’t prove calibration, but it tells you when the forest is unsure.

How often it rained, by share of trees voting rainAll 27,256 test days. Top: observed rain rate per bin against the diagonal of a perfectly calibrated share. Bottom: days in each bin.
How often it rained, by share of trees voting rain0 %25 %50 %75 %100 %calibratedit rained 2 %it rained 90.3 %11,6940–103,73110–202,66720–301,90830–402,09240–501,93850–601,25160–7087270–8073180–9037290–100daysshare of trees voting rain (%)

With under 10% of votes it rained 2% of the time; with 90% or more, 90.3%. In all ten bins it rained less than the middle of the bin. Most days (11,694) sit in the lowest bin: the forest is rarely sure it will rain.

Where each model draws the lineBackground: probability of rain tomorrow. Dots: real test days, filled if it rained.
Where each model draws the line1939587797-1.2-0.60.00.71.3mean relative humidity (%)pressure change vs. yesterday (kPa)
  • < 10 %
  • 10–25 %
  • 25–50 %
  • 50–75 %
  • ≥ 75 %

AUC on the test years with just these two variables: one tree 0.664, 10 trees 0.726, 200 trees 0.734. The single tree draws hard staircases; the forest smooths them by averaging.

Which feature matters? The impurity trap

Every scikit-learn forest ships with feature_importances_, and it’s the first thing people plot. That measure adds up how much each feature reduces impurity at the trees’ splits, computed on the training data. To test it, I added a column of random numbers with no relation at all to rain.

The same 16 features, ranked two waysSwitch the measure and watch each row move to its place in the other ranking.
  1. max temperature0.0580.0086 ± 0.0005
  2. min temperature0.0490.0043 ± 0.0003
  3. rain today0.1830.0983 ± 0.0015
  4. rain yesterday0.0570.0042 ± 0.0004
  5. wind speed0.0500.0066 ± 0.0005
  6. wind east-west0.0480.0079 ± 0.0003
  7. wind north-south0.1250.0396 ± 0.0010
  8. radiation0.0540.0051 ± 0.0004
  9. humidity0.0560.0044 ± 0.0004
  10. pressure0.0510.0065 ± 0.0003
  11. pressure change0.0600.0191 ± 0.0008
  12. dew point0.0420.0030 ± 0.0002
  13. day of year (sine)0.0380.0013 ± 0.0003
  14. day of year (cosine)0.0380.0029 ± 0.0004
  15. latitude0.0490.0302 ± 0.0011
  16. random noise0.041-0.0002 ± 0.0001

By impurity, 13 of the 16 features sit between 0.038 and 0.058 and the random noise ranks 14th. By permutation only 4 features drop AUC by more than 0.01, and the noise is last.

By impurity, the noise ranked above both day-of-year features, with an importance of 0.041. Permutation importance, which shuffles one column in the test years and measures how much worse the model gets, puts it last and at zero. The scikit-learn guide warns about both causes: impurity is computed on the training data and favors features with many distinct values, and a continuous random column has as many distinct values as rows.

With all 16 features in view something else shows up: by impurity, 13 of the 16 sit between 0.038 and 0.058, so they all seem to matter a bit; by permutation, only 4 drop AUC by more than 0.01 (today’s rain, the north-south wind, latitude and pressure change). Impurity doesn’t just inflate the noise: it spreads importance almost evenly across everything.

Permutation isn’t a causal ranking either. It measures how much this model depends on each column in this data, and when two features are alike (today’s and yesterday’s rain, pressure and its change) shuffling one leaves the other to cover the gap, so both can come out lower than they weigh together.

When not to use Random Forest

When you have to predict beyond what it saw

A regression tree answers with the average of the examples in a leaf, so a forest can never predict a value higher than the highest it saw in training. To see it without ambiguity I set up a controlled case: a forest and a linear regression learn tomorrow’s maximum temperature in Santiago only from April to September between 1984 and 2012, and then predict the summers of 2016 to 2026. That is 5,307 days to fit and 962 to score, with no day in common: on top of the change of season comes the jump in years. Both models learn without the day of the year and without latitude, because in summer those features fall outside the range they saw.

The forest stays inside what it saw, the line keeps goingTrained on April–September of 1984-2012, asked about December–February of 2016-2026. Santiago, 300 summer days drawn at random.
The forest stays inside what it saw, the line keeps going151520202525303035354040hottest day seen in training · 33.0 °Cactual max temperature (°C)predicted (°C)

perfect prediction

Hottest training day: 33.0 °C. The summers averaged 29.63 °C; the forest predicted 26.01 °C on average and never above 28.9 °C, the line 29.01 °C. Mean absolute error: forest 3.76 °C, line 1.86 °C.

Realistic case, every season: hottest day each city saw in training (red) and the highest the forest predicted from 2016 (green)
  • La Serenasaw 31.2 · max predicted 30.25 °C · 0 test days above
  • Valparaísosaw 23.1 · max predicted 22.37 °C · 0 test days above
  • Santiagosaw 38.4 · max predicted 34.43 °C · 0 test days above
  • Concepciónsaw 34.9 · max predicted 31.09 °C · 2 test days above
  • Temucosaw 39.5 · max predicted 34.28 °C · 5 test days above
  • Puerto Monttsaw 36.0 · max predicted 27.96 °C · 0 test days above
  • Punta Arenassaw 22.2 · max predicted 19.91 °C · 1 test day above

The forest predicted 26.01 °C on average for summers that averaged 29.63 °C, and its highest value was 28.9 °C. It had seen days of up to 33.0 °C, but averaging leaves kept it well below that maximum: the hot days got flattened. The line kept rising with the day’s variables and missed by a little less than half as much, 1.86 °C of mean absolute error versus 3.76 °C for the forest.

In the realistic case, with every month and the test years from 2016, the forest won (1.50 °C versus 1.70 °C), and almost no test day exceeded its city’s training maximum: 5 in Temuco, 2 in Concepción, 1 in Punta Arenas. On those days the forest’s mean error was 4.7 to 8.6 °C depending on the city, and the line’s was even larger in all three. The line’s advantage shows when almost everything you predict falls outside the range, as in the controlled case; with a few isolated records, both models fail.

Corrected on 16 September 2026. The first version of this controlled case merged the training and test years before splitting the seasons. The fit ended up with winters from 2016-2026 inside it, and the “summer no model had seen” included summers from 1984-2012: the figures published then, 3.43 °C for the forest and 1.81 °C for the line, came from a contaminated experiment. With the periods kept apart the forest’s ceiling is still there and the gap between the two models is slightly wider.

When the model has to travel

A forest without a depth limit stores every node, and with a lot of data it grows fast. I measured the serialized size with pickle and prediction latency on a single thread:

ForestSizeNodesOne row (p50)27,256 rows at onceAUC
10 trees, no limit12.3 MB153,5700.48 ms0.019 s0.848
100 trees, no limit122.9 MB1,535,6241.73 ms0.182 s0.878
100 trees, depth 1226.0 MB325,0841.74 ms0.103 s0.880
400 trees, depth 12104.1 MB1,299,3105.73 ms0.408 s0.882

Two readings. Limiting depth to 12 made the 100-tree forest 4.7 times lighter without losing any AUC, so the default setting, which lets trees grow to the end, is expensive to ship and to hold in memory. And predicting row by row costs far more than in batch: 1.74 ms for one row versus 0.103 s for 27 thousand, about 3.8 microseconds per row. The cost of a single call grows with the number of trees, from 0.48 ms with 10 to 5.73 ms with 400, because scikit-learn dispatches each tree separately even when the query is a single row.

Where a Random Forest lives in a real system

A random forest is almost never the product. It’s a piece between the feature table and the decision, and in practice it shows up in three places:

  • Batch scoring. A nightly job reads the table of customers, accounts or sensors, computes a probability per row and stores it. This is where the forest shines: the 100-tree, depth-12 forest processed the batch at 3.8 microseconds per row, so a million rows would take about four seconds on one thread if that pace holds. A batch dilutes the fixed cost of each call, although trees and depth still weigh in: 400 trees took four times as long as 100.
  • Behind an API, carefully. If the decision is made online (approving a payment, flagging an email), every request pays the per-call cost and the model has to sit in memory on every replica. There it pays to limit depth, cut the number of trees to where the curve stops rising, or batch requests.
  • As a baseline and a feature detector. Before trying something more complex, a forest with default settings gives an honest number in minutes. To learn which features it depends on, use permutation importance on evaluation data, not the built-in one.

The data pattern it suits is the table: rows with numeric columns (and categorical ones once encoded as numbers, because scikit-learn doesn’t accept text), a few hundred thousand examples, non-linear relationships. It needs less preparation than other models: features don’t need scaling. For images, audio or raw text it’s not the tool. For detecting anomalies without labels there’s a relative with a similar name and different mechanics, the Random Cut Forest, which I measured in Flink on streaming events.

When I’d choose it: tabular data, a good result without much tuning, training across several CPUs, and batch prediction. When I wouldn’t: if you need to extrapolate beyond the seen range, if the model has to be small and answer in microseconds per request, or if every point of accuracy matters and there’s time to tune histogram boosting, which matched or slightly beat the forest on this same data in a fraction of the time on a single CPU. And with small tables it’s worth looking beyond trees: TabPFN, a model that doesn’t even train on the table, beat tuned XGBoost on all fourteen tables I measured.

Sources

  • Breiman, L. (2001). “Random Forests”. Machine Learning, 45(1), 5–32. DOI 10.1023/A:1010933404324.
  • Ho, T. K. (1995). “Random decision forests”. Proceedings of the 3rd International Conference on Document Analysis and Recognition, pp. 278–282. DOI 10.1109/ICDAR.1995.598994.
  • scikit-learn user guide, “Ensembles: Gradient boosting, random forests, bagging, voting, stacking”: the note that its implementation averages probabilities instead of counting votes, and the warning about impurity-based importance.
  • NASA POWER, Daily API and data sources: where the daily data comes from. Precipitation is PRECTOTCORR, which the API itself defines as MERRA-2 bias-corrected total precipitation.
  • Open-Meteo, Historical Weather API: the best_match selection, which blends IFS HRES, ERA5 and ERA5-Land, and which made me drop that source.

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.