17°
Portada del artículo: Naive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edges
Machine learningAlgorithmsPythonData

Naive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edges

What Naive Bayes is, what the independence assumption means and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Fitting on 74,145 rows took 0.008 s and the model weighs 1.3 KB, at AUC 0.822; on binned variables it reached 0.844, with no conclusive difference from logistic regression. 62 % of its probabilities fell below 0.01 or above 0.99, and calibrating it took log loss from 1.09 to 0.39.

Efrain Garay 15 September 2026

Playing summary

Fitting this model on 74,145 rows of weather took 0.008 seconds, and the serialized artifact weighs 1.3 KB: two numbers per variable and class. With that it reached AUC 0.822 on the 2016-2026 test period, and measuring the variables in bins, 0.844, with no conclusive difference from a logistic regression. The price shows up elsewhere: 62 % of its probabilities landed below 0.01 or above 0.99.

This is the ninth post in the series, with the same daily NASA POWER data from seven Chilean cities I used for Random Forest, linear regression, K-Means, logistic regression, decision trees, gradient boosting, SVM and KNN. The question is still whether at least 1 mm of rain will fall tomorrow. I trained on 1984-2012 and evaluated on the 2016-2026 test period, in a container with 8 CPUs and scikit-learn 1.7.2. I chose the representation of the variables, the number of bins and the smoothing by fitting on 1984-2009 (up to 30 December, because the label of the 31st is rain on 1 January 2010) and scoring on 2010-2012. The comparison models use the settings from their own posts, and the 2016-2026 curves and examples are descriptive.

In 47 seconds, narrated: the whole model weighs 1.3 kilobytes and was fitted in eight milliseconds; how it looks at each variable separately inside each class; what happens when a column is copied seven times; why six out of ten days get a probability pinned to zero or to one; and how it behaves with only thirty training rows. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

What is Naive Bayes?

Bayes’ theorem flips the question: instead of modeling the probability of rain given the day, it models how days look inside each class and then turns the calculation around. The problem is that describing fifteen variables jointly needs enormous amounts of data. The naive version solves it by assuming that, within each class, the variables are independent: then it is enough to describe each variable on its own and multiply. It has roots in the probabilistic document-indexing idea Maron and Kuhns published in 1960.

How Naive Bayes sees two variablesThe same 200 days and two standardized variables as the SVM and KNN posts. The loop goes through the days, the bell fitted to dry days, the bell fitted to rainy days, and the line where the probability crosses 0.5.
dry next dayrain next dayprobability 0.5AUC 2010-12 0.795humidity (standardized)today's rain, log scale (standardized)

The model fits one mean and one variance per variable and class, so each bell comes out with its axes aligned: it assumes the two variables are independent within each class. In these days they are not, with a correlation of 0.40 among dry days and 0.42 among rainy ones. Even so, on 2010-2012 it reached AUC 0.795 against 0.803 for a logistic regression on the same two variables.

With two variables, humidity and today’s rain, and the same 200 days as the earlier posts, the model fits one bell per class, with axes aligned because it assumes the two variables are unrelated within the class. In those days they are related, with a correlation of 0.40 among dry days and 0.42 among rainy ones. Even so, on 2010-2012 it reached AUC 0.795 against 0.803 for a logistic regression on the same two variables.

One prediction, opened into addends

Since everything is a multiplication, taking logarithms turns it into a sum: the log-odds of rain is the prior term plus one term per variable.

One prediction, opened into addendsLog-odds of rain for Santiago on May 28, 2010, as the prior term plus one term per variable. Positive pushes towards rain, negative towards a dry day.
prior (base rate)−0.99today's rain+4.29pressure−2.90north-south wind+2.00yesterday's rain−0.88radiation+0.84humidity−0.76latitude−0.63day of year (cos)+0.46pressure change−0.40min temp+0.36wind−0.24east-west wind+0.18dew point+0.16max temp−0.04day of year (sin)−0.02total+1.42 · probability of rain 81 %

Naive Bayes adds one log-likelihood ratio per variable, which is why a prediction can be read term by term. That day the rain already fallen pushed +4.29 towards rain and pressure −2.90 against it; starting from a prior of −0.99, the terms add up to 1.42, which is a probability of 81 %. The model's own log-odds is 1.42, the same number. It rained the next day.

For a Santiago day in May 2010, the rain already fallen contributed +4.29 and pressure −2.90; starting from a prior of −0.99, the fifteen terms add up to 1.42, which is a probability of 81 %. The log-odds the model itself returns is 1.42, the same number: the explanation is not an approximation, it is the calculation. It rained the next day.

The assumption that does not hold

In the 66,466 rows of 1984-2009, within the dry-next-day class, 10.5 % of variable pairs exceed 0.5 in absolute correlation: radiation and day of year 0.81, humidity and pressure 0.80, maximum temperature and humidity −0.71. Within rainy days, 7.6 %.

What that costs is clearest when exaggerated: I copied humidity inside the table, so the model counts the same evidence several times.

Counting the same evidence twiceHumidity copied 0, 1, 3 and 7 extra times. Naive Bayes and logistic regression fit on 1984-2009 and scored on 2010-2012.
AUClog loss0.800.820.840120.821+01.26+00.818+11.37+10.814+31.64+30.805+72.29+7extra copies of humidity
  • Naive Bayes
  • logistic

Naive Bayes multiplies one likelihood per column, so a copied column counts its evidence again. With 7 extra copies its AUC fell from 0.821 to 0.805, log loss rose from 1.26 to 2.29, and the average absolute log-odds went from 8.5 to 16.5: the model got far more confident about the same day. Logistic regression, which fits the columns together, stayed at AUC 0.850 and log loss 0.400.

With seven extra copies, AUC on 2010-2012 fell from 0.821 to 0.805 and log loss rose from 1.26 to 2.29. The average absolute log-odds went from 8.5 to 16.5: the model became twice as confident on the same information. Logistic regression, which fits the columns together, stayed at 0.850. It is what Domingos and Pazzani showed in 1997: Naive Bayes can classify well even when its probabilities are wrong, because the ranking survives the broken assumption far better than the number. Hand and Yu reviewed in 2001 why it keeps working so often.

The bell against the data

The Gaussian version also assumes that each variable, within its class, follows a bell. Rain does not.

The bell against the dataDays of 1984-2009 followed by rain and days followed by dry weather, with the normal curve GaussianNB fits to each class: same mean, same variance.
033today's rain (mm)04log(1 + rain today)1797humidity (%)
  • rain next day
  • dry next day
  • fitted normal

Rain today is not bell-shaped: 33 % of the rainy-next-day column sits in the first bin and the tail runs far to the right, with a skew of 3.63. The normal fitted to it, mean 5.44 and variance 71.6, spreads mass below zero, where no day exists. Taking the logarithm brings the shape closer to a bell, and humidity already looks like one. This is the assumption that the binned version of the model drops.

In 1984-2009, 34 % of days have exactly zero millimeters and rain has a skew of 6.0. The normal fitted to days followed by rain has mean 5.4 and variance 71.6, so it spreads mass below zero, where no day exists. I tried three ways to fix it, choosing on 2010-2012: taking the logarithm of the two rain columns raised AUC from 0.821 to 0.827; a quantile transform to a normal shape, to 0.825; and cutting every variable into 40 quantile bins and counting frequencies with CategoricalNB, with smoothing alpha = 10 chosen in the same grid, to 0.843. With that last one, refit on 1984-2012, the 2016-2026 test period gave 0.844 against 0.822 for the raw Gaussian. Leaving the bell behind cost 26.7 KB instead of 1.3 KB.

No scaling needed, but the smoothing does look at the units

The smoothing does look at the unitsGaussianNB fits one mean and one variance per variable and class, so rescaling a column changes nothing by itself. But var_smoothing adds epsilon = 1e-9 × the largest variance in the table to every variance.
0.60.70.80.821kPaε 3.5e-70.821Pa (×1,000)ε 1.7e-20.675×1,000,000ε 1.7e40.615×1,000,000,000ε 1.7e10pressure written as · epsilon added to every variance

With pressure in kPa, epsilon was 3.5e-7 against a smallest variance of 0.140, and AUC on 2010-2012 was 0.821; in Pa it was 0.821. Multiplied by a million, epsilon reached 1.7e4, far above every other variance, every variable was flattened to almost the same bell, and AUC fell to 0.675; times a billion, 0.615. Standardizing all the variables gave 0.821, the same as kPa.

One mean and one variance per variable and class change nothing if a column is multiplied by a thousand: standardizing every variable gave the same 0.821 as the raw data, and writing pressure in pascals, 0.8205. The catch is var_smoothing, which adds to every variance an epsilon equal to 1e-9 times the largest variance in the table. With pressure multiplied by a million, that epsilon reached 17,167, far above the smallest variance in the table, 0.140: every variable was flattened into almost the same bell and AUC fell to 0.675. Times a billion, 0.615.

That same parameter, swept from 1e-12 to 0.1 on validation, barely moved the AUC up to 1e-4 (0.820); at 1e-3 it dropped to 0.813 and at 0.1 to 0.796, while log loss fell from 1.26 to 0.53, because flattening the bells makes the model less confident. I kept the default: the best of the sweep, 1e-7, improved AUC by 7e-7.

Probabilities at the edges

It ranks well and lies about how muchReliability on the 2016-2026 test period: for each tenth of predicted probability, the real share of rainy days. The diagonal is perfect agreement.
0.00.00.50.51.01.0days below 0.01 or above 0.9961.9 %Gaussian Naive Bayesloss 1.092.2 %calibrated (isotonic)loss 0.390.2 %logistic regressionloss 0.3814.6 %boostingloss 0.33

Multiplying fifteen likelihoods pushes the result to the edges: 61.9 % of test days got a probability below 0.01 or above 0.99, and the log loss was 1.09 against 0.378 for logistic regression, even though their AUCs are 0.822 and 0.847. An isotonic calibration fit on 2010-2012 left the ranking untouched, AUC 0.822, and brought the log loss down to 0.391. For a decision you need the calibrated version.

Multiplying fifteen likelihoods, several of them correlated, pushes the result to the extremes: 61.9 % of test days got a probability below 0.01 or above 0.99. Log loss was 1.09 against 0.38 for logistic regression, even though their AUCs are 0.822 and 0.847. With the binned version, 1.01.

Calibrating separately fixes it without touching the ranking. Fitting the calibration on 2010-2012 over a model trained on 1984-2009, a sigmoid left log loss at 0.402 and an isotonic one at 0.391, at AUC 0.823 and 0.822; over the binned version, the isotonic reached 0.383, close to the 0.378 of logistic regression. This is what Niculescu-Mizil and Caruana measured in 2005 across model families. Changing only the prior, using the 2010-2012 rain rate instead of the 1984-2012 one, barely helped: 1.07.

With few rows it is hard to beat

With few rows, few parameters hold up betterAUC on the 2016-2026 test period against training rows, drawn at random from 1984-2012. Median of up to 20 draws, with the interquartile range shaded. Descriptive: no choice was made with this curve.
301001,00010,00074,1450.50.60.70.80.9training rows (log)
  • Gaussian Naive Bayes
  • Naive Bayes on bins
  • logistic regression
  • boosting, 594 rounds

With 30 rows, Gaussian Naive Bayes reached a median AUC of 0.795 and logistic regression 0.790, while the boosting model stayed at 0.500: with that little data it could not even order the days. From 300 rows on, logistic regression is ahead, and with all 74,145 the boosting model reaches 0.883 while Gaussian Naive Bayes stays at 0.822. Naive Bayes fits two numbers per variable and class, so it has little to estimate and little room to improve.

Naive Bayes estimates two numbers per variable and class, so it has little to estimate. With 30 training rows, its median AUC on the test period was 0.795 and logistic regression’s 0.790, while the boosting model stayed at 0.500: with that little data it could not even order the days. With 100 rows, 0.817 against 0.809 and 0.801. From 300 on, logistic regression leads and the boosting model pulls away, up to 0.883 with all the rows while the Gaussian stays at 0.822. It is the comparison Ng and Jordan formalized in 2001: the generative model reaches its ceiling sooner, and that ceiling is lower.

Learning one year at a time

One year of data already gets thereThe model is trained with partial_fit, one year at a time from 1984 to 2012, never looking at earlier years again. After each year, AUC on the 2016-2026 test period (descriptive).
0.8120.8180.8240.83019841991199820052012years seenAUC 2016-26

After 1984 alone, 2,562 rows, it already scored 0.826; with all 74,145 rows it ended at 0.822, and the best year along the way was 0.826. Each yearly update took a median of 1.21 ms, because it only adds up counts, sums and sums of squares. The streamed model matches a single fit on all the rows: the largest difference in predicted probability is 5e-9.

Fitting is counts, sums and sums of squares, so it can be done in pieces. I trained with partial_fit one year at a time, never revisiting the earlier ones: after 1984 alone, 2,562 rows, it already scored 0.826, above the final 0.822 with all 74,145. Each update took a median of 1.2 ms, and the model trained in pieces matches one fitted in a single pass: the largest difference in predicted probability is 5e-9.

Naive Bayes, logistic regression, KNN, forest and boosting

The same 74,145 rows, one threadMedian of 3 fits, except the forest (one measurement). AUC measured on the 2016-2026 test period. Switch scenario to see the cost of scoring.
Gaussian Naive Bayes · AUC 0.8220.0081 s
KNN, k = 100 · AUC 0.8700.0142 s
Logistic regression · AUC 0.8470.0496 s
Naive Bayes on bins · AUC 0.8440.0581 s
Tree, depth 7 · AUC 0.8620.3703 s
Boosting, 594 rounds · AUC 0.8832.6182 s
Forest, 200 trees · AUC 0.88223.9211 s

At a quarter of real time.

Gaussian Naive Bayes0.0023 s
Naive Bayes on bins0.0151 s
Tree, depth 70.0015 s
Logistic regression0.0016 s
Forest, 200 trees0.3248 s
Boosting, 594 rounds0.5834 s
KNN, k = 1005.7856 s

The 27,256 rows, in real time.

Serialized, Gaussian Naive Bayes takes 1.3 KB, the binned version 26.7 KB, logistic regression 1.8 KB, boosting 2.1 MB and the forest 124 MB. Scoring one day took the Gaussian 0.38 ms and boosting 3.2 ms.

With the same rows, I resampled the city-year blocks of 2016-2026 2,000 times to compare AUC in pairs. The binned version landed 0.003 below logistic regression, with a 95 % interval from −0.008 to 0.001, the only one that crosses zero: on ranking ability, the difference is not conclusive. Against the depth-7 tree it was 0.018 below, against KNN 0.026 and against boosting 0.039 (from −0.044 to −0.035). Over the raw Gaussian it was 0.021 above.

Where Naive Bayes lives in a real system

Where Naive Bayes livesCounts and sums per class, an artifact of a few kilobytes, and a mandatory detour through calibration.
Where Naive Bayes livestable74,145 rowscounts and sumsper classartifact1.3 KBnew batchpartial_fitonline filter1 row · 0.38 mscalibrationbefore decidingbaselineAUC 0.822

Measured on one thread with all 74,145 rows: Gaussian Naive Bayes fits in 0.008 s, weighs 1.3 KB serialized, scores one row in 0.38 ms and the whole test period in 0.002 s, at AUC 0.822 and log loss 1.09. Calibrated on 2010-2012 the log loss drops to 0.391. Logistic regression: 0.050 s, AUC 0.847, log loss 0.378. Boosting: 2.6 s, 2.1 MB, AUC 0.883.

  • Online filters and triage. Scoring a row takes 0.38 ms and the whole model fits in 1.3 KB: it works where decisions must be fast and cheap, as in mail filters, one of the earliest academic applications being Sahami and coauthors in 1998.
  • Data arriving in pieces. partial_fit updates the model with each batch without keeping the previous ones.
  • Very few examples. With dozens of rows it already ranks, where models with more parameters have not started.
  • Calibration is mandatory if the number is used. Uncalibrated, its probabilities live at the edges.
  • A baseline. It is the reference any more expensive model has to justify itself against.

When I would choose it: as the first measurement of a new problem, for small sets, for text with many sparse columns, or when the model must be tiny and updatable. When not: when a credible probability is needed without an extra calibration step, when the variables are strongly correlated with each other, or when there is plenty of data and the ceiling matters. On this data, with all 74,145 rows, a boosting model landed 0.039 of AUC above it with less than a third of its log loss.

Sources

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.