
Naive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edges
What Naive Bayes is, what the independence assumption means and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Fitting on 74,145 rows took 0.008 s and the model weighs 1.3 KB, at AUC 0.822; on binned variables it reached 0.844, with no conclusive difference from logistic regression. 62 % of its probabilities fell below 0.01 or above 0.99, and calibrating it took log loss from 1.09 to 0.39.
Fitting this model on 74,145 rows of weather took 0.008 seconds, and the serialized artifact weighs 1.3 KB: two numbers per variable and class. With that it reached AUC 0.822 on the 2016-2026 test period, and measuring the variables in bins, 0.844, with no conclusive difference from a logistic regression. The price shows up elsewhere: 62 % of its probabilities landed below 0.01 or above 0.99.
This is the ninth post in the series, with the same daily NASA POWER data from seven Chilean cities I used for Random Forest, linear regression, K-Means, logistic regression, decision trees, gradient boosting, SVM and KNN. The question is still whether at least 1 mm of rain will fall tomorrow. I trained on 1984-2012 and evaluated on the 2016-2026 test period, in a container with 8 CPUs and scikit-learn 1.7.2. I chose the representation of the variables, the number of bins and the smoothing by fitting on 1984-2009 (up to 30 December, because the label of the 31st is rain on 1 January 2010) and scoring on 2010-2012. The comparison models use the settings from their own posts, and the 2016-2026 curves and examples are descriptive.
What is Naive Bayes?
Bayes’ theorem flips the question: instead of modeling the probability of rain given the day, it models how days look inside each class and then turns the calculation around. The problem is that describing fifteen variables jointly needs enormous amounts of data. The naive version solves it by assuming that, within each class, the variables are independent: then it is enough to describe each variable on its own and multiply. It has roots in the probabilistic document-indexing idea Maron and Kuhns published in 1960.
The model fits one mean and one variance per variable and class, so each bell comes out with its axes aligned: it assumes the two variables are independent within each class. In these days they are not, with a correlation of 0.40 among dry days and 0.42 among rainy ones. Even so, on 2010-2012 it reached AUC 0.795 against 0.803 for a logistic regression on the same two variables.
With two variables, humidity and today’s rain, and the same 200 days as the earlier posts, the model fits one bell per class, with axes aligned because it assumes the two variables are unrelated within the class. In those days they are related, with a correlation of 0.40 among dry days and 0.42 among rainy ones. Even so, on 2010-2012 it reached AUC 0.795 against 0.803 for a logistic regression on the same two variables.
One prediction, opened into addends
Since everything is a multiplication, taking logarithms turns it into a sum: the log-odds of rain is the prior term plus one term per variable.
Naive Bayes adds one log-likelihood ratio per variable, which is why a prediction can be read term by term. That day the rain already fallen pushed +4.29 towards rain and pressure −2.90 against it; starting from a prior of −0.99, the terms add up to 1.42, which is a probability of 81 %. The model's own log-odds is 1.42, the same number. It rained the next day.
For a Santiago day in May 2010, the rain already fallen contributed +4.29 and pressure −2.90; starting from a prior of −0.99, the fifteen terms add up to 1.42, which is a probability of 81 %. The log-odds the model itself returns is 1.42, the same number: the explanation is not an approximation, it is the calculation. It rained the next day.
The assumption that does not hold
In the 66,466 rows of 1984-2009, within the dry-next-day class, 10.5 % of variable pairs exceed 0.5 in absolute correlation: radiation and day of year 0.81, humidity and pressure 0.80, maximum temperature and humidity −0.71. Within rainy days, 7.6 %.
What that costs is clearest when exaggerated: I copied humidity inside the table, so the model counts the same evidence several times.
- Naive Bayes
- logistic
Naive Bayes multiplies one likelihood per column, so a copied column counts its evidence again. With 7 extra copies its AUC fell from 0.821 to 0.805, log loss rose from 1.26 to 2.29, and the average absolute log-odds went from 8.5 to 16.5: the model got far more confident about the same day. Logistic regression, which fits the columns together, stayed at AUC 0.850 and log loss 0.400.
With seven extra copies, AUC on 2010-2012 fell from 0.821 to 0.805 and log loss rose from 1.26 to 2.29. The average absolute log-odds went from 8.5 to 16.5: the model became twice as confident on the same information. Logistic regression, which fits the columns together, stayed at 0.850. It is what Domingos and Pazzani showed in 1997: Naive Bayes can classify well even when its probabilities are wrong, because the ranking survives the broken assumption far better than the number. Hand and Yu reviewed in 2001 why it keeps working so often.
The bell against the data
The Gaussian version also assumes that each variable, within its class, follows a bell. Rain does not.
- rain next day
- dry next day
- fitted normal
Rain today is not bell-shaped: 33 % of the rainy-next-day column sits in the first bin and the tail runs far to the right, with a skew of 3.63. The normal fitted to it, mean 5.44 and variance 71.6, spreads mass below zero, where no day exists. Taking the logarithm brings the shape closer to a bell, and humidity already looks like one. This is the assumption that the binned version of the model drops.
In 1984-2009, 34 % of days have exactly zero millimeters and rain has a skew of 6.0. The normal fitted to days followed by rain has mean 5.4 and variance 71.6, so it spreads mass below zero, where no day exists. I tried three ways to fix it, choosing on 2010-2012: taking the logarithm of the two rain columns raised AUC from 0.821 to 0.827; a quantile transform to a normal shape, to 0.825; and cutting every variable into 40 quantile bins and counting frequencies with CategoricalNB, with smoothing alpha = 10 chosen in the same grid, to 0.843. With that last one, refit on 1984-2012, the 2016-2026 test period gave 0.844 against 0.822 for the raw Gaussian. Leaving the bell behind cost 26.7 KB instead of 1.3 KB.
No scaling needed, but the smoothing does look at the units
With pressure in kPa, epsilon was 3.5e-7 against a smallest variance of 0.140, and AUC on 2010-2012 was 0.821; in Pa it was 0.821. Multiplied by a million, epsilon reached 1.7e4, far above every other variance, every variable was flattened to almost the same bell, and AUC fell to 0.675; times a billion, 0.615. Standardizing all the variables gave 0.821, the same as kPa.
One mean and one variance per variable and class change nothing if a column is multiplied by a thousand: standardizing every variable gave the same 0.821 as the raw data, and writing pressure in pascals, 0.8205. The catch is var_smoothing, which adds to every variance an epsilon equal to 1e-9 times the largest variance in the table. With pressure multiplied by a million, that epsilon reached 17,167, far above the smallest variance in the table, 0.140: every variable was flattened into almost the same bell and AUC fell to 0.675. Times a billion, 0.615.
That same parameter, swept from 1e-12 to 0.1 on validation, barely moved the AUC up to 1e-4 (0.820); at 1e-3 it dropped to 0.813 and at 0.1 to 0.796, while log loss fell from 1.26 to 0.53, because flattening the bells makes the model less confident. I kept the default: the best of the sweep, 1e-7, improved AUC by 7e-7.
Probabilities at the edges
Multiplying fifteen likelihoods pushes the result to the edges: 61.9 % of test days got a probability below 0.01 or above 0.99, and the log loss was 1.09 against 0.378 for logistic regression, even though their AUCs are 0.822 and 0.847. An isotonic calibration fit on 2010-2012 left the ranking untouched, AUC 0.822, and brought the log loss down to 0.391. For a decision you need the calibrated version.
Multiplying fifteen likelihoods, several of them correlated, pushes the result to the extremes: 61.9 % of test days got a probability below 0.01 or above 0.99. Log loss was 1.09 against 0.38 for logistic regression, even though their AUCs are 0.822 and 0.847. With the binned version, 1.01.
Calibrating separately fixes it without touching the ranking. Fitting the calibration on 2010-2012 over a model trained on 1984-2009, a sigmoid left log loss at 0.402 and an isotonic one at 0.391, at AUC 0.823 and 0.822; over the binned version, the isotonic reached 0.383, close to the 0.378 of logistic regression. This is what Niculescu-Mizil and Caruana measured in 2005 across model families. Changing only the prior, using the 2010-2012 rain rate instead of the 1984-2012 one, barely helped: 1.07.
With few rows it is hard to beat
- Gaussian Naive Bayes
- Naive Bayes on bins
- logistic regression
- boosting, 594 rounds
With 30 rows, Gaussian Naive Bayes reached a median AUC of 0.795 and logistic regression 0.790, while the boosting model stayed at 0.500: with that little data it could not even order the days. From 300 rows on, logistic regression is ahead, and with all 74,145 the boosting model reaches 0.883 while Gaussian Naive Bayes stays at 0.822. Naive Bayes fits two numbers per variable and class, so it has little to estimate and little room to improve.
Naive Bayes estimates two numbers per variable and class, so it has little to estimate. With 30 training rows, its median AUC on the test period was 0.795 and logistic regression’s 0.790, while the boosting model stayed at 0.500: with that little data it could not even order the days. With 100 rows, 0.817 against 0.809 and 0.801. From 300 on, logistic regression leads and the boosting model pulls away, up to 0.883 with all the rows while the Gaussian stays at 0.822. It is the comparison Ng and Jordan formalized in 2001: the generative model reaches its ceiling sooner, and that ceiling is lower.
Learning one year at a time
After 1984 alone, 2,562 rows, it already scored 0.826; with all 74,145 rows it ended at 0.822, and the best year along the way was 0.826. Each yearly update took a median of 1.21 ms, because it only adds up counts, sums and sums of squares. The streamed model matches a single fit on all the rows: the largest difference in predicted probability is 5e-9.
Fitting is counts, sums and sums of squares, so it can be done in pieces. I trained with partial_fit one year at a time, never revisiting the earlier ones: after 1984 alone, 2,562 rows, it already scored 0.826, above the final 0.822 with all 74,145. Each update took a median of 1.2 ms, and the model trained in pieces matches one fitted in a single pass: the largest difference in predicted probability is 5e-9.
Naive Bayes, logistic regression, KNN, forest and boosting
At a quarter of real time.
The 27,256 rows, in real time.
Serialized, Gaussian Naive Bayes takes 1.3 KB, the binned version 26.7 KB, logistic regression 1.8 KB, boosting 2.1 MB and the forest 124 MB. Scoring one day took the Gaussian 0.38 ms and boosting 3.2 ms.
With the same rows, I resampled the city-year blocks of 2016-2026 2,000 times to compare AUC in pairs. The binned version landed 0.003 below logistic regression, with a 95 % interval from −0.008 to 0.001, the only one that crosses zero: on ranking ability, the difference is not conclusive. Against the depth-7 tree it was 0.018 below, against KNN 0.026 and against boosting 0.039 (from −0.044 to −0.035). Over the raw Gaussian it was 0.021 above.
Where Naive Bayes lives in a real system
Measured on one thread with all 74,145 rows: Gaussian Naive Bayes fits in 0.008 s, weighs 1.3 KB serialized, scores one row in 0.38 ms and the whole test period in 0.002 s, at AUC 0.822 and log loss 1.09. Calibrated on 2010-2012 the log loss drops to 0.391. Logistic regression: 0.050 s, AUC 0.847, log loss 0.378. Boosting: 2.6 s, 2.1 MB, AUC 0.883.
- Online filters and triage. Scoring a row takes 0.38 ms and the whole model fits in 1.3 KB: it works where decisions must be fast and cheap, as in mail filters, one of the earliest academic applications being Sahami and coauthors in 1998.
- Data arriving in pieces.
partial_fitupdates the model with each batch without keeping the previous ones. - Very few examples. With dozens of rows it already ranks, where models with more parameters have not started.
- Calibration is mandatory if the number is used. Uncalibrated, its probabilities live at the edges.
- A baseline. It is the reference any more expensive model has to justify itself against.
When I would choose it: as the first measurement of a new problem, for small sets, for text with many sparse columns, or when the model must be tiny and updatable. When not: when a credible probability is needed without an extra calibration step, when the variables are strongly correlated with each other, or when there is plenty of data and the ceiling matters. On this data, with all 74,145 rows, a boosting model landed 0.039 of AUC above it with less than a third of its log loss.
Sources
- Maron, M. E. and Kuhns, J. L. (1960). “On Relevance, Probabilistic Indexing and Information Retrieval”. Journal of the ACM, 7(3), 216–244. DOI 10.1145/321033.321035.
- Domingos, P. and Pazzani, M. (1997). “On the Optimality of the Simple Bayesian Classifier under Zero-One Loss”. Machine Learning, 29, 103–130. DOI 10.1023/A:1007413511361.
- Hand, D. J. and Yu, K. (2001). “Idiot’s Bayes — Not So Stupid After All?”. International Statistical Review, 69(3), 385–398. DOI 10.1111/j.1751-5823.2001.tb00465.x.
- Ng, A. Y. and Jordan, M. I. (2001). “On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes”. Advances in Neural Information Processing Systems 14.
- Niculescu-Mizil, A. and Caruana, R. (2005). “Predicting Good Probabilities with Supervised Learning”. ICML ‘05, 625–632. DOI 10.1145/1102351.1102430.
- Sahami, M., Dumais, S., Heckerman, D. and Horvitz, E. (1998). “A Bayesian Approach to Filtering Junk E-Mail”. AAAI Workshop on Learning for Text Categorization.
- scikit-learn 1.7, Naive Bayes.
- NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.