
Logistic regression explained: 16 numbers to decide whether it rains tomorrow
What logistic regression is, how it turns a sum into a probability and when it misleads, measured on daily weather from seven Chilean cities between 1984 and 2026. A straight line gives negative probabilities on 9% of days; at a 0.5 threshold logistic regression catches only half of the rain; and it promises more rain than falls.
A logistic regression fits in 16 numbers: 15 weights and a constant. With them I decided whether it would rain tomorrow in seven Chilean cities, and it ranked days almost as well as a 200-tree forest with 1.5 million nodes. But with the default threshold it warned about only half of the rain, and when it said “30%” it rained less than that.
This is the fourth part of the series, with the same NASA POWER daily data I used for Random Forest, linear regression and K-Means. The question is the one the forest answered: will it rain at least 1 mm tomorrow? I trained on 74,145 city-day rows from 1984 to 2012 and evaluated on 27,256 from 2016 to 2026, inside a container with 8 CPUs.
What is logistic regression?
Logistic regression estimates a probability. It adds up the day’s variables, each with its weight, just like a linear regression. The difference is at the end: that sum goes through the sigmoid, an S-shaped curve that turns any number into a value between 0 and 1. A very negative sum ends near 0, a very positive one near 1, and a sum of 0 lands at 0.5.
The weights are not chosen by least squares but by maximizing the likelihood: the weights that make what actually happened most probable. scikit-learn adds a mild L2 penalty on the weights by default; on this data, weakening it to C = 10 did not change the metrics at four decimals. Cox framed it as regression for binary sequences in 1958, and Berkson had already used the logistic function in bioassay in 1944.
Why a straight line is not enough
The temptation is to fit a linear regression to a column of 0s and 1s. It half works, and the chart shows where it breaks.
- observed rate
- straight line on 0/1
- logistic regression
With all fifteen variables, the straight line gave a negative probability on 9.1 % of 2016-2026 days and one above 1 on 0.4 %, from -0.10 to 2.28. Its ranking of days was almost as good (AUC 0.847 on the raw output, against 0.847), but a negative number is not a probability.
With humidity alone, the real rain rate climbs slowly and then shoots up on the most humid days. The line cannot bend: it drops below zero on dry days. With all fifteen variables, it gave negative probabilities on 9.1% of 2016-2026 days and 2.28 in the worst case. For ranking days it was almost as good as logistic regression (AUC 0.847 for both, computed on the unclipped output), but a negative number cannot be used as a probability.
What the weights say
Each logistic weight reads as an odds ratio. The odds are the probability of rain divided by the probability of no rain; a weight turned into an odds ratio says by how much they multiply when the variable rises and the rest stay fixed.
- rain today×2.19
- wind N-S (cos)×1.94
- latitude×0.58
- max wind×0.76
- min today×0.83
- dew point×1.18
- day of year (cos)×0.86
- pressure change×0.88
- rain yesterday×1.14
- pressure×1.12
- radiation×1.11
- max today×0.92 · unstable sign
- humidity×1.06 · unstable sign
- day of year (sin)×1.04
- wind E-W×1.02 · unstable sign
the same weights in real units (odds multiplied by)
- ×6.29wind direction 0° instead of 180°
- ×2.135 mm more rain today
- ×1.36pressure 1 kPa lower than yesterday
- ×1.0310 more points of humidity
- ×0.815 °C higher min temperature
Humidity, whose rain rate clearly rises on its own, gets a weight whose interval crosses 1 (0.90 to 1.22) probably because other variables carry similar information (in the linear regression post its VIF was 33). The weights describe this model with these variables, not causes.
With standardized variables, today’s rain multiplied the odds of rain tomorrow by 2.19 per standard deviation, and the north-south component of wind direction by 1.94. In real units: 5 mm more rain today multiplied them by 2.13; pressure 1 kPa lower than yesterday, by 1.36; and a direction near 0° instead of 180°, by 6.29. In the usual meteorological convention, 0° is wind coming from the north, but I did not verify which convention NASA POWER uses, so I keep it as a direction and not a cause.
Humidity is the interesting case. On its own, the rain rate clearly rises as it increases, as the previous chart shows; inside the model, 10 more points barely multiplied the odds by 1.03, and across 200 resamples its interval crossed 1 (the weight changed sign in 20.5% of them). Probably other variables carry similar information: in linear regression I measured that humidity can be explained almost entirely by the others (VIF of 33). A weight is the effect of moving one variable with the others fixed, not its importance on its own.
Cost decides the threshold
Logistic regression returns a probability. Deciding “warn about rain” needs a threshold, and the default one, 0.5, does not have to be the right one.
threshold 0.05
threshold 0.10
threshold 0.15
threshold 0.20
threshold 0.25
threshold 0.30
threshold 0.35
threshold 0.40
threshold 0.45
threshold 0.50
threshold 0.55
threshold 0.60
threshold 0.65
threshold 0.70
threshold 0.75
threshold 0.80
threshold 0.85
threshold 0.90
threshold 0.95
threshold 0.49
threshold 0.23
threshold 0.13
Threshold 0.49 chosen on 2010-2012. In 2016-2026: recall 50 %, precision 58 %, alarm on 17 % of days. Cost per day 0.175, against 0.203 never warning and 0.797 always warning.
Threshold 0.23 chosen on 2010-2012. In 2016-2026: recall 80 %, precision 45 %, alarm on 36 % of days. Cost per day 0.318, against 0.608 never warning and 0.797 always warning.
Threshold 0.13 chosen on 2010-2012. In 2016-2026: recall 92 %, precision 35 %, alarm on 53 % of days. Cost per day 0.497, against 2.026 never warning and 0.797 always warning.
At 0.5 the model warns on few days and catches 48 % of the rain, with 82.6 % accuracy; never warning already scores 79.7 %. Accuracy is the wrong lens when rain is one day in five: the threshold that makes sense depends on what a missed rain costs.
At 0.5, the model warned on few days: it caught 48.1% of rainy days with 58.7% precision and 82.6% accuracy. That sounds good until you see that never warning already reaches 79.7% accuracy, because it rains one day in five. That is why accuracy misleads with imbalanced classes.
To choose the threshold without looking at the test data, I trained up to 2009 and chose it on 2010-2012. If missing rain costs the same as a false alarm, the threshold was 0.49 and the 2016-2026 cost fell from 0.203 per day (never warning) to 0.175. If missing costs three times more, the threshold dropped to 0.23: it caught 80.2% of rainy days, with alarms on 36% of days. If it costs ten times more, to 0.13: it caught 92.5%, warning on more than half of the days, and still cost less than always warning (0.497 against 0.797).
Does a 30% chance of rain come true 30% of the time?
A probability is only useful if it means what it says. That is calibration, and it is measured by grouping days by predicted probability and counting how many ended with rain. Brier proposed a way to score such forecasts in 1950, precisely for weather.
AUC 0.847 · Brier 0.119 · Brier skill vs. always predicting the training rain rate 0.279 · PR-AUC 0.573
AUC 0.848 · Brier 0.156 · Brier skill vs. always predicting the training rain rate 0.057 · PR-AUC 0.571
AUC 0.882 · Brier 0.106 · Brier skill vs. always predicting the training rain rate 0.357 · PR-AUC 0.660
logistic trained up to 2009: observed rain rate against mean predicted probability
- observed
- mean predicted
The logistic regression predicted 24.1 % rain on average and it rained on 20.3 % of days. Among the 11,010 days it gave under 10%, it rained on 2.6 %; among the 874 days it gave 70 to 80%, on 56.9 %. It learned from 1984-2012, when rain was more frequent. With class_weight='balanced' it catches more rain but its probabilities stop meaning what they say (average 38.5 %). The forest is the best calibrated of the three.
Logistic regression promised more rain than fell. On average it predicted 24.1% and it rained on 20.3% of days. On the 11,010 days it gave under 10%, it rained on 2.6%; on the 874 days it gave 70 to 80%, on 56.9%. Part of the gap is about the period: it learned from 1984-2012, when it rained on 26.6% of days. The model trained up to 2009 predicted more rain than observed in every year from 2016 to 2026.
Balancing the classes with class_weight='balanced' is a common recipe to catch more rain, and it works: balanced accuracy rose from 0.70 to 0.78. But its mean probability was 38.5% and, on the days it gave 50 to 60%, it rained on 33%. If the probabilities are going to be used, that forces a recalibration. The forest ended better calibrated than both, as the Random Forest post had already shown.
Which variables survive the penalty
Regularization penalizes large weights. L2 (ridge) shrinks them; L1 (lasso, Tibshirani, 1996) can set them exactly to zero, which makes it a way to select variables.
C = 1015 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 115 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.315 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.115 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.0314 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.0112 variables alive · log loss 0.3780 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.00311 variables alive · log loss 0.3772 · AUC 0.848
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.0018 variables alive · log loss 0.3781 · AUC 0.848
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.00035 variables alive · log loss 0.3935 · AUC 0.842
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.00012 variables alive · log loss 0.4621 · AUC 0.802
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 1015 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 115 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.315 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.115 variables alive · log loss 0.3782 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.0315 variables alive · log loss 0.3781 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.0115 variables alive · log loss 0.3780 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.00315 variables alive · log loss 0.3778 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.00115 variables alive · log loss 0.3777 · AUC 0.847
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.000315 variables alive · log loss 0.3800 · AUC 0.846
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
C = 0.000115 variables alive · log loss 0.3891 · AUC 0.844
- max
- min
- rain today
- rain yest.
- wind
- wind E-W
- wind N-S
- radiation
- humidity
- pressure
- press. change
- dew point
- day (sin)
- day (cos)
- latitude
With L1 and C = 0.0001 only rain today and wind N-S remain, with AUC 0.802. With C = 0.001, 8 of 15 variables already give log loss 0.3781, practically the same as all 15 with a weak L1 (C = 10) (0.3782). C was not chosen on these test figures; they only show the path.
With L1 and the strongest penalty I tried, only two variables were left: today’s rain and the north-south component of wind, with AUC 0.802. With 8 variables, the 2016-2026 loss was already practically the same as with all 15 and a weak L1 (C = 10). On this data almost half the columns are spare for ranking days; I did not choose the penalty from these figures, they only show the path.
Scale before solving
Logistic regression has no closed-form solution: it is solved iteratively, and the feature scale issue from linear regression comes back.
Four times slower than real time.
In real time.
All reached very similar performance (AUC 0.847 to 0.848). Unscaled, lbfgs, scikit-learn's default, needed 819 iterations and 33 times more time than with scaled variables.
In that one-thread measurement, with scaled variables, every solver finished in under 0.21 seconds. Unscaled, lbfgs went from 26 to 819 iterations and from 0.045 to 1.51 seconds, and saga from 0.20 to 1.29. They reached very similar performance (log loss 0.377 to 0.378). Unscaled, the variables have very different ranges, such as pressure in kilopascals next to the sine of the day of year, which tends to slow down first-order methods such as lbfgs and saga.
Logistic regression or Random Forest
The 200-tree forest ranked days better: AUC 0.882 against 0.847, and PR-AUC 0.660 against 0.573. It was also better calibrated (Brier skill of 0.357 against 0.279 relative to always predicting the training rain rate). What that difference costs shows in everything else.
At half of real time.
In a separate run on 8 threads, the forest took 3.26 s. The 16 logistic parameters take 355 bytes as JSON; the serialized forest, 124 MB.
On one thread, logistic regression trained in 0.049 seconds and the forest in 23.8, about 490 times more. Scoring one row took 0.39 ms with scikit-learn and 0.003 ms with a direct NumPy operation, against 3.1 ms for the forest. And logistic regression learns fast: with 2,000 rows it already had AUC 0.840, 0.007 from its value with all 74,145; the forest rose from 0.866 to 0.882 with more data.
Where logistic regression lives in a real system
The 16 logistic parameters fit in 0.4 KB of JSON; the pickled 200-tree forest takes 124.3 MB. One row scored with plain NumPy took 0.0029 ms (0.394 ms through scikit-learn), matching scikit-learn to within 10⁻¹⁵.
- Risk scores. Credit, fraud or clinical risk: one probability per case, weights that can be audited, and a threshold set by the business according to what each mistake costs.
- Filters and alerts. Spam, rain or failure alerts: trained offline and scored online with a sum and a sigmoid, without loading a heavy model.
- Baseline. Before a forest or a neural network, a logistic regression trained in milliseconds tells you how much the rest adds. Here, the forest beat it by 0.035 AUC.
- Calibration monitoring. The loop that is almost always missing: comparing the mean probability with what happened. Here it would have caught the gap from the first year.
When I would choose it: when I need fast probabilities, auditable weights and a threshold adjustable by cost. When I would not: with strongly non-linear relationships or strong interactions, where a forest ranks better, or when I cannot monitor calibration over time.
Sources
- Cox, D. R. (1958). “The Regression Analysis of Binary Sequences”. Journal of the Royal Statistical Society: Series B, 20(2), 215–232. DOI 10.1111/j.2517-6161.1958.tb00292.x.
- Berkson, J. (1944). “Application of the Logistic Function to Bio-Assay”. Journal of the American Statistical Association, 39(227), 357–365. DOI 10.1080/01621459.1944.10500699.
- Brier, G. W. (1950). “Verification of Forecasts Expressed in Terms of Probability”. Monthly Weather Review, 78(1), 1–3. DOI 10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2.
- Niculescu-Mizil, A. and Caruana, R. (2005). “Predicting good probabilities with supervised learning”. Proceedings of the 22nd International Conference on Machine Learning, 625–632. DOI 10.1145/1102351.1102430.
- Tibshirani, R. (1996). “Regression Shrinkage and Selection via the Lasso”. Journal of the Royal Statistical Society: Series B, 58(1), 267–288. DOI 10.1111/j.2517-6161.1996.tb02080.x.
- scikit-learn,
LogisticRegression(solvers,class_weightand regularization) and the calibration guide. - NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.