
Linear regression explained: the line that measured the heat in Chile, and what breaks it
What linear regression is, how its coefficients are computed and when it stops working, measured on daily data from seven Chilean cities. The max temperature rises 0.33 °C per decade in Temuco and falls in Valparaíso, unscaled gradient descent does not reach a useful solution, 5% of rows with a misplaced decimal point multiply the error by 4.7, and the line trains about 1,100 times faster than a random forest.
In the Random Forest post there was an awkward moment for the forest: trained on Santiago’s cold months, it could not predict the summer, and a plain linear regression missed by a little less than half as much. The line kept rising because that is the only thing it knows how to do.
This time the line is the star. I used the same NASA POWER daily data for seven Chilean cities, from La Serena to Punta Arenas, and measured inside a container with CPU and memory limits what usually stays in theory: how the coefficients are computed, what happens if the features are not scaled, when the coefficients stop being trustworthy, what a mistyped value does, and how much you pay against a forest.
What is linear regression?
Linear regression predicts a number as a weighted sum of features plus a constant. With one feature it is a line; with fifteen it is the same thing in fifteen dimensions. The weights are chosen so the sum of squared errors is as small as possible, the least squares method Legendre published in 1805 and Gauss justified in 1809.
Here I use it to predict tomorrow’s max temperature from fifteen features of the day: temperatures, rain, wind, radiation, humidity, pressure, dew point, the day of the year and the city’s latitude.
Largest weights: max today +4.69 and day of year (cos) +0.95 °C per standard deviation. Least squares and the normal equation give the same coefficients down to 3×10⁻¹³.
Read this way, a linear regression is a single neuron with identity activation, trained with squared error. Each weight says how much its feature adds or subtracts, which is why it is among the easiest models to inspect, as long as the weights can be trusted. Further down you will see when they cannot.
The slope: how much temperature rises in each city
The oldest use of regression is one x and one y. I averaged the daily max temperature for each year between 1981 and 2025 and fitted a line against the year. The slope is the trend, and the 95% interval summarizes its uncertainty under the assumptions of the conventional fit.
A dashed line means the conventional 95% interval includes zero. Intervals assume independent years and are not corrected for autocorrelation or for testing seven cities; with a Bonferroni threshold only the conventional p-values of Temuco and Valparaíso stay below it. The source switches to GEOS-IT for the most recent months.
Temuco rises 0.332 °C per decade (interval between 0.161 and 0.503), Santiago 0.193 and Puerto Montt 0.167. Valparaíso falls 0.095 °C per decade. La Serena, Concepción and Punta Arenas have intervals that include zero.
Those intervals assume years are independent, and they are not entirely: in Temuco one year’s residual resembles the next (autocorrelation of 0.37), which makes the interval narrower than it should be. I also tested seven cities at once; applying the stricter Bonferroni threshold, only Temuco’s and Valparaíso’s conventional p-values stay below that threshold, Temuco’s with the autocorrelation caveat. Santiago and Puerto Montt stay as hints, not conclusions.
Valparaíso’s negative slope has the same sign as the coastal cooling Falvey and Garreaud described in 2009 for central and northern Chile: −0.2 °C per decade at coastal stations between 1979 and 2006, while the Andes warmed. But my data are different (the daily max of a reanalysis, not stations), cover another period, and Valparaíso’s grid cell probably includes ocean. And of the three coastal cities in that latitude range, La Serena (+0.043) and Concepción (+0.067) show no such cooling. The sign matches in one; that neither confirms the pattern nor rules out a data product artifact.
One more caveat: the NASA POWER series combines MERRA-2 for the historical record with GEOS-IT for recent months, so part of the final trend may come from the source change.
How are the coefficients computed?
To predict tomorrow’s max temperature I trained on 74,145 city-day rows from 1984 to 2012 and evaluated on 27,256 rows from 2016 to August 30, 2026. I left out 2013 to 2015 so training and test are not back to back in time. I use that same evaluation period for every comparison in the post, so the figures describe that period; they are not a final test kept untouched. It is the same table as the forest post, with the same caveats about the source.
There are three ways to reach the weights. The normal equation builds the 16 by 16 matrix AᵀA and solves that system. scikit-learn’s LinearRegression uses scipy.linalg.lstsq, which works on the data without forming that matrix and is more numerically stable. And gradient descent moves the weights little by little in the direction that lowers the error.
Forty times slower than real time.
At twice real time.
The normal equation and lstsq differed by at most 3×10⁻¹³ in the coefficients. Unscaled gradient descent took about 27 times longer than scaled and ended with a useless error.
The normal equation was the fastest, 1.2 ms versus 11.9 ms. Part of that lead is that I timed only solving the system, while LinearRegression also validates and centers the data. Nor does it make it the one to use: forming AᵀA squares the condition number of the data, and AᵀA’s reached 4.1×10⁷ here, from the mix of scales (pressure near 100, sines between −1 and 1) and the features that resemble each other. With this number it was enough; as it grows, that route loses precision before lstsq does.
Do you need to scale the features?
For least squares, no: scaled or unscaled gave exactly the same test error, 1.697 °C. For gradient descent it changes everything. The scikit-learn guide says it plainly: stochastic gradient descent is sensitive to feature scaling.
zoom on the scaled run, linear scale
- raw features
- scaled features
- least squares 4.56
Raw, the lowest error in 60 passes was 1.6×10²³; with tolerance 10⁻⁴ the fit stopped after 614 iterations and missed evaluation days by 9.8×10¹¹ °C on average. Scaled, it stopped after 25 iterations at 1.692 °C, close to least squares.
The reason is the step size. Pressure comes in kilopascals, between 87 and 104, humidity reaches 100, and the sine of the day of year goes from −1 to 1. A step that works for one feature is huge for the other, and the error bounces instead of falling. With SGDRegressor (tolerance 10⁻⁴, at most 1,000 passes, everything else default), the unscaled fit stopped after 614 iterations, far from the minimum, with a mean error of 9.8×10¹¹ °C on the evaluation period. Scaled, it stopped after 25 iterations with 1.692 °C. Another learning rate might have rescued the unscaled version; what the experiment shows is that this usual configuration does not.
When coefficients stop being trustworthy: collinearity
If two features carry almost the same information, the line can split the weight between them in many ways with a similar error. The variance inflation factor (VIF) measures how much each feature can be explained by the others: humidity reaches 33, dew point 23 and today’s max 15.
To see the effect, I fitted the line 200 times on random subsamples of 3,000 rows and looked at where each coefficient landed, with least squares and with Ridge, which penalizes large weights. It is a stability test, not an interval for the final model: the subsamples are small and ignore that rows come grouped by city and date.
- humidityVIF 33.217 % with the minority sign
- dew pointVIF 23.346 % with the minority sign
- max today (off scale: 4.69 / 2.82)VIF 15.4
- min todayVIF 8.410 % with the minority sign
- day of year (cos)VIF 6.0
- radiationVIF 5.923 % with the minority sign
- latitudeVIF 4.6
- pressureVIF 4.21 % with the minority sign
- wind N-SVIF 2.0
- wind speedVIF 1.8
- rain todayVIF 1.617 % with the minority sign
- wind E-WVIF 1.5
- day of year (sin)VIF 1.4
- rain yesterdayVIF 1.35 % with the minority sign
- pressure changeVIF 1.2
With least squares and 3,000 rows, dew point takes the minority sign in 46 % of subsamples; with Ridge, 2 features ever flip. With full-size subsamples (74,145 rows) only dew point still flips (28 %). On those 3,000-row fits Ridge missed by 1.79 °C versus 1.70 °C; trained on the full table, 1.698 versus 1.696 °C.
With least squares and 3,000 rows, the dew point came out with the minority sign in 46% of subsamples, radiation in 22.5%, humidity in 16.5% and the min temperature in 9.5%, plus features with coefficients near zero, such as today’s rain (16.5%). With Ridge, none of the high-VIF features flipped; only today’s and yesterday’s rain, with coefficients very close to zero, kept doing it (11% and 0.5%).
Two nuances change the reading. First, much of that instability comes from small samples: I repeated the test with 50 subsamples of full size, 74,145 rows, and only the dew point kept flipping (28%); every other feature always kept the same sign. Second, Ridge is not free: with α = 300, picked by hand and not by validation, those 3,000-row subsamples missed by 1.79 °C on average versus 1.70 °C for least squares. Trained on the full table, where that α weighs much less, Ridge landed at 1.698 °C, virtually the same as least squares (1.697).
The practical rule holds: if you are going to read a coefficient as “this feature raises the temperature”, first check whether its sign survives a resample of the real size.
A misplaced decimal point: outliers and Huber
Least squares squares every error, so a very distant value weighs a lot. To measure it, I multiplied a share of the training max temperatures by 10, as if someone had typed 243 instead of 24.3, and compared least squares with Huber regression: its loss is quadratic near zero and linear in the tails, with a threshold that depends on a scale the model estimates itself, and scikit-learn’s implementation adds a very small L2 penalty by default.
- least squares
- Huber
- 0 % · 0 rows · slope 0.93 vs 0.95 · MAE with 15 features 1.70 vs 1.69 °C
- 0.1 % · 74 rows · slope 0.95 vs 0.95 · MAE with 15 features 1.69 vs 1.69 °C
- 1 % · 741 rows · slope 1.00 vs 0.95 · MAE with 15 features 2.12 vs 1.69 °C
- 5 % · 3,707 rows · slope 1.39 vs 0.95 · MAE with 15 features 7.89 vs 1.69 °C
With 5 % of the rows corrupted, the least squares slope goes from 0.93 to 1.39; Huber stays at 0.95. With all 15 features, least squares error rises to 7.89 °C and Huber stays at 1.69 °C.
With 74 corrupted rows (0.1%) the error barely changed. With 741 (1%), the least squares error rose from 1.70 to 2.12 °C. With 3,707 (5%) it reached 7.89 °C, 4.7 times as much. Huber stayed at 1.69 °C in all four cases. It is a single realization, with one draw per level and deliberately extreme errors: it shows the difference in behavior, not an estimate of how much bad data NASA POWER contains.
What the line leaves unexplained: residuals
The 1.70 °C mean error is an average, and averages hide things. A residual is the actual max minus the predicted one; looking at them by group tells you where the line fails systematically.
by city: mean ± one standard deviation
- La Serena+0.32 ± 1.86
- Valparaíso−0.54 ± 1.07
- Santiago+0.15 ± 3.06
- Concepción+0.28 ± 2.43
- Temuco+0.49 ± 2.89
- Puerto Montt+0.16 ± 2.13
- Punta Arenas+0.05 ± 1.73
Overall mean +0.13 °C and standard deviation 2.28 °C; 4.1% of days miss by more than 5 °C. Monthly means only move between −0.08 and +0.38 °C: in the means, cities vary more than months, which does not rule out seasonal patterns in the spread. Adding the city as a feature lowered the error to 1.65 °C.
In Valparaíso the line predicts 0.54 °C too high on average, with little spread, and in Temuco 0.49 °C too low. There were two possible explanations: a city-specific effect latitude does not capture, or drift between periods (Temuco is the city warming the most and Valparaíso the one cooling). I tested it by adding the city as a feature. The mean error dropped from 1.70 to 1.65 °C and Valparaíso’s bias almost disappeared, from −0.54 to −0.11 °C, but Temuco stayed at +0.39 °C. In Valparaíso it was mostly the city; what remains in Temuco is consistent with drift over time, although this test does not isolate it.
The spread is not even either: 3.06 °C in Santiago versus 1.07 °C in Valparaíso. A single prediction interval for every city could end up poorly calibrated in some of them; its coverage per city would need checking before publishing it.
Linear regression or Random Forest
On the same table, a 200-tree forest missed by 1.50 °C and the line by 1.70 °C. The forest wins on accuracy. What that difference costs is measured in time and size.
At 30% of real time.
In real time. A thousand queries were not run: it is the one-row median times a thousand.
In this run the line trained about 1,100 times faster and answered about 14 times faster. The line has 15 coefficients and an intercept; the forest, 3,011,352 nodes, which are not comparable units of memory.
The line fits in one line of code and can be audited weight by weight. The forest captures relationships the line does not see, but its predictions stay within what it saw, while the line continues beyond the range. In the controlled case of the previous post that gave it the edge (1.86 °C versus 3.76 °C on summers no model had seen), but it is no guarantee: on the few evaluation days that exceeded their city’s training maximum, the line missed by more than the forest in all three cities where it happened.
Corrected on 16 September 2026. That controlled case merged the training and test years before splitting the seasons; redone with the periods kept apart, the figures went from 1.81 and 3.43 °C to 1.86 and 3.76 °C. The line keeps its edge.
Where a linear regression lives in a real system
- Trends and reports. When the question is “how much does this change per year”, the slope with its interval is the answer, as in the weather scoreboard I maintain, where I compare forecasts against what happened.
- Baseline. Before a complex model, a line trained in milliseconds tells you how much everything else adds. Here the forest improved on it by 0.2 °C.
- Where every microsecond counts. A prediction is a dot product: it can be computed in a database, a browser or a microcontroller without loading anything. The 0.26 ms I measured include going through scikit-learn; the bare dot product, which I did not time separately, should cost far less.
- Where the decision has to be explained. One weight per feature can be shown to whoever has to approve the model, as long as collinearity has been checked first.
When I would choose it: roughly linear relationships, a need to interpret, to answer very fast or to extend a trend carefully. When I would not: relationships with thresholds or strong interactions, data with gross errors left uncleaned (or, in that case, Huber instead of least squares), and when every tenth of a degree is worth more than simplicity.
Sources
- Legendre, A.-M. (1805). Nouvelles méthodes pour la détermination des orbites des comètes. Paris. Appendix “Sur la méthode des moindres quarrés”. Scan on the Internet Archive.
- Gauss, C. F. (1809). Theoria motus corporum coelestium in sectionibus conicis solem ambientium. Hamburg. Scan on the Internet Archive.
- Huber, P. J. (1964). “Robust Estimation of a Location Parameter”. The Annals of Mathematical Statistics, 35(1), 73–101. DOI 10.1214/aoms/1177703732.
- Hoerl, A. E. and Kennard, R. W. (1970). “Ridge Regression: Biased Estimation for Nonorthogonal Problems”. Technometrics, 12(1), 55–67. DOI 10.1080/00401706.1970.10488634.
- Falvey, M. and Garreaud, R. D. (2009). “Regional cooling in a warming world: Recent temperature trends in the southeast Pacific and along the west coast of subtropical South America (1979–2006)”. Journal of Geophysical Research: Atmospheres, 114(D4). DOI 10.1029/2008JD010519.
- scikit-learn,
LinearRegression(implemented withscipy.linalg.lstsq), the stochastic gradient descent guide, with its warning about feature scaling, andHuberRegressor. - NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.