17°
Portada del artículo: Linear regression explained: the line that measured the heat in Chile, and what breaks it
Machine learningAlgorithmsPythonData

Linear regression explained: the line that measured the heat in Chile, and what breaks it

What linear regression is, how its coefficients are computed and when it stops working, measured on daily data from seven Chilean cities. The max temperature rises 0.33 °C per decade in Temuco and falls in Valparaíso, unscaled gradient descent does not reach a useful solution, 5% of rows with a misplaced decimal point multiply the error by 4.7, and the line trains about 1,100 times faster than a random forest.

Efrain Garay 16 September 2026

Playing summary

In the Random Forest post there was an awkward moment for the forest: trained on Santiago’s cold months, it could not predict the summer, and a plain linear regression missed by a little less than half as much. The line kept rising because that is the only thing it knows how to do.

This time the line is the star. I used the same NASA POWER daily data for seven Chilean cities, from La Serena to Punta Arenas, and measured inside a container with CPU and memory limits what usually stays in theory: how the coefficients are computed, what happens if the features are not scaled, when the coefficients stop being trustworthy, what a mistyped value does, and how much you pay against a forest.

In 57 seconds and narrated: on Santiago summers no model had seen, the line missed by 1.86 °C and the random forest by 3.76 °C; the yearly max rises 0.33 °C per decade in Temuco and falls in Valparaíso; unscaled, gradient descent ends with an error of 9.8×10¹¹ °C; with 3,000 rows, the dew point comes out with the minority sign in 46% of samples; and with 5% of rows carrying a misplaced decimal point, least squares rises to 7.89 °C while Huber stays at 1.69 °C. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

What is linear regression?

Linear regression predicts a number as a weighted sum of features plus a constant. With one feature it is a line; with fifteen it is the same thing in fifteen dimensions. The weights are chosen so the sum of squared errors is as small as possible, the least squares method Legendre published in 1805 and Gauss justified in 1809.

Here I use it to predict tomorrow’s max temperature from fifteen features of the day: temperatures, rain, wind, radiation, humidity, pressure, dew point, the day of the year and the city’s latitude.

A linear regression is a weighted sumEdge width is the standardized coefficient (mean of 200 resamples); green adds, violet subtracts, dashed grey flips sign in more than 20% of resamples. Below: three ways to compute almost the same weights.
A linear regression is a weighted summax today+4.69day of year (cos)+0.95latitude+0.86pressure change+0.44day of year (sin)+0.35pressure−0.27wind N-S−0.25humidity−0.23min today−0.18wind E-W+0.18wind speed−0.12radiation+0.08rain yesterday−0.06rain today+0.05dew point+0.02Σweighted summax tomorrowhow the weights are computedleast squares (lstsq)12 msnormal equation1.2 msSGD, scaled89 ms

Largest weights: max today +4.69 and day of year (cos) +0.95 °C per standard deviation. Least squares and the normal equation give the same coefficients down to 3×10⁻¹³.

Read this way, a linear regression is a single neuron with identity activation, trained with squared error. Each weight says how much its feature adds or subtracts, which is why it is among the easiest models to inspect, as long as the weights can be trusted. Further down you will see when they cannot.

The slope: how much temperature rises in each city

The oldest use of regression is one x and one y. I averaged the daily max temperature for each year between 1981 and 2025 and fitted a line against the year. The slope is the trend, and the 95% interval summarizes its uncertainty under the assumptions of the conventional fit.

The slope, city by cityYearly mean of the daily max temperature, 1981–2025, with the least squares line and the band of slopes inside the 95% interval.
La Serena+0.043 °C/decade
[-0.060; 0.145] · R² 0.02 · interval crosses zero
21°23°19812025
Valparaíso−0.095 °C/decade
[-0.155; -0.035] · R² 0.19
16°17°19812025
Santiago+0.193 °C/decade
[0.035; 0.350] · R² 0.12
22°24°19812025
Concepción+0.067 °C/decade
[-0.049; 0.183] · R² 0.03 · interval crosses zero
17°19°19812025
Temuco+0.332 °C/decade
[0.161; 0.503] · R² 0.26
16°19°19812025
Puerto Montt+0.167 °C/decade
[0.042; 0.292] · R² 0.14
14°16°19812025
Punta Arenas+0.017 °C/decade
[-0.072; 0.107] · R² 0.00 · interval crosses zero
8°9°19812025

A dashed line means the conventional 95% interval includes zero. Intervals assume independent years and are not corrected for autocorrelation or for testing seven cities; with a Bonferroni threshold only the conventional p-values of Temuco and Valparaíso stay below it. The source switches to GEOS-IT for the most recent months.

Temuco rises 0.332 °C per decade (interval between 0.161 and 0.503), Santiago 0.193 and Puerto Montt 0.167. Valparaíso falls 0.095 °C per decade. La Serena, Concepción and Punta Arenas have intervals that include zero.

Those intervals assume years are independent, and they are not entirely: in Temuco one year’s residual resembles the next (autocorrelation of 0.37), which makes the interval narrower than it should be. I also tested seven cities at once; applying the stricter Bonferroni threshold, only Temuco’s and Valparaíso’s conventional p-values stay below that threshold, Temuco’s with the autocorrelation caveat. Santiago and Puerto Montt stay as hints, not conclusions.

Valparaíso’s negative slope has the same sign as the coastal cooling Falvey and Garreaud described in 2009 for central and northern Chile: −0.2 °C per decade at coastal stations between 1979 and 2006, while the Andes warmed. But my data are different (the daily max of a reanalysis, not stations), cover another period, and Valparaíso’s grid cell probably includes ocean. And of the three coastal cities in that latitude range, La Serena (+0.043) and Concepción (+0.067) show no such cooling. The sign matches in one; that neither confirms the pattern nor rules out a data product artifact.

One more caveat: the NASA POWER series combines MERRA-2 for the historical record with GEOS-IT for recent months, so part of the final trend may come from the source change.

How are the coefficients computed?

To predict tomorrow’s max temperature I trained on 74,145 city-day rows from 1984 to 2012 and evaluated on 27,256 rows from 2016 to August 30, 2026. I left out 2013 to 2015 so training and test are not back to back in time. I use that same evaluation period for every comparison in the post, so the figures describe that period; they are not a final test kept untouched. It is the same table as the forest post, with the same caveats about the source.

There are three ways to reach the weights. The normal equation builds the 16 by 16 matrix AᵀA and solves that system. scikit-learn’s LinearRegression uses scipy.linalg.lstsq, which works on the data without forming that matrix and is more numerically stable. And gradient descent moves the weights little by little in the direction that lowers the error.

How long each way of computing almost the same weights takesMedian of five fits (three for gradient descent) on 74,145 rows and 15 features, in the container.
Normal equation (solve only)0.0012 s
lstsq (scikit-learn)0.0119 s
Gradient descent, scaled0.089 s

Forty times slower than real time.

Gradient descent, scaled0.089 s
Gradient descent, unscaled2.359 s

At twice real time.

The normal equation and lstsq differed by at most 3×10⁻¹³ in the coefficients. Unscaled gradient descent took about 27 times longer than scaled and ended with a useless error.

The normal equation was the fastest, 1.2 ms versus 11.9 ms. Part of that lead is that I timed only solving the system, while LinearRegression also validates and centers the data. Nor does it make it the one to use: forming AᵀA squares the condition number of the data, and AᵀA’s reached 4.1×10⁷ here, from the mix of scales (pressure near 100, sines between −1 and 1) and the features that resemble each other. With this number it was enough; as it grows, that route loses precision before lstsq does.

Do you need to scale the features?

For least squares, no: scaled or unscaled gave exactly the same test error, 1.697 °C. For gradient descent it changes everything. The scikit-learn guide says it plainly: stochastic gradient descent is sensitive to feature scaling.

Gradient descent, with and without scalingTraining mean squared error after each pass over the 74,145 rows, log scale, from separate 60-pass partial_fit runs. The dashed line is the closed-form least squares solution; the iteration and error figures below come from the fit() runs.
Gradient descent, with and without scaling10⁰10⁵10¹⁰10¹⁵10²⁰10²⁵1204060epochsraw featuresscaled features

zoom on the scaled run, linear scale

least squares 4.564.5624.5744.5861204060
  • raw features
  • scaled features
  • least squares 4.56

Raw, the lowest error in 60 passes was 1.6×10²³; with tolerance 10⁻⁴ the fit stopped after 614 iterations and missed evaluation days by 9.8×10¹¹ °C on average. Scaled, it stopped after 25 iterations at 1.692 °C, close to least squares.

The reason is the step size. Pressure comes in kilopascals, between 87 and 104, humidity reaches 100, and the sine of the day of year goes from −1 to 1. A step that works for one feature is huge for the other, and the error bounces instead of falling. With SGDRegressor (tolerance 10⁻⁴, at most 1,000 passes, everything else default), the unscaled fit stopped after 614 iterations, far from the minimum, with a mean error of 9.8×10¹¹ °C on the evaluation period. Scaled, it stopped after 25 iterations with 1.692 °C. Another learning rate might have rescued the unscaled version; what the experiment shows is that this usual configuration does not.

When coefficients stop being trustworthy: collinearity

If two features carry almost the same information, the line can split the weight between them in many ways with a similar error. The variance inflation factor (VIF) measures how much each feature can be explained by the others: humidity reaches 33, dew point 23 and today’s max 15.

To see the effect, I fitted the line 200 times on random subsamples of 3,000 rows and looked at where each coefficient landed, with least squares and with Ridge, which penalizes large weights. It is a stability test, not an interval for the final model: the subsamples are small and ignore that rows come grouped by city and date.

The coefficient whose sign flipsStandardized coefficient across 200 random subsamples of 3,000 rows: bar from the 5th to the 95th percentile, dot at the mean. A stability test, not an interval for the final model.
least squaresRidge (α = 300, picked by hand)
  • humidityVIF 33.217 % with the minority sign
  • dew pointVIF 23.346 % with the minority sign
  • max today (off scale: 4.69 / 2.82)VIF 15.4
  • min todayVIF 8.410 % with the minority sign
  • day of year (cos)VIF 6.0
  • radiationVIF 5.923 % with the minority sign
  • latitudeVIF 4.6
  • pressureVIF 4.21 % with the minority sign
  • wind N-SVIF 2.0
  • wind speedVIF 1.8
  • rain todayVIF 1.617 % with the minority sign
  • wind E-WVIF 1.5
  • day of year (sin)VIF 1.4
  • rain yesterdayVIF 1.35 % with the minority sign
  • pressure changeVIF 1.2

With least squares and 3,000 rows, dew point takes the minority sign in 46 % of subsamples; with Ridge, 2 features ever flip. With full-size subsamples (74,145 rows) only dew point still flips (28 %). On those 3,000-row fits Ridge missed by 1.79 °C versus 1.70 °C; trained on the full table, 1.698 versus 1.696 °C.

With least squares and 3,000 rows, the dew point came out with the minority sign in 46% of subsamples, radiation in 22.5%, humidity in 16.5% and the min temperature in 9.5%, plus features with coefficients near zero, such as today’s rain (16.5%). With Ridge, none of the high-VIF features flipped; only today’s and yesterday’s rain, with coefficients very close to zero, kept doing it (11% and 0.5%).

Two nuances change the reading. First, much of that instability comes from small samples: I repeated the test with 50 subsamples of full size, 74,145 rows, and only the dew point kept flipping (28%); every other feature always kept the same sign. Second, Ridge is not free: with α = 300, picked by hand and not by validation, those 3,000-row subsamples missed by 1.79 °C on average versus 1.70 °C for least squares. Trained on the full table, where that α weighs much less, Ridge landed at 1.698 °C, virtually the same as least squares (1.697).

The practical rule holds: if you are going to read a coefficient as “this feature raises the temperature”, first check whether its sign survives a resample of the real size.

A misplaced decimal point: outliers and Huber

Least squares squares every error, so a very distant value weighs a lot. To measure it, I multiplied a share of the training max temperatures by 10, as if someone had typed 243 instead of 24.3, and compared least squares with Huber regression: its loss is quadratic near zero and linear in the tails, with a threshold that depends on a scale the model estimates itself, and scikit-learn’s implementation adds a very small L2 penalty by default.

The line that runs after the bad dataToday's max against tomorrow's max. Background: a sample of 300 clean rows. Lines: fitted on all 74,145 training rows after multiplying the indicated share of targets by 10.
The line that runs after the bad data0055101015152020252530303535today's max (°C)tomorrow's max (°C)
  • least squares
  • Huber
  • 0 % · 0 rows · slope 0.93 vs 0.95 · MAE with 15 features 1.70 vs 1.69 °C
  • 0.1 % · 74 rows · slope 0.95 vs 0.95 · MAE with 15 features 1.69 vs 1.69 °C
  • 1 % · 741 rows · slope 1.00 vs 0.95 · MAE with 15 features 2.12 vs 1.69 °C
  • 5 % · 3,707 rows · slope 1.39 vs 0.95 · MAE with 15 features 7.89 vs 1.69 °C

With 5 % of the rows corrupted, the least squares slope goes from 0.93 to 1.39; Huber stays at 0.95. With all 15 features, least squares error rises to 7.89 °C and Huber stays at 1.69 °C.

With 74 corrupted rows (0.1%) the error barely changed. With 741 (1%), the least squares error rose from 1.70 to 2.12 °C. With 3,707 (5%) it reached 7.89 °C, 4.7 times as much. Huber stayed at 1.69 °C in all four cases. It is a single realization, with one draw per level and deliberately extreme errors: it shows the difference in behavior, not an estimate of how much bad data NASA POWER contains.

What the line leaves unexplained: residuals

The 1.70 °C mean error is an average, and averages hide things. A residual is the actual max minus the predicted one; looking at them by group tells you where the line fails systematically.

What the line leaves unexplainedResiduals (actual minus predicted) over the 27,256 evaluation city-day rows; positive means the line predicted too cold.
What the line leaves unexplainedmean +0.13-10-50510residual (°C)

by city: mean ± one standard deviation

  • La Serena+0.32 ± 1.86
  • Valparaíso−0.54 ± 1.07
  • Santiago+0.15 ± 3.06
  • Concepción+0.28 ± 2.43
  • Temuco+0.49 ± 2.89
  • Puerto Montt+0.16 ± 2.13
  • Punta Arenas+0.05 ± 1.73

Overall mean +0.13 °C and standard deviation 2.28 °C; 4.1% of days miss by more than 5 °C. Monthly means only move between −0.08 and +0.38 °C: in the means, cities vary more than months, which does not rule out seasonal patterns in the spread. Adding the city as a feature lowered the error to 1.65 °C.

In Valparaíso the line predicts 0.54 °C too high on average, with little spread, and in Temuco 0.49 °C too low. There were two possible explanations: a city-specific effect latitude does not capture, or drift between periods (Temuco is the city warming the most and Valparaíso the one cooling). I tested it by adding the city as a feature. The mean error dropped from 1.70 to 1.65 °C and Valparaíso’s bias almost disappeared, from −0.54 to −0.11 °C, but Temuco stayed at +0.39 °C. In Valparaíso it was mostly the city; what remains in Temuco is consistent with drift over time, although this test does not isolate it.

The spread is not even either: 3.06 °C in Santiago versus 1.07 °C in Valparaíso. A single prediction interval for every city could end up poorly calibrated in some of them; its coverage per city would need checking before publishing it.

Linear regression or Random Forest

On the same table, a 200-tree forest missed by 1.50 °C and the line by 1.70 °C. The forest wins on accuracy. What that difference costs is measured in time and size.

Line against forest: training and answeringMedian of five fits of the line and a single measurement of the 200-tree forest on 8 CPUs; one-row latency on one thread.
Linear regression0.012 s
Random Forest, 200 trees13.52 s

At 30% of real time.

Linear regression0.26 s
Random Forest, 200 trees3.63 s

In real time. A thousand queries were not run: it is the one-row median times a thousand.

In this run the line trained about 1,100 times faster and answered about 14 times faster. The line has 15 coefficients and an intercept; the forest, 3,011,352 nodes, which are not comparable units of memory.

The line fits in one line of code and can be audited weight by weight. The forest captures relationships the line does not see, but its predictions stay within what it saw, while the line continues beyond the range. In the controlled case of the previous post that gave it the edge (1.86 °C versus 3.76 °C on summers no model had seen), but it is no guarantee: on the few evaluation days that exceeded their city’s training maximum, the line missed by more than the forest in all three cities where it happened.

Corrected on 16 September 2026. That controlled case merged the training and test years before splitting the seasons; redone with the periods kept apart, the figures went from 1.81 and 3.43 °C to 1.86 and 3.76 °C. The line keeps its edge.

Where a linear regression lives in a real system

  • Trends and reports. When the question is “how much does this change per year”, the slope with its interval is the answer, as in the weather scoreboard I maintain, where I compare forecasts against what happened.
  • Baseline. Before a complex model, a line trained in milliseconds tells you how much everything else adds. Here the forest improved on it by 0.2 °C.
  • Where every microsecond counts. A prediction is a dot product: it can be computed in a database, a browser or a microcontroller without loading anything. The 0.26 ms I measured include going through scikit-learn; the bare dot product, which I did not time separately, should cost far less.
  • Where the decision has to be explained. One weight per feature can be shown to whoever has to approve the model, as long as collinearity has been checked first.

When I would choose it: roughly linear relationships, a need to interpret, to answer very fast or to extend a trend carefully. When I would not: relationships with thresholds or strong interactions, data with gross errors left uncleaned (or, in that case, Huber instead of least squares), and when every tenth of a degree is worth more than simplicity.

Sources

  • Legendre, A.-M. (1805). Nouvelles méthodes pour la détermination des orbites des comètes. Paris. Appendix “Sur la méthode des moindres quarrés”. Scan on the Internet Archive.
  • Gauss, C. F. (1809). Theoria motus corporum coelestium in sectionibus conicis solem ambientium. Hamburg. Scan on the Internet Archive.
  • Huber, P. J. (1964). “Robust Estimation of a Location Parameter”. The Annals of Mathematical Statistics, 35(1), 73–101. DOI 10.1214/aoms/1177703732.
  • Hoerl, A. E. and Kennard, R. W. (1970). “Ridge Regression: Biased Estimation for Nonorthogonal Problems”. Technometrics, 12(1), 55–67. DOI 10.1080/00401706.1970.10488634.
  • Falvey, M. and Garreaud, R. D. (2009). “Regional cooling in a warming world: Recent temperature trends in the southeast Pacific and along the west coast of subtropical South America (1979–2006)”. Journal of Geophysical Research: Atmospheres, 114(D4). DOI 10.1029/2008JD010519.
  • scikit-learn, LinearRegression (implemented with scipy.linalg.lstsq), the stochastic gradient descent guide, with its warning about feature scaling, and HuberRegressor.
  • NASA POWER, Daily API and data sources.

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.