17°
Portada del artículo: SVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 seconds
Machine learningAlgorithmsPythonData

SVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 seconds

What a support vector machine is, what the margin, C and gamma do, and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Unscaled, switching pressure from kPa to Pa dropped validation AUC from 0.835 to 0.555; on all 74,145 rows the RBF kernel took 51 s and, on the 2016-2026 test set, still ranked below a boosting model that trained in 2.6 s.

Efrain Garay 15 September 2026

Playing summary

An SVM with an RBF kernel, on unscaled variables, scored AUC 0.835 on a validation period at predicting tomorrow’s rain. I wrote the same atmospheric pressure in pascals instead of kilopascals, a thousand times larger, and the same model fell to 0.555, close to chance. With the variables standardized, it scored 0.845 in both cases.

This is the seventh part of the series, with the same NASA POWER daily data from seven Chilean cities I used for Random Forest, linear regression, K-Means, logistic regression, the decision tree and gradient boosting. The question is still whether it will rain at least 1 mm tomorrow. I trained on 1984-2012 and evaluated on the 2016-2026 test split, inside a container with 8 CPUs and scikit-learn 1.7.2. I chose C and gamma by fitting on 1984-2009 and scoring on 2010-2012. Because a single kernel fit on every row takes about a minute and choosing C and gamma takes dozens of fits, several tests use fixed subsamples of 10,000 or 20,000 days, and every comparison says which rows each model saw.

In 49 seconds and narrated: the same unscaled SVM collapses when one variable changes units; what the margin and support vectors are; how gamma bends the boundary until it memorizes; how fast training cost grows when using every row; and how to approximate the kernel with Nystroem plus a linear model. Muted by default: turn the sound on in the controls.Watch it in the reel viewer →

What is an SVM?

A support vector machine looks for a boundary between two classes, days followed by rain and days followed by dry weather, and among all possible boundaries picks the one that leaves the most room on each side. That room is the margin. Classes almost never separate cleanly, so some days are allowed inside the margin or on the wrong side, with a penalty set by the parameter C. This is the formulation by Corinna Cortes and Vladimir Vapnik (1995).

The days that sit on the margin, inside it or on the wrong side are the support vectors. Only they define the boundary: delete a day that is far away and the boundary does not change.

The margin rests on a few daysA linear SVM on two standardized variables (humidity and today's rain) over 200 training days, half with rain the next day. Solid line: the decision. Dashed: the margin. Rings: support vectors. The loop goes through C = 0.01, 1 and 100.
C = 0.01 · 164 SVC = 1 · 118 SVC = 100 · 117 SVhumidity (standardized)today's rain, log scale (standardized)
  • rain tomorrow
  • dry tomorrow
  • SV

With C = 0.01 the margin is 3.43 wide and 164 of the 200 days are support vectors; with C = 1, 1.51 and 118; with C = 100, 1.34 and 117. The days hold the line only if they sit on or inside the margin: on these two variables the classes overlap, so more than half of the days end up there. Scored on 2010-2012 with the same two variables, AUC goes from 0.802 to 0.814.

With two variables, humidity and today’s rain, and 200 training days (100 followed by rain and 100 not), C = 0.01 left a margin of 3.43 and 164 support vectors. With C = 1 the margin narrowed to 1.51 with 118; with C = 100, 1.34 with 117. On these two variables the classes overlap so much that more than half of the days stayed inside the margin even with a large C. On 2010-2012, with the same two variables, AUC went from 0.802 to 0.814.

The kernel bends the boundary

A straight boundary is not always enough. The kernel trick, which Boser, Guyon and Vapnik used in 1992 for maximum-margin classifiers, lets the SVM be computed as if the variables lived in a space with many more dimensions without building it. The RBF kernel measures how similar two days are from their distance, and gamma decides how fast that similarity drops.

The kernel bends the boundaryThe same 200 days and two variables with an RBF kernel, C = 1. The loop goes through gamma 0.03, 0.3, 3 and 30. Each label gives AUC on those 200 days and on all of 2010-2012.
gamma 0.03AUC train 0.844 · 2010-12 0.807gamma 0.3AUC train 0.828 · 2010-12 0.801gamma 3AUC train 0.877 · 2010-12 0.791gamma 30AUC train 0.943 · 2010-12 0.765humidity (standardized)today's rain, log scale (standardized)

With gamma 0.03 the boundary is almost straight: AUC 0.844 on the 200 days and 0.807 on 2010-2012. With gamma 30 it wraps around small groups of days: 0.943 on the 200 days it learned from and 0.765 on 2010-2012. A larger gamma makes each support vector's influence reach less far, so the boundary follows individual days instead of the trend.

With gamma 0.03 the boundary comes out almost straight: AUC 0.844 on the 200 days and 0.807 on 2010-2012. With gamma 30 it wraps around small groups of days: 0.943 on the days it learned from and 0.765 on 2010-2012. It is overfitting, like the unlimited decision tree, with a different shape.

Scaling changes the answer

If the kernel measures distances, units matter. I tested on 10,000 days of 1984-2009 and scored on 2010-2012.

Scaling changes the answer10,000 days of 1984-2009, AUC on 2010-2012. Same data in both tabs; only the unit of one column changes.
RBF, raw
0.835
RBF, StandardScaler
0.845
RBF, MinMaxScaler
0.852
LinearSVC, raw
0.851
LinearSVC, StandardScaler
0.850
0.50.9
RBF, raw
0.555
RBF, StandardScaler
0.845
RBF, MinMaxScaler
0.852
LinearSVC, raw
0.813
LinearSVC, StandardScaler
0.850
0.50.9

With pressure in kPa, the raw RBF SVM scored 0.835 and the standardized one 0.845. Writing the same pressure in Pa, a thousand times larger, sank the raw RBF to 0.555, close to chance: the distance between days became almost only a pressure difference. The scaled versions did not move (0.845 and 0.852), and the raw LinearSVC fell to 0.813.

With pressure in kPa, as NASA POWER delivers it, the unscaled RBF SVM scored 0.835, standardized 0.845 and with MinMaxScaler 0.852. No variable in this table dominates the distance by orders of magnitude (standard deviations range from 0.4 to 19), so the difference was small. With pressure in Pa the distance between two days became almost only their pressure difference, and the unscaled SVM fell to 0.555. The scaled versions did not move, and the unscaled linear SVM dropped from 0.851 to 0.813.

In the logistic regression post, scaling changed how many iterations the fit needed to converge. In an SVM, with or without a kernel, it changes the solution.

C and gamma are chosen on a grid

C and gamma are chosen on a grid10,000 days of 1984-2009. Each cell: AUC and share of training days kept as support vectors. Ring: chosen on 2010-2012.
C \ gamma0.00030.0010.0030.010.030.10.310,0000.8590.88538 %0.8590.89037 %0.8480.89636 %0.8380.92736 %0.7980.97334 %0.7721.00033 %0.7901.00042 %1,0000.8570.87940 %0.8600.88638 %0.8580.89137 %0.8470.90736 %0.8260.94736 %0.7750.99536 %0.7901.00042 %1000.8500.87441 %0.8580.88039 %0.8600.88738 %0.8510.89337 %0.8420.91937 %0.8100.97538 %0.7930.99943 %100.8410.86642 %0.8520.87541 %0.8580.88140 %0.8560.88638 %0.8490.89638 %0.8330.93939 %0.8040.99046 %10.8400.86047 %0.8410.86344 %0.8510.87442 %0.8570.88040 %0.8550.88539 %0.8420.90040 %0.8380.94748 %0.10.8390.85952 %0.8400.85950 %0.8410.86147 %0.8510.87144 %0.8530.87542 %0.8500.87943 %0.8480.89651 %

The best cell on 2010-2012 was C = 1,000 with gamma = 0.001: AUC 0.860 with 38% of the days as support vectors. Refit on 10,000 days of 1984-2012, it scored 0.865 on 2016-2026. The top-right corner (large C and gamma) reaches AUC 1.000 on its own training days and falls on 2010-2012. The chosen cell is not on the edge of the grid.

I tried 42 combinations of C and gamma on 10,000 days of 1984-2009. The best on 2010-2012 was C = 1,000 with gamma = 0.001, with AUC 0.860; C = 100 with gamma = 0.003 came within 0.0003, and the same choice came out when repeating the best row and column on 20,000 days. The first two grids I built left the choice on the edge, so I widened them until it no longer was. The corner of large C and gamma reached AUC 1.000 on its own training days and 0.790 on 2010-2012. Refit on 10,000 days of 1984-2012, the chosen combination scored 0.865 on 2016-2026.

Cost climbs faster than the rows

scikit-learn builds SVC on LIBSVM, which evaluates kernel similarities between pairs of examples during optimization. I timed the fit on one thread over nested subsamples, with the SVM at its defaults (C = 1, gamma = ‘scale’) and the other models on the same rows.

The cost of an RBF SVM climbs faster than the rowsFit time on one thread against training rows, both on log scales, same nested rows for every model.
0.001 s0.01 s0.1 s1.0 s10.0 s100 s2k5k10k20k40k74ktraining rows (log)predicted 47.3 s
  • SVM RBF
  • forest 200
  • boosting 594
  • LinearSVC
  • logistic

AUC on 2016-2026 by rows

2k0.8440.852
5k0.8530.863
10k0.8510.872
20k0.8500.877
40k0.8480.880
74k0.8490.883

Between 5,000 and 40,000 rows the SVM's fit time grew with a log-log slope of 2.02: doubling the rows multiplied the time by about 4.1. That slope predicted 47.3 s for all 74,145 rows; it took 50.8 s. On those same rows boosting trained in 2.6 s and logistic regression in 0.06 s.

Between 5,000 and 40,000 days, time grew with a log-log slope of 2.02: doubling the rows multiplied the time by four. That slope predicted 47 s for all 74,145 rows; it took 51 s. On those rows the SVM kept 27,707 support vectors, took 17 s to score the 27,256 test days and scored AUC 0.849 on 2016-2026, at its defaults rather than the grid’s C and gamma. Boosting trained in 2.6 s and reached 0.883.

Going from 2,000 to 74,145 days moved the SVM from 0.844 to 0.849; boosting, from 0.852 to 0.883, and the forest, from 0.866 to 0.882. Eight threads did not help: with 40,000 days it took 13.6 s on one and on eight, because LIBSVM trains on a single thread. Raising the kernel cache from 200 to 2,000 MB did not speed it up either (16.3 s).

Every prediction goes through every support vector

Every prediction goes through every support vectorAn RBF SVM scores a new day by comparing it with each support vector (48 drawn here). Table: 20,000 training rows, gamma 0.001, one thread. C was chosen earlier on the grid; test AUC per C is shown, not used to choose.
8,767 support vectors · C = 1
CSVms/daysizeAUC 2016-26
0.0110,7120.731.5 MB0.843
0.109,8830.711.4 MB0.843
18,7670.671.2 MB0.846
108,3880.661.2 MB0.857
1008,0660.661.1 MB0.862
1,0007,8130.641.1 MB0.865

With C = 0.01 the model kept 10,712 support vectors, 54% of the training rows, and took 0.73 ms per day; with C = 1,000, 7,813 and 0.64 ms. Serialized size follows the support vectors, from 1.5 MB to 1.1 MB. A logistic regression scores a day with one dot product.

With 20,000 days and gamma 0.001, raising C from 0.01 to 1,000 cut support vectors from 10,712 (54% of the days) to 7,813 (39%). Latency per day went from 0.73 to 0.64 ms and the serialized model from 1.5 to 1.1 MB, following the support vectors. In every case more than 98% of the support vectors sat at the upper bound of their penalty, that is, on or inside the margin, including the misclassified ones: they are the days that are hard to separate, not a selection of neat examples.

A linear SVM is not a logistic regression

Without a kernel, the SVM is a linear model with a different loss. On all 74,145 standardized rows, LinearSVC, which runs on LIBLINEAR and whose C I chose on 2010-2012, scored AUC 0.849 in 0.08 s and logistic regression 0.847 in 0.05 s. Of the 15 coefficients, 13 had the same sign in both models, and the three largest weights matched: today’s rain, the north-south wind component and latitude.

The trap is using SVC(kernel="linear"), which goes through LIBSVM. With 20,000 days it took 3.0 s against 0.02 s for LinearSVC, for a similar but not identical linear model: LinearSVC uses a different loss and penalizes the intercept. The time of SVC(kernel="linear") keeps climbing like the curve above.

An SVM does not return probabilities

An SVM returns a score, not a probability

decision_function on 2016-2026 · light: all days · dark: days with rain tomorrow

-3.506.6
mean predictedobserved rain0.373mean predictedobserved rain0.367mean predictedobserved rain0.352mean predictedobserved rain0.380

probability=True ran libsvm's internal 5-fold Platt scaling and multiplied training time from 5.8 to 31.8 s. It averaged 24 % rain against 20 % observed (log loss 0.373), and on 31 test days predict and predict_proba > 0.5 disagreed. A sigmoid fit on 2010-2012 gave log loss 0.367, an isotonic map 0.352, and logistic regression on the same rows 0.380.

decision_function is a signed score, proportional to the distance to the boundary, not a probability. probability=True adds Platt scaling with an internal five-fold cross-validation: on 20,000 days of 1984-2009 it took training from 5.8 to 31.8 s. It predicted 24.2% rain on average against 20.3% observed on 2016-2026, with log loss 0.373, and on 31 days of 2016-2026 predict and predict_proba greater than 0.5 disagreed, which the scikit-learn documentation warns about.

Calibrating separately on 2010-2012 worked better: on 2016-2026 a sigmoid gave log loss 0.367 and an isotonic calibration 0.352, both with 21.5% predicted rain. On the same rows, logistic regression gave 0.380 and boosting 0.343.

class_weight="balanced" gave more weight to rainy days. With 20,000 days, the unweighted SVM caught 49% of 2016-2026 rainy days and the balanced one 83%, with AUC 0.865 and 0.872. I then lowered the unweighted SVM’s threshold until, on 2010-2012, it caught as many as the balanced one (84%), and froze it: on 2016-2026 it caught 83% with a balanced accuracy of 0.793, almost the same as the balanced model (0.795). In this experiment, moving the threshold nearly reproduced the recall and balanced accuracy of class_weight, but not its AUC gain, from 0.865 to 0.872.

A bridge for many rows: approximate the kernel

If the exact kernel cannot see every row, it can be approximated. Nystroem, the idea of Williams and Seeger (2001), builds new variables from a sample of days; the random Fourier features of Rahimi and Recht (2007) approximate the RBF with sines and cosines. After that, a linear model is enough.

A bridge for many rows: approximate the kernelAUC on 2016-2026 against fit time on one thread (log). Exact RBF SVM on 10,000 and 20,000 rows; Nystroem and random Fourier features with 100 to 3,000 components plus LinearSVC on all 74,145 rows.
1101001,0000.8600.8750.890fit time, one thread (s, log)10k20k1003001000300010030010003000
  • exact RBF
  • Nystroem + LinearSVC
  • random Fourier + LinearSVC

fit time with 8 threads

Nystroem 100: 4.4 → 4.3 sNystroem 300: 24.9 → 23.1 sNystroem 1,000: 66.0 → 84.3 sNystroem 3,000: 269.1 → 197.8 s

The exact kernel on 20,000 rows scored 0.865 in 5.9 s. Nystroem with 300 components on all 74,145 rows reached 0.879 in 24.9 s, and the best Nystroem, 1,000 components, 0.879. The approximations on 74,145 rows ranked better than the exact kernel on 20,000. With 8 threads Nystroem 1,000 took 84.3 s.

With the same gamma, the linear model’s C chosen on 2010-2012 for each size (3,000 components reuse the 1,000 choice) and the sizes 100, 300, 1,000 and 3,000 that I fixed before measuring and report as diagnostics, Nystroem with 300 components on all 74,145 rows reached AUC 0.879 on 2016-2026 in 25 s, more than the exact SVM on 20,000 days (0.865 in 5.9 s). Going to 1,000 components improved AUC by only 0.0005 and 3,000 added nothing, at a cost of 66 and 269 s. With random Fourier features, 3,000 components reached 0.880 in 249 s. In this comparison, the approximations trained on 74,145 rows ranked better than the exact kernel trained on 20,000; the design does not allow attributing the difference only to the number of rows. Among the sizes whose C I tuned, except for Nystroem with 100 components, C = 1,000 won, the top value I tried, and validation was still creeping up there: a larger C might gain a little more.

For numbers: SVR

The same idea works for predicting tomorrow’s max temperature. An SVR with an RBF kernel, with gamma, C and epsilon chosen on 2010-2012 (gamma 0.01, C = 1,000 at the top value I tried, and epsilon 0.5 °C), trained on 10,000 days, missed by 1.49 °C on average on 2016-2026; linear regression on all rows, by 1.70 °C, and boosting, by 1.44 °C.

In the series’ extrapolation test, training on Santiago’s April to September months and predicting the 2016-2026 summers, the SVR predicted up to 35.73 °C, above the 32.97 °C of the warmest day it saw, and missed by 2.43 °C on average. Boosting, which never went past 29.56 °C, missed by 5.48 °C; linear regression, by 1.72 °C. In this seasonal cut the SVR went past the maximum it saw and boosting did not; that does not show that an SVR extrapolates in general. It is a seasonal cut, not a clean out-of-range test: summer changes several variables at once.

SVM, boosting, forest, Nystroem and logistic regression

The same 20,000 rows: training on one threadOn one thread, after a warm-up fit: median of 3 fits for logistic regression, boosting and Nystroem; a single measurement for the SVM and the forest. AUC measured on the 2016-2026 test set.
Logistic regression · AUC 0.8470.011 s
Boosting, 594 rounds · AUC 0.8771.18 s
SVM RBF · AUC 0.8655.436 s
Forest, 200 trees · AUC 0.8785.776 s
Nystroem 1,000 + LinearSVC · AUC 0.87715.058 s

At half of real time.

Scoring one day took 0.64 ms for the SVM, 0.39 ms for logistic regression and 3.2 ms for boosting; serialized, the SVM takes 1.1 MB, boosting 2.1 MB and the forest 34 MB.

On the same 20,000 rows, I resampled the city-year blocks of 2016-2026 2,000 times to compare AUC in pairs. The SVM came out below boosting by 0.012 (95% interval: −0.016 to −0.008) and below Nystroem with a linear model by 0.013; it came out above logistic regression by 0.018. None of the three intervals crosses zero.

Where an SVM lives in a real system

Where an SVM livesThe scaler travels inside the artifact, and the support vectors set the cost of every prediction.
Where an SVM livestable20,000 rowsscalerinside the artifacttrainkeeps the SVsartifact1.1 MBbatch scoring27,256 days · 4.85 sonline API1 day · 0.64 msmany rows: Nystroem+ linear model

Measured on one thread with the same 20,000 rows: the RBF SVM trains in 5.44 s, takes 1.1 MB serialized, scores one day in 0.64 ms and the 27,256 test days in 4.85 s (AUC 0.865). Nystroem with 1,000 components plus LinearSVC: 15.06 s, 0.72 ms, AUC 0.877. Boosting: 1.18 s, 3.16 ms, AUC 0.877. Logistic regression: 0.011 s, AUC 0.847.

  • Medium tables of numeric variables. Thousands to a few tens of thousands of rows, where the quadratic cost still fits and the boundary is not linear.
  • The scaler inside the artifact. The StandardScaler has to travel with the model, in the same pipeline. If a source changes units, the SVM changes its answer without warning.
  • An API with few support vectors. Latency per row grows with them: here 0.64 ms with 7,813.
  • With many rows, Nystroem plus a linear model. Here it did better than the exact kernel on a subsample.

When I would choose it: for medium tables where a smooth boundary is enough and scaling can be handled carefully. When not: with hundreds of thousands of rows, or when C and gamma must be searched over tens of thousands, where training and prediction cost grow with the support vectors; when probabilities are needed without separate calibration; or when variables have units that may change. On this data a boosting model scored a higher test AUC, trained faster and had a lower log loss.

Sources

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.