11°

~/topics/machine-learning

Machine learning algorithms, measured

Eleven algorithms on the same weather, the same split and the same container.

All topics

Most explanations of these algorithms use toy data that never shows where they break. Here all eleven run on the same table — daily weather from seven Chilean cities since 1984 — with the same fit through 2009, the same 2010-2012 validation and the same test period from 2016, inside a container with fixed CPU and memory. That makes the numbers comparable, and it makes the negative results publishable: the Transformer landed below three simpler models, and the forest cannot extrapolate. One caveat about that test period: no model used it to pick its parameters, but after eleven pieces I have read it eleven times, so it is better called a frozen historical benchmark than an untouched holdout. A "which algorithm wins" conclusion would need years nobody has looked at yet.

12 pieces

Portada: How file compression actually works today (and how I beat zlib by 2.86% with graph theory)ArticleHow file compression actually works today (and how I beat zlib by 2.86% with graph theory)I translated zlib from C to Rust with c2rust, proved the translation is bit-exact across more than 15,000 cases, then replaced zlib's heuristic with a shortest path in a graph. Result: 2.86% smaller than zlib -9, verified byte for byte against the real C implementation. Along the way: how lossless compression actually works today, from LZ77 and Huffman to why zopfli and PPMd don't break Shannon's limit either.#Compression#Rust#CPortada: Neural network explained: it matched boosting with 2,177 weights, and changing the seed moved its AUC by 0.004TutorialNeural network explained: it matched boosting with 2,177 weights, and changing the seed moved its AUC by 0.004What a neural network is, what each layer does and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. One hidden layer of 128 units, 2,177 weights and 81 KB, reached AUC 0.883 on the 2016-2026 test period: the same figure as a 2.1 MB boosting model, at 0.40 ms per row against 3.09. Training it cost five times more, and repeating the training with only the seed changed moved the validation AUC between 0.881 and 0.884.#Machine learning#Algorithms#PythonPortada: Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwardsTutorialTransformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwardsWhat a Transformer is, what attention does and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. The chosen network has 9,633 parameters and 49 KB, and landed below boosting, the forest and a one-layer network. The ranking is not the lesson: among the six architectures that competed, the validation order inverted on the test period.#Machine learning#Algorithms#PythonPortada: KNN explained: training took 0.007 s, predicting 5.5 s, and with k = 1 it memorized what it sawTutorialKNN explained: training took 0.007 s, predicting 5.5 s, and with k = 1 it memorized what it sawWhat k-nearest neighbors is, why you must scale, how to choose k and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Fitting on the 74,145 rows took 0.007 s and scoring the 2016-2026 test period 5.5 s on one thread; with k = 1 it scored AUC 1.000 on its own training days and 0.69 on validation. With k = 100 it reached 0.870, below a boosting model and above logistic regression.#Machine learning#Algorithms#PythonPortada: Naive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edgesTutorialNaive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edgesWhat Naive Bayes is, what the independence assumption means and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Fitting on 74,145 rows took 0.008 s and the model weighs 1.3 KB, at AUC 0.822; on binned variables it reached 0.844, with no conclusive difference from logistic regression. 62 % of its probabilities fell below 0.01 or above 0.99, and calibrating it took log loss from 1.09 to 0.39.#Machine learning#Algorithms#PythonPortada: SVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 secondsTutorialSVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 secondsWhat a support vector machine is, what the margin, C and gamma do, and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Unscaled, switching pressure from kPa to Pa dropped validation AUC from 0.835 to 0.555; on all 74,145 rows the RBF kernel took 51 s and, on the 2016-2026 test set, still ranked below a boosting model that trained in 2.6 s.#Machine learning#Algorithms#PythonPortada: Decision tree explained: 15 nodes you can read, 8,524 leaves that memorizeTutorialDecision tree explained: 15 nodes you can read, 8,524 leaves that memorizeWhat a decision tree is, how it picks its questions and when it memorizes, measured on daily weather from seven Chilean cities between 1984 and 2026. With no depth limit it got everything it saw right and dropped to AUC 0.689 on new days; at depth 7, chosen on validation, it reached 0.862.#Machine learning#Algorithms#PythonPortada: Gradient boosting explained: 594 chained trees that train 9 times faster than a forestTutorialGradient boosting explained: 594 chained trees that train 9 times faster than a forestWhat gradient boosting is, how each tree corrects the previous ones and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. With rate 1.0 it destabilized and after 3,000 rounds ended at AUC 0.681; with rate 0.03 and 594 rounds chosen on validation it reached 0.883, with no conclusive difference from a 200-tree forest.#Machine learning#Algorithms#PythonPortada: K-Means explained: I asked Chile's weather for four kinds of day and it found SantiagoTutorialK-Means explained: I asked Chile's weather for four kinds of day and it found SantiagoWhat K-Means is, how Lloyd's algorithm works and when it misleads, measured on daily weather from seven Chilean cities between 1984 and 2026. Unscaled, humidity accounts for 76% of the separation; with k = 8, only 18% of random starts end within 0.1% of the best inertia found; and no criterion agrees on how many groups there are.#Machine learning#Algorithms#PythonPortada: Random Forest explained: how a forest of trees votes, measured on 42 years of Chilean weatherTutorialRandom Forest explained: how a forest of trees votes, measured on 42 years of Chilean weatherWhat Random Forest is, how bootstrap sampling and feature subsampling work, and why its trees train in parallel instead of in a chain. I measured it predicting rain in seven Chilean cities: 7.5 times faster on 8 CPUs, a plateau past 100 trees, a noise column that the default importance ranks above real features, and a ceiling when extrapolating temperatures.#Machine learning#Algorithms#PythonPortada: Linear regression explained: the line that measured the heat in Chile, and what breaks itTutorialLinear regression explained: the line that measured the heat in Chile, and what breaks itWhat linear regression is, how its coefficients are computed and when it stops working, measured on daily data from seven Chilean cities. The max temperature rises 0.33 °C per decade in Temuco and falls in Valparaíso, unscaled gradient descent does not reach a useful solution, 5% of rows with a misplaced decimal point multiply the error by 4.7, and the line trains about 1,100 times faster than a random forest.#Machine learning#Algorithms#PythonPortada: Logistic regression explained: 16 numbers to decide whether it rains tomorrowTutorialLogistic regression explained: 16 numbers to decide whether it rains tomorrowWhat logistic regression is, how it turns a sum into a probability and when it misleads, measured on daily weather from seven Chilean cities between 1984 and 2026. A straight line gives negative probabilities on 9% of days; at a 0.5 threshold logistic regression catches only half of the rain; and it promises more rain than falls.#Machine learning#Algorithms#Python

All writing