18°
← back to the blog

Reels

40 · swipe or use the arrows

  1. 59 s

    01 / 40

    A model that does not write, it only decides: I had Intern-Decision predict the rain without training it

    In 59 seconds, with narration: an oracle that does not speak, it only raises a card with a number. What broke when installing it, the 19.9 ms per query, the AUC of 0.823 predicting rain without training, and the 0.62 ceiling on its probabilities. Every figure is a measured one. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  2. 1:18

    03 / 40

    How file compression actually works today (and how I beat zlib by 2.86% with graph theory)

    In 78 seconds, narrated: DEFLATE's parse as a shortest path in a graph (44 bits from zlib's real output vs 36 bits from the brute-force-verified optimum in the example), the Huffman tree that iterates (-8.5% on a real file), and the real result across 29 files: -2.86% against zlib, verified byte for byte. Muted by default: turn sound on in the controls.

    Read the articleDirect link

  3. 1:12

    04 / 40

    Rust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly three times more, and mv between disks loses the dates

    In 72 seconds, narrated: Rust Coreutils (uutils) 0.12 against GNU coreutils 9.12. The 653 in its notes is the best result per test out of six CI runs, and it did not reproduce in my container; in cp, mv, rm, install, ls, sort and du only 3 tests fail. On speed, uutils is slower in 91 of 157 cells (starting a process costs almost three times more), but wins in 45. And mv across disks loses the modification time. Muted by default: turn sound on in the controls.

    Read the articleDirect link

  4. 1:15

    05 / 40

    Dragonfly 2.0 vs Redis 8 and Valkey 9: with 2 CPUs it wins big, with 4 it depends on how you configure the others

    In 75 seconds, narrated: Dragonfly 2.0, Redis 8 and Valkey 9 with the same CPU budget. With 2 CPUs and no pipelining Dragonfly beats stock Redis (342,601 against 172,081 operations per second with io_uring) and uses 8 to 15 percent less memory per key, but with 4 CPUs and no pipelining the bench cannot rank the three, and with pipeline 16 Redis and Valkey with io-threads win (4.21 and 4.04 million against 2.73). One workload, closed-loop bench. Muted by default: turn sound on in the controls.

    Read the articleDirect link

  5. 1:07

    06 / 40

    CUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thing

    In 67 seconds and narrated: the same Mandelbrot kernel in cuda-oxide, CUDA written in Rust, and in CUDA C++ on an RTX 4070 Ti SUPER. With default options Rust comes out 3.2% faster, but the two images differ in 27,510 pixels because each toolchain fuses different operations into FMA. With FMA off the outputs match bit for bit and Rust is 7% slower; the cause of the 3% is not established. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  6. 1:09

    07 / 40

    Wild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneck

    In 69 seconds and narrated: Wild against mold linking ripgrep and cargo, median of 20 repetitions. Wild won all 20 of each project, by 1.9 ms on ripgrep and about 8 ms on cargo. With rust-lld as the default linker, linking is 2 to 5% of a release rebuild of 2.2 to 2.6 s, and close to 1% with mold or Wild. It includes a mistake in my benchmark: a fake linker passed with -B never ran. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  7. 46 s

    08 / 40

    Neural network explained: it matched boosting with 2,177 weights, and changing the seed moved its AUC by 0.004

    In 46 seconds, narrated: one hidden layer of 128 units matched a boosting model on the test period; why with no hidden layer the network is exactly a logistic regression; what happens when it grows too much; how much the result moves when only the seed changes; and what it gains in size and latency per row. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  8. 46 s

    09 / 40

    Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwards

    In 46 seconds, narrated: why the Transformer landed below three simpler models; the 9,633 weights the rule picked and the 0.007 of AUC that separate it from boosting; how the ranking of eleven architectures flipped between validation and the test period; what happened when the attention was removed; and the layer that shields it from a unit change by leaving one column mute. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  9. 47 s

    10 / 40

    KNN explained: training took 0.007 s, predicting 5.5 s, and with k = 1 it memorized what it saw

    In 47 seconds, narrated: fitting KNN took seven milliseconds and scoring the test set five and a half seconds; how it decides with its neighbors and why a single one memorizes; how unscaled humidity takes over the distance; why the search trees were slower than exhaustive search; and where it landed against logistic regression and boosting. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  10. 47 s

    11 / 40

    Naive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edges

    In 47 seconds, narrated: the whole model weighs 1.3 kilobytes and was fitted in eight milliseconds; how it looks at each variable separately inside each class; what happens when a column is copied seven times; why six out of ten days get a probability pinned to zero or to one; and how it behaves with only thirty training rows. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  11. 49 s

    12 / 40

    SVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 seconds

    In 49 seconds and narrated: the same unscaled SVM collapses when one variable changes units; what the margin and support vectors are; how gamma bends the boundary until it memorizes; how fast training cost grows when using every row; and how to approximate the kernel with Nystroem plus a linear model. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  12. 42 s

    13 / 40

    Decision tree explained: 15 nodes you can read, 8,524 leaves that memorize

    In 42 seconds and narrated: a decision tree with no limit reached AUC 1.000 on the days it learned from and dropped to 0.689 on new days; it asks one variable at a time; unchecked it grew more than 8,500 leaves, and at depth 7, chosen on validation, it did better; the first question barely changes, but the predictions disagree on a third of the days; and for numbers it predicts in steps, flat beyond what it saw. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  13. 1:04

    14 / 40

    ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always win

    In 64 seconds and narrated: the same Qwen3.8-27B on the same 16 GB GPU, with ExLlamaV3 and llama.cpp getting the same tokens. With MTP on code, 103.1 tokens per second against 79.5; in an 11,000-token agentic turn the lead shrinks to 8.62 against 9.29 seconds; and on quality ExLlamaV3 wins with 13 GB files and loses with 15 GB ones. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  14. 51 s

    15 / 40

    Gradient boosting explained: 594 chained trees that train 9 times faster than a forest

    In 51 seconds and narrated: with learning rate 1.0, gradient boosting destabilized and after 3,000 rounds ended at AUC 0.681, below a decision tree with no depth limit; it adds small trees in a chain and each one fits what the previous ones missed; with rate 0.03, validation chose 594 rounds and it reached 0.883; against a 200-tree forest the difference was not conclusive, but on one thread it trained 9 times faster and weighs 2.1 MB against 124 MB; and with one-question trees it needed almost 20,000 rounds. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  15. 49 s

    16 / 40

    K-Means explained: I asked Chile's weather for four kinds of day and it found Santiago

    In 49 seconds and narrated: without being told the cities, K-Means put 92% of Santiago's days in one group; Lloyd's algorithm assigns and moves centroids; unscaled, humidity accounts for 76% of the separation; inertia shows no elbow and the silhouette stays within sampling noise across k = 3, 5 and 6; and with k = 8, only 18% of random starts end close to the best inertia, against 38% for k-means++. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  16. 1:06

    17 / 40

    Random Forest explained: how a forest of trees votes, measured on 42 years of Chilean weather

    In 64 seconds and narrated: the viral Random Forest diagram draws the trees in a chain, and that is boosting. On eight cores, 200 trees drop from 26.1 to 3.5 seconds; AUC climbs from 0.69 with one tree to 0.88 with a hundred; a column of pure noise looks important with the default measure; and trained on cold months from 1984 to 2012, the forest never predicted above 28.9 °C in the summers of 2016 to 2026. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  17. 57 s

    18 / 40

    Linear regression explained: the line that measured the heat in Chile, and what breaks it

    In 57 seconds and narrated: on Santiago summers no model had seen, the line missed by 1.86 °C and the random forest by 3.76 °C; the yearly max rises 0.33 °C per decade in Temuco and falls in Valparaíso; unscaled, gradient descent ends with an error of 9.8×10¹¹ °C; with 3,000 rows, the dew point comes out with the minority sign in 46% of samples; and with 5% of rows carrying a misplaced decimal point, least squares rises to 7.89 °C while Huber stays at 1.69 °C. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  18. 43 s

    19 / 40

    Logistic regression explained: 16 numbers to decide whether it rains tomorrow

    In 43 seconds and narrated: a straight line fitted to days with and without rain gives negative probabilities on 9% of days; logistic regression passes the sum of the variables through an S-shaped curve; at a 0.5 threshold it warns about only half of the rain; if missing rain costs three times a false alarm, it catches 80%; and it promises 24% rain when it rains 20%. Muted by default: turn the sound on in the controls.

    Read the articleDirect link

  19. 52 s

    20 / 40

    A chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the prompt

    In 52 seconds and narrated: the gatekeeper sorts the messages, an injection is cut off thirteen times sooner than a full answer, the verifier strikes out the invented figure, and the cost, which depends on how you build it: about 992 dollars per million messages paying as you go, or the 22-dollar monthly plan this blog uses. Every figure comes from what was measured. Muted by default: turn the sound on in the controls.

    Read the articleDirect link