11°

~/topics/performance

Measuring performance without fooling yourself

The numbers are the easy part. Making them mean something is not.

All topics

Most performance comparisons measure something other than what they claim: a corpus that fits in cache, a bench with the wrong ground truth, a 2% difference sold as a win when the spread across runs is 14%. In these pieces the method takes up as much room as the result, and the measurements that had to be thrown out get published too.

27 pieces

Portada: A model that does not write, it only decides: I had Intern-Decision predict the rain without training itArticleA model that does not write, it only decides: I had Intern-Decision predict the rain without training itIntern-Decision came out on 26 September 2026: 0.8B to 4B models that do not generate text and return one probability per question in a single forward pass. I installed it the same day, ran into three failures along the way, one of which made it between 2 and 3.2 times slower, and had it predict tomorrow's rain over 27,256 days in seven Chilean cities without training it. The 2B reaches an AUC of 0.823, almost the same as a Naive Bayes trained on 74,145 rows, and stays below logistic regression.#Models#Benchmarks#GPUPortada: AWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloudArticleAWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloudI tested Strands Harness, AWS's agent runtime, against a bare LangGraph agent: same 3 tools, same gatekeeper, same 20-turn conversation. Against a local Qwen3.8-27B with a 16k context window it spent 1.82 times more tokens. Against the same job on a cloud model, with temperature finally matched between the two, it spent 88.8% less. And three rounds of adversarial audit had to correct my own mistakes along the way: 11 extra tools, a turn limit that lied, and a cache double-count that inflated the first number 4x.#Agents#LangGraph#AWSPortada: How file compression actually works today (and how I beat zlib by 2.86% with graph theory)ArticleHow file compression actually works today (and how I beat zlib by 2.86% with graph theory)I translated zlib from C to Rust with c2rust, proved the translation is bit-exact across more than 15,000 cases, then replaced zlib's heuristic with a shortest path in a graph. Result: 2.86% smaller than zlib -9, verified byte for byte against the real C implementation. Along the way: how lossless compression actually works today, from LZ77 and Huffman to why zopfli and PPMd don't break Shannon's limit either.#Compression#Rust#CPortada: Rust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly three times more, and mv between disks loses the datesArticleRust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly three times more, and mv between disks loses the datesI measured Rust Coreutils (uutils) 0.12 against GNU coreutils 9.12 in containers with limits: the GNU test suite, 157 commands with their output compared, and process start-up. uutils is slower in 91 of 157 cells, takes 2.8 times longer to launch 5,000 processes, and mv across file systems loses the modification dates.#Rust#Coreutils#GNUPortada: Dragonfly 2.0 vs Redis 8 and Valkey 9: with 2 CPUs it wins big, with 4 it depends on how you configure the othersArticleDragonfly 2.0 vs Redis 8 and Valkey 9: with 2 CPUs it wins big, with 4 it depends on how you configure the othersI measured Dragonfly 2.0, Redis 8.10 and Valkey 9.1 in containers with 2 and 4 CPUs. With io-threads on, Redis and Valkey catch up with Dragonfly without pipelining and beat it at pipeline 16; with 2 CPUs and no pipelining, Dragonfly wins by 59% to 99% over stock Redis. It uses 8 to 15% less memory per key.#Dragonfly#Redis#ValkeyPortada: CUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thingArticleCUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thingI built cuda-oxide, NVIDIA's backend for writing CUDA kernels in Rust, and timed the same kernel against CUDA C++. On Mandelbrot Rust comes out 3.2% faster, but the two programs differ in 27,510 pixels; with the arithmetic matched bit for bit, Rust ends up 7% slower.#Rust#CUDA#GPUPortada: Wild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneckArticleWild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneckI measured Wild 0.10.0, mold 2.42.1, rust-lld and GNU ld linking ripgrep and cargo. Wild won all 20 repetitions of each project, by 1.9 ms and about 8 ms. With either one the link is close to 1% of the rebuild. And on the way I almost published false numbers twice.#Rust#Linkers#BenchmarksPortada: Neural network explained: it matched boosting with 2,177 weights, and changing the seed moved its AUC by 0.004TutorialNeural network explained: it matched boosting with 2,177 weights, and changing the seed moved its AUC by 0.004What a neural network is, what each layer does and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. One hidden layer of 128 units, 2,177 weights and 81 KB, reached AUC 0.883 on the 2016-2026 test period: the same figure as a 2.1 MB boosting model, at 0.40 ms per row against 3.09. Training it cost five times more, and repeating the training with only the seed changed moved the validation AUC between 0.881 and 0.884.#Machine learning#Algorithms#PythonPortada: Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwardsTutorialTransformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwardsWhat a Transformer is, what attention does and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. The chosen network has 9,633 parameters and 49 KB, and landed below boosting, the forest and a one-layer network. The ranking is not the lesson: among the six architectures that competed, the validation order inverted on the test period.#Machine learning#Algorithms#PythonPortada: KNN explained: training took 0.007 s, predicting 5.5 s, and with k = 1 it memorized what it sawTutorialKNN explained: training took 0.007 s, predicting 5.5 s, and with k = 1 it memorized what it sawWhat k-nearest neighbors is, why you must scale, how to choose k and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Fitting on the 74,145 rows took 0.007 s and scoring the 2016-2026 test period 5.5 s on one thread; with k = 1 it scored AUC 1.000 on its own training days and 0.69 on validation. With k = 100 it reached 0.870, below a boosting model and above logistic regression.#Machine learning#Algorithms#PythonPortada: Naive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edgesTutorialNaive Bayes explained: it trained in 0.008 s and weighs 1.3 KB, but 62 % of its probabilities landed at the edgesWhat Naive Bayes is, what the independence assumption means and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Fitting on 74,145 rows took 0.008 s and the model weighs 1.3 KB, at AUC 0.822; on binned variables it reached 0.844, with no conclusive difference from logistic regression. 62 % of its probabilities fell below 0.01 or above 0.99, and calibrating it took log loss from 1.09 to 0.39.#Machine learning#Algorithms#PythonPortada: SVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 secondsTutorialSVM explained: pressure in Pa sank AUC to 0.555 and the exact kernel took 51 secondsWhat a support vector machine is, what the margin, C and gamma do, and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. Unscaled, switching pressure from kPa to Pa dropped validation AUC from 0.835 to 0.555; on all 74,145 rows the RBF kernel took 51 s and, on the 2016-2026 test set, still ranked below a boosting model that trained in 2.6 s.#Machine learning#Algorithms#PythonPortada: Decision tree explained: 15 nodes you can read, 8,524 leaves that memorizeTutorialDecision tree explained: 15 nodes you can read, 8,524 leaves that memorizeWhat a decision tree is, how it picks its questions and when it memorizes, measured on daily weather from seven Chilean cities between 1984 and 2026. With no depth limit it got everything it saw right and dropped to AUC 0.689 on new days; at depth 7, chosen on validation, it reached 0.862.#Machine learning#Algorithms#PythonPortada: Gradient boosting explained: 594 chained trees that train 9 times faster than a forestTutorialGradient boosting explained: 594 chained trees that train 9 times faster than a forestWhat gradient boosting is, how each tree corrects the previous ones and when it fails, measured on daily weather from seven Chilean cities between 1984 and 2026. With rate 1.0 it destabilized and after 3,000 rounds ended at AUC 0.681; with rate 0.03 and 594 rounds chosen on validation it reached 0.883, with no conclusive difference from a 200-tree forest.#Machine learning#Algorithms#PythonPortada: K-Means explained: I asked Chile's weather for four kinds of day and it found SantiagoTutorialK-Means explained: I asked Chile's weather for four kinds of day and it found SantiagoWhat K-Means is, how Lloyd's algorithm works and when it misleads, measured on daily weather from seven Chilean cities between 1984 and 2026. Unscaled, humidity accounts for 76% of the separation; with k = 8, only 18% of random starts end within 0.1% of the best inertia found; and no criterion agrees on how many groups there are.#Machine learning#Algorithms#PythonPortada: Random Forest explained: how a forest of trees votes, measured on 42 years of Chilean weatherTutorialRandom Forest explained: how a forest of trees votes, measured on 42 years of Chilean weatherWhat Random Forest is, how bootstrap sampling and feature subsampling work, and why its trees train in parallel instead of in a chain. I measured it predicting rain in seven Chilean cities: 7.5 times faster on 8 CPUs, a plateau past 100 trees, a noise column that the default importance ranks above real features, and a ceiling when extrapolating temperatures.#Machine learning#Algorithms#PythonPortada: Linear regression explained: the line that measured the heat in Chile, and what breaks itTutorialLinear regression explained: the line that measured the heat in Chile, and what breaks itWhat linear regression is, how its coefficients are computed and when it stops working, measured on daily data from seven Chilean cities. The max temperature rises 0.33 °C per decade in Temuco and falls in Valparaíso, unscaled gradient descent does not reach a useful solution, 5% of rows with a misplaced decimal point multiply the error by 4.7, and the line trains about 1,100 times faster than a random forest.#Machine learning#Algorithms#PythonPortada: Logistic regression explained: 16 numbers to decide whether it rains tomorrowTutorialLogistic regression explained: 16 numbers to decide whether it rains tomorrowWhat logistic regression is, how it turns a sum into a probability and when it misleads, measured on daily weather from seven Chilean cities between 1984 and 2026. A straight line gives negative probabilities on 9% of days; at a 0.5 threshold logistic regression catches only half of the rain; and it promises more rain than falls.#Machine learning#Algorithms#PythonPortada: A chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the promptArticleA chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the promptTutorial: LangGraph chat on Amazon Bedrock AgentCore with an S3 Knowledge Base and guardrails, tested on a 64-message chat. Costs and a local option.#AI#AWS#SecurityPortada: I measured Pingora 0.9.0 six times slower than nginx. The culprit was one line of my codeArticleI measured Pingora 0.9.0 six times slower than nginx. The culprit was one line of my codeI wrote a minimal reverse proxy with Pingora 0.9.0 and in a 2 vCPU container it measured 21 thousand requests per second against nginx's 126 thousand: six times slower. The culprit was not Pingora but a blocking getaddrinfo my upstream_peer ran on every request. Fixed, the real gap is 1.70x, and perf quantifies it: 1.70x in cycles per request, 119.4 thousand against 70.4 thousand.#Rust#Benchmarks#NetworkingPortada: Hard rules against Random Cut Forest on the same stream: the model adds 13 points of recall and costs 7 times the throughputArticleHard rules against Random Cut Forest on the same stream: the model adds 13 points of recall and costs 7 times the throughputI ran Apache Flink with Kafka and measured three deterministic rules against Amazon's Random Cut Forest over exactly the same payment stream, on a four-core box with no GPU. The rules reach 0.814 recall at 173 thousand events per second; the model reaches 0.943 but drops the pipeline to 23 thousand; together they hit 0.982.#Streaming#Benchmarks#DataPortada: Bun rewrote itself from Zig to Rust in eleven days. I measured what they published, and what they did notArticleBun rewrote itself from Zig to Rust in eleven days. I measured what they published, and what they did notBun ported 535,000 lines from Zig to Rust and published three concrete figures: start-up, requests per second, and binary size. Version 1.3.14 was the last Zig build and 1.4.0 the first Rust one, so you can download both and measure the same thing with only the binary changing. Start-up beats it comfortably, binary size falls short, and for memory and idle CPU the announcement gives no number at all: those two I measure with nothing official to compare against.#Rust#ArquitecturaPortada: I translated DOOM from C to Rust without writing a line, then used the model for what it is actually good atArticleI translated DOOM from C to Rust without writing a line, then used the model for what it is actually good atI needed to move a program from C to Rust and my first idea was to ask a model. A tool already existed that does it in 23 seconds. The translation matched the original bit for bit across the 11,113 frames compared, except one that traces back to an uninitialised-memory bug already present in 1993's DOOM, and the model ended up doing something else entirely: textures, per-pixel relief and continuous lighting on top of the original engine.#Rust#C#Code translationPortada: Go 1.27 brings portable SIMD: it ties with NumPy out of cache and loses inside itArticleGo 1.27 brings portable SIMD: it ties with NumPy out of cache and loses inside itI measured Go 1.27's experimental simd package against NumPy on a real task. They tie when the corpus does not fit in cache, and NumPy wins by 2.4 times when it does. Along the way I nearly published two false comparisons, and those are the useful part.#Go#Benchmarks#PythonPortada: TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteenArticleTabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteenThe claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without an account.#Models#Benchmarks#GPUPortada: I put the same agent to fix real bugs with three local engines: 284B, 27B and 27B in two bitsArticleI put the same agent to fix real bugs with three local engines: 284B, 27B and 27B in two bitsThe DeepSeek Harness has only been tested against the official API. I wired it to three local engines and gave them real SWE-bench Verified instances. Two tie on fixes, one is eleven times faster, and the one with 284 billion parameters barely leaves the starting line.#Agents#Models#GPUPortada: Does Qwen3.8-27B fit in 16 GB of VRAM? I measured it, and found out my own benchmark was lyingArticleDoes Qwen3.8-27B fit in 16 GB of VRAM? I measured it, and found out my own benchmark was lyingEveryone repeats that Alibaba's new model needs about 15 GB in 4 bits. My GPU has 16 and the file weighs 17. I measured tokens per second, real VRAM and the CPU/GPU split across three context sizes, and the model ended up correcting me.#Models#GPU#Benchmarks

All writing