18°

~/blog/articulo

Articles

Analysis and measurements of my own: networks, low-level work, hardware and systems.

All writing

Portada: A model that does not write, it only decides: I had Intern-Decision predict the rain without training itArticleA model that does not write, it only decides: I had Intern-Decision predict the rain without training itIntern-Decision came out on 26 September 2026: 0.8B to 4B models that do not generate text and return one probability per question in a single forward pass. I installed it the same day, ran into three failures along the way, one of which made it between 2 and 3.2 times slower, and had it predict tomorrow's rain over 27,256 days in seven Chilean cities without training it. The 2B reaches an AUC of 0.823, almost the same as a Naive Bayes trained on 74,145 rows, and stays below logistic regression.#Models#Benchmarks#GPUPortada: AWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloudArticleAWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloudI tested Strands Harness, AWS's agent runtime, against a bare LangGraph agent: same 3 tools, same gatekeeper, same 20-turn conversation. Against a local Qwen3.8-27B with a 16k context window it spent 1.82 times more tokens. Against the same job on a cloud model, with temperature finally matched between the two, it spent 88.8% less. And three rounds of adversarial audit had to correct my own mistakes along the way: 11 extra tools, a turn limit that lied, and a cache double-count that inflated the first number 4x.#Agents#LangGraph#AWSPortada: How file compression actually works today (and how I beat zlib by 2.86% with graph theory)ArticleHow file compression actually works today (and how I beat zlib by 2.86% with graph theory)I translated zlib from C to Rust with c2rust, proved the translation is bit-exact across more than 15,000 cases, then replaced zlib's heuristic with a shortest path in a graph. Result: 2.86% smaller than zlib -9, verified byte for byte against the real C implementation. Along the way: how lossless compression actually works today, from LZ77 and Huffman to why zopfli and PPMd don't break Shannon's limit either.#Compression#Rust#CPortada: Rust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly three times more, and mv between disks loses the datesArticleRust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly three times more, and mv between disks loses the datesI measured Rust Coreutils (uutils) 0.12 against GNU coreutils 9.12 in containers with limits: the GNU test suite, 157 commands with their output compared, and process start-up. uutils is slower in 91 of 157 cells, takes 2.8 times longer to launch 5,000 processes, and mv across file systems loses the modification dates.#Rust#Coreutils#GNUPortada: Dragonfly 2.0 vs Redis 8 and Valkey 9: with 2 CPUs it wins big, with 4 it depends on how you configure the othersArticleDragonfly 2.0 vs Redis 8 and Valkey 9: with 2 CPUs it wins big, with 4 it depends on how you configure the othersI measured Dragonfly 2.0, Redis 8.10 and Valkey 9.1 in containers with 2 and 4 CPUs. With io-threads on, Redis and Valkey catch up with Dragonfly without pipelining and beat it at pipeline 16; with 2 CPUs and no pipelining, Dragonfly wins by 59% to 99% over stock Redis. It uses 8 to 15% less memory per key.#Dragonfly#Redis#ValkeyPortada: CUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thingArticleCUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thingI built cuda-oxide, NVIDIA's backend for writing CUDA kernels in Rust, and timed the same kernel against CUDA C++. On Mandelbrot Rust comes out 3.2% faster, but the two programs differ in 27,510 pixels; with the arithmetic matched bit for bit, Rust ends up 7% slower.#Rust#CUDA#GPUPortada: Wild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneckArticleWild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneckI measured Wild 0.10.0, mold 2.42.1, rust-lld and GNU ld linking ripgrep and cargo. Wild won all 20 repetitions of each project, by 1.9 ms and about 8 ms. With either one the link is close to 1% of the rebuild. And on the way I almost published false numbers twice.#Rust#Linkers#BenchmarksPortada: ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always winArticleExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always winI gave both engines exactly the same input tokens inside a container with pod limits. ExLlamaV3 decodes faster and handled 131k tokens of context; with 13 GB files it loses less quality, with 15 GB files it loses to GGUF. In an agentic turn, prefill eats almost all of the lead.#ExLlamaV3#llama.cpp#QuantizationPortada: A chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the promptArticleA chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the promptTutorial: LangGraph chat on Amazon Bedrock AgentCore with an S3 Knowledge Base and guardrails, tested on a 64-message chat. Costs and a local option.#AI#AWS#SecurityPortada: I measured Pingora 0.9.0 six times slower than nginx. The culprit was one line of my codeArticleI measured Pingora 0.9.0 six times slower than nginx. The culprit was one line of my codeI wrote a minimal reverse proxy with Pingora 0.9.0 and in a 2 vCPU container it measured 21 thousand requests per second against nginx's 126 thousand: six times slower. The culprit was not Pingora but a blocking getaddrinfo my upstream_peer ran on every request. Fixed, the real gap is 1.70x, and perf quantifies it: 1.70x in cycles per request, 119.4 thousand against 70.4 thousand.#Rust#Benchmarks#NetworkingPortada: Hard rules against Random Cut Forest on the same stream: the model adds 13 points of recall and costs 7 times the throughputArticleHard rules against Random Cut Forest on the same stream: the model adds 13 points of recall and costs 7 times the throughputI ran Apache Flink with Kafka and measured three deterministic rules against Amazon's Random Cut Forest over exactly the same payment stream, on a four-core box with no GPU. The rules reach 0.814 recall at 173 thousand events per second; the model reaches 0.943 but drops the pipeline to 23 thousand; together they hit 0.982.#Streaming#Benchmarks#DataPortada: I measured 320 Chilean sites: half already use post-quantum cryptography, and almost none of them chose toArticleI measured 320 Chilean sites: half already use post-quantum cryptography, and almost none of them chose toI probed the 320 most visited .cl domains: 145 negotiate X25519MLKEM768 (47.7%), but only 9 of the 128 with their own infrastructure. Measured cost: 0 ms.#Post-quantum cryptography#TLS#ChilePortada: I built a fully post-quantum handshake, then went to see which of my servers could do itArticleI built a fully post-quantum handshake, then went to see which of my servers could do itRustls turned ML-DSA certificates on by default, but only for private hierarchies. I built an mTLS where the key exchange, the server signature and the client identity are all post-quantum. I measured what the chain weighs, which version of each server supports it, and corrected a number I had got wrong: the private key is not nineteen times bigger, it is 128 bytes.#Post-quantum cryptography#TLS#mTLSPortada: pg_anon found my database's emails, but not the RUTArticlepg_anon found my database's emails, but not the RUTI tried pg_anon, the Russian tool that masks personal data in Postgres. With stock rules it caught 1 of 8 PII columns on a Chilean database; with my own rules, all 8. I measured recall, integrity and speed against pg_dump.#PostgreSQL#Personal data#AnonymizationPortada: More bits where it hurts: I quantized a Qwen3.8-27B to fit my 4070 TiArticleMore bits where it hurts: I quantized a Qwen3.8-27B to fit my 4070 TiQ4 overflows 16 GB and Q3 gives up quality. A llama.cpp patch spreads the bits layer by layer: I measured a Qwen3.8-27B at 94% agreement with an 8-bit reference that does fit on a 4070 Ti.#Quantization#llama.cpp#GGUFPortada: Bun rewrote itself from Zig to Rust in eleven days. I measured what they published, and what they did notArticleBun rewrote itself from Zig to Rust in eleven days. I measured what they published, and what they did notBun ported 535,000 lines from Zig to Rust and published three concrete figures: start-up, requests per second, and binary size. Version 1.3.14 was the last Zig build and 1.4.0 the first Rust one, so you can download both and measure the same thing with only the binary changing. Start-up beats it comfortably, binary size falls short, and for memory and idle CPU the announcement gives no number at all: those two I measure with nothing official to compare against.#Rust#ArquitecturaPortada: I built a public sandbox for Chile's Open Finance System, and Keycloak alone did not get thereArticleI built a public sandbox for Chile's Open Finance System, and Keycloak alone did not get thereChile has an open finance regulation and no public sandbox to test it against. I built a complete unofficial one, faithful to the Technical Annex: participant directory, bank, authorization server and requesting application, with mandatory PAR, an encrypted request object, mTLS and RAR reaching the token. Keycloak 26 covers much of it through configuration, but the consent model Chile picked does not exist in the product: it had to be written from scratch.#Seguridad#Arquitectura#ChilePortada: The seven GEO rules: I reproduced the only test that exists, then audited my own 72 pagesArticleThe seven GEO rules: I reproduced the only test that exists, then audited my own 72 pagesA Russian study tested the seven rules the GEO industry keeps repeating against 73 pages that six engines had cited, and not one survived. I rebuilt its counts and reproduced all seven p-values exactly. Then I took the detector to my own 72 pages, and there something appeared that none of the rules looks at: four different pages declaring the same entity.#SEO#GEO#StatisticsPortada: I translated DOOM from C to Rust without writing a line, then used the model for what it is actually good atArticleI translated DOOM from C to Rust without writing a line, then used the model for what it is actually good atI needed to move a program from C to Rust and my first idea was to ask a model. A tool already existed that does it in 23 seconds. The translation matched the original bit for bit across the 11,113 frames compared, except one that traces back to an uninitialised-memory bug already present in 1993's DOOM, and the model ended up doing something else entirely: textures, per-pixel relief and continuous lighting on top of the original engine.#Rust#C#Code translationPortada: Go 1.27 brings portable SIMD: it ties with NumPy out of cache and loses inside itArticleGo 1.27 brings portable SIMD: it ties with NumPy out of cache and loses inside itI measured Go 1.27's experimental simd package against NumPy on a real task. They tie when the corpus does not fit in cache, and NumPy wins by 2.4 times when it does. Along the way I nearly published two false comparisons, and those are the useful part.#Go#Benchmarks#PythonPortada: TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteenArticleTabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteenThe claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without an account.#Models#Benchmarks#GPUPortada: I put the same agent to fix real bugs with three local engines: 284B, 27B and 27B in two bitsArticleI put the same agent to fix real bugs with three local engines: 284B, 27B and 27B in two bitsThe DeepSeek Harness has only been tested against the official API. I wired it to three local engines and gave them real SWE-bench Verified instances. Two tie on fixes, one is eleven times faster, and the one with 284 billion parameters barely leaves the starting line.#Agents#Models#GPUPortada: Your architecture diagram is lying: I tested both ways of writing it as codeArticleYour architecture diagram is lying: I tested both ways of writing it as codeA visual C4 editor showed up on Habr and cannot be self-hosted. I tested the two alternatives that can: the Structurizr CLI generates all three levels of a real system in 1 second from a single file; C4-PlantUML reaches the same result without installing anything, at the cost of tripling the model.#Architecture#ToolsPortada: The chip that guards your keys is broken: I audited all four machines in my swarmArticleThe chip that guards your keys is broken: I audited all four machines in my swarmTwo flaws rated CVSS 8.5 and 8.3 in firmware TPM affect Ryzen 3000 through Ryzen 9000 and Intel up to Ultra 200. I checked my four machines: two are affected, one has had a patch since July and the other was abandoned by its vendor in 2022.#Security#Hardware#LinuxPortada: PhantomRelay: why your automation gives itself away before the first HTTP byteArticlePhantomRelay: why your automation gives itself away before the first HTTP byteA Node HTTP client is distinguishable from real Chrome in the TLS handshake, before sending a single header. How I built a browser relay with a Rust addon over BoringSSL, cost-based escalation and nearly 900 tests.#Systems#NetworksPortada: Does Qwen3.8-27B fit in 16 GB of VRAM? I measured it, and found out my own benchmark was lyingArticleDoes Qwen3.8-27B fit in 16 GB of VRAM? I measured it, and found out my own benchmark was lyingEveryone repeats that Alibaba's new model needs about 15 GB in 4 bits. My GPU has 16 and the file weighs 17. I measured tokens per second, real VRAM and the CPU/GPU split across three context sizes, and the model ended up correcting me.#Models#GPU#BenchmarksPortada: The OSI model explained: the 7 layers and encapsulationArticleThe OSI model explained: the 7 layers and encapsulationWhat the OSI model is, what each of its 7 layers is for, and how a piece of data gets encapsulated into segments, packets, frames and bits as it travels across the network.#Networks#Low level