11°

~/topics/local-models

Models on my own hardware

What actually fits in a 16 GB card, and what gets lost along the way.

All topics

Almost everything published about open models is measured on datacenter cards. The question here is a different one: if the file weighs more than the memory available, does it still run, and at what cost? These pieces share a bench — the same 16 GB GPU, the same frozen inputs — so the numbers can be placed side by side instead of read in isolation.

11 pieces

Portada: A model that does not write, it only decides: I had Intern-Decision predict the rain without training itArticleA model that does not write, it only decides: I had Intern-Decision predict the rain without training itIntern-Decision came out on 26 September 2026: 0.8B to 4B models that do not generate text and return one probability per question in a single forward pass. I installed it the same day, ran into three failures along the way, one of which made it between 2 and 3.2 times slower, and had it predict tomorrow's rain over 27,256 days in seven Chilean cities without training it. The 2B reaches an AUC of 0.823, almost the same as a Naive Bayes trained on 74,145 rows, and stays below logistic regression.#Models#Benchmarks#GPUPortada: AWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloudArticleAWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloudI tested Strands Harness, AWS's agent runtime, against a bare LangGraph agent: same 3 tools, same gatekeeper, same 20-turn conversation. Against a local Qwen3.8-27B with a 16k context window it spent 1.82 times more tokens. Against the same job on a cloud model, with temperature finally matched between the two, it spent 88.8% less. And three rounds of adversarial audit had to correct my own mistakes along the way: 11 extra tools, a turn limit that lied, and a cache double-count that inflated the first number 4x.#Agents#LangGraph#AWSPortada: CUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thingArticleCUDA in Rust vs CUDA C++: Rust wins by 3.2% until you check that they do not compute the same thingI built cuda-oxide, NVIDIA's backend for writing CUDA kernels in Rust, and timed the same kernel against CUDA C++. On Mandelbrot Rust comes out 3.2% faster, but the two programs differ in 27,510 pixels; with the arithmetic matched bit for bit, Rust ends up 7% slower.#Rust#CUDA#GPUPortada: ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always winArticleExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always winI gave both engines exactly the same input tokens inside a container with pod limits. ExLlamaV3 decodes faster and handled 131k tokens of context; with 13 GB files it loses less quality, with 15 GB files it loses to GGUF. In an agentic turn, prefill eats almost all of the lead.#ExLlamaV3#llama.cpp#QuantizationPortada: A chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the promptArticleA chat on Amazon Bedrock AgentCore for my blog: I broke it with 64 messages and rebuilt it with rules outside the promptTutorial: LangGraph chat on Amazon Bedrock AgentCore with an S3 Knowledge Base and guardrails, tested on a 64-message chat. Costs and a local option.#AI#AWS#SecurityPortada: More bits where it hurts: I quantized a Qwen3.8-27B to fit my 4070 TiArticleMore bits where it hurts: I quantized a Qwen3.8-27B to fit my 4070 TiQ4 overflows 16 GB and Q3 gives up quality. A llama.cpp patch spreads the bits layer by layer: I measured a Qwen3.8-27B at 94% agreement with an 8-bit reference that does fit on a 4070 Ti.#Quantization#llama.cpp#GGUFPortada: TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteenArticleTabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteenThe claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without an account.#Models#Benchmarks#GPUPortada: I ran AIFS 2.0, ECMWF's weather model, on my desktop card and landed 0.45 °C from the official forecastTutorialI ran AIFS 2.0, ECMWF's weather model, on my desktop card and landed 0.45 °C from the official forecastECMWF published the weights of its AI forecast model. It weighs under a gigabyte and runs on a desktop card: 48 hours of forecast in 108 seconds and 5.77 GB of memory, or 33 minutes if you have no GPU. I installed it step by step, compared it against the same centre's operational physical model, and wrote down the two stumbles that nearly made me publish nonsense.#Models#GPU#WeatherPortada: I put the same agent to fix real bugs with three local engines: 284B, 27B and 27B in two bitsArticleI put the same agent to fix real bugs with three local engines: 284B, 27B and 27B in two bitsThe DeepSeek Harness has only been tested against the official API. I wired it to three local engines and gave them real SWE-bench Verified instances. Two tie on fixes, one is eleven times faster, and the one with 284 billion parameters barely leaves the starting line.#Agents#Models#GPUPortada: Pocket TTS: I cloned my voice on a 2019 CPU, no GPU and no cloudTutorialPocket TTS: I cloned my voice on a 2019 CPU, no GPU and no cloudA 209 MB voice model that runs on CPU, clones a voice from 25 seconds of reference audio and allows commercial use with attribution. Step-by-step install, metrics measured on a 2019 i5, and my honest take next to ElevenLabs and friends.#AI#AudioPortada: Does Qwen3.8-27B fit in 16 GB of VRAM? I measured it, and found out my own benchmark was lyingArticleDoes Qwen3.8-27B fit in 16 GB of VRAM? I measured it, and found out my own benchmark was lyingEveryone repeats that Alibaba's new model needs about 15 GB in 4 bits. My GPU has 16 and the file weighs 17. I measured tokens per second, real VRAM and the CPU/GPU split across three context sizes, and the model ended up correcting me.#Models#GPU#Benchmarks

All writing