Reels · 14 of 40
Reel · 1:04
ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always win
In 64 seconds and narrated: the same Qwen3.8-27B on the same 16 GB GPU, with ExLlamaV3 and llama.cpp getting the same tokens. With MTP on code, 103.1 tokens per second against 79.5; in an 11,000-token agentic turn the lead shrinks to 8.62 against 9.29 seconds; and on quality ExLlamaV3 wins with 13 GB files and loses with 15 GB ones. Muted by default: turn the sound on in the controls.
- Length
- 1:04
- Published
This reel sums up ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always win, where the method, the tables and what did not work are.