17°

Reels · 14 of 40

1:04

Reel · 1:04

ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always win

In 64 seconds and narrated: the same Qwen3.8-27B on the same 16 GB GPU, with ExLlamaV3 and llama.cpp getting the same tokens. With MTP on code, 103.1 tokens per second against 79.5; in an 11,000-token agentic turn the lead shrinks to 8.62 against 9.29 seconds; and on quality ExLlamaV3 wins with 13 GB files and loses with 15 GB ones. Muted by default: turn the sound on in the controls.

Length
1:04
Published

ExLlamaV3llama.cppCuantizaciónGPUModelos de lenguaje

This reel sums up ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always win, where the method, the tables and what did not work are.

Read the articleOpen in the viewer