
ExLlamaV3 vs llama.cpp with Qwen3.8-27B on 16 GB: faster decoding, but it does not always win
I gave both engines exactly the same input tokens inside a container with pod limits. ExLlamaV3 decodes faster and handled 131k tokens of context; with 13 GB files it loses less quality, with 15 GB files it loses to GGUF. In an agentic turn, prefill eats almost all of the lead.
A month ago I measured whether Qwen3.8-27B fit on my 16 GB RTX 4070 Ti SUPER, and then I quantized the model to the size of the card with llama.cpp. Always the same engine. On September 13 ExLlamaV3 1.5.0 came out, and its performance table starts at 24 GB cards. Mine is not there.
So I gave both engines exactly the same input tokens, inside the same container. ExLlamaV3 decoded faster in everything I measured and handled 131k tokens of context, where llama.cpp reached 64k, the largest size I tried. On quality it wins with 13 GB files and loses with 15 GB files. And in a real agentic turn, prefill eats almost all of the speed lead.
What I compared, and why it is fair
ExLlamaV3 uses its own quantization format, EXL3, and does not load GGUF. You cannot measure “the same file on two engines”: the closest thing is to pick files of the same size and see what each engine does with its own format.
- ExLlamaV3 1.5.0 with the 3.00 bits-per-weight EXL3 that turboderp publishes: 12.9 GiB on disk, MTP head included. For the quality comparison at a larger size, also the 3.50-bit one (14.3 GiB).
- llama.cpp with a 12.6 GiB
Q3_K_M, quantized with the same importance matrix as the previous post, and the 15.0 GiB GGUF that post custom-quantized to fit the card. For MTP it uses a separate 1.3 GB draft.
Both receive the same token ids. I took the prompt battery ExLlamaV3 ships to measure speculative decoding (code, agents with tools, translation, creative writing), tokenized it once and sent those same numbers to llama-server. With a 16k context, 15 of the 22 prompts fit; the 7 longest agentic ones are left out on both engines.
Same input, same limits, same card. What changes is the engine, its quantization format and the shape of its MTP.
All speed measurements ran inside a container with pod limits: 8 CPUs, 24 GB of RAM and the GPU handed over through CDI. I also repeated llama.cpp directly on the host and the difference did not exceed 0.6%: the workload lives on the GPU. The prompt battery ran three times per engine and no category varied more than 0.3% between repetitions, except one I explain below. The agentic turn, five times per configuration. The context sweep is one run per size on ExLlamaV3 and two internal llama-bench repetitions on llama.cpp. Quality and ExLlamaV3 peak VRAM with MTP were measured on the host: qbench is deterministic and GPU memory does not depend on the container limits.
Installation, with the stumbles
Installation took more work than measuring, and almost none of it was ExLlamaV3’s fault.
- Installing torch from the PyTorch index,
uvgave up twice withoperation timed outwhile fetching NVIDIA packages, for both CUDA 13.2 and 12.8. I did not find the cause: later, the same 218 MB package downloaded in full at 11 MiB/s withcurl. I ended up installing torch 2.11.0 from PyPI with a longer network timeout. - The release ships wheels per CUDA and torch combination. My script built the name from variables that came out empty and produced
exllamav3-1.5.0+.torch.0, a nameuvrejects. With the full name,+cu132.torch2.11.0, it installed fine. exllamav3does not expose__version__. My check failed even though everything was installed correctly.- To measure quality with qbench, ExLlamaV3’s tool, I had to build
llama-cpp-pythonwith CUDA (352 s, there are no recent wheels), installtransformersandgguf, and fix the wikitext dataset name, which the newhuggingface_hubno longer accepts without theSalesforce/prefix. - Inside the container three more showed up.
llama-servercould not findlibcudart.so.13, which the host resolves throughldconfig. torch crashed looking up the name of a user the container does not have. And ExLlamaV3 would not start on the bare Fedora 43 image because Triton needs a C compiler to build its kernels: I addedgcc.
The measuring tool gets measured too
Inside the container, eval/spec_decode.py produced speeds in two categories that did not hold up, the same across three consecutive processes: 51 tokens per second with MTP on an agentic prompt the host handled at 89 with the same tokens, and 69.8 on trivial repetition, which the runner below measured at 115.6. Running that prompt alone, the container dropped even further: 18.5, and as low as 14.4 in other runs. CPU or memory limits, /dev/shm, the cuBLAS library, text output and switching the Triton cache did not change it. I did not find the cause. What I did see: within one process, the first pass over that prompt is slow and the following ones are not. That is why the ExLlamaV3 battery below comes from a runner of mine that repeats every prompt three times in the same process: the first repetition gave 51, the other two 90.9.
Speed: the same tokens, two engines
Without speculative decoding, ExLlamaV3 generated between 40.7 and 43.1 tokens per second depending on the category, and llama.cpp between 36.0 and 38.0. On a 4,000-token prefill they were almost equal: about 1,490 against 1,530 tokens per second.
With MTP, both with a 3-token window:
- ExLlamaV3 · code103.1
- llama.cpp · code79.5
- ExLlamaV3 · creative writing75.2
- llama.cpp · creative writing56.2
- ExLlamaV3 · translation75.2
- llama.cpp · translation60.2
- ExLlamaV3 · trivial repetition115.6
- llama.cpp · trivial repetition89.8
RTX 4070 Ti SUPER in a container with 8 CPUs and 24 GB. Prompts from ExLlamaV3 1.5.0 eval/spec_decode.py, same token ids on both engines, greedy, up to 1024 new tokens. Generated tokens over generation time, median of three repetitions.
The MTP head proposes, the model verifies. A bit over twice as fast (41 → 90.9 tok/s) and the text that comes out is identical.
ExLlamaV3 won all eight categories. And not because its draft guesses better: llama.cpp accepted 91% of what it proposed on the agentic curl prompt and ExLlamaV3 78%; on translation with reasoning, 71% against 57%. ExLlamaV3 is faster while guessing less. The arithmetic explains it: on curl, llama.cpp gets 3.7 tokens per verification round and takes about 45 ms per round; ExLlamaV3 gets 3.3 tokens, but each round takes about 37 ms. I derive the rounds assuming three proposed tokens per round, which is the configured window.
llama.cpp accepts more per round, but each round takes almost 9 ms longer. In the same half second ExLlamaV3 fits two more rounds, and that is where its lead comes from. Rounds and milliseconds per round are derived from the counters, assuming three proposed tokens per round.
An empty Triton cache ruins the measurement
ExLlamaV3 compiles Triton kernels when they are not in the cache, and that compilation lands inside the measured time. On the host, with an empty Triton cache, the curl prompt with MTP gave 38 tokens per second; with a full cache, 89.
Real speed. Same answer, same draft acceptance.
Real speed. Container with 8 CPUs and 24 GB.
Compiling kernels or a slow first pass can double the time of the same answer. That is why the battery figures are the median of three repetitions.
Another odd detail: with different Triton caches the generated text changed, just enough to move draft acceptance from 2.34 to 2.46 tokens per round. My hypothesis is that autotuning picked different kernels, but I did not verify it. That is why the battery figures are the median of three repetitions: with a slow first pass, like the 51 on curl, the median lands on one of the two fast ones. With three samples it is not a measure of variability; it only keeps one anomalous pass from defining the number.
The agentic turn: where the lead shrinks
The agentic prompts carry 11,000 to 14,000 tokens of history and tools, and the answer is a tool call of 60 to 97 tokens. There, prefill rules, not decoding.
At double speed. Container with 8 CPUs and 24 GB.
At double speed. Container with 8 CPUs and 24 GB.
Decoding multiplies, but the full turn drops only 6 to 8%: almost all the time goes into reading the context.
On that prompt, llama.cpp took 7.51 s to reach the first token and ExLlamaV3 7.87 s. On the curl one, with 13,703 tokens, 9.44 against 10.11 s. Reading long context, llama.cpp is faster. ExLlamaV3 wins back ground while writing: with window 4 it decoded at 117.9 tokens per second against 85.9 for llama.cpp. For the full turn it takes less time on the code prompt (8.62 against 9.29 s) and the same on the curl one (11.33 against 11.34 s). Careful reading that as an exact tie: the answers are not the same length. On curl, ExLlamaV3 wrote 97 tokens and llama.cpp 81, so in the same time it produced 20% more; on code it wrote 60 against 64.
Quality: how far each one drifts from the 8-bit model
I measured KL divergence against a Q8_0 of the same model with qbench, over 16 blocks of 512 wikitext tokens. It is a small sample, 8,192 tokens, and the reference is 8-bit, not the original model.
With files of about 13 GB:
- ExLlamaV3 3.00 bpw (12.9 GiB): mean divergence 0.0435, median 0.0164.
- llama.cpp Q3_K_M (12.6 GiB): mean divergence 0.0554, median 0.0238.
With files of about 15 GB:
- ExLlamaV3 3.50 bpw (14.3 GiB): mean divergence 0.0234, median 0.0089.
- llama.cpp custom 15 GB (15.0 GiB): mean divergence 0.0159, median 0.0075.
- ExLlamaV3 3.00 bpw · 12.9 GiB · mean 0.0435
- llama.cpp Q3_K_M · 12.6 GiB · mean 0.0554
reference confidence
- ExLlamaV3 3.50 bpw · 14.3 GiB · mean 0.0234
- llama.cpp custom 15 GB · 15.0 GiB · mean 0.0159
reference confidence
At 13 GB ExLlamaV3 is closer in four of five bands; only on near-certain tokens does the GGUF edge ahead in the mean. At 15 GB the custom GGUF wins every band but the last one. 8,192 tokens of wikitext: a small sample.
The comparison is matched by file size, not by bits per weight. With 13 GB files, ExLlamaV3 stays closer to the reference than the Q3_K_M in every confidence band, except for the mean on the tokens the 8-bit model predicts with more than 95% confidence. With 15 GB files it flips: the custom GGUF wins, with 700 MiB more on disk. With 15 GiB of model plus the 1.3 GB draft, the arithmetic leaves no room for MTP on 16 GB, and in my tests it could not create a 4,000-token context with llama-bench; the previous post loaded it with 8,000 from llama-cli.
There is a counting detail that looks like a mistake and is not. In its layers, the 3-bit EXL3 uses 3.02 bits per weight and the Q3_K_M 3.84, yet the EXL3 file is larger. Part of it is the MTP head it bundles; I did not break down the rest.
Why perplexity does not match the previous post
qbench gave the Q3_K_M a perplexity of 9.86. In the quantization post I published 7.19 for the same file. They do not contradict each other: they are different tools. I measured it again with llama-perplexity, the tool from the previous post, and got 7.1856. qbench scores the full 512-token window over 16 blocks, and llama-perplexity only the second half of each window over all of wikitext.
The 16 GB budget
- ExLlamaV3 + MTP · 16k: MTP head included
- llama.cpp + MTP · 16k: separate 1.3 GB draft
With MTP, llama.cpp scrapes the ceiling: at a 32k context it no longer fit. ExLlamaV3 leaves about 2 GB free out of the 15,945 MiB the card reports as usable.
Without MTP, the question is how much context fits:
- ExLlamaV3 with a 4-bit cache ran a 131,072-token prefill, and with 130,816 tokens of context it still generated 31.4 tokens per second. At 32k it did 39.4. With an fp16 cache it loaded 32k (37.8 tokens per second), but not 64k.
- llama.cpp with a q8_0 cache reached 64k tokens, the largest size I tried (28.9 tokens per second). With an f16 cache it ran 32k (34.0) and could not create a 64k context.
- ExLlamaV3 · Q4 cache:43.5 → 31.4
- ExLlamaV3 · fp16 cache:43.6 → 37.8 · 66k does not load
- llama.cpp · q8_0 cache:37.9 → 28.9 · not tested beyond 66k
- llama.cpp · f16 cache:38.3 → 34.0 · 66k does not load
With its cache quantized to 4 bits, ExLlamaV3 still decodes 31.4 tokens per second at 131k. llama.cpp with a q8_0 cache reaches 64k, the largest size I tried. Both fp16/f16 caches stop at 32k.
My take
For a single 16 GB NVIDIA GPU, ExLlamaV3 is today the fastest thing I have measured for this model: it decodes faster in every category, with and without MTP, and handled 131k tokens of context. I would not say it for every use. If the job is an agent that rereads its history every turn, with 10,000 tokens per round, the real-time difference is small, and llama.cpp reads long context faster. And if what you want is the highest quality that fits on the card without MTP, the custom 15 GB GGUF drifts less from the 8-bit model than the EXL3 of similar size.
It is not a pure engine comparison either. The engine, the quantization format and the shape of MTP all change at once. What I measure is what you get by picking one or the other with the same disk budget.
When to use it and when not
- Yes: an NVIDIA GPU on Linux or Windows, long answers (code, translation, writing), MTP on 16 GB with headroom.
- No: without an NVIDIA GPU, because ExLlamaV3 needs CUDA (its CPU offload is only for the experts of MoE models); agents where prefill dominates; when you want the best possible quality at 15 GB without MTP; when you need the GGUF ecosystem (Ollama, LM Studio, one file that runs anywhere).
Sources
- ExLlamaV3 v1.5.0, release notes, September 13, 2026.
- turboderp/Qwen3.8-27B-exl3, 3.00bpw branch.
- eval/spec_decode.py and eval/qbench.py from ExLlamaV3 v1.5.0.
- llama.cpp, commit 458681e, the
llama-serverused for the battery. - llama.cpp, PR #15550, the
--target-bpwdial, built from commit 325319b. - llama-cpp-python v0.3.35.
- Qwen/Qwen3.8-27B.
Comments
No comments yet. The first one is yours.