11°

Reels · 2 of 16

48 s

Reel · 48 s

More bits where it hurts: I quantized a Qwen3.8-27B to fit my 4070 Ti

In 45 seconds: why the model that fits is the dumber one, how the dial spreads the bits layer by layer, and the 94% fidelity that fits in 16 GB where Q4 doesn't.

Length
48 s
Published

Cuantizaciónllama.cppGGUFGPUModelos de lenguaje

This reel sums up More bits where it hurts: I quantized a Qwen3.8-27B to fit my 4070 Ti, where the method, the tables and what did not work are.

Read the articleOpen in the viewer