Your Quantized LLM Is Not Slow Because of the Quantization

Wait 5 sec.

The SymptomI spent months building a 2-bit quantization scheme for Qwen3. The model went from 8 GB to 2.6 GB, a 4.5x reduction. Then I measured throughput.It was barely faster than the FP16 baseline.