Regular MLX QMM (quantized matmul) stages operands (dequant) to FP16 before matrix multiplication. But if you modify MLX to stage to FP8 instead, you can use the 2x faster FP8 matrix path on the new Apple Silicon M6 hardware! The same idea works on the M5 family: stage 4-bit affine weights to int8 before the matmul and you get a similar prefill boost (M5 and M6). Benefits apply to prefill (+40% prefill Qwen3-8B or +50% Qwen3.8-27B). Decode is bandwidth limited so not much difference on decode side and better to let it use default FP16 staging on decode side. It's an experimental fork, not upstream (mlx#4627) and quality was measured: worst-case perplexity increase under ~1%. FP8 performance improvement needs an M6 on macOS 27 with a deployment-target-27 build; int8 works on M5 and later. This was quite an interesting research project and it helped me get a deeper understanding of the math behind LLMs on the hardware side. 😄 All details, code and evidence available on my blog: https://precisit.com/en/blog/apple-matrix-formats/ The MLX fork itself with FP8 + Int8 QMM patch: https://github.com/precisit/mlx/tree/staged-8bit-qmm Disclosure: the blog is from Precisit, where I work.   submitted by   /u/Brilliant-Hall1387 [link]   [comments]