GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M

Wait 5 sec.

Using llama.cpp Ubuntu Vulkan prebuilt binary and the Acemagic miniPC S3A using an AMD Ryzen 7 6800H is a high-performance 8-core, 16-thread mobile processor launched on January 4, 2022, built on the 6nm Zen 3+ architecture loaded with 64GB of DDR5 RAM. It features a 3.2 GHz base clock, a 4.7 GHz boost clock, a 45W TDP, and powerful integrated (iGPU) Radeon 680M graphics. Tested Models Based on the benchmark commands and llama-bench output labels: GLM-4.7-Flash-MXFP4_MOE.gguf (Reported: deepseek2 30B.A3B MXFP4 MoE | 15.79 GiB) GLM-4.7-Flash-UD-Q4_K_XL.gguf (Reported: deepseek2 30B.A3B Q4_K - Medium | 16.31 GiB) GLM-4.7-Flash-Q4_K_M.gguf (Reported: deepseek2 30B.A3B Q4_K - Medium | 17.05 GiB) Note: The filename contains GLM-4.7, but llama-bench reads the internal GGUF header and reports deepseek2 30B.A3B. The benchmark data corresponds to a ~30B parameter MoE architecture. Average Performance Results Model Filename Reported Name Size Avg Prompt Processing (pp512) t/s Avg Token Gen (tg128) t/s GLM-4.7-Flash-MXFP4_MOE.gguf deepseek2 30B.A3B MXFP4 MoE 15.79 GiB 258.36 t/s 11.66 t/s GLM-4.7-Flash-Q4_K_M.gguf deepseek2 30B.A3B Q4_K - Medium 17.05 GiB 218.22 t/s 12.09 t/s GLM-4.7-Flash-UD-Q4_K_XL.gguf deepseek2 30B.A3B Q4_K - Medium 16.31 GiB 160.31 t/s 13.13 t/s (Values are arithmetic means of 3 runs. fa on = Flash Attention enabled) Summary Analysis 🔹 Hardware & Memory Context Device: AMD Radeon Graphics (RADV REMBRANDT) Integrated GPU Architecture: UMA (Unified Memory Access) with fp16: 1, bf16: 0, fp4: 0 Implication: The models (~16–17 GB) exceed typical iGPU VRAM, forcing offloading to system RAM. Performance is heavily bound by system memory bandwidth (~50–65 GB/s DDR5) and PCIe/NB link latency. The fp4: 0 flag confirms native FP4 compute is unsupported, so MXFP4 is emulated or converted at runtime. 🔹 Prompt Processing (pp512) vs Generation (tg128) Trade-off Format PP Speed TG Speed Best Use Case MXFP4 MoE 🥇 Fastest (258 t/s) 🥉 Slowest (11.66 t/s) Long context windows, RAG, document processing Q4_K_M 🥈 Balanced (218 t/s) 🥈 Balanced (12.09 t/s) General-purpose chat, mixed workloads Q4_K_XL 🥔 Slowest (160 t/s) 🥇 Fastest (13.13 t/s) Fast response generation, streaming UIs Why MXFP4 excels in PP: Despite lacking native FP4 support, the MoE structure and extreme quantization drastically reduce active compute and memory reads during attention scoring. Flash Attention further optimizes cache locality for prompt parsing. Why Q4_K_XL leads in TG: Generation is purely memory-bandwidth bound. The Q4_K_XL quantization layout appears better optimized for the RADV driver's memory prefetching, yielding ~13% faster token streaming than Q4_K_M and ~12% over MXFP4. 🔹 Consistency & Stability All runs show extremely tight standard deviations (±0.02–0.06 t/s for TG), indicating stable thermal/power delivery and no background interference. Outlier: Q4_K_XL's first run showed high PP variance (±17.76 t/s), likely due to cold cache/memory allocation overhead. Subsequent runs stabilized (±1.11 and ±1.54), typical of VM/page cache warmup. 🔹 Recommendations For Chat/Streaming: Use Q4_K_XL. Slightly slower prompt processing is negligible in typical conversational turns, but faster TG improves perceived latency. For RAG/Long Context: Use MXFP4_MOE. The ~60% PP speed boost dramatically reduces wait times for context loading, with minor TG impact being acceptable for batched or paused workflows.   submitted by   /u/tabletuser_blogspot [link]   [comments]