Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

Wait 5 sec.

TL;DR: Halogen v0.14.0 with its native .hgn weights is the fastest, followed by gufo and CIRU. Halogen is closed source and runs in Docker. gufo is open source and loads 4x faster from cold. gufo is also fastest to first token on follow-ups (1.6-1.9 s against 2.6-3.3 s). I benchmarked different engines for Qwen 3.8-Flash-Next on an AMD Strix Halo (ASUS ROG Flow Z13 GZ302 with Ryzen AI Max+ and 128 GB of RAM). All the engines were run via LlamaStash (My own orchestrator tool, the tool does not add any overhead) at 70 W TDP on performance profile on Arch Linux. Here are the results of the benchmark: First was a screening round at 64k context with 50% filled (32k prompt). Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval Halogen 0.14.0 Halogen native (.hgn) 1.9 min 1,045 39.3 84% 14/14 gufo (ROCm 7.2.4) UD-Q4_K_XL 2.1 min 1,033 32.9 74% 14/14 gufo (ROCm 10.0) UD-Q4_K_XL 2.1 min 1,009 33.1 72% 14/14 CIRU (MTP 3) CIRU IU4 2.3 min 814 33.8 72% 14/14 CIRU (n-gram + MTP 3) CIRU IU4 2.3 min 808 31.9 64% 14/14 Halogen 0.14.0 UD-Q4_K_XL 2.4 min 1,030 28.1 84% 14/14 CIRU (MTP 6) CIRU IU4 2.7 min 796 28.6 47% 14/14 strixllama (llama.cpp fork) UD-Q4_K_XL 2.7 min 716 29.7 71% 14/14 rdna-boosts (llama.cpp fork) UD-Q4_K_XL 2.7 min 788 22.0 52% 14/14 llama.cpp, Unsloth build (Vulkan) UD-Q4_K_XL 3.7 min 314 28.4 56% 14/14 llama.cpp, Unsloth build (ROCm) UD-Q4_K_XL 4.8 min 285 20.5 50% 14/14 CIRU UD-Q4_K_XL + Unsloth MTP head did not finish 884 11.8 0% - 3-turn time is the time to first token plus the time to write 1,000 tokens, added up over the first question and 2 follow-ups. Output length varies a lot with sampling, so this compares the engines on equal output. MTP accept is the share of drafted tokens the model kept. Retrieval is how many of 7 exact values from the prompt the model got right, over both runs. Qwen model card sampling, 2 runs each. Top 3 engines from the screening round got a 128k context (50% and 75% filled) and 256k context (50% filled) run. 128k context, 50% filled (64k prompt) Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval Halogen 0.14.0 Halogen native (.hgn) 2.4 min 1,148 40.4 84% 16/16 gufo (ROCm 7.2.4) UD-Q4_K_XL 2.7 min 1,047 32.3 72% 16/16 CIRU (MTP 3) CIRU IU4 3.3 min 808 30.2 64% 16/16 128k context, 75% filled (97k prompt) Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval Halogen 0.14.0 Halogen native (.hgn) 2.9 min 1,103 39.3 85% 16/16 gufo (ROCm 7.2.4) UD-Q4_K_XL 3.3 min 1,025 30.7 70% 16/16 CIRU (MTP 3) CIRU IU4 4.1 min 789 27.8 62% 16/16 256k context, 50% filled (130k prompt) Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval Halogen 0.14.0 Halogen native (.hgn) 3.4 min 1,093 38.9 83% 14/16 gufo (ROCm 7.2.4) UD-Q4_K_XL 4.0 min 997 27.9 70% 14/16 CIRU (MTP 3) CIRU IU4 5.0 min 767 25.6 60% 15/16 Retrieval here is 8 values per run. All 5 misses are the same answer: the right glyphs for UNICODE_SPINNER, with the leading & dropped. End-to-end, 10 Aider polyglot Python exercises run through pi -p (128k window, 50% cell), graded by their own tests: Engine Passed Total time Halogen 0.14.0 10/10 21.5 min CIRU (MTP 3) 10/10 24.6 min gufo (ROCm 7.2.4) 10/10 36.0 min   submitted by   /u/deepu105 [link]   [comments]