Update: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark

Wait 5 sec.

Hi folks, I've spent the last couple of months experimenting with Strix Halo and previously I released a proof of concept I called llama-halo-hybrid. I've continued updating it and it now performs very well. The idea is that you can take an R9700, or similar, and place dense parts of the model, KV, and some of the layers on the GPU and let the APU take the rest of the model. You can add the extra GPU through a PCIe extender (framework desktop), Occulink, or a thunderbolt dock depending on which machine you have. Detailed notes along with code in the repo on github. I'm not selling anything, this is all 100% open, MIT-licensed. It breaks 60+ tok/s decode and 2000+ tok/s prefill, supporting full 256k context. This is not some custom inference engine that requires a custom quant to run. This is llama.cpp modified to run whatever you want, albeit mostly tuned for Qwen and GLM families. After continuing to tinker with it, it now performs better than DGX Spark (albeit cheaper) running Qwen-3.8-flash-next and slightly better yet with the Swift-1.5 variant. Most of my testing was done with the Q4/Q4_K_XL models to balance size and quality. Note - if you are just using Strix Halo by itself, this is probably not the right tool. Check out gufo, which looks very promising. https://github.com/sixvolts/llama-halo-hybrid I would love any feedback you all have and happy to investigate tuning for different "sidecar" GPUs other than the R9700 if there's demand and I can get my hands on one.   submitted by   /u/darklordfireape [link]   [comments]