Advice needed on a budget hybrid build for Qwen3.8-Flash-Next at 4-bit

Wait 5 sec.

After seeing how good the cloud models are getting, I feel like this is something we cannot let big tech hold over us so deiced to build a budget local box. After testing about 20 open models, Qwen3.8-Flash-Next (medium reasoning) was the only one that passed my task without inventing config options when used with a harness that forced doc lookups. So the box is built around that model. GLM-5.3-Flash performed even better but it's too big for my budget. Planned build (Netherlands prices): Ryzen 5 9600, about €200 MSI B850 Gaming Plus MAX WiFi, about €170 2×48 GB DDR5-5600, €1,199–1,549. Two sticks only to avoid the four-stick speed penalty. Used RTX 3090, about €1,150–1,500 Case, 850 W PSU and NVMe I already own That's about 79 GB of the model in RAM (experts plus the 28.8 GB n-gram table) and about 20.6 GB on the card. Questions: Will 6 Zen 5 cores hold back generation with 40 MoE layers on the CPU? At 96 GB with about 79 GB of mode will 17 GB be enough for the OS, a sandbox container and an embedding model? Should I use --mlock? Is DDR5-6000 worth it over 5600? Has anyone run Unsloth's MTP branch with experts on the CPU? What speedup did you get, and does it break the prompt cache on the DeltaNet layers? Is anything wrong with a used 3090 here? Also has anyone tried the Arc Pro B60 (€772 new) workable on Vulkan or SYCL with this model yet? It's so much cheaper but I am worried becase of the software. Super exciting to work on it but I am really inexperienced so this would be my first build. Does it make sense?   submitted by   /u/Confident-Truth3607 [link]   [comments]