I've been messing with MoE offloading for a while. Setup: Qwen3-Next-80B-A3B Q4_K_M (48.5 GB), RTX 3090, only 1/4 of the experts kept in VRAM, the rest read from NVMe when the router asks for them. When the router picks an expert that isn't in VRAM you can either wait for the SSD read or use the next best expert that's already on the GPU. Substituting everything wrecks quality (+5.7% ppl in my emulation tests). Waiting only for the router's top pick and for experts with gate weight >= 0.15, and substituting the rest, brought it down to +0.35%. In the real engine that rule costs more like +1.2%. Decode tok/s on a rented 3090 box (NVMe ~5.7 GB/s, 16 threads), both using about 15.7 GB of VRAM: | free RAM | stock llama.cpp (--n-cpu-moe 34) | patched | |---|---|---| | plenty | 72.9 | 108.4 | | ~32 GB | 65.8 | 97.8 | | ~16 GB | 31.8 | 94.5 | At 16 GB it still did 89 tok/s when reading every miss straight from the SSD. Perplexity was 1.6% higher than stock on the same text. On GSM8K (500 problems) it lost 1.8 points vs waiting for every expert, on HumanEval no real difference. Things to know before trying it: - it's a research patch, not a polished fork. It builds a benchmark tool, I haven't tested llama-server or llama-cli with it - only Qwen3-Next, only Linux + CUDA, one sequence at a time - the 16 GB case was simulated by locking RAM on a bigger machine - the table is decode speed while feeding real text through the model. In actual greedy generation it did 64-74 tok/s Code, run scripts and raw logs: https://github.com/SOCIALPINE/moe-miss-substitution Paper with the details, including what didn't work: https://doi.org/10.21203/rs.3.rs-11268552/v1 Has anyone tried something like this, or have numbers from slower SSDs? Curious how much the SSD matters.   submitted by   /u/Zestyclose_Reality15 [link]   [comments]