Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)

Wait 5 sec.

(Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!) About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could actually help speed up CPU/GPU offloading. Original post: Tried predicting which MoE experts get used next token to speed up CPU/GPU offload Well, quick confession first. I actually shelved that project shortly after. The reason? The speeds I was getting back then were kinda fake. My engine was aggressively pruning experts based on their router weights, basically dropping cold experts to get better performance. Sure, the numbers looked great, but doing that on an already quantized model was hurting output quality and coherence. Didn't really like that tradeoff, so I abandoned it and never released it. Fast forward to recently, and Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) drops. I downloaded the 68GB GSQ-RCO IQ2_XS build, hoping to run it on my daily driver. That's when I decided to revisit the idea, but this time without cutting corners. My setup GPU: RTX 3060 12GB RAM: 16GB DDR4, single channel (~19 GB/s bandwidth) OS: CachyOS / Arch Linux Storage: Mid-tier NVMe SSD (~2.1 GB/s read) Software: llama.cpp, built with CUDA And if you've tried running a 68GB MoE on a 16GB RAM machine with stock llama.cpp, you probably know how painful it gets. I'm talking 1.4–2.1 tok/s, with over 1,500 major page faults per token in some runs. Linux ends up constantly pulling model data from the SSD because there's simply not enough memory to keep the working set around. Then engines like Strata started showing up with claims of around 40 tok/s on consumer hardware. Pretty impressive, but there's a catch for people with less RAM. Some of these approaches rely on keeping around 24 GiB of experts pinned in memory using mlock. If you've only got 16GB RAM, you're obviously not doing that. Depending on the setup, you either run into OOM issues or end up with terrible performance. So I went back to my original idea and started implementing it properly as an optional feature inside llama.cpp: --moe-direct-io The goal this time was simple. No dropping experts, no sacrificing output quality, and bit-exact output compared to stock. The numbers All tests below were run with MemoryMax=6G using a cgroup. Model: Qwen 3.8 Flash Next IQ2_XS (68GB) Hardware: RTX 3060 12GB + 16GB DDR4 RAM Engine / mode Decode speed Major page faults per token SSD I/O Output Stock llama.cpp (mmap) 1.41–2.12 tok/s 1,140–1,565 208–312 MB/token Coherent Our engine (blocking, demand-only) 0.73 tok/s 0 ~206 MB/token Bit-exact Our engine (--moe-direct-io + prefetch) 20.14–21.13 tok/s ~0 (+1 across 32 tokens!) Sequential streaming Bit-exact to stock The blocking version is actually slower than stock, which makes sense. It's basically waiting on disk reads without doing much to hide the latency. The prefetching version is where things get interesting. Once the cache warms up, it sustains 20–21 tok/s, with some runs hitting 24+ tok/s. That's roughly a 10–15x speedup over stock llama.cpp on the same machine. And no, we're not getting those numbers by dropping experts. The output is bit-exact to stock. That's the part I'm most excited about, honestly. Being able to run a model this large on a 16GB machine without the usual page-fault nightmare is pretty much what I wanted to achieve with the original project. What's still rough It's not all perfect yet. There are a few things we're still working on. 1. Cold starts are noticeably slower Right now, the slots start empty (-1), so the first request on a new topic can start around 3.5–4.5 tok/s before ramping up to 20+ tok/s as the hot working set settles. We're working on offline hot-profile seeding so it can start with a useful working set instead of learning everything from scratch. 2. Prompt processing is slow Feeding a prompt of 512+ tokens can touch a huge number of experts in a short period. That puts a lot of pressure on the 72 slots per layer and causes the prefill stage to struggle. We're working on micro-batching prompt chunks (-ub 32) to help with this. 3. Speculative decoding gets weird with SSD offloading We found that standard MTP speculation can actually make things slower. Verifying 2–3 tokens can require loading the combined set of experts needed for those tokens from disk, which eats into the gains. Right now, confidence-gated speculation (min-p 0.8) or suffix prompt lookup seems more promising for this kind of setup. Anyway, that's where the project is at right now. Still plenty to improve, especially prefill and cold starts, but getting 20+ tok/s out of this setup without pruning experts is a pretty big deal for me. Happy to answer questions or get into the io_uring and slot-remapping implementation details if anyone's interested. (The second half of this post was written with some help from Claude.)   submitted by   /u/zyxciss [link]   [comments]