Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

Wait 5 sec.

I wanted to see how far I could push a fairly ordinary laptop with a huge MoE model. Turns out, Qwen3.8 Flash Next 176B can run on: RTX 3080 Laptop — 16GB VRAM 32GB system RAM SSD No 128GB/256GB RAM workstation and no multi-GPU setup. I’m running it with TensorSharp, my open-source local LLM inference engine: TensorSharp on GitHub The interesting part for me wasn't simply getting a 176B model to load. I wanted to make a model much larger than both available VRAM and RAM actually usable. The approach is basically: Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD. Rather than treating SSD as a last-resort swap space, TensorSharp coordinates the different memory/storage tiers around MoE execution and tries to keep the right experts/data in the right tier at the right time. I previously benchmarked TensorSharp against llama.cpp and got very encouraging results. This time I wanted to compare it with Strata, since Strata's approach to running large models with constrained memory is particularly interesting. Here are the results from the attached benchmark: Measurement TensorSharp Strata Decode tokens/s 11.09 (9.22–14.02) 10.24 (9.37–10.46) Whole-process time 16.54s (14.95–19.31) 62.15s (59.76–66.89) Device-wide GPU peak 14,832.5 MiB 15,729 MiB OS peak working set 19.74 GiB 18.51 GiB The decode throughput is fairly close: 11.09 vs. 10.24 tok/s. What surprised me more was the end-to-end result: 16.54s vs. 62.15s in this test. I think this points to an interesting direction for local LLM inference. For huge sparse MoE models, the question may not simply be: “Do I have enough RAM/VRAM to fit this model?” but rather: “How efficiently can the runtime coordinate VRAM, RAM, SSD, caching, and expert activation?” With the right quantization and memory hierarchy, you can apparently do some pretty ridiculous things on consumer hardware. I’d be especially interested if anyone here has tried the same model with llama.cpp, Strata, or another MoE/offloading implementation. It would be great to compare results on similar hardware.   submitted by   /u/fuzhongkai [link]   [comments]