Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

Wait 5 sec.

Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game. Hardware: RTX 5070 12GB 32GB DDR4-2400 Ryzen 5 5600GT PCIe Gen3 Windows The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently. (In the video its around 16 minutes for 10k tokens and 10.41 tok/s Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set. Demos: https://www.youtube.com/watch?v=cOPumMlyj_4 https://www.youtube.com/watch?v=rc-uTjVpXM8 In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know   submitted by   /u/ayobluestarr [link]   [comments]