I got Qwen3.8-27B running at ~18 tok/sec decode & ~500-600 tok/sec prefill (at 64k context) on just 12GB VRAM + 8GB RAM, using ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3_S quant (~11GB). This recipe also works with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3_S quant, achieving similar speeds (about a 7% loss). This is the best Qwen3.8-27B quant I’ve tested so far (and I’ve tried everything), and for it to fit in such limited RAM/VRAM is wild. GSQ-RCO quantization is magic, it performs very close to the full precision weights in all of my testing. The reason it fits at all is Qwen3.8 is hybrid, so only 16 of the 64 layers need KV cache. With q4_0 for cache the full 64k is only about 1.1GB instead of 4GB for f16. I'm on a 9900X + 4070S 12GB + 32GB RAM for reference, using stock llama.cpp. Settings are -ngl 58 -ot token_embd=CPU -ctk q4_0 -ctv q4_0 -c 64000. Full build + serve scripts and all my numbers are here if you’d like to reproduce yourselves: https://github.com/bodhi37/Qwen3.8-27B-12GBVRAM-Recipe   submitted by   /u/bodhi371 [link]   [comments]