I'm running Qwen 3.8 27B Q4 XS with ~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM. Build llama.cpp with Vulkan (ROCm uses more VRAM and has leak issues): cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then: llama-server \ --model Qwen3.8-27B-UD-IQ4_XS.gguf \ --mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \ --n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \ --batch-size 2048 --ubatch-size 512 \ --flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \ --load-mode none --fit off \ --cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \ --no-context-shift --jinja --reasoning-format deepseek \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \ --threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080   submitted by   /u/Haunting-Stretch8069 [link]   [comments]