Improve token per second without touching quant

Wait 5 sec.

Spent the past month tweaking and experimenting with many different numbers to achieve 30tps. Hardware: -Rx6700xt 12gb vram (AMD) -2x16 ddr4 3200 ram -r5 5600x -llama.cpp vulkan sdk -window11 (no wsl switching since im not used to the environment) Im looking for any improvements to achieve maybe 40tps? without touching quant at all, q8 and q4_k_xl remains. I've experiment with threads at 6 is the best out of (4,8,12) Other than that im not sure on what to improve to achieve higher tps, any advices? appreciate it. ctx remains 100,000 |[tiel-coder-35b] model = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\Tiel-Coder-35B-A3B-MTP-UD-Q4_K_XL.gguf mmproj = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\mmproj-BF16.gguf no-mmproj-offload = on c = 100000 parallel = 1 flash-attn = on cache-type-k = q8_0 cache-type-v = q8_0 load-mode = dio fit = on fit-target = 256 # MiB n-gpu-layers = 99 n-cpu-moe = 28 batch-size = 2048 ubatch-size = 512 threads = 6 prio = 2 prio-batch = 2 spec-type = draft-mtp spec-draft-n-max = 3 spec-draft-p-min = 0.75 jinja = on chat-template-file = C:\Users\brain\Qwen-Fixed-Chat-Templates\chat_template.jinja temp = 0.6 top-p = 0.95 top-k = 20 min-p = 0.0 presence-penalty = 0.0 alias = tiel-coder-35b   submitted by   /u/Loose_Doubt367 [link]   [comments]