Which speculative decoding are you using and why?

Wait 5 sec.

I have 2 working models at the moment Qwen3.8-Flash-Next-UD-Q4_K_XL and Qwen3.8-27B-UD-Q5_K_XL. For Flash model I'm using these params spec-draft-device: Vulkan0 spec-draft-n-max: 4 spec-draft-ngl: 99 spec-draft-p-min: 0.5 spec-type: draft-mtp,ngram-map-k And for 27B one I use these spec-draft-device: Vulkan0 spec-draft-n-max: 4 spec-draft-ngl: 99 spec-draft-p-min: 0.5 spec-type: draft-dflash,ngram-map-k I don't know if ngram-* like decodings make sense at all or maybe they make it worse. I didn't see much difference on my working tasks and you can't use llama-bench with speculative models to check it without bias. My hardware: AMD Ryzen AI 9 HX 470 w/ Radeon 890M, 96Gb DDR5 RAM, llama-server Vulkan backend from Nathan   submitted by   /u/cradlemann [link]   [comments]