Hi all, having searched google and asked several AIs without success, maybe s.o. can help. I have 2 x dgx spark in a cluster. deepseek-flash is running successful. Now I want to test qwen3.8-flash-next from unsloth in the mtp version. Following start-command runs into a dgx-stall: ./llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:Q8_0 -ngl 999 -ngld 999 --load-mode none --fit off -fa on --host 0.0.0.0 --port 8090 --ctx-size 256000 --parallel 1 --chat-template-kwargs '{"preserve_thinking": true}' -sm layer --cache-ram 0 --spec-type draft-dspark --spec-draft-n-max 2 --reasoning on --seed 3407 --temp 1.0 --top-p 0.95 --top_k 20 --min_p 0.0 --presence_penalty 0.0 --repeat_penalty 1.0 --rpc 10.10.188.10:50052 There are two mtp files: mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf mtp-Qwen3.8-Flash-Next-Q8_0.gguf I can't figure out how to use them, start the server properly. Help appreciated!   submitted by   /u/Impossible_Art9151 [link]   [comments]