Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

Wait 5 sec.

https://github.com/ggml-org/llama.cpp/pull/27694 Now merged. Update your llama if you haven't done so yet. Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling. Main gain seems to be on prose generation. Tests above ran with thinking off, ngram-mod off.   submitted by   /u/Dreeew84 [link]   [comments]