I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length. The results and code are below: Results (same hardware, my megakernel vs llama.cpp with MTP): - Writing code: 140 tok/s vs 73 (1.9x faster) - Coding with a 4K-token file in context: 95 tok/s vs 56 (1.7x faster) - Prompt processing (prefill): ~1,600 tok/s vs ~1,100 (1.4x faster) The speedup comes from running each whole speculative-decoding cycle in one kernel launch, so checking 4-5 drafted tokens costs about the same as checking one. It's an OpenAI-compatible server, so it works as a drop-in replacement for the llama server. Caveats and limitations: it's built for one model (unsloth's Q4_K_M) on 3090s only. The outputs match llama.cpp's apart from rare rounding near-ties. KVCache is fp16 up to ~80k context, then q8. Code and full benchmarks: https://github.com/L-Forster/open-jet/tree/master/megakernel   submitted by   /u/Adorable_Weakness_39 [link]   [comments]