This is a follow-up to my post from yesterday (https://www.reddit.com/r/LocalLLaMA/comments/1x2erdj/qwen3827b_on_a_single_3090_140_toks_on_code_with/) detailing the CUDA megakernel. Thank you all for the positive feedback and PRs! RECAP: The megakernel is a CUDA engine for Qwen3.8-27B that runs 1.4-1.9x faster than llama.cpp on a single 3090. The most common questions were about accuracy, the baseline and context size, so here are the numbers. Code and benchmarks: https://github.com/L-Forster/open-jet/tree/master/megakernel Firstly, I benchmarked the engine accuracy against llama.cpp. No measurable loss in quality: Decode accuracy vs llama.cpp Context Positions MK top token MK KL LC top token LC KL 1K 13,299 99.08% 0.0009 98.68% 0.0019 30K 1,999 99.65% 0.0007 99.35% 0.0009 96K 1,999 99.60% 0.0008 99.45% 0.0014 MK = megakernel, LC = llama.cpp batched. Top token = how often the top token matches the reference. KL = mean KL divergence in nats. Perplexity (1K run only): megakernel 2.4274, llama.cpp batched 2.4283, reference 2.4263. Log-loss vs the reference: +0.0005 +/- 0.0004 at 1K, +0.0018 +/- 0.001 at 30K, +0.0012 +/- 0.001 at 96K. Not significant. The 30K and 96K rows are one sequence of code each & the 1K row is code and prose. Speculative decoding is exact: greedy output is the main model's greedy choice, and sampling keeps the same distribution as without drafts. Speed (same 3090, megakernel vs llama.cpp with MTP) Writing code: 140 vs 73 tok/s Reasoning mode: 78 vs 51 tok/s With 30K tokens of code in context: 73 vs 52 tok/s Prompt processing (prefill): ~1,600 vs ~1,100 tok/s With speculative decoding off they are about equal (42.9 vs 40.4). The gain comes from checking 4-5 drafted tokens per pass for about the cost of one. llama.cpp ran at 2 drafts, the megakernel at 4. Since the first post Two PRs have been merged: the loader now checks every tensor's type, and the server warms up before accepting requests. Model download fixed after unsloth removed the Q4_K_M from their repo. It also loads the Q4_K_M from lmstudio-community and mradermacher. openjet setup shows which engine each model will use. Quantisation support This table details the quantisation support, with unsloth Q4_K_M being the native model and what I benchmarked, lmstudio and mradermacher being semi-compatible, not fully benchmarked, tentative results show similar speed to native after type conversions. Other quants fall back to llama.cpp Qwen3.8-27B file Runs on unsloth's original Q4_K_M (what openjet setup downloads) megakernel Q4_K_M from lmstudio-community, mradermacher megakernel, semi-compatible Other quants (unsloth UD-*, bartowski, GSQ-RCO and others) llama.cpp Limits Single GPU. RTX 30/40/50 series with 24 GB or more. MK is tuned on a 3090 only. 147K context runs on a 3090 (fp16 KV cache to ~80K, using q8 above). No task-level pass-rate benchmark yet. Contributions that would help Tunes for specific hardware: RTX 4090, 5090 and other 24 GB+ cards. It has only been tuned on a 3090. Support for other quantisations of Qwen3.8-27B. Please include benchmarks with the PR: your GPU, and bench/bench.py or bench/code_bench.py numbers before and after. How to contribute: see CONTRIBUTING.md in the repo. Comment the quant you use for Qwen3.8-27B so I know what to add support for next! (also planning for Qwen 4 when that drops).   submitted by   /u/Adorable_Weakness_39 [link]   [comments]