We published three Qwen3.8-Flash-Next REAP builds for Apple Silicon. Rather than call one “best,” here is the decision boundary from the exact artifact configs and published oMLX runs. All three are qwen4_exp derivatives with 48 transformer layers, top-10 routing, a retained 512-expert native MTP predictor, and a 262k configured context. The repository names retain BF16 for release identity, but the inspected configs are mixed-precision packages — not uniform BF16 checkpoints. RESULTS AT 4K CONTEXT Build HumanEval LiveCodeBench PP tok/s TG tok/s Peak memory REAP-384 oQ5e 97/100 no-thinking; 96/100 thinking 61/100 no-thinking; 62/100 thinking 679.1 53.9 72.8 GB REAP-384 oQ4e 91/100 no-thinking; 95/100 thinking 55/100 no-thinking; 56/100 thinking 689.5 59.3 61.5 GB REAP-288 oQ4e 91/100 no-thinking 56/100 no-thinking 692.3 48.7 46.0 GB Hardware: M4 Max, 40-core GPU, 128 GB unified memory. PP = prompt processing; TG = text generation. JUNDOT OQ4E COMMUNITY CONTEXT The Jundot oQ4e MTP package is a useful 512-expert, 4-bit community baseline. It is context, not a winner-ranking comparison: expert count, quantization, runtime, batch size, and some host details differ. Shared LiveCodeBench context Hardware Package Batch Sample Score Mensa REAP-384 oQ5e, thinking M4 Max 40c / 128 GB 5-bit 4 100/1,054 62.0% Jundot oQ4e MTP, thinking M4 Max 40c / 128 GB 4-bit 1 100/1,054 55.0% These use the same benchmark family, sample size, and host class, but they are still not a controlled head-to-head. In particular, do not read the table as a universal score delta or winner claim. Published 4k throughput context Chip / RAM Package PP tok/s TG tok/s Peak memory Mensa REAP-384 oQ5e M4 Max 40c / 128 GB 5-bit 679.1 53.9 72.8 GB Jundot oQ4e MTP M5 Max 40c / 128 GB 4-bit 1,286 62.5 72.3 GB The throughput rows are deployment snapshots, not a speed ranking. Jundot's record uses a newer M5 Max and a different 512-expert, 4-bit package. Jundot model card https://huggingface.co/Jundot/Qwen3.8-Flash-Next-oQ4e-mtp WHICH ONE SHOULD YOU USE? If your priority is... Pick Why Strongest measured coding result in this set REAP-384 oQ5e Highest published HumanEval and LiveCodeBench results here. A 384-expert build with lower memory use and faster generation REAP-384 oQ4e Used 11.3 GB less peak memory and generated 5.4 more TG tok/s than the 384 oQ5e snapshot. Maximum memory headroom REAP-288 oQ4e Its published footprint was 26.8 GB below 384 oQ5e and 15.5 GB below 384 oQ4e. The 288 oQ4e keeps the same 48-layer/MTP structure while reducing the target-expert bank to 288. MODEL CARDS REAP-384 oQ5e https://huggingface.co/mensaprodigy/Qwen3.8-Flash-Next-REAP-384-oQ5e-BF16-MTP-PLE REAP-384 oQ4e https://huggingface.co/mensaprodigy/Qwen3.8-Flash-Next-REAP-384-oQ4e-BF16-MTP-PLE REAP-288 oQ4e https://huggingface.co/mensaprodigy/Qwen3.8-Flash-Next-REAP-288-oQ4e-BF16-MTP-PLE IMPORTANT CAVEAT The accuracy figures are 100-question stratified samples. Throughput is workload-, host-, and configuration-specific. HumanEval and LiveCodeBench are kept separate; neither should be treated as a combined score or a claim of a universal leaderboard position. The model cards include the source oMLX records, runtime settings, context length, batch size, KV policy, and MTP state. Feedback on additional reproducible evaluation suites is welcome.   submitted by   /u/MensaProdigy [link]   [comments]