Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context

Wait 5 sec.

After my Splash M1 port the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP) would take forever, so I took antirez's ds4 ( DwarfStar ), which already runs it, and spent the time making it fast on the M1. Same laptop: M1 Max, 32 GPU cores, 64 GB. What I care about is not the peak but how little it falls in long agent sessions, where local models usually die. TL;DR, MTP on: One chat grown to its full context: Q2_0 (35 GiB) decodes at 44 tok/s at 4K, 38.8 at 256K, 35.4 at 398K. Prefill 328 to 292 tok/s. Five prompts on a loop for 5 minutes: 43.2 tok/s at the start, 43.1 at the end. GPU 73 to 75 C, no throttling, 0.70 J per token. Real OpenCode sessions: 162 requests, context up to 371K. Median decode 44 tok/s under 64K, 38 at 128 to 256K. IQ3_XXS (44 GiB, better answers): 35.2 tok/s at 4K, 31.7 at 259K. Against my Splash port of Qwen3.8-27B: at 128K+ in OpenCode, 14 tok/s vs 38; prefill 4 to 7x faster. For scale: a DGX Spark with the popular vLLM NVFP4 + MTP recipe does 36.8 tok/s on prose, 45.8 on code. Bonus: DeepSeek V4 Flash (81 GiB) on the same 64 GB Mac, streamed from the SSD: 12 tok/s at 4K, 9 at 127K. https://preview.redd.it/6qs2jrbeqath1.png?width=2080&format=png&auto=webp&s=5f03c22100c174b992e2de4d6f0729c50eb3218a How it works Stock ds4 could not open ISTA's files files; with a patch to load them it did 21 to 23 tok/s, 17.7 at 262K. What changed: Loading the files. Metal decoders for every quant type in ISTA's mixed-precision GGUFs, each checked bit-exact against the CPU. The n-gram table is 95 GiB in BF16, so the fork reads the 16 rows a token needs from ISTA's 28.8 GB IQ4_NL shard on the SSD. The MTP head is missing from ISTA's release, so it is grafted from a separate 1.5 GB GGUF. Decode kernels for the M1. Every token reads about 4 GB of weights, so it is all about matvec kernels: half2 tricks for Q2_0, typed block pointers (5 to 15%), and 2- and 3-token variants that read the weights once for a whole MTP verify. MTP on code went from 24 to 45 tok/s. Why the curve stays flat. The attention layers only attend to blocks an indexer picks, which should be cheap, but on the M1 the indexer's scalar scorer took 10.7 ms of a 42.5 ms token at 128K. A bit-identical vector scorer for every Apple GPU, a coalesced top-k select and an attention loop that loads four keys per round took decode at 128K from 23.8 to 33.8 tok/s and at 262K from 17.7 to 29.7, byte-identical output. Keeping the GPU busy. The n-gram SSD read now overlaps the first layer, and the sampler went from 1.34 to 0.13 ms per token. Prefill. Small tiles for leftover expert tokens (182 to 223 tok/s at 300 tokens), a 4-at-a-time Q6_K decoder, weight prefetch and register tiles from my Splash port (276 to 327 at 16K). Agent turns that do not replay. Three of four layers are recurrent and cannot be rewound, so a retried turn meant prefilling everything again. The server now keeps the state at the last turn marker: 109 s to 0.3 s on a 31K prompt. Past 262K. YaRN factor 2: needles 20/20 at 400K, 18/20 at 524K, same NLL as from position zero. Quality: every kernel is tested against the reference path, and score_official gives the same NLL before and after (0.45044 vs 0.45045). Benchmarks One chat to the full context (each turn adds ~30K tokens of C source, reasoning xhigh, 800 tokens out): Context Q2_0 decode Q2_0 prefill IQ3_XXS decode IQ3_XXS prefill 4K 44.0 tok/s 328 tok/s 35.2 tok/s 241 tok/s 128K 43.0 320 34.2 278 256K 38.8 313 31.7 (259K) 258 (259K) 398K 35.4 292 Sustained load (npanj's five Splash prompts for 5 minutes, xhigh, macmon, fans on a curve to full speed at 80 C): https://preview.redd.it/y8q63vkarath1.png?width=1760&format=png&auto=webp&s=dcd4fee4f7fda2e31426d4c6ca337f6afa465735 Decode First / last quarter GPU while generating Energy per token Q2_0 43.2 tok/s 43.2 / 43.1 73.1 C IQ3_XXS 35.0 tok/s 35.1 / 35.0 73.9 C OpenCode, against Splash 27B. The same Three.js galaxy task as in my last post, reasoning medium; for Q2_0 a second task on top pushed the context to 371K. https://preview.redd.it/16s8lhggrath1.png?width=1950&format=png&auto=webp&s=4c937d6f4b23e3d7a1608a386541275bcb2a5fa2 Decode per turn (median) Qwen3.8-27B, Splash M1 Flash-Next Q2_0 Flash-Next IQ3_XXS 0-64K 30.6 tok/s 44.7 tok/s 34.2 tok/s 64-128K 19.8 42.7 34.6 128-192K 14.2 38.1 32.7 Prefill of new tokens 79 down to 42 334 down to 302 282 down to 260 On short prompts they are even (41.5 vs 43.2 tok/s). In a session they are not: the 27B is dense and its attention reads the whole cache, so a 30K tool result at 128K takes it over 10 minutes to read and Flash-Next about 90 seconds. Different models, and I am not claiming a 2-bit Flash-Next answers as well as a 4-bit 27B; the 27B also fits a 32 GB Mac, Flash-Next needs 64. DGX Spark. vLLM NVFP4 + MTP on one Spark: 36.8 tok/s on prose, 45.8 on code. The M1 Max: 43.2 on mixed prompts, 44 to 45 on code. The Spark still prefills 5x faster and runs 4-bit weights, and tuned INT4 builds there report more, but I did not expect this to be close. DeepSeek V4 Flash (81 GiB Q2, experts streamed from the SSD, 128K context): 11.8 tok/s for 5 minutes at 66.5 C, 12.1 to 9.1 tok/s up to 127K, prefill 34 to 94. Fine for chat, slow for agents. Along the way I found and fixed a few bugs in ds4 itself; they are upstream too, and the PRs to antirez will go gradually, one at a time. Try it dstar is one launcher for every model ds4 runs: it picks the context that fits your Mac, turns on MTP and vision, streams models larger than your RAM from the SSD. Prebuilt release (macOS 15 or newer, nothing to compile; installs into ~/dstar and puts dstar on your PATH): curl -fsSL https://raw.githubusercontent.com/paperniuk/ds4/m1-flash-next/install.sh | bash dstar doctor # what your Mac fits dstar pull # Qwen3.8-Flash-Next Q2_0 + MTP + vision, 67 GB (--quant iq3 for IQ3_XXS) dstar serve # OpenAI/Anthropic compatible server on http://127.0.0.1:8010/v1 dstar opencode # provider block for OpenCode dstar models # DeepSeek V4 / V4.1 Flash, GLM and the rest; dstar serve deepseek, or any GGUF path Or from source: git clone https://github.com/paperniuk/ds4.git && cd ds4 && make, then the same commands as ./dstar from the repo folder. 64 GB and up, for now. Flash-Next needs a 64 GB Mac. 32 and 48 GB are in the plans, but until then my Splash M1 port or Original splash (M3+) with Qwen3.8-27B (21 GB) is the better choice there. GPU memory limit. macOS lets the GPU wire ~48 GiB of a 64 GB Mac by default, enough for Q2_0 up to 131K. For longer contexts: sudo sysctl iogpu.wired_limit_mb=57344 (Q2_0 at 262K/400K) or 61440 (IQ3_XXS at 262K, Q2_0 at 524K). It resets at reboot, and dstar serve prints the exact line when the context you ask for does not fit. Not only M1/M2. Almost nothing is M1-only, so on M3, M4 and M5 this could be one of the fastest ways to run Qwen3.8-Flash-Next right now, especially on 64 GB Macs. Untested though; dstar doctor plus one dstar chat run from your Mac would be a great report. What this says about the M1 Max. To me this is the real result: a 2021 laptop runs a ~126B-parameter MoE (6.7B active) at 35 to 44 tok/s across 400K tokens and decodes like NVIDIA's 2025 AI box. And it is not the ceiling. Every token reads about 4.2 GB, so plain decode at 35.5 tok/s uses ~150 GB/s of the M1 Max's 400 (the Spark has 273); single big matvecs reach 200 to 300 GB/s, the rest is lost to ~1000 small dispatches per token. Prefill at 327 tok/s is ~4 TFLOPS of the ~10 TFLOPS fp16 matrix peak (9.9 measured). Plenty of headroom left on a five-year-old chip. Links: The fork: https://github.com/paperniuk/ds4 The prebuilt release: https://github.com/paperniuk/ds4/releases/latest ds4 (DwarfStar) by antirez: https://github.com/antirez/ds4 The GSQ-RCO quants by ISTA-DASLab: https://huggingface.co/ISTA-DASLab npanj's benchmark prompts: https://github.com/npanj/splash-plus DGX Spark numbers: https://madeye.github.io/qwen38-flash-next-on-dgx-spark/ The engine is antirez's and the quants are ISTA-DASLab's; this is an unofficial fork for Apple Silicon.   submitted by   /u/Erp4759 [link]   [comments]