Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra

Wait 5 sec.

https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality. This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it! Measurement This pack Swift llama.cpp IQ3_XXS Prefill · 25k prompt 3,191 tok/s 1,427 tok/s Prefill · 95k prompt 2,928 tok/s 1,307 tok/s Decode · after 4k prompt 113.7 tok/s 62.7 tok/s Decode · after 95k prompt 81.1 tok/s 44.7 tok/s Top-1 agreement with Swift BF16 91.0% 84.1% 91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy. Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.   submitted by   /u/DankpawsDev [link]   [comments]