I got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming

Wait 5 sec.

Bit of a side project I wanted to share. The metrics are 17.5tks decode, 360tks prompt processing based testing against my normal ai usage. I tested a couple of new things others haven’t done (at least that I’ve seen). Setup a carousel buffer for streaming in experts for prompt processing which got my pp +30% tks. Tried a second external ssd to get parallel reads which got my +15% on both prompt processing and decode. Plus a long tail of small efficiency gains. I also setup a system where by you can have a chat application make a call to the server and effectively kick out a coding run (which is kept alive until after the chat then continues). Good if you run long coding jobs , but want to chat inbetween. Probably useful for all setups where you want to save on local caching memory. I also noticed there is still a lot of gains to be made. I make this statement as there is still a lot of essentially free time on decode where the gpu is waiting for experts to stream in. There’s also work that could be done for an optimised kernel on metal. I also think the way things are going with Qwen (flash next being a precursor to 4), we’re gonna see a lot more efficiencies we can take advantage of like the ngram table and the cheap hybrid attention caching. I’m really liking qwen flash next .. the coding is actually very good. I’m quite surprised in fact I’m leaving it on during the workday to do large jobs. The chat, decode would be technically fast enough IMO but not really with qwen. The actual issue qwen spends so long thinking, so the decode hurts. Anyone else working on this? I’d love to compare notes. Yes I’ve heard of strata it does look pretty sic. https://github.com/skeggsguy/Flash-next-ssd   submitted by   /u/turtleninja99 [link]   [comments]