Looking for some assistance /ideation. I am running qwen flash next q4 in my Mac mini m5 64gb. QFN doesn’t fit so this is done by having as many experts hot in cache as possible and streaming in the rest from ssd. I’m getting 17.5 tks decode and 390 tks pp. Have done a bunch of optimisations including a carousel buffering system for the prompt processing which essentially loads faster than the GPU can prompt process in most cases. I feel like I have mostly maxed out this lane. The decode part 27% of the time is still the gpu waiting for experts to stream in from the ssd (see photo). The biggest unlock is really getting the gpu working more. I’m already doing mtp. Hot cache hit rate is 75% Some ideas I already have - Im already lookahead guess fetching the following layers experts, can I expand this more successfully. Current fetch accuracy is 72% - Use a seperate staging buffer for lookahead guess experts ahead (so I’m not evicting hot experts as guesses come in) … im learning a lot right now. Feel free to ask questions for clarifying. GitHub for reference. https://github.com/skeggsguy/Flash-next-ssd   submitted by   /u/turtleninja99 [link]   [comments]