OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!

Wait 5 sec.

I'm genuinely shocked! I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX! Even a couple of weeks ago, I wasn't able to get it to run. I just tried the latest commit for fun, and it worked! It feels like some kind of sorcery to be able to run a 100GB model with 58GB allocated to GPU! I can even run up to 130k context window! Here are stats after running a short session with PI. Requests: 13 Total Prefill Tokens: 321,309 Cached Tokens: 292,209 Cache Efficiency: 90.9% Prompt Processing (excl. cached): 140.0 tok/s Token Generation: 15.8 tok/s Here are the settings I used on oMLX: Memory guard: Aggressive Hot Cache Limit (In-Memory Cache): 2GB Lightning MTP: on MoE Expert Offload: on RESIDENT EXPERTS: 50%   submitted by   /u/chibop1 [link]   [comments]