Possible to run Qwen3.8-Flash-Next-4bit on Mac with 64GB?

Wait 5 sec.

I have m3Max 64GB, and I can allocate up to 58GB to GPU. Is that possible to run Qwen3.8-Flash-Next-4bit? I'd like to avoid lower quant than 4bit. I know 58GB can't load everything, but I think there are now some engines that let you load only some parts of the model and stream from SSD? If so with what engine: Strata, omlx? Also roughly how much context can I expect to run, and at what speed? Thanks!   submitted by   /u/chibop1 [link]   [comments]