I put together some text-only Qwen3.5 2B, 4B and 9B MLX 4-bit packages (including abliterated variants) perfect for local use — have at 'em. :)

Wait 5 sec.

Hey folks — I put together some text-only MLX 4-bit packages of Qwen3.5 great for local inference because nothing else quite exists in those specific configurations/sizes. Sharing them here in case they’re useful to anyone else running models on Apple Silicon. Brief summary: There are six packages: 2B, 4B and 9B, each in the original Qwen version and the corresponding Huihui abliterated version. They’re text-only (vision-removed) and 4-bit, which makes them incredibly svelte (Only 1GB for the 2B version!) and great for running passively with resources to spare. I've got them currently working on my writing software Minstrel doing summarization tasks and making contextual decisions (more on this later!), but they're of course free for everyone to use. Hopefully someone finds them useful! Models and download links on Hugging Face Note: These build on existing upstream models and community conversions, so my work here was mostly the text-only packaging and MLX conversion. The Huihui 4B GGUF source was already text-only, so all it needed was conversion. Credit to Qwen, Huihui, and the community conversion authors linked in each model card. Let me know if you run into issues. ✌️   submitted by   /u/sachasayan [link]   [comments]