LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM

Wait 5 sec.

https://github.com/woct0rdho/transformers5-qwen3.5-recipe An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed. On Strix Halo it trains context chunk size 2048 at 9.5 s/it. That's 200 token/s. There is still room to optimize, compared to > 1600 token/s PP we've achieved, and the common sense that LoRA training takes 2-3x work of PP. Since transformers 5.18, initial support for modern GGUF has been merged, and we can expect more work in this direction. Spoiler: In the torch-ggml-ops repo there is something called GGTensile. Basically it's Tensile-like asm-level optimization on MMQ kernels. We already see it's faster than HIP in many cases. I'll make a new post when I have something to show on this. I guess I'll skip DeepSeek-V4.1, unless someone can quantize or prune it to < 125 GiB. &#32; submitted by &#32; /u/woct0rdho [link] &#32; [comments]