Hi, Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant. Made with love and precision (read the model card) on my 4 RTX 3090, I believe it's the best compromise you can actually get on 48Gb of VRAM. Running it with vLLM and MTP on 2 RTX 3090 in my workflows and it's working flawlessly at almost 100 tok/s decode (with MTP=5), while being less an overthinker than the OG 3.8 27B :) Around 21.6Gb used per card (vLLM launched with 128k context and FP16 KV cache). https://huggingface.co/TacGibs/Swift-1.5-Qwen3.8-27B-W8A16 Hope it'll be usefull to some else !   submitted by   /u/TacGibs [link]   [comments]