What's the best setup for Qwen3.8 27b for a 16 gig VRAM?

Wait 5 sec.

Hello guys! I have been experimenting with qwen 3.8 for a long time and I hadn't been able to get reasonable speed. I am on an 5060Ti 16 GB, and although this gpu can game, I am aware that AI demands more than 16 GB. I am on a Fedora 44, AMD Ryzen 9600x and 16 GB system ram (16 GB system ram and 16 GB vram, totaling to 32 GB) and I would like to use llamacpp, although I would use any other tool if I could if it meant faster speed. I am looking for a large context. Atleast 128k context. The first question is, what quantization to pick? In my experience Q3 UD was satisfying, but I am looking for uncensored model. In my experience, MTP has never lived up to its hype for me (and I don't know why!?), which is why I am thoroughly lost on making a good setup after an honest week of experimentation, which is why I have resorted to ask here as a last resort.   submitted by   /u/SultanGreat [link]   [comments]