Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards

Wait 5 sec.

Hello all! Thank you to everyone that contributed their feedback and notes for LlamAmpere the last time I posted. I've continued to chip away at improvements for the 3090 crowd -- this release is modestly faster (3-4%), with a few hundred MB smaller runtime (if you're using YaRN, you can now support up to ~340K ctx). But, it is mostly targeted to the 12GB card crowd, who had not been the focus of the previous two releases. In addition to a series of configurable runtime refinements + adjustments (compact MTP caches, 16 bit activations, redundant overhead items removed), I've also added a new variant of KVaRN (Staged + Journaled KVaRN), that reduced KLD vs paper faithful + competitive versions by ~40%. 4/4 is my new recommended default (significantly better performance vs a 8/4 KV cache) for larger cards, but for people looking to get maximum ctx, the 3/3 bit (0.001 nats KLD) and 3/2 (0.0024 nats) are still very solid (the KLD from 3/2 is smaller than the performance drop you see going from 4 XL to 4 S quants). They support ctx lengths of 205k to 230k for the model tested, respectively, at 11 GB (significantly more if you're using 100% of GPU in a headless config). The model tested was 2.3bpw fusion of swift-1.5-uncensored and mirai's 2.5 bpw model. It retained 85% of the BF16's performance on LiveCode Bench across 7 runs on 3060/80/ti cards. On 3080/ti it averaged ~65-70 tps (thanks to MTP) on the evals. By no means is the model lossless, and I won't pretend that it is like a certain other model did. But using it personally in hermes for a couple days, and running it through coding evals + pi testing, it is genuinely usable for standard agentic tasks. Subjectively, I'd put it somewhere between the BF16 versions of qwen3.5 and 3.6, which is pretty good for a model that fits in 12GB with context! LlamAmpere is MIT license as before: https://github.com/JakeATX/llamAmpere Models (4 XS-M, 2.3bpw, etc) based on swift 1.5 are here: https://huggingface.co/collections/jakeatx/qwen38-27b-models (llamAmpere supports EXL if you prefer that family, but I am working on improving prefill + decode kernels, so still not quite first class speeds yet in the 3-5 bpw range) As before, please share your results, bugs, etc. They will get added to the backlog. Enjoy!   submitted by   /u/Brief-Tap-6616 [link]   [comments]