https://huggingface.co/Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit Qwen3.8-27B CODER — IQ4_XS imatrix · 24 GB card fit · ~262k context · MTP draft Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context. It built the importance matrix from code-heavy text and kept the MTP head at Q8_0, so --spec-type draft-mtp works without a separate draft model. The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context. File File Type Size Inside Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf IQ4_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8_0, token embeddings at Q4_K Measured speed (real use, not a synthetic benchmark) These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself: Context already in the window Requests Decode, median Decode, range 65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s 131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s MTP draft acceptance: the median is 85% (the middle half of requests falls between 75% and 94%). That works out to about 2.7 tokens per decode step at draft depth 2. Prefill: 387 t/s for a cold 108k-token prompt; 175–183 t/s for about 4.5k new tokens added at 147k–156k depth. VRAM: 24.2 of 24.6 GB in use at 245,760 tokens of context, with a q4_1 KV cache and the vision projector on the CPU. Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot. Run it with llama.cpp llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \ -c 262144 -np 1 -ngl 99 --flash-attn on \ --cache-type-k q4_1 --cache-type-v q4_1 \ --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \ --jinja --reasoning on --reasoning-format deepseek \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \ -b 2048 -ub 512 --cache-reuse 256 Context: -c 262144 is what fits next to the weights on a 24 GB card with a q4_1 KV cache. The model's native window is 262,144 tokens. On a smaller card, lower -c first. Speculative decoding: --spec-type draft-mtp drafts with the MTP layer inside this file, so no separate draft model is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235. Sampling: these are Qwen's recommended settings, and they are also stored in the file. Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB). Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context. How LexiPanel made it Conversion: the source weights were converted to a BF16 GGUF with the MTP head included. Importance matrix: computed from about 300k tokens (570 chunks) of code-heavy calibration text. Three quarters is Python source (the standard library and installed packages). The rest is technical documentation, READMEs and license texts, the kind of text a coding agent's context fills with. Quantization: llama-quantize from llama.cpp b11182 made the IQ4_XS file with that matrix. The MTP head stays at Q8_0 so its drafts stay accurate, and the token embeddings are Q4_K. Fitting the card: LexiPanel's Fit planner chose the mix, quality first, for one 24 GB card at 262144 tokens of context. It took the best quality that card could afford at that context, not the smallest file. Credits and license Quantization, importance matrix, MTP draft setup and card fit: LexiPanel. Tools: llama.cpp. Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility. Qwen3.8-27B CODER — IQ4\_XS imatrix · 24 GB card fit · \~262k context · MTP draft Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context. It built the importance matrix from code-heavy text and kept the MTP head at Q8_0, so --spec-type draft-mtp works without a separate draft model. The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context. File File Type Size Inside Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf IQ4_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8_0, token embeddings at Q4_K Measured speed (real use, not a synthetic benchmark) These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself: Context already in the window Requests Decode, median Decode, range 65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s 131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s MTP draft acceptance: the median is 85% (the middle half of requests falls between 75% and 94%). That works out to about 2.7 tokens per decode step at draft depth 2. Prefill: 387 t/s for a cold 108k-token prompt; 175–183 t/s for about 4.5k new tokens added at 147k–156k depth. VRAM: 24.2 of 24.6 GB in use at 245,760 tokens of context, with a q4_1 KV cache and the vision projector on the CPU. Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot. Run it with llama.cpp llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \ -c 262144 -np 1 -ngl 99 --flash-attn on \ --cache-type-k q4_1 --cache-type-v q4_1 \ --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \ --jinja --reasoning on --reasoning-format deepseek \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \ -b 2048 -ub 512 --cache-reuse 256 Context: -c 262144 is what fits next to the weights on a 24 GB card with a q4_1 KV cache. The model's native window is 262,144 tokens. On a smaller card, lower -c first. Speculative decoding: --spec-type draft-mtp drafts with the MTP layer inside this file, so no separate draft model is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235. Sampling: these are Qwen's recommended settings, and they are also stored in the file. Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB). Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context. How LexiPanel made it Conversion: the source weights were converted to a BF16 GGUF with the MTP head included. Importance matrix: computed from about 300k tokens (570 chunks) of code-heavy calibration text. Three quarters is Python source (the standard library and installed packages). The rest is technical documentation, READMEs and license texts, the kind of text a coding agent's context fills with. Quantization: llama-quantize from llama.cpp b11182 made the IQ4_XS file with that matrix. The MTP head stays at Q8_0 so its drafts stay accurate, and the token embeddings are Q4_K. Fitting the card: LexiPanel's Fit planner chose the mix, quality first, for one 24 GB card at 262144 tokens of context. It took the best quality that card could afford at that context, not the smallest file. Credits and license Quantization, importance matrix, MTP draft setup and card fit: LexiPanel. Tools: llama.cpp. Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.   submitted by   /u/W61k3r [link]   [comments]