ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (GGUF) High-Precision i1-Q5_K_M with True 131k Context on Consumer 24GB GPUs Most benchmarks in the community measure generation speed at trivial context depths (2k to 8k tokens). However, running a 27B parameter model at high quantization precision (Q5_K_M) across 131,072 tokens (128k) on a single consumer 24GB GPU without overflowing into slow system RAM or sacrificing attention fidelity is a fundamentally different challenge. I am releasing ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an optimized quantization suite built with a dedicated calibration imatrix, custom asymmetric tensor mapping, and native llamAmpere hardware acceleration. Hugging Face Model Card: https://huggingface.co/bjivanovich/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-GGUF Available Quants: i1-Q8_0, i1-Q6_K, i1-Q5_K_M (Primary), i1-Q4_K_M, plus mmproj-BF16.gguf for multimodal vision. The 131k Context & Q5 Precision Challenge on 24GB VRAM On a standard 24GB card (RTX 3090 / 4090): Weight Footprint: A standard 27B model at Q5_K_M occupies ~19.2 GB of raw weights. Context Memory at 131,072 Tokens: Standard FP16 KV cache for 131k tokens requires >24 GB on its own, making full-context inference impossible without dropping precision down to severe Q3/Q2 compromises or offloading layers to CPU RAM. MTP Quantization Pitfall: Standard community quants compress the Multi-Token Prediction draft block (blk.64) uniformly. At Q5 or Q4, this degrades draft accuracy, causing speculative acceptance to plunge from 85% down to ~55%, destroying generation speed. Our Architecture: Asymmetric Tensor Mapping + llamAmpere KV Compression To solve this, we applied an asymmetric layer-by-layer quantization layout calibrated on a custom domain-rich dataset (imatrix_atx_uncensored.dat): MTP Speculative Head (blk.64) Isolated at Q8_0: Guarantees near-lossless draft predictions, increasing acceptance rates to 76% - 88% (averaging 3.3 to 3.7 verified tokens per generation round). Attention Layers (attn_q, attn_k, attn_v, attn_output) Protected at Q6_K / Q8_0: Prevents attention drift and catastrophic reasoning decay at 64k, 96k, and 128k+ token horizons. FFN Layers (ffn_gate, ffn_up, ffn_down) at Q5_K_M: Absorbs standard compression without degrading semantic coherence. Unified Turbo KV Cache (-ctk turbo5 -ctv turbo4 or -ctk q8_0 -ctv turbo3): Compresses the 131,072 KV cache down to just ~3.5 to 4.2 GB of VRAM, allowing the entire Q5 model + full 131k context window to reside 100% inside the 24GB VRAM envelope. GPU Memory Footprint & Resource Breakdown (RTX 3090 24GB) Total VRAM Allocated: 23.4 GB / 24.0 GB (100% GPU offload, -ngl 99, 0 layers in CPU RAM). Model Weights (Q5_K_M Asymmetric): ~19.2 GB. KV Cache (131,072 tokens, Unified Turbo4/5): ~3.8 GB. System RAM Cache (--cache-ram 4096): 4.0 GB RAM dedicated to multi-session prompt state preservation. CUDA Compute Architecture: Ampere SM86 with FlashAttention-2 (-fa on) and hardware Tensor Core MMA fused kernels. Direct Benchmark Comparison: Standard Swift-1.5 Q5 vs ATX-Swift-1.5 Q5 Tested on Single NVIDIA RTX 3090 (24GB) with llamAmpere under Deep Context (~80,000 to 98,000 active tokens) | Measured Metric | Standard Swift-1.5 Q5 (mradermacher) | ATX-Swift-1.5 Q5 (Our Quant) | Real Delta | | Sustained Speed (~80k-98k ctx) | 44.94 to 45.79 t/s (Tasks 740, 19637) | 50.05 to 53.72 t/s (Tasks 0, 100, 155) | +5.5 to +8.0 t/s (+13% to +17%) | | Burst Generation Peaks (tg_3s) | 45.6 to 52.3 t/s | 58.10 to 63.05 t/s | +10.7 t/s higher peak bursts | | MTP Draft Acceptance Rate | 63.7% to 65.1% (Tasks 740, 19637) | 75.2% to 81.7% (Tasks 0, 155) | +11.5% to +16.6% higher accuracy | | Mean Draft Length (mean len) | 2.91 to 2.95 tokens / round | 3.31 to 4.10 tokens / round | Up to +1.1 tokens / verification step | | Compute Time per Token | 21.84 to 22.25 ms / token | 18.62 to 19.45 ms / token | ~3 ms lower latency per token | | KV Cache Precision Evaluated 1| -ctk q8_0 -ctv turbo3 (3-bit V) | -ctk turbo5 -ctv turbo4 (4-bit V, higher precision) | ATX wins in speed despite higher KV fidelity | Real Execution Log Excerpts Standard Swift-1.5 Q5 (Symmetric Quantization) Task 740 (Context: 97,182 tokens | Generated: 1,130 tokens): eval time = 25123.57 ms / 1130 tokens (22.25 ms per token, 44.94 tokens per second) draft acceptance = 0.63746 (742 accepted / 1164 generated), mean len = 2.91 Task 19637 (Context: 92,075 tokens | Generated: 834 tokens): eval time = 18192.26 ms / 834 tokens (21.84 ms per token, 45.79 tokens per second) draft acceptance = 0.65130 (551 accepted / 846 generated), mean len = 2.95 ATX-Swift-1.5 Q5 (Asymmetric Custom Tensor Mapping) Task 155 (Context: 84,099 tokens | Generated: 3,478 tokens): eval time = 67619.68 ms / 3478 tokens (19.45 ms per token, 51.42 tokens per second) Burst Peaks: tg_3s = 60.18 t/s and tg_3s = 63.05 t/s draft acceptance = 0.73096 (2543 accepted / 3479 generated), mean len = 3.72 Task 100 (Context: 97,847 tokens | Generated: 525 tokens): eval time = 10175.17 ms / 525 tokens (19.42 ms per token, 51.50 tokens per second) Burst Peak: tg_3s = 58.10 t/s draft acceptance = 0.76939 (367 accepted / 477 generated), mean len = 3.31 Task 0 (Context: 80,016 tokens | Generated: 456 tokens): eval time = 8470.07 ms / 456 tokens (18.62 ms per token, 53.72 tokens per second) draft acceptance = 0.81710 (344 accepted / 421 generated), mean len = 4.10 Technical Takeaway for the Post Why ATX is ~15% faster under identical deep context: In standard quants, compressing the speculative head (blk.64) to Q5 causes ~36% of proposed draft tokens to fail rejection sampling, reducing throughput to ~45 t/s. In ATX-Swift-1.5, isolating blk.64 at Q8_0 increases draft accuracy from ~64% to ~77%+, delivering 3.31 to 3.72 verified tokens per round and raising sustained generation speed past 51.5 t/s (with burst peaks over 63 t/s). Optimized Execution Script (llamAmpere) Make sure to pass the explicit MTP vocabulary shortlist (atx_65536.txt). This restricts speculative draft projections to the top 65,536 power-of-two tokens, aligning perfectly with NVIDIA Ampere Tensor Cores and preventing a 73% compute penalty: cd D:\llamAmpere $env:GGML_Q8_TURBO3_MMA_FUSED = "1" .\build-sm86\bin\Release\llama-server.exe -m "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-i1-Q5_K_M.gguf" -ngl 99 -c 131072 -b 2048 -ub 512 -t 20 -tb 8 -fa on -ctk turbo5 -ctv turbo4 --kv-unified --prio 3 --parallel 1 --jinja --fit off --cache-prompt --cache-ram 4096 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.1 --spec-draft-type-k q8_0 --spec-draft-type-v q8_0 --spec-draft-vocab-map "D:\llamAmpere\docs\mtp-vocab\atx_65536.txt" --reasoning-format none --temp 0.2 --top-p 0.90 --top-k 40 --min-p 0.05 --repeat-penalty 1.08 --repeat-last-n 256 --alias "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-Q5_K_M" --host 127.0.0.1 --port 8080   submitted by   /u/bjivanovich [link]   [comments]