Qwen3.8-27B (Q6_K_XL) 110+ TPS at 256k context on a single RTX 5090, with a KV buffer decoupled from context size

Wait 5 sec.

I'm sharing this project (honestly 2nd time) for anyone who wants to run ultra-long context tasks with high-precision quantizations, especially for heavy coding. TL;DR: In vanilla llama.cpp, -c N allocates a physical KV buffer for all N tokens up front, including promprts and KV caches. In my fork (focus-llama) the logical context stays at 256k, but the physical KV buffer is capped (--kv-cache-size, ~85k cells). Older chunks are offloaded to an external store and pulled back on demand. This runs a 27B Q6 model with 256k logical context on one 5090, at roughly 100–120 t/s depending on how full the context is. How to? Vanilla llama.cpp allocates a contiguous KV buffer for the full -c (VRAM), and every decode step attends over all tokens currently in context, so per-token cost grows roughly linearly with context length (only the attention part; the weights matmul is constant). The two problems are separate: the allocation wastes VRAM, and the growing context slows decoding. Initial speed is around 120 t/s, dropping to 60 t/s as context grows. focus-llama attacks both: the physical buffer is capped (~85k cells) so VRAM is bounded, and since the number of resident cells can't exceed the buffer, per-step attention cost is bounded by the buffer size instead of the logical context length. It based on 2 techniques declarative attention and skill.state, introduced by google deepmind. I've spent the past two weeks ironing out bugs, and now that it has stabilized, I'm honestly blown away. On a single RTX 5090 + Qwen3.8-27B UD-Q6_K_XL, adaptive MTP speculative decoding), it shows average 110+ t/s and 256k logical context with a ~85k physical buffer. Configuration is somewhat tricky (focus-memory: kv cache store also required) but, You'll see the MAGIC in action: context usage stays capped at around 18–30%, token generation speeds remain consistently high, and you'll never hit full-day compaction pauses when running coding harnesses like Cline or Qwen Code. Link? https://github.com/edwardyoon/focus-llama/blob/master/README.md My conf: -m /home/edwardyoon/my_model/qwen3.8/Qwen3.8-27B-UD-Q6_K_XL.gguf \ --alias qwen27b \ -ngl 99 \ -b 1024 \ -ub 1024 \ -c 200000 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --da-auto \ --kv-unified \ --da-min-ctx 2048 \ --da-chunk-tokens 4096 \ --fm-offload \ --kv-offload-threshold 36864 \ --kv-offload-holes \ --kv-cache-size 85536 \ --kv-retain-tokens 6000 \ --sparse-gate-threshold 60 \ --focus-memory-host http://192.168.219.124:3900 \ --focus-memory-token focus-memory-local \ --spec-type draft-mtp-adaptive \ --spec-draft-n-max 6 \ --spec-draft-ngl all \ --spec-draft-type-k q8_0 \ --spec-draft-type-v q8_0 \ few journal logs: llama-server[827038]: da_sparse[VEC]: SPARSE - gather bound 5120 of 15872 KV rows (32.3% of KV) … n_gen = 361, tg = 119.44 t/s, tg_3s = 119.77 t/s n_gen = 729, tg = 120.98 t/s, tg_3s = 122.53 t/s n_gen = 1083, tg = 119.71 t/s, tg_3s = 117.18 t/s n_gen = 1497, tg = 123.98 t/s, tg_3s = 136.71 t/s n_gen = 1819, tg = 120.50 t/s, tg_3s = 106.61 t/s n_gen = 2157, tg = 119.06 t/s, tg_3s = 111.85 t/s n_gen = 2509, tg = 118.67 t/s, tg_3s = 116.32 t/s n_gen = 2837, tg = 117.37 t/s, tg_3s = 108.35 t/s n_gen = 3234, tg = 118.99 t/s, tg_3s = 131.94 t/s n_gen = 3643, tg = 120.65 t/s, tg_3s = 135.62 t/s n_gen = 3983, tg = 119.98 t/s, tg_3s = 113.25 t/s n_gen = 4350, tg = 120.09 t/s, tg_3s = 121.34 t/s n_gen = 4713, tg = 120.08 t/s, tg_3s = 119.89 t/s n_gen = 5150, tg = 121.80 t/s, tg_3s = 144.11 t/s n_gen = 5527, tg = 122.04 t/s, tg_3s = 125.37 t/s n_gen = 5866, tg = 121.45 t/s, tg_3s = 112.57 t/s n_gen = 6291, tg = 122.59 t/s, tg_3s = 140.83 t/s   submitted by   /u/Ok-Shower7286 [link]   [comments]