TL;DR: Qwen3.8 27B on a single AMD Radeon AI PRO R9700 (32 GB, 300 W), with speculative decoding and our 3-bit weights (not a blanket 3-bit quant, see below), now keeps 569,878 tokens of reusable cache in its new coding mode. Every request gets 262,144 tokens of context, and two full-length requests fit at once. Coding agents resend the whole conversation on every turn; now only the new part is read. Later agent turns start up to 12× sooner, whole agent sessions finish about 6× faster, and a 258K-token document the server has already read comes back in 2.7 s instead of 134 s. Decode speed and accuracy are within noise of our previous release. Prefer 4-bit? MXFP4 is still one command away. "3-bit? Pass." Fair. Here's what it actually is It's not a blanket 3-bit quant: Only the large projection matrices are 3-bit: MLP, attention and the recurrent (Gated DeltaNet) projections, 24.3B of the 27B parameters. Each block of 128 weights gets its own scale, about 3.1 bits per weight in total. Everything else is not 3-bit: the embeddings, the output head, the norms and the other recurrent-layer parameters come from our MXFP4 release. The matrices are rotated before quantizing: a rotation spreads out the outliers that usually wreck low-bit weights. They're then GPTQ-calibrated on ~293K tokens of permissively licensed data. Token generation runs with 8-bit (FP8) activations. Reading the prompt uses 4-bit activations on the rotated matrices (the "A4" in W3A4). The accuracy numbers below include both. Everything is published: weights, calibration sources and per-shard hashes are on Hugging Face. What you trade, same card, default 65K mode: MXFP4 (run-mxfp4.sh) 3-bit (run-3bit.sh) Decode, single stream 156.1 tok/s Aggregate, 8 requests 428.0 tok/s KV cache on the card 174,634 tokens Longest request 200,000 (220,000 tested), one at a time Prefix cache for agents 200,000 tokens, one request GSM8K 5-shot (1,319) 95.68 HumanEval pass@1 (164) 95.12 MMLU-Pro subset (1,400) 62.57 We also paired the two per question, both with the FP8 cache. Only the MMLU-Pro gap was beyond noise (−2.9 points, 95 % CI −4.8 to −0.9). hat's knowledge recall, the usual cost of fewer bits; math and code stayed within noise. So 2–3 points of MMLU-Pro buy you 15 % more speed, 2.25× the cache and 262K context per request. If knowledge recall matters most to you, run MXFP4: it's also still there, it's still our most accurate option, and it uses the same download. The 3-bit weights are a 9.55 GB add-on, so you can run both and judge on your own work. You asked for it on r/ROCm: more context In our r/ROCm threads you asked for more context. Well, now you got more than half a million tokens of cache on one card, with multi-hour accuracy runs on every configuration (GSM8K, HumanEval, MMLU-Pro, needle tests), at essentially the same speed as our previous release. Previous release (2 Oct) This release, --mode long-kv4 Reusable (prefix-cached) tokens in the 4-bit mode 0 (no prefix cache) Tokens per request 262,144 Full-length requests at once 1 Decode, single stream 174.3 tok/s Aggregate, 8 requests 462.4 tok/s Time to first token, short prompts (p50) 86.7 ms GSM8K / HumanEval / MMLU-Pro 95.53 / 92.07 / 60.57 Our previous release's FP8 --mode long already had a prefix cache, holding 281,665 tokens. The new 4-bit coding mode caches about twice as much. Run it git clone https://github.com/Eliovp-BV/paiton-vllm-plugin && cd paiton-vllm-plugin # set the PAITON_* model paths as shown in the README (already set up? just `git pull`) bash models/Qwen3.8-MXFP4-DFlash2/run-3bit.sh --mode long-kv4 Point your agent at http://127.0.0.1:18982/v1. The model name is Qwen3.8, any API key works, and the context window is 262,144 tokens. Already running our 3-bit weights? No new download. The launcher pins the new container and Docker pulls it on first start. Every mode is in the README. No new engine This is stock vLLM with a plugin: the same OpenAI-compatible server, the same API, and the same tools and workflows you already use. Nothing to relearn, nothing to migrate. We will keep publishing new ready-built containers that follow upstream vLLM. Each one goes through the same full validation (speed, accuracy, long-context) before it ships. You get upstream improvements without building anything yourself. One command and you're serving. What it does for coding agents Without a prefix cache, the server rereads the whole prompt on every turn: at 250K tokens that is about two minutes before the first token. Now finished requests stay cached until the space is needed, turn 20 reads only what changed, and in our sessions 91–92 % of prompt tokens came from the cache. Workload (3-bit, thinking off) Previous 4-bit mode (no prefix cache) --mode long-kv4 Faster 20-turn conversation growing from 50K to 253K tokens, whole session 1,237 s 204 s 6.1× … average wait for the first token, turns 2–20 63.6 s 7.8 s 8× 3 agents sharing a 100K-token repo, 5 turns each, whole session 655 s 111 s 5.9× … average wait for the first token, later turns 84 s 7.0 s 12× New question about a 258K-token document already read 134 s 2.7 s 49× The cached answer to the 258K-token document matched the cold read exactly. Speed details (BetterBench, full 20-pass run): decode within 0.5 % single-stream; +3.5 % with 8 requests (−2.5 % at 4); time to first token unchanged. A new long prompt reads about 3 % slower, and short prompts 7–13 % slower (a fraction of a second). That's the price of ending each prefill step where a later request can pick up from the cache. Accuracy (greedy, paired per question): the coding mode is within noise of the previous release, and the 512K mode is within noise of the coding mode. MXFP4 Previous release (3-bit) --mode long-kv4 --mode long-512k GSM8K 5-shot (1,319) 95.68 95.53 95.45 HumanEval pass@1 (164) 95.12 92.07 92.68 MMLU-Pro subset (1,400) 62.57 60.57 59.93 Opt-in: 524,288 tokens in one request --mode long-512k extends the position range with the model's official long-context scaling. That scaling applies to every request in this mode, so use it only when a single request needs more than 262K. It found 4/4 planted facts at 300K and 4/4 at 500K tokens. A cold read takes 158 s at 300K and about 6 minutes at 500K. Follow-up questions take 2.8–4.4 s from the cache. The cache holds 594,290 tokens, and weighted decode is within 0.6 % (in one run the chat category was 16 % slower). Unchanged MXFP4 (run-mxfp4.sh) is still there and still the most accurate option. --vision works in the default 65K mode and in --mode long (images with up to 245K tokens of context). The previous release is one --image flag away (see the README). Caveats, honestly System RAM: the coding and 512K modes pin 2.4 GiB of system RAM for the embedding table, and 16 GB of RAM is enough (tested). Below about 13.5 GiB of total RAM, --mode long-kv4 keeps the table on the GPU and caches 451,879 tokens. You still get 262K per request and the prefix cache. --mode long-512k needs the RAM. Cold reads: the first read of a 258K-token prompt still takes over two minutes. The cache helps from the second request on, as long as the start of the prompt stays the same (same system prompt, no timestamp at the top). usage.prompt_tokens_details.cached_tokens shows every hit. Vision isn't in the coding or 512K modes yet. Use --mode long --vision for images with long context. RAM/SSD cache tier: the experimental system-RAM tier for the prefix cache (--host-cache-gib, plus an SSD tier behind it) spills even more context off the card. For now it works with the FP8 --mode long; support for the 4-bit coding mode lands in the next release. The coding mode's 569,878-token GPU cache works today. The agent-session numbers and the accuracy columns were measured on pre-release builds of these configurations; the README has every number. Our testbench is ancient, slow CPU, limited and slow memory (16GB), so your results will most probably be even better! What's next: more GPUs More R9700s are on the way, and we're going after tensor parallelism next (and other models). Others are working on multi-card setups too, so expect some healthy competition on that front. Good for everyone running AMD at home! More: Paiton   submitted by   /u/evp-cloud [link]   [comments]