I've been tinkering with GLM-5.3-Flash on two 64GB CMP 170HX cards for a while, and the setup is finally stable enough that I figured I'd share it. I also compared it against the Qwen3.8-Flash-Next setup I've been using on the same machine: AWQ INT4 + FP8 PLE on vLLM. Besides PP/TG benchmarks, I hooked both models up to DSH and gave them the same small coding/agent tasks to see how raw inference speed translated into actual task completion time. A few caveats up front: this is not an apples-to-apples quant comparison GLM and Qwen are using different engines and different speculative decoding setups speculative decode speed depends heavily on acceptance rate and generated text the coding tasks are just a few practical examples, not a serious benchmark suite when I mention “Strata-style” below, I mean the HBM-first/full-residency approach I previously used with Strata, not that this is running Strata itself Repo and playable demos: GitHub: https://github.com/Flun/glm53-flash-cmp170hx-exl3 Hardware CPU Ryzen 5 5600X RAM 80GB DDR4 GPU 2x CMP 170HX 64GB GPU arch SM80 PCIe Gen2 x8 GPU P2P unavailable OS Ubuntu 24.04 Both models were tested on the same machine. For the agent tests I used DSH as the harness. GLM-5.3-Flash setup Target model: turboderp/GLM-5.3-Flash-exl3 3.05bpw Engine: ExLlamaV3 1.5.4 Current setup: GLM-5.3-Flash EXL3 3.05bpw ~125.2GB / 116.6GiB target weights target fully resident across the two 64GB cards k_hcfuse DFlash2 EXL3 6bpw DFlash2 K7 Q8 KV cache 384K context in actual use max request budget around 392,960 tokens GLM-5.3-Flash itself is a 320B-total / ~18B-active MoE model. The important part here is that the target weights stay resident in HBM. I'm not continuously streaming experts from system RAM during decode. About the 3.05bpw quality This was probably the part I cared about most. At first glance, “3.05bpw” sounds like a pretty aggressive quant, especially compared to the UD Q4 variants people commonly use. But EXL3 isn't simply “make every tensor 3-bit”. It uses a trellis-based quantization scheme with different bit allocation depending on the tensor. The 3.05bpw number is an average target bitrate. The published quant configs for this family also keep more sensitive parts at higher precision. For example, lm_head remains at 6-bit in the 3.05bpw branch. Looking at the same model family, the 4.05bpw build has been inspected with something roughly like: routed experts: K4 attention: K6 shared experts: K6 dense MLP: K5 lm_head: K6 embedding / norms / router: native So the general idea is to compress the huge routed-expert portion more aggressively while spending more bits on the smaller/more sensitive paths. That makes quite a bit of sense for a MoE model like this, because most of the storage is in the expert weights. Rough comparison with the common UD quants Quant Size Top-1 agreement vs BF16 Mean KLD UD-IQ3_XXS 120.37GB 81.63% 0.28377 EXL3 3.05bpw (my target) 125.18GB ~93.05% (estimated from published 3.0bpw results) ~0.050 (estimated from published 3.0bpw results) UD-IQ4_XS 156.82GB 88.18% 0.11665 UD-Q4_K_XL 199.71GB 92.22% 0.04929 UD-Q5_K_XL 240.31GB 94.35% 0.02705 For reference, public GLM-5.3-Flash GGUF fidelity numbers look roughly like this: One thing worth pointing out is that something named Q4_K_XL is not literally “4 bits per parameter across the entire model”. At ~200GB for a 320B model, it's a mixed-precision quant with an effective average bitrate much higher than 4bpw. There is also a published GLM-5.3-Flash EXL3 3.0bpw fidelity test using 51,175 held-out next-token positions that reported: Top-1 agreement: ~93.0% Mean KLD: ~0.0505 Numerically, that's in roughly the same neighborhood as the published UD-Q4_K_XL result. That said, I would not claim that “EXL3 3bpw is better than UD-Q4_K_XL” from those numbers alone. They were not measured through the exact same evaluation pipeline/corpus, and the public 3.0bpw artifact isn't the exact same quant I'm running either. My takeaway is simply that 3bpw-class EXL3 can preserve a surprising amount of fidelity for its size, and it doesn't behave like a naive 3-bit quant. For my use case, getting the target down to ~116.6GiB while still retaining usable coding/agent quality was the main reason this setup was interesting. DFlash2 6bpw is only the drafter Just to avoid confusion: the 3.05bpw model is the actual GLM target. The 6bpw DFlash2 model is only the speculative drafter. The drafter proposes tokens, and the GLM target verifies them. So this is not some kind of “3.05bpw + 6bpw averaged quality” setup. The draft quant mostly affects draft speed, VRAM use, and acceptance efficiency. HBM-first / “Strata-style” part This is where I borrowed an idea from a Strata setup I had used previously. Again, this does not run Strata. What I mean by “Strata-style” is simply: keep as much of the model permanently resident in HBM as possible, and avoid runtime CPU↔GPU weight traffic The GLM target itself fits across the two cards, so I leave the target fully resident. Instead of offloading experts, I focused on reducing the memory used by the parts that scale with context: KV cache and the speculative drafter. For this particular machine that made more sense to me than constantly moving weights over PCIe. DFlash2 + 384K context I originally used GLM's MTP d2 path. Later I switched to DFlash2. Instead of keeping the original BF16 incoai/GLM-5.3-Flash-DFlash2 drafter, I converted it to an ExLlamaV3-compatible EXL3 6bpw build. The resulting draft weights are about 0.96GiB. The bigger problem at long context was actually the draft KV cache. If the target is running 384K and the drafter also grows a 384K KV cache, VRAM disappears quickly. So I changed the drafter side to use a fixed SWA window plus a GPU ring cache. The target still sees the full 384K context and keeps its full target KV. Only the drafter's KV storage is kept inside a bounded ring. That's what lets the current setup run: DFlash2 K7 + Q8 KV + 384K target context without growing the draft cache to the full target length. The implementation and validation tests are in the repo. Cold start I also measured from a cold compile/start until the API was actually ready. Model Ready time GLM-5.3-Flash EXL3 ~1m 04s Qwen3.8 Flash Next / vLLM ~3m 50s This isn't really a model-size comparison. The Qwen vLLM setup has quite a bit more startup work: PP workers distributed runtime model placement MTP PLE GDN Triton compilation memory profiling KV allocation The ExLlamaV3 GLM path is comparatively static. Inference benchmarks These are the numbers from my dashboard workload. Again, especially for speculative decode, I wouldn't treat these as universal model speeds. Acceptance rate and generated text matter a lot. GLM-5.3-Flash / DFlash2 K7 / Q8 Input PP Decode 8K 1,529 tok/s 95.6 tok/s 40K 1,624 90.7 73K 1,647 90.6 106K 1,650 91.0 131K 1,611 94.4 385K 1,535 90.1 DFlash acceptance on this particular workload was mostly around 87%. With the older MTP d2 path, the same dashboard workload was generally in the ~60 tok/s range. Switching to DFlash2 K7 brought it to around ~90 tok/s here. Qwen3.8 Flash Next / vLLM The Qwen setup is: AWQ INT4 + FP8 PLE / PP2 / MTP3 Input PP Decode 8K 5,513 tok/s 129.5 tok/s 40K 5,628 123.6 73K 5,466 141.9 106K 5,307 141.5 131K 5,183 159.9 252K 4,706 149.4 So on raw throughput, Qwen is clearly faster. At roughly 131K: PP: ~5.18K vs ~1.61K decode: ~160 vs ~94 tok/s No argument there. The interesting part for me was what happened once I actually let both models do multi-step coding work. Why I stopped at 384K for now I tested roughly 385K input and PP was still around 1.5K tok/s. The problem wasn't PP collapsing. It was simply wall-clock time. Prefilling ~385K from scratch already takes about 4 minutes. Even if I can make 1M fit, doing a full 1M cold prefill at this speed isn't particularly attractive for normal use. So I'm currently leaving the service at 384K Q8. I still want to see if I can get 1M working eventually, mostly for the technical exercise. Small agent tests Originally I was only going to make both models build Tetris and stop there. Both were connected to DSH and got the same request. 1. Tetris Prompt: Build a playable Tetris game for the web. Model Completion time GLM-5.3 4m 40s Qwen3.8 5m 30s Both produced working versions, and honestly the difference wasn't dramatic enough to be very interesting. So I added two more tasks. https://reddit.com/link/1x0b1ws/video/1sn8netal4uh1/player 2. AI mini PC landing page Exact prompt given to both: Build a polished single-file HTML landing page for an AI mini PC with a dark theme, specs, performance charts, pricing, FAQ, and smooth scroll animations, using no external libraries. Model Completion time GLM-5.3 10m 05s Qwen3.8 14m 40s https://reddit.com/link/1x0b1ws/video/pdglw7kbl4uh1/player 3. Vampire-Survivors-style game Exact prompt: Build a single-file HTML vampire-survivors-style game with WASD movement, auto-attacks, enemy waves, XP, 3-choice level-up upgrades, HP, game over, and restart, using no external libraries. Model Completion time GLM-5.3 4m 40s Qwen3.8 24m 10s This one had a much larger difference than I expected. There is one obvious caveat: the Qwen version added sound, while the GLM version did not. The prompt didn't ask for sound, so I didn't go back and ask GLM to add it afterward. I wanted to leave both runs as the result of the same one-shot prompt. https://reddit.com/link/1x0b1ws/video/4hl6g2bcl4uh1/player Task completion times Task GLM-5.3 Qwen3.8 Tetris 4m 40s 5m 30s Landing page 10m 05s 14m 40s Vampire-style game 4m 40s 24m 10s I wouldn't read too much into three examples. This definitely isn't evidence that GLM is “5x better at coding” or anything like that. What I found interesting is simply that raw tok/s and end-to-end agent completion time didn't track each other very well. Qwen has much higher PP and decode throughput, but on these particular tasks GLM often finished sooner. For agent work, planning, number of retries, file rereads, edits, and how close the first implementation is to working all matter too. So I think raw inference speed and actual task completion time are worth looking at separately. Current state The GLM service I'm using now is: GLM-5.3-Flash EXL3 3.05bpw DFlash2 EXL3 6bpw K7 Q8 KV 384K context** The main thing I like about this configuration is the memory/quality tradeoff. The target fits in ~116.6GiB of HBM, stays resident, and the public 3bpw-class EXL3 fidelity results suggest the quant is holding up much better than I would have expected from the bitrate alone. The runtime side is basically an HBM-first setup: keep target weights resident, then save memory on the drafter/KV side rather than moving experts back and forth during decode. Full config, conversion scripts, ring-cache changes, benchmark code and raw results are here: GitHub: https://github.com/Flun/glm53-flash-cmp170hx-exl3 If anyone is running GLM-5.3-Flash on other weird 128GB-class GPU setups, I'd be interested in seeing what numbers you're getting too. Next thing I want to try is 1M context, although at that point prefill time is probably the bigger problem than just making it fit.   submitted by   /u/Prudent_Appearance71 [link]   [comments]