Strata 0.1.40.1 left ~10 GB of VRAM unused on my 4-GPU rig — I thought it was a bug. It isn't.

Wait 5 sec.

Running Qwen3.8 Flash-Next (125B MoE, IQ3_XXS) on the Strata engine across a mixed rig — RTX 5060 Ti + 3090 + 2x 3060, 262k context. After upgrading 0.1.39 -> 0.1.40.1, nvidia-smi showed a lot of VRAM just sitting there unused. My first thought: is Strata leaving VRAM on the table (a bug)? 1) v0.1.40.1 leaves ~10.7 GiB of VRAM unused (live nvidia-smi) GPU Total (MiB) Used Free RTX 5060 Ti 16,311 15,514 337 RTX 3060 12,288 5,580 6,332 RTX 3090 24,576 23,799 328 RTX 3060 12,288 7,912 4,000 Total 65,463 52,805 10,997 ~51.6 GiB used, ~10.7 GiB free of 63.9 GiB — and the free VRAM sits mostly on the two 3060s (6.3 + 4.0 GiB). So Strata really is not filling the cards. Bug? 2) It's caching ~24% fewer experts The model is fixed: 48 MoE layers x 512 experts = 24,576 cacheable experts (plus 48 always-on shared experts). "Resident experts" is how many Strata keeps on-GPU. Engine Resident experts (GSQ-RCO) Resident experts (orca) 0.1.36 23,354 - 0.1.38 23,170 22,131 0.1.39 23,168 22,145 0.1.40.1 17,515 18,586 That's -24% (GSQ-RCO) and -16% (orca) resident experts on 0.1.40.1 — which is where the free VRAM comes from. 3) ...but the cache hit rate barely moved, and generation actually improved Engine Cache hit rate Gen tok/s 0.1.36 99.7% 65.4 0.1.38 99.9% 69.2 0.1.39 99.7% ~87 0.1.40.1 98.1% 104 Hit rate dropped ~1.6 points while resident experts dropped 24%. Decode went up. Why it is not a bug (according to Flash Next) The experts Strata dropped are cold — the profile ranks experts by routed mass, and the tail carries