Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp

Wait 5 sec.

This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab. LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no account and no install. The new part: two engines that stream mixture-of-experts weights from disk while they generate. That lets a tab run models bigger than the machine's RAM. Live: https://localmind.naklitechie.com · Code (MIT): https://github.com/NakliTechie/LocalMind All numbers are from one MacBook M4 Pro (24 GB) in Chrome. How it works On first load the GGUF is copied into OPFS, the browser's private file system. Dense weights, routers and the KV cache go to the GPU. Routed experts stay on disk. A pool of workers reads them on demand with sync access handles into a GPU slot cache (LRU, two layers of prefetch). The trunk kernels are hand-written WGSL that follow llama.cpp's graphs. That lets me test against llama.cpp on the exact same GGUF. Gemma 4 26B-A4B (Google's QAT Q4_0, 14.4 GB) Same output as llama.cpp b9830 Metal: the live site's chat replies were character-identical on 9/9 test conversations (capped at 64 tokens). 15/16 fresh prompts matched token for token. The 16th split on a 0.00009-nat near tie, where llama.cpp's own two attention paths also disagree. Memory: the Chrome GPU process sits at 6.9 GB with a 4 GB expert cache. About 8.6 GB of experts stay on disk. Speed: 23.6 tok/s decode, 55 tok/s prompt processing. llama.cpp Metal does 70.6 and 204 on the same Mac, so the tab is ~3× slower at decode. Per token: ~22.5 ms GPU compute, ~11 ms routing round trips, ~8–13 ms SSD reads. First load from the site: 11.5 min (14.4 GB download). After that: 1.6 s. Qwen3.6 35B-A3B (unsloth Q8_0, 36.9 GB, on a 24 GB Mac) — experimental The file is bigger than the machine's memory. The GPU process measured 7.3 GB with a 4 GB expert cache. Live site: 9.9 tok/s decode, 2.2 s to first token. First load is 36 min (download plus the OPFS copy). Output matches llama.cpp Metal 8/8 on 4- and 16-layer cuts. On the full model it matches llama.cpp CPU 5/8; the other 3 swap near-tie tokens. I can't run the full file on llama.cpp Metal on this Mac, so full-model parity is still open. Per token (~99 ms): ~23 ms GPU compute, ~39 ms routing round trips, ~35 ms expert reads from the SSD. Moving routing onto the GPU gave no gain (10.3 vs 10.3 tok/s): the misses are experts nobody predicted. Also Gemma 4 E2B can keep its 1.2 GB per-layer embedding table on disk: GPU process 4.27 → 2.07 GB, identical output, 3–8% slower decode. It's a setting, off by default. The whole app is one index.html again (854 KB with brotli). Engines, workers and the disk tier are rolled into it, and the tab builds them from blob URLs. The disk tier is also a standalone library: diskformer.js. Prior art As far as I can find (searched 6 Oct 2026), no earlier browser engine reads weights from disk during generation. wllama and LlamaWeb stream from OPFS only at load. Pooled runs Qwen3.6-35B-A3B in a browser with experts paged from system RAM. On-demand disk reads exist in native runtimes: llama.cpp's --moe-stream PR (#25294) and Google's LiteRT-LM for Gemma's per-layer embeddings. Corrections welcome. The Gemma 4 E2B kernels are webml-community's (Xenova and the Transformers.js team). My part there is the disk path. Limits Chrome or Edge with WebGPU. Tested on one M4 Pro 24 GB only; 8 and 16 GB machines are untested. Not faster than native: llama.cpp is ~3× faster on Gemma 26B. The point is that a tab can run these at all, with the same output. Parity covers greedy decoding, the prompts listed above, and 64 tokens each. I haven't tried llama.cpp's expert-offload flags (-ot exps=CPU) for comparison. If you have an NVIDIA/AMD GPU or a 32–64 GB Mac, I'd like your tok/s numbers. A bigger expert cache should move the Qwen3.6 number the most.   submitted by   /u/naklitechie [link]   [comments]