A few days ago I posted about sf-ds4-1flash, my fork of antirez's ds4 stripped down to one model (DeepSeek V4.1 Flash) and one backend (Metal). It streams the 341 GB Q2 GGUF from the SSD on a 128 GB M5 Max. Since then two fixes landed. Both are macOS gotchas anyone streaming weights from disk can hit, so here they are. 1. F_NOCACHE doesn't fully mean "no cache" on APFS Prefill reads expert weights with F_NOCACHE so they don't fill the page cache. Turns out on APFS, if a read doesn't start on a 16 KB page boundary, part of it still ends up cached. Experts sit every ~3 MB in the file, so they almost never start on a boundary. Quick test, 20 random expert reads: 982 of 3731 pages stayed cached. Same reads starting on a boundary: 0. Those "uncached" bytes pushed the decode weights out of RAM, so the first token after a long prompt had to page them back in. Fix: read the partial first page on its own, then the rest from the boundary. Same bytes, same addresses. first token after an 8K prompt: 1.23 s → 0.15 s appending 1.5K tokens to a chat: +12.5% 2. Don't wait for the command buffer, poll a mailbox On every layer of every token, the CPU needs the expert ids the GPU just picked. The command buffer's "Completed" status shows up ~48 µs after the GPU is done. If the router kernel writes the ids into a shared buffer and the CPU spins on it, they show up in ~20 µs. Tiny, but it happens on every layer of every token. decode: +5.5% at 2K context, +6.2% at 8K Side finding: a GPU write becomes visible to the CPU when its command buffer ends, not when its kernel does. Where it stands vs upstream ds4 (same GGUF, same flags, ds4's own bench): prefill: +45% to +99% decode: ~30 tok/s at 2K context, ~26 at 32K (+62% to +78%) first token after a prefill: 0.15-0.2 s vs 2-3 s Output is still bit-identical: logits are checked bit for bit after every change, and greedy output matches ds4 token for token. Caveat: one machine only. These numbers are from a cool Mac; after an hour of heavy use decode settles around 24 tok/s. If you have another Mac with lots of RAM and a fast SSD, I'd love to see your numbers. Repo: https://github.com/Chida82/sf-ds4-1flash Huge thanks to antirez and the llama.cpp/GGML folks. None of this exists without them.   submitted by   /u/Chida82 [link]   [comments]