A 341 GB DeepSeek on a 128 GB Mac: 2x decode and first token in 0.2 s instead of 2.5 s, streaming from two SSDs, same tokens as stock ds4. The trick wasn't a kernel: I made the codebase small enough for an agent

Wait 5 sec.

DeepSeek V4.1 Flash at Q2 is 341 GB on disk: 152 GB of weights plus 189 GB of Engram tables. My Mac is an M5 Max with 128 GB. It runs anyway, because ds4 streams the experts from SSD, and stock ds4 gave me around 12 tok/s on the CLI. Usable. But I had the feeling the SSD wasn't the only thing holding it back, so I started poking at it with a coding agent, and that turned into something bigger than I planned. The first thing I noticed is that every session began the same way: the agent re-reading huge chunks of ds4.c (85k lines, three model families, three GPU backends) to figure out which 2% of it my model actually goes through. Most of what it read was about hardware I don't own and models I don't run. So I deleted all of it. Not #ifdef, deleted. ds4.c is 34k lines now, the whole tree 150k instead of 278k. Only DeepSeek V4.1 Flash, only Metal. That changed the economics of trying things. An optimization attempt that used to cost me an afternoon of the agent wandering around now costs maybe an hour, so I tried a lot more of them and measured every single one instead of picking the three I believed in. One rule the whole time: output doesn't change. Every change has to produce the same tokens as upstream ds4 on the same GGUF (greedy, ten prompts), and pass an A/B/B/A bench against the previous build with the logits compared bit for bit. No KV quant, no approximate kernels. If it's faster but a logit moved, it doesn't go in. This is where it landed, internal SSD only, same GGUF, same flags, upstream's own bench (ds4 → fork): - generation, ctx 2048 (128 tokens, first one included): 12.1 → 24.4 tok/s - steady decode, ctx 2048: 15.7 → 25.7 tok/s - steady decode, ctx 32768: 15.7 → 22.0 tok/s - prefill, 16k → 32k context: 404 → 636 tok/s - first token after a prefill: 2.1–3.0 s → 0.3–1.1 s None of it is clever. Decode layers get committed to the GPU without waiting for each other. Expert reads are split across a thread pool and the cache slabs sit in a Metal residency set. A handful of kernel fusions per decode token. Prefill reads the next layer's experts while the current layer computes. Individually each one is a small diff you can read in a few minutes. Together they double the speed, and you get all of this with the Mac as it is, nothing to buy. Then I got curious about the SSD part. Streaming is bound by read bandwidth and a Mac has exactly one internal drive, so I put a byte-identical copy of the GGUF on an external Thunderbolt 5 SSD and made prefill read part of every layer from each drive at the same time. The engine checks the copy against the model at every start (about 7 s) and refuses to run if anything differs, because I don't trust myself to keep two 341 GB files in sync by hand. Internal SSD only → with the external copy: - 3.5K-token prompt, time to first token: 13.2 s → 11.2 s - 10K-token prompt: 29.0 s → 25.0 s - +1.5K tokens appended to a 5.3K chat: 8.6 s → 7.0 s - first token after an 8K context: 1.38 s → 0.18 s Decode doesn't change, it never reads the copy. My enclosure also runs the drive at PCIe 4.0 x4, about half of the internal SSD, so a better enclosure should do better than this. Nice side effect: the KV cache can write to the external drive, so the soldered internal SSD takes zero writes while the model runs. Again: this part is optional, the table above it is the one that matters for most people. About staying in sync with ds4, because that was my main worry: the fork never renames the ds4_* files, every cut is marked in the source at the exact spot, and git merge upstream/main with rerere replays the conflict resolutions. After each merge the parity check tells me if the tokens still match. So far antirez's fixes have kept flowing in without drama. Not everything worked. I tried to bring the two-SSD trick back into upstream ds4 through its mmap path: bit-exact, but prefill got 13–38% slower and I still don't know why, so no PR for now. Twenty-odd other ideas were measured and dropped. I keep all of them in a "rejected ideas" table in the repo with the numbers, mostly so the agent (and I) stop re-proposing the same thing every other week. There's a growing trend of single-model inference engines, and ds4 itself started that way. This is just that idea pushed a bit further, one model and one backend, and at least here it holds up: faster, still correct, still merging upstream. I've done four of these forks, one per model; the procedure is in a separate repo (StarForge) and has nothing DeepSeek- or Metal-specific in it. Repo: github.com/Chida82/sf-ds4-1flash. The README and speed-bench/perf-record.md have the conditions behind every number. A few things I'd like to hear opinions on: Does "per-model fork that merges upstream" scale past a handful of forks, or is it just fragmentation with extra steps? Is bit-exact the right bar? I left speed on the table by refusing KV quant. Would you take 10% more for a slightly different token? If you have a 96–128 GB Apple Silicon Mac, I'd love to see stock ds4 and this side by side on your machine. One machine is an anecdote.   submitted by   /u/Chida82 [link]   [comments]