I've had one pull request merged into ds4 (DwarfStar), a tiny one. There are a few more still waiting in the queue. I’m not complaining. Antirez says it clearly in the README: with coding agents everyone can tune the engine for their hardware and model and he can’t review everything. That made me think. If the plan is that everyone applies their patches using an agent then the real cost of a patch isn’t just the code change. It’s also how tokens the agent has to read before it knows what it’s actually touching. The ds4 codebase runs DeepSeek, GLM and Qwen on Metal, CUDA and ROCm—all in a 85k-line file. I'm running Qwen3.8 Flash Next on an M5 Max with 128GB RAM. Everything else in that file is noise for me and for my agent.. Every time the agent runs it has to re-read all of it. So I ripped it out. I didn’t just ifdef it. I deleted it. The ds4.c file went from 85k lines down to 45k. Now the entire code tree fits inside a context window. Metal is the production backend now. The CPU path is kept as a reference for tests. My guess was that making the codebase smaller would make optimizing cheaper and safer. Here's what happened: Q2: decode speeds up by 9–13% prefill improves by % (up to 64k context) and MTP goes from 75.8 to 86.7 tok/s Q4: prefill gains 2–11% MTP rises from 77.8 to 85.9 tok/s Output stays bit-exact compared to stock ds4 at every step. No KV cache quantization. No approximate kernels. Every change must pass a parity check— GGUF, greedy decoding identical tokens—plus an interleaved A/B benchmark against the previous build. The smaller codebase also let me go through the PRs in ds4. I tested them against my version of the model and ported the ones that worked. Twenty commits were adopted. Around thirty were dropped. The results are in the repo. I also added SSD streaming for the experts. It matches a resident run token-for-token. On a simulated 48GB machine Q2 runs at 27 tok/s. With MTP it reaches around 35 tok/s. The fork still keeps up with upstream. It runs git merge upstream/main with rerere plus the parity check. So antirez’s fixes keep flowing in. The whole process—what to delete, what to keep how to sync—lives in a repo called StarForge. I have four of these "children," one for each model. Nothing in StarForge depends on Qwen or Metal. If you want a cut-down ds4 tailored to your model or to CUDA just clone it and run the checklist with your agent. Repo: sf-q3-8flash with tables in the README. This setup uses one machine and one model. If you’re on Apple Silicon I’d love to see your numbers, ideally side by side, with stock ds4.   submitted by   /u/Chida82 [link]   [comments]