DeepSeek V4.1 Flash on a single DGX Spark: 113.6 GB VQ base + 40 MB domain sidecars, 74–82% top-1 agreement vs original

Wait 5 sec.

I’ve been working on YoungAi, a native C/CUDA inference engine that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). The original weights are ~510 GB. I deploy it as three files: ① Base GGUF — 113.6 GB, universal, zero-corpus. Quantized once from official weights. ② Domain sidecar — ~40 MB per domain. Solved once per domain, then frozen. ③ Post-training file — experimental, re-solved nightly. Delete it to roll back. Each routed expert row is scaled by g_base × s_sidecar × s_posttrain, and the router gets a bias Δb_sidecar. The base alone is a complete model; sidecars just add tiny scaling/bias without changing kernels. TL;DR Single DGX Spark, 113.6 GB resident + ~40 MB sidecar. 5 domains: finance, code, law, medicine, science. Top-1 agreement vs original improves +2.8 to +3.7 points with a domain sidecar. Speculative decode: 43 tok/s on a real 14.1k-token Agent request (greedy). Prefill: 1,055 tok/s on 12.5k prompt; 671 tok/s on 106.7k prompt. English WikiText-2 does not regress when any domain sidecar is attached (it actually goes up). Core implementation ideas Base (VQ-8 + per-layer shared codebook). Every 8 consecutive weights in an expert row become one 12-bit (or 13-bit) codebook index, multiplied by a single per-row gain. Codebooks are trained per layer and shared across all 384 experts and three matrices. Codebooks are stored in FP8 (E4M3). 13-bit layers use a “12+1” bit-plane layout for 128-byte cache line alignment. The base is zero-corpus: it never sees domain data. Domain sidecar (“anti-solver”). For each domain, I solve a multiplicative gain per output channel of every expert’s down projection, plus a router bias per expert. Objective: reproduce the original model’s MoE block output on domain text, layer by layer, using the engine’s own prefill hooks. Gains are stored in FP4 with lattice-aware Gauss-Seidel. The sidecar is ~40 MB and adds ~0.6 MB read per decoded token. Post-training file (experimental). Turn “the model should write a, not b” into a linear equation on last-layer expert gains, then solve with conjugate gradient. On the training request, decision points flip from 55% to 88%, but it does not generalize across trading days yet. Multi-domain real metrics All metrics are teacher-forced against the original DeepSeek V4.1 Flash (official PyTorch code, full precision). Higher top-1 / Σmin is better; lower KL / PPL ratio is better. Domain (judgment slice) Base only top-1 + domain sidecar top-1 Σmin (median / p5) Avg KL PPL ratio Finance (8,192 tok) 71.73% 74.57% 0.745 (0.810 / 0.283) 0.512 1.267 Code (15,360 tok) 78.61% 82.26% 0.803 (0.872 / 0.402) 0.303 1.229 Law (15,360 tok) 72.90% 76.36% 0.760 (0.834 / 0.272) 0.454 1.251 Medicine (15,360 tok) 69.08% 72.82% 0.739 (0.767 / 0.325) 0.465 1.280 Science (15,360 tok) 71.65% 74.93% 0.753 (0.788 / 0.347) 0.422 1.173 English WikiText-2 (512 tok) 78.52% 80.66–82.81% (any sidecar) 0.787–0.801 0.566–0.619 1.564–1.658 English row shows that domain sidecars don’t hurt general ability; all five sidecars actually improve it slightly. Speed on one DGX Spark Scenario Prefill Decode 12.5k-token prompt 1,055 tok/s — Real 14.1k-token Agent request 940 tok/s — 106.7k-token prompt via server 671 tok/s (159 s TTFT) — Short prompt, pure greedy — 30.5–30.7 tok/s 14.1k-token Agent request, pure decode — 28.9–29.4 tok/s Same request, speculative (default) — 43.0 tok/s (3.04 tok/round) Unseen 9.2k prompt, speculative — 40.0 tok/s 51k context, pure decode — 27.5 tok/s Decode is memory-bound: ~6.3 GB read per token. GB10 measured bandwidth is ~235 GB/s, so the wall is ~37 tok/s; we hit ~32.5 ms, or 82% of the wall. Honest limitations Post-training (③) is a working mechanism, not a product yet. It flips specified decisions on the solving request but does not transfer to held-out days (55% → 55%). Five domains only. Sidecars are evaluated teacher-forced on held-out text, not yet end-to-end. CUDA only, validated only on DGX Spark. No Metal. Speculative decoding only kicks in for greedy; sampling requests fall back to pure decode. Source code (engine, quantizer, solver) is not public yet. Feedback welcome Are these top-1 agreement / Σmin numbers useful for real workloads? Is the VQ-8 + per-layer codebook + sidecar gain approach reasonable? What benchmarks or integration points would you want to see next? Model card and weights: https://huggingface.co/wenzhouwu/YoungAi-DeepSeek-V4.1-Flash This is not an official DeepSeek release. If this kind of post isn’t appropriate here, let me know and I’ll move or remove it. Thanks!   submitted by   /u/Physical_Toe_2499 [link]   [comments]