Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.

Wait 5 sec.

​ ​I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster. ​Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive. ​Performance Benchmarks ​Coding Generation / Decode: Sustaining ~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks. ​Prefill Throughput: ~750–910 tok/s (optimized via modified NCCL topology and dual-switch setup). ​Concurrency & Stress Testing: Handling multiple concurrent user requests smoothly without dropping token generation rates or starving KV cache memory. ​Context / Tool Bench: Stable multi-hundred-thousand token context runs agentic workfows with multiple 500k compaction. ​Compute: 16x GB10 Cluster Nodes ​Connectivity: Dual MikroTik Switch (CRS804-4DDQ) using 4x 400G-to-4x100G breakout cables. ​Runtime: Customized gb10-vllm stack using dspark / Inferact/Kimi-K3-DSpark wrappers with custom MLA/KV kernels. ​I attached a short clip showing real-time token streaming, coding output. ​GitHub & Setup Files: All runtime patches, config files, and build scripts are on my GitHub: 👉 https://github.com/ciprianveg/gb10-vllm   submitted by   /u/ciprianveg [link]   [comments]