​ I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster. Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive. Performance Benchmarks Coding Generation / Decode: Sustaining ~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks. Prefill Throughput: ~750–910 tok/s (optimized via modified NCCL topology and dual-switch setup). Concurrency & Stress Testing: Handling multiple concurrent user requests smoothly without dropping token generation rates or starving KV cache memory. Context / Tool Bench: Stable multi-hundred-thousand token context runs agentic workfows with multiple 500k compaction. Compute: 16x GB10 Cluster Nodes Connectivity: Dual MikroTik Switch (CRS804-4DDQ) using 4x 400G-to-4x100G breakout cables. Runtime: Customized gb10-vllm stack using dspark / Inferact/Kimi-K3-DSpark wrappers with custom MLA/KV kernels. I attached a short clip showing real-time token streaming, coding output. GitHub & Setup Files: All runtime patches, config files, and build scripts are on my GitHub: 👉 https://github.com/ciprianveg/gb10-vllm   submitted by   /u/ciprianveg [link]   [comments]