Details, full tables and raw data: https://github.com/ggml-org/llama.cpp/discussions/30071 Repo with binaries and docker images: https://github.com/lukolszewski/llama.cpp-multigpu I run Qwen3.8-Flash-Next on six 3090s over plain PCIe (no NVLink, some cards on x4 and x2 lanes), five sessions of 262k each. Stock llama.cpp got slower the deeper the context went and fell apart with several sessions decoding at once: 2.3 t/s per session at 5x250k. So I spent September fixing it. The patches sit on top of upstream df03399b8 and ship as tarballs (CUDA 12.9 for V100 to 5090, CUDA 13.4 for Ampere+) and ghcr images. Same GGUF, same llama-server, everything switched on by env vars. Same model (unsloth UD-Q4_K_XL), same command line, 5 slots x 262k, q8_0 KV, layer split. Tokens/s, upstream -> patched: workload ctx 6x3090 (mine) 6x4090 (rented) prefill, 1 session 250k 263 -> 2111 (8x) 744 -> 7403 (10x) decode, 1 session 250k 10.2 -> 33.7 (3.3x) 21.0 -> 48.7 (2.3x) decode, 5 sessions, each 250k 2.3 -> 27.3 (10.8x) not run -> 30.8 decode, 1 session 5k 38.5 -> 45.9 62.2 -> 62.8 The point is the shape: patched prefill is flat from 5k to 250k and decode barely drops, while upstream halves every 50k or so. At 5k with one user there is nothing to gain. The 10.8x is against a 2.3 t/s baseline, so do not quote that one. llama.cpp-multigpu is a temporary performance fork (until upstream catches up). Long-context decode is fixed for everyone, including single GPU; the multi-GPU part is for layer split over PCIe and is off unless you turn it on. What each patch does is in the repo. MTP: tried it, it was slower in most cases on this box, and the base commit predates upstream's MTP for this model anyway, so it is not included. N-gram lookup speculation instead: 2-2.5x on code rewrites and refactoring, 1.5x on code explanation, nothing on prose, and it switches itself off beyond two active users so the multi-user numbers do not suffer. The benchmarks above ran with it off. Caveats: tested with one model, CUDA only, written for slow PCIe, may regress NVLink or single-GPU boxes if you turn the multi-GPU switches on. Mixed prefill plus decode is better than upstream but still the weak spot and to be improved. The code was written with an LLM and validated by measurement and output checks (needle tests, temp-0 output identical), not by review, so I am not opening upstream PRs from it; each change is one commit and anyone can pick up any piece.   submitted by   /u/flynth92 [link]   [comments]