A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

Wait 5 sec.

I forked Strata and worked with Claude code with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V100 SXM2 GPUs, the CPU-GPU BW advertised as 150GB/s and the nvidia drivers do allow unified memory access. The machine I have has 4 x 16GB GPUs, llama.cpp had like terrible results before I started this journey, it produced 130tk/s prefill and 15tk/s decode. So I was fighting Claude the whole afternoon, beating it with facts and logic, like FP16 instead of BF16, memory management, expert caching on GPU, better NVLink usage, Tensor Core utilization rather than CUDA core ops. Claude was great at iterating, executing nsight nsys to debug time gaps. Prompt reading: 7,350 tok/s peak, still 7,090 tok/s on a 252K-token prompt (35 s) Generation: ~113 tok/s peak (JSON), ~100 on code, ~84 on prose (MTP speculative decoding) Follow-up at 252K depth: first token after 0.26 s, 60 tok/s All 72 GiB of experts page-locked in RAM across both sockets; GPUs pull from NVLink 2.0 at ~70 GB/s each I will try to contribute back some of the changes, but I suspect Strata will remain consume-hw-first inference engine, and that is totally ok, Niko1221 did a great job The forked repo: github.com/eelgaev/Strata-AC922   submitted by   /u/okoyl3 [link]   [comments]